# Vespa Blog

We Make AI Work

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Deploy, Discover, Inspect, Observe: A Summer Spent Making a Public Vespa MCP Server

DevFeed: [Deploy, Discover, Inspect, Observe: A Summer Spent Making a Public Vespa MCP Server](<https://devfeed.tech/articles/deploy-discover-inspect-observe-a-summer-spent-making-a-public-vespa-mcp-server-12795.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/public-mcp-interns/>)

Author: eivinbingen oystein viktor mfstort

Published: 2026-09-13T00:00:00Z

Content type: article

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [Model Context Protocol (MCP)](<https://devfeed.tech/topics/model-context-protocol-mcp.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>)

Tags: [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-assistants](<https://devfeed.tech/tags/ai-assistants.md>), [claude](<https://devfeed.tech/tags/claude.md>), [cli](<https://devfeed.tech/tags/cli.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [codex](<https://devfeed.tech/tags/codex.md>), [internships](<https://devfeed.tech/tags/internships.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>)

### AI overview

This article describes the construction of a standalone, publicly hosted Vespa Cloud MCP server. It explains how MCP connects AI assistants and language models to external systems through resources, tools, and prompts, and discusses evaluating MCP usage against terminal access and Vespa CLI access.

### Source excerpt

We built a standalone Vespa Cloud MCP server as a summer interns project

## Vespa Newsletter, September 2026

DevFeed: [Vespa Newsletter, September 2026](<https://devfeed.tech/articles/vespa-newsletter-september-2026-12801.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/vespa-newsletter-sept-2026/>)

Author: Bonnie Chase

Published: 2026-09-08T00:00:00Z

Content type: news

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [ann](<https://devfeed.tech/topics/ann.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Provisioning](<https://devfeed.tech/topics/provisioning.md>), [Algorithm](<https://devfeed.tech/topics/algorithm.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Graphs](<https://devfeed.tech/topics/graphs.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [algorithm](<https://devfeed.tech/tags/algorithm.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [features](<https://devfeed.tech/tags/features.md>), [graph](<https://devfeed.tech/tags/graph.md>), [latency](<https://devfeed.tech/tags/latency.md>), [newsletter](<https://devfeed.tech/tags/newsletter.md>), [product](<https://devfeed.tech/tags/product.md>), [provisioning](<https://devfeed.tech/tags/provisioning.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [september-2026](<https://devfeed.tech/tags/september-2026.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [updates](<https://devfeed.tech/tags/updates.md>)

### AI overview

The September 2026 Vespa newsletter announces updates including time-constrained ANN search, sub-query ranking support, flexible provisioning, new rank features, and telemetry export. It also introduces Vespa.ai Live, an in-person community meetup focused on retrieval and ranking systems.

### Source excerpt

Advances in Vespa include time-constrained ANN search, sub-query ranking support, flexible provisioning, new rank features and telemetry export

## The Human Context Advantage: Takeaways from the Ai4 Keynote with Andrew Ng, Geoffrey Hinton, and Fei-Fei Li

DevFeed: [The Human Context Advantage: Takeaways from the Ai4 Keynote with Andrew Ng, Geoffrey Hinton, and Fei-Fei Li](<https://devfeed.tech/articles/the-human-context-advantage-takeaways-from-the-ai4-keynote-with-andrew-ng-geoffrey-hinton-and-fei-fei-li-12798.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/the-human-context-advantage/>)

Author: Bonnie Chase

Published: 2026-08-10T00:00:00Z

Content type: opinion

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Strategy](<https://devfeed.tech/topics/ai-strategy.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [business](<https://devfeed.tech/tags/business.md>), [developers](<https://devfeed.tech/tags/developers.md>), [education](<https://devfeed.tech/tags/education.md>), [future-of-work](<https://devfeed.tech/tags/future-of-work.md>), [genai](<https://devfeed.tech/tags/genai.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [keynote](<https://devfeed.tech/tags/keynote.md>), [rag](<https://devfeed.tech/tags/rag.md>), [work](<https://devfeed.tech/tags/work.md>)

### AI overview

The article reflects on an Ai4 keynote featuring Andrew Ng, Geoffrey Hinton, and Fei-Fei Li. It argues that humans retain a context advantage because they understand organizational goals, customers, relationships, constraints, and unwritten rules that current AI systems lack. The article suggests that enterprise AI success will depend on delivering relevant context at the right time, while AI changes jobs by automating tasks and broadening developers' responsibilities rather than simply eliminating work.

### Source excerpt

From jobs and education to open models and AI infrastructure, the AI4 keynote made one thing clear: competitive advantage won't come from AI alone, but from connecting it to the right context.

## Your agent wants to search like a 2010 quant

DevFeed: [Your agent wants to search like a 2010 quant](<https://devfeed.tech/articles/your-agent-wants-to-search-like-a-2010-quant-12802.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/your-agent-wants-to-search-like-a-2010-quant/>)

Author: Jon Bratseth

Published: 2026-07-07T00:00:00Z

Content type: opinion

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [AI search](<https://devfeed.tech/topics/ai-search.md>), [information retrieval](<https://devfeed.tech/topics/information-retrieval.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Google Search](<https://devfeed.tech/topics/google-search.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [bm25](<https://devfeed.tech/tags/bm25.md>), [genai](<https://devfeed.tech/tags/genai.md>), [google-search](<https://devfeed.tech/tags/google-search.md>), [hybrid-search](<https://devfeed.tech/tags/hybrid-search.md>), [information-retrieval](<https://devfeed.tech/tags/information-retrieval.md>), [rag](<https://devfeed.tech/tags/rag.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [search](<https://devfeed.tech/tags/search.md>)

### AI overview

The article argues that AI agents should retrieve information with more control and sophistication than ordinary human search users. It describes a progression from vector retrieval to hybrid search using methods such as BM25 and machine-learned ranking, and presents search as code as a possible next stage.

### Source excerpt

The idea of empowering AI agents to retrieve information like a professional is going mainstream.

## Re-autoresearching MSMARCO BM25, on Vespa

DevFeed: [Re-autoresearching MSMARCO BM25, on Vespa](<https://devfeed.tech/articles/re-autoresearching-msmarco-bm25-on-vespa-12796.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/re-autoresearching-msmarco-bm25-on-vespa/>)

Author: andreer thomas

Published: 2026-05-29T00:00:00Z

Content type: article

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [Python](<https://devfeed.tech/topics/python.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [generalization in machine learning](<https://devfeed.tech/topics/generalization-in-machine-learning.md>), [pandas](<https://devfeed.tech/topics/pandas.md>), [Google Search](<https://devfeed.tech/topics/google-search.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [bm25](<https://devfeed.tech/tags/bm25.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [google-search](<https://devfeed.tech/tags/google-search.md>), [information-retrieval](<https://devfeed.tech/tags/information-retrieval.md>), [openai](<https://devfeed.tech/tags/openai.md>), [pandas](<https://devfeed.tech/tags/pandas.md>), [python](<https://devfeed.tech/tags/python.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

### AI overview

This article reproduces an MSMARCO BM25 autoresearch experiment in Vespa. It compares LLM-driven Python reranking with an approach restricted to existing Vespa rank features and reports a comparable improvement on a 650,000-passage subset, with better generalization to the full dataset.

### Source excerpt

BM25 is having a moment. We reproduce Doug Turnbull's MSMARCO autoresearch experiment in Vespa and get a comparable MRR@10 lift from existing rank features -- with twice the generalization to full MSMARCO.

## Vespa Newsletter, May 2026

DevFeed: [Vespa Newsletter, May 2026](<https://devfeed.tech/articles/vespa-newsletter-may-2026-12800.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/vespa-newsletter-may-2026/>)

Author: Bonnie Chase

Published: 2026-05-27T00:00:00Z

Content type: news

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [AI search](<https://devfeed.tech/topics/ai-search.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [code productivity](<https://devfeed.tech/topics/code-productivity.md>), [configuration](<https://devfeed.tech/topics/configuration.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [developer-productivity](<https://devfeed.tech/tags/developer-productivity.md>), [embedding](<https://devfeed.tech/tags/embedding.md>), [newsletter](<https://devfeed.tech/tags/newsletter.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [search](<https://devfeed.tech/tags/search.md>), [vector-search](<https://devfeed.tech/tags/vector-search.md>)

### AI overview

The May 2026 Vespa newsletter announces updates for retrieval and ranking systems, including improved ranking flexibility, embedding integrations with Voyage AI, OpenAI, and Mistral AI, Vespa Cloud dashboards and backups, maintenance controls, custom resource tags, agent skills, and new query and array-field capabilities.

### Source excerpt

Advances in Vespa include finer control over deployments, smarter ranking, richer embedding integrations, and more scalable vector search.

## Scaling a Vespa Application: Feeding Fast and Furiously

DevFeed: [Scaling a Vespa Application: Feeding Fast and Furiously](<https://devfeed.tech/articles/scaling-a-vespa-application-feeding-fast-and-furiously-12797.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/scaling-a-vespa-application-feeding-fast-and-furiously/>)

Author: Kai Borgen

Published: 2026-04-28T00:00:00Z

Content type: tutorial

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [AI search](<https://devfeed.tech/topics/ai-search.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [information retrieval](<https://devfeed.tech/topics/information-retrieval.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Homebrew](<https://devfeed.tech/topics/homebrew.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>), [onnx](<https://devfeed.tech/topics/onnx.md>), [optimum](<https://devfeed.tech/topics/optimum.md>), [XML](<https://devfeed.tech/topics/xml.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-search](<https://devfeed.tech/tags/ai-search.md>), [cli](<https://devfeed.tech/tags/cli.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [genai](<https://devfeed.tech/tags/genai.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [information-retrieval](<https://devfeed.tech/tags/information-retrieval.md>), [install](<https://devfeed.tech/tags/install.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [onnx](<https://devfeed.tech/tags/onnx.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [performance](<https://devfeed.tech/tags/performance.md>), [rag](<https://devfeed.tech/tags/rag.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [tensors](<https://devfeed.tech/tags/tensors.md>), [tutorial](<https://devfeed.tech/tags/tutorial.md>)

### AI overview

This tutorial demonstrates how to scale a Vespa application while feeding the full MS_marco passages dataset. It covers preparing the dataset, configuring access, deploying a sample application, and using scaling and metrics to improve feed throughput and performance.

### Source excerpt

A tutorial on how to scale the resources in a Vespa application to increase feed throughput. Using the metrics dashboard for informed and optimised scaling.

## The Vespa Cloud Metrics Dashboard

DevFeed: [The Vespa Cloud Metrics Dashboard](<https://devfeed.tech/articles/the-vespa-cloud-metrics-dashboard-12799.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/the-vespa-cloud-metrics-dashboard/>)

Author: Bjørn Meland

Published: 2026-04-24T00:00:00Z

Content type: tutorial

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [HTTP](<https://devfeed.tech/topics/http.md>), [payload](<https://devfeed.tech/topics/payload.md>), [Network](<https://devfeed.tech/topics/network.md>)

Tags: [guide](<https://devfeed.tech/tags/guide.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [http](<https://devfeed.tech/tags/http.md>), [latency](<https://devfeed.tech/tags/latency.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [network](<https://devfeed.tech/tags/network.md>), [payload](<https://devfeed.tech/tags/payload.md>), [performance](<https://devfeed.tech/tags/performance.md>)

### AI overview

A guide to using the Vespa Cloud metrics dashboard to investigate production issues. It presents a workflow that moves from system health to latency bottlenecks and resource utilization, then highlights health indicators and annotations added in the latest revision.

### Source excerpt

A guide to the Vespa Cloud metrics dashboard -- how to move from symptom to bottleneck to action, and what's new in the latest revision.

## Using Large ONNX Models with External Data in Vespa Embedders

DevFeed: [Using Large ONNX Models with External Data in Vespa Embedders](<https://devfeed.tech/articles/using-large-onnx-models-with-external-data-in-vespa-embedders-12794.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/onnx-external-data-in-vespa-embedders/>)

Author: bjorncs thomas

Published: 2026-03-27T00:00:00Z

Content type: tutorial

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [onnx](<https://devfeed.tech/topics/onnx.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [authentication](<https://devfeed.tech/tags/authentication.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [embedding](<https://devfeed.tech/tags/embedding.md>), [files](<https://devfeed.tech/tags/files.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [models](<https://devfeed.tech/tags/models.md>), [onnx](<https://devfeed.tech/tags/onnx.md>), [protocol](<https://devfeed.tech/tags/protocol.md>), [serialization-format](<https://devfeed.tech/tags/serialization-format.md>)

### AI overview

Vespa embedders now support large ONNX models whose weights are stored in external data files. Starting with Vespa 8.544, Vespa automatically downloads referenced external files when loading URL-based models, with support for private models through propagated authentication tokens. The feature is limited to embedders and supported model references.

### Source excerpt

Many ONNX models exceed the 2GB protobuf limit and store weights in external data files. Vespa now supports these models for embedders.

## Asymmetric Retrieval: Spend on Docs, Embed your Queries for Free

DevFeed: [Asymmetric Retrieval: Spend on Docs, Embed your Queries for Free](<https://devfeed.tech/articles/asymmetric-retrieval-spend-on-docs-embed-your-queries-for-free-12793.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/asymmetric-retrieval-spend-on-docs-queries-for-free/>)

Author: thomas bjorncs

Published: 2026-03-10T00:00:00Z

Content type: article

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [container](<https://devfeed.tech/tags/container.md>), [cost](<https://devfeed.tech/tags/cost.md>), [embedding](<https://devfeed.tech/tags/embedding.md>), [genai](<https://devfeed.tech/tags/genai.md>), [latency](<https://devfeed.tech/tags/latency.md>), [local](<https://devfeed.tech/tags/local.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [quality](<https://devfeed.tech/tags/quality.md>), [rag](<https://devfeed.tech/tags/rag.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [vector](<https://devfeed.tech/tags/vector.md>), [voyage-ai](<https://devfeed.tech/tags/voyage-ai.md>)

### AI overview

The article presents asymmetric retrieval using Voyage AI models with Vespa. Documents are embedded once with a high-quality, API-based model, while frequently submitted queries are embedded with a small model running locally. Because the models share a compatible vector space, this approach can eliminate recurring query-embedding API costs, reduce latency, and allow query models to be upgraded independently.

### Source excerpt

Documents are embedded once -- worth the spend for maximum quality. Queries hit you on every request. This is what drives your cost at scale. Asymmetric retrieval with Voyage AI and Vespa. Real numbers, real config.