Skip to content
research

Vector DB Benchmarks Are Lying to You

Every vector database claims to be the fastest and most accurate, but their benchmarks are deeply flawed. Discover the quadrillion-calculation problem that hid the truth, and the new 24TB dataset that finally exposes it.

Aki Tanaka
Vector DB Benchmarks Are Lying to You

The Benchmark Mirage: Why Every Vendor Wins

Choosing a reliable vector database for your project often feels like navigating a hall of mirrors. Every vendor’s benchmark page proudly showcases their product at the top, painting a contradictory picture that creates industry-wide confusion and erodes trust in published performance metrics. This widespread phenomenon isn't necessarily a symptom of malicious intent or deliberate deception.

Instead, it underscores the extreme difficulty of honest vector search benchmarking. Accurately evaluating these complex systems demands immense computational resources and meticulous methodology, compelling many vendors to take shortcuts. These often lead to incomplete or "half-baked" results, making objective comparison nearly impossible for users seeking genuine performance insights.

Central to this benchmarking challenge is recall, the most critical metric for accuracy. Recall precisely quantifies how many of a query's true nearest neighbors an engine actually retrieves from a dataset. To measure this metric reliably, one must first compute these true nearest neighbors exactly—not approximately—establishing a foundational "ground truth" for comparison.

Achieving this exact ground truth is computationally prohibitive for large-scale deployments. Consider a dataset of 10 billion vectors processed with 120,000 queries: this scenario necessitates over a quadrillion (10^15) exact distance computations. This astronomical computational barrier explains why no one had previously undertaken such an exhaustive, precise calculation, fostering deep skepticism around vendor claims and leaving the industry without truly reliable, independent benchmarks.

A Quadrillion Reasons to Be Skeptical

Measuring a vector database’s true performance hinges on recall, the proportion of actual nearest neighbors an engine retrieves for a given query. This metric presents a fundamental catch-22: to accurately assess recall, one must first possess the "ground truth"—the precise set of nearest neighbors calculated without approximation. Without this definitive answer, any recall measurement is inherently speculative.

Establishing this ground truth is computationally immense. Consider a dataset of 10 billion vectors requiring 120,000 queries; determining the exact nearest neighbors necessitates over a quadrillion (10^15) distance computations. This astronomical scale represents a computational barrier that has historically rendered comprehensive, unbiased benchmarking impractical for the vast majority.

Such prohibitive costs have directly prevented the industry from developing truly neutral, large-scale benchmark datasets. Consequently, vendors have been compelled to create their own, often smaller, benchmarks, inadvertently perpetuating the cycle of self-serving performance claims. This historical barrier explains the persistent challenge in obtaining objective vector database evaluations.

The 24 Terabyte Dataset That Changes Everything

Qdrant dramatically shifted the landscape by undertaking the computationally expensive task of establishing ground truth. Previously, calculating the exact nearest neighbors for 10 billion vectors involved over a quadrillion distance computations, an insurmountable barrier. But Qdrant invested heavily, burning through massive GPU compute to create a 24-terabyte dataset of 10 billion vectors, now publicly available. This monumental effort provides the first precisely computed ground truth for Internet-scale vector search; explore it further at Internet-Scale Knowledge Retrieval: A Novel Vector Search Dataset at 10B Scale - Hugging Face.

This new clarity reveals stark performance differences between engines. When tested against a slice of this rigorous dataset, Elasticsearch struggled to surpass 95% accuracy, regardless of configuration. In contrast, both Qdrant and Milvus consistently achieved 98% accuracy, demonstrating their superior ability to retrieve true nearest neighbors.

These results highlight distinct operational profiles. Elasticsearch remains the fastest way to be roughly right, offering speed for less critical recall demands. But Qdrant emerges as a clear leader for applications demanding both high precision and rapid retrieval, delivering the most precise results at competitive speeds. This dataset finally enables objective, apples-to-apples comparisons.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Your Action Plan: How to Choose a Vector DB Now

Qdrant’s monumental effort, computing the ground truth for a 10-billion-vector, 24-terabyte dataset, shifts the entire landscape. This publicly available resource, born from quadrillions of distance computations, finally offers a verifiable baseline for measuring recall—a feat previously deemed too expensive and complex. This development far outweighs any individual benchmark result.

Remain acutely skeptical of vendor-published claims. Every vector database vendor still presents their product as superior on their own benchmark pages, a natural consequence of benchmarking's inherent difficulty. Instead of relying on these curated narratives, empower yourself with this new data.

Leverage these unprecedented tools to rigorously test databases against your specific production workload. This means evaluating performance for your unique query patterns and data characteristics. The objective is not to declare a single 'winner,' but to meticulously understand the real-world trade-offs.

Consider the complex interplay between search speed, recall accuracy, and operational cost for your application. For instance, while Milvus and Qdrant might both achieve 98% accuracy, their speed profiles could differ significantly. Your informed decision will hinge on balancing these critical factors for optimal system design.

Frequently Asked Questions

Why are most vector database benchmarks unreliable?

They are unreliable because accurately measuring performance metrics like 'recall' is computationally expensive. Many vendors take shortcuts, leading to biased, 'half-baked' results where they appear to outperform competitors.

What is 'recall' in a vector database context?

Recall measures the accuracy of a search. It's the percentage of the true nearest neighbors that the database successfully retrieves for a given query. A 98% recall means the engine found 98 out of the 100 actual closest vectors.

What is the significance of Qdrant's 10-billion-vector dataset?

This 24TB dataset is the first of its scale to include the pre-computed 'ground truth' of exact nearest neighbors. It provides a public, unbiased foundation for accurately benchmarking different vector databases against each other, solving a major industry problem.

Which vector database is best according to the new benchmark?

There's no single 'best' for all cases. The tests show different strengths: Elasticsearch is fast for 'roughly right' results (~95% accuracy), while Qdrant and Milvus can achieve higher precision (~98% accuracy), with Qdrant noted for being both quick and precise.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$500 · AI tools & software only

For builders

This page is doing a job for someone else’s tool.

AI agents read it. Buyers land on it. It answers in eight languages and over MCP. Your tool can have one like it — live in 24 hours.