Why DynamoDB Became the Bottleneck
Perplexity’s search API processes a demanding workload: each request fetches approximately 100 to 120 page keys, typically in batches of 10 to 20. Each record averages about 50 KB, leading to a substantial data volume per query. This access pattern, combined with a rapidly expanding web index, quickly exposed the limitations of a usage-based billing model.
DynamoDB’s cost structure bills every byte read or written. As Perplexity’s data index grew and production traffic intensified, these byte-based read costs scaled linearly, making the managed database increasingly expensive. This economic pressure became a significant driver for Perplexity to re-evaluate its database strategy.
Beyond cost, DynamoDB's managed nature imposed critical control limits. Perplexity could not tune fundamental database behaviors to optimize for its specific access patterns. Engineers lacked the ability to dictate:
- Partition placement on specific machines
- Local memory allocated for caching
- Which replica responded to a read request
This lack of granular control, particularly over replica selection, meant that one slow replica could stall an entire batch read, significantly impacting tail latency for users. Perplexity required a more flexible and performant solution.
The Slowest Replica Set the Pace
Batch reads in DynamoDB amplified latency. A single search request, fetching 100 to 120 page keys in batches of 10 to 20, meant that one slow replica could delay the entire request, even if other records arrived quickly. This phenomenon, where the slowest operation dictates overall performance, is known as tail latency.
Perplexity addressed this by building CobbleDB, a specialized hot key-value store. Developed in Rust, CobbleDB runs on RocksDB, serving data directly from local NVMe storage. Keys are logically grouped by partition, ensuring efficient data locality. This architecture granted Perplexity granular control over caching and data placement, capabilities absent in managed services.
A stateless router orchestrates CobbleDB’s reads. It sends parallel requests to multiple replicas and employs hedged reads: if one replica lags, the router immediately sends the same read to another replica. This aggressive strategy minimizes the impact of slow nodes, significantly reducing tail latency for batch operations. This approach dropped median batch read latency from 31.4 milliseconds to 5.6 milliseconds, and P99 latency from 123 milliseconds to just 24 milliseconds—a nearly five-fold improvement.
The 5x Win—and What the Numbers Mean
Perplexity’s shift to CobbleDB dramatically improved production batch-read latency. The median response time plummeted from 31.4 ms on DynamoDB to just 5.6 ms. Even more striking, the P99 tail latency — which previously held up entire search requests — dropped from 123 ms to approximately 24 ms, marking a roughly fivefold speedup.
This performance boost came with a significant economic advantage. Perplexity’s internal cost model projected CobbleDB to be at least 20% cheaper than DynamoDB across all commitment tiers. Synthetic tests further validated CobbleDB’s robustness, demonstrating stable throughput up to 500,000 requests per second without performance degradation.
While these results are impressive, they reflect Perplexity’s specific workload and operational model. Their unique search API, characterized by batch reads of 100-120 page keys each averaging 50 KB, directly benefited from CobbleDB’s tailored architecture, including its use of hedged reads to mitigate tail latency.
This success does not universally guarantee that a custom-built database will outperform managed services for every team or use case. Perplexity’s bespoke solution was optimized for their precise needs, leveraging two engineers and AI agents to craft a system uniquely suited to their scale and cost drivers. For more details on the architecture, read about CobbleDB: Rebuilding AI Search Storage for Lower Latency and Cost.
Enjoying this? Get one like it in your inbox each morning.
one email a day · unsubscribe in two clicks · no third-party tracking
Two Engineers, AI Agents, and the Hidden Trade-Off
CobbleDB’s rapid development highlights a new frontier in engineering. Roughly 40,000 lines of Rust code were delivered in about two months by two engineers, working in tandem with AI coding agents. This swift execution underscores the potential for AI-augmented development.
The division of labor was critical. AI agents handled repetitive tasks like writing tests, implementing fixes, generating observability hooks, producing documentation, and tracking CI/CD follow-ups. Meanwhile, the human engineers retained ownership of strategic elements:
- Architecture design
- Code review
- Deployment gates
- Production decisions
This collaboration allowed the small team to move with unusual speed, focusing human expertise on high-leverage problems.
However, replacing a managed service like DynamoDB introduces a significant operational trade-off. Perplexity now directly assumes responsibility for hardware failures, data backups, and ensuring continuous reliability. While this model offers unparalleled control and performance for specialized, hyperscale workloads, it demands a level of operational maturity and resource commitment most startups cannot afford. CobbleDB’s success is a testament to bespoke engineering, but it comes with the hidden cost of increased operational burden.
Frequently Asked Questions
What is CobbleDB?
CobbleDB is Perplexity’s in-house key-value store, built for serving search data with lower latency and cost than its previous DynamoDB setup.
How did CobbleDB reduce latency?
It uses RocksDB on local NVMe and parallel replica reads, with hedged reads that retry a slow request against another replica.
How much faster was CobbleDB than DynamoDB?
Reported median batch-read latency fell from 31.4 ms to 5.6 ms, while P99 latency fell from 123 ms to about 24 ms.
Did AI agents build CobbleDB autonomously?
No. Agents assisted with tasks such as tests, fixes, monitoring, and documentation, while engineers designed the system, reviewed changes, and controlled production releases.

