Forty-six seconds. That's how long it took Mintlify's documentation assistant to spin up a new conversation in early 2026. A sandbox environment would initialize, files would clone, permissions would get set—all the necessary choreography of isolated compute. In isolation, the latency seemed defensible. Maybe even reasonable.
Then the engineers did the math.
With 30,000 conversations happening daily, those 46 seconds weren't just an inconvenience. They were a projected $70,000-per-year infrastructure problem, assuming growth to 850,000 monthly conversations and Daytona's per-second sandbox pricing. The kind of number that makes you rethink your entire architecture.
So Mintlify rebuilt it. On March 24, 2026, the company published details of ChromaFs, a virtual filesystem layered atop their existing Chroma vector database that cuts session creation to roughly 100 milliseconds. The shift represented more than a performance optimization—it surfaced a broader realization rippling through developer tooling: AI agents now consume documentation at a scale approaching human traffic, with Mintlify's own data showing agents accounting for 45.3% of documentation requests. And they prefer filesystems to embeddings.
The Sandbox Tax
Traditional sandboxing made sense, once. When documentation assistants needed to execute code or manipulate files in isolated, stateful environments, spinning up VMs or containers was the prudent choice. But Mintlify's assistant didn't actually need execution capability. It needed fast, read-only access to structured knowledge.
The team's projections were sobering. At 850,000 monthly conversations—well within reach given current growth—sandbox provisioning alone would incur over $70,000 annually in marginal compute costs, based on Daytona's per-second pricing model referenced in their engineering post. That's real money for what amounts to waiting.
The obvious alternative—static retrieval-augmented generation with top-k chunk retrieval—had its own constraints. When answers spanned multiple documentation pages or required exact string matching (think grep-style searches for error codes, function signatures, configuration keys), chunk-based retrieval fragmented results. Reranking could help, sure, but it added latency and didn't solve the fundamental mismatch: agents were being asked to work with a knowledge format optimized for embedding similarity, not for the way they naturally navigate information.
It's a problem more platforms are encountering. Agents traverse directories, pipe grep output, cat entire files when context windows permit. They work with filesystems daily in their core workflow—navigating codebases, reading configs, checking logs. Presenting documentation as a virtual filesystem aligns with their existing mental model. Top-k chunk retrieval, by contrast, forces them into a narrower interaction pattern: ask a question, receive N disconnected snippets, synthesize. Not impossible. Just... awkward.
A Filesystem That Isn't There
ChromaFs layers a POSIX-like interface over Chroma's vector store. When an agent issues a grep, cat, ls, find, or cd command through Vercel Labs' just-bash shell emulation, ChromaFs intercepts it and translates the request into Chroma queries. A gzipped JSON path tree provides fast directory operations. Large artifacts—OpenAPI specifications, for instance—use lazy pointers to avoid bloating the in-memory tree. The system is strictly read-only; attempts to write return an EROFS (read-only filesystem) error, the digital equivalent of a polite but firm no.
Access control happens at the path level. Before constructing the tree, ChromaFs prunes paths based on tenant permissions, ensuring each user's virtual filesystem reflects only the documentation they're authorized to see. The architecture exploits Chroma's existing indexing while preserving the file/directory metaphor that agents handle natively.
No real files exist. The "filesystem" is an API surface shaped like one.
The results, according to Mintlify's March 2026 post, are stark. Session creation dropped from roughly 46 seconds to around 100 milliseconds—a 460x improvement. Marginal per-conversation compute cost effectively rounds to zero, since ChromaFs reuses the already-provisioned vector database. At scale—hundreds of thousands of users, over 30,000 conversations daily—the approach eliminates what you might call the sandbox tax entirely.
When Machines Become Your Primary Readers

The architectural shift coincided with a pattern emerging in Mintlify's traffic data that's hard to ignore. In March 2026, the company analyzed 30 days of requests across all Mintlify-powered documentation sites: 357.6 million of the roughly 790 million total requests came from AI coding agents.
That's 45.3 percent. Near parity with traditional browser traffic.
Claude Code accounted for 55.8 percent of that AI volume (199.4 million requests), with Cursor at 39.8 percent (142.3 million). The figures, published April 3, 2026, come with a caveat Mintlify was quick to note: not all agents identify via user-agent strings. OpenAI's Codex, for example, lacks a consistent user-agent, suggesting the true share of AI-driven documentation consumption is likely higher still. Perhaps significantly higher.
This isn't a Mintlify anomaly. Stack Overflow's 2025 AI survey showed 84 percent of developers either using or planning to use AI tools. Postman's 2025 State of the API report found 41 percent of respondents using AI to generate API documentation. DX's Q4 2025 impact report, covering 266 companies, estimated roughly 22 percent of code was AI-authored by late 2025.
Documentation platforms, in other words, are now serving two audiences—one human, one algorithmic. And the algorithmic one is growing faster.
When documentation serves agents at that scale, retrieval patterns shift in subtle but important ways. Agents traverse directories, pipe grep output, cat entire files when context windows allow. They work with filesystems daily. Presenting documentation as a virtual filesystem aligns with their existing workflow. Top-k chunk retrieval forces them into something less natural.
Beyond the Chunk
Mintlify's approach doesn't eliminate vector search—it reframes how agents access the indexed knowledge. The Chroma database still holds embeddings, still supports semantic queries. But instead of returning ranked chunks, ChromaFs lets agents query, filter, and navigate the corpus using familiar shell commands. A grep for a specific error code becomes a targeted Chroma query, not a semantic similarity match that might surface tangentially related content. A find operation traverses the virtual directory tree backed by the database's path index.
The shift reflects broader skepticism around static RAG as a one-size-fits-all solution, something that's been percolating in developer communities for months. Research published throughout 2025 highlighted RAG pain points: chunking strategies that fragment code blocks, reranking sensitivity, version drift in evolving documentation. A January 2025 survey on agentic RAG catalogued iterative tool-calling loops, planning, and retrieval strategies designed to move beyond simple top-k fetch. Version-aware RAG for evolving docs appeared in October 2025. Planning-driven chunk compression surfaced in December.
Practitioner commentary has been, well, blunt. Reddit threads in late 2025 and early 2026 describe RAG hallucinations persisting despite careful chunking and reranking. The issue isn't necessarily the vector database itself—it's the interface. When answers require exact matches, span multiple pages, or depend on hierarchical structure, chunk-based retrieval introduces friction.
Agents already have a solution. They grep, cat, and cd their way through problems every day.
The ecosystem is responding. Letta (formerly MemGPT) introduced a filesystem abstraction for agent memory in July 2025, later extending it to git-backed "context repositories." The Model Context Protocol (MCP), donated to the Linux Foundation's Agentic AI Foundation in December 2025, provides a standardized way to expose structured knowledge—including filesystem-like endpoints—to agents. Microsoft's Copilot Studio announced MCP support in March 2025. OpenAI's Agents SDK documentation includes MCP integration examples. The pieces are assembling.
Design Principles for an Agent-First World

Mintlify's architecture change signals a tactical evolution, not a wholesale rejection of retrieval-augmented generation. The company still relies on Chroma for indexing and search. What's changed is the abstraction layer: instead of serving chunks, they serve a virtual filesystem backed by the same underlying database.
It's a small conceptual shift with large operational consequences. One hundred-millisecond sessions instead of 46-second waits. Near-zero marginal cost instead of projected six-figure annual compute overhead.
The broader lesson may be less about filesystems specifically and more about interface design in an agent-first world. As AI coding tools drive toward half of documentation traffic—and Mintlify's March data suggests that threshold is approaching—the incentive to optimize for machine-readable structure intensifies. Mintlify co-developed llms.txt and llms-full.txt with Anthropic, formats designed to give agents concise, structured overviews of a documentation corpus. (Adoption remains contested. Vendor posts from August 2025 claim active crawling by major LLMs, while independent analyses from October 2025 and January 2026 dispute widespread uptake. The truth, as usual, is probably somewhere in between.)
GitBook, another documentation platform, has rolled out AI-native features including an embedded assistant, Ask AI insights, and llms.txt support through 2025. ReadMe offers an AI Agent for content creation and a GitHub AI Writer that proposes doc updates from pull request diffs. The competitive dynamic favors platforms that can serve both audiences—developers reading in browsers, agents scraping at scale—without the performance or cost penalties of per-session isolation.
For teams building their own documentation or retrieval systems, the ChromaFs approach suggests a few design principles worth considering. First, evaluate whether your agent interactions require true execution environments or simply structured read access. If the latter, virtual abstractions over existing databases may be cheaper and faster than sandboxing. Second, expose knowledge in formats agents already use—filesystems, structured APIs, MCP-compatible endpoints—rather than forcing them into embedding-centric workflows. Third, measure your AI agent traffic. If it's approaching human levels (as Mintlify's March data suggests), treat machine readability as a first-class design constraint, not an afterthought.
The EU AI Act's general-purpose AI obligations, phasing in through August 2026, add another dimension: access control, auditability, and copyright compliance for documentation feeding assistants become regulatory concerns, not just engineering ones. ChromaFs's path-level pruning before tree construction hints at how these requirements might be implemented at the retrieval layer—users see only what they're authorized to access, baked into the virtual filesystem itself. It's early days, but the regulatory pressure is real.
What It Means

Perhaps the most durable insight here is that agents and traditional RAG aren't incompatible. Mintlify didn't tear out Chroma; they built ChromaFs on top of it. The vector database still powers search, still handles embeddings. The filesystem is just a more natural interface for the tools developers—and the AI agents writing code alongside them—already use daily.
Whether that pattern spreads depends on how much of documentation traffic shifts from browsers to agents. If Mintlify's March numbers hold, the shift is already underway. Forty-six seconds was unacceptable. One hundred milliseconds is table stakes. And the companies figuring that out first are the ones building for an audience that's half-human, half-machine, and entirely unforgiving of latency.
