When OpenClaw's creator revealed last month that his agent fleet had torched $1.3 million in API costs—in a single month, no less—the figure landed somewhere between jaw-dropping and inevitable. OpenAI picked up the tab, which was generous. But the incident laid bare something most people building with AI coding agents already knew: the economics are brutal.
603 billion tokens. 7.6 million requests. Thirty days.
The numbers crystallized what had been whispered in Slack channels and GitHub issues for months. These tools work, sure. Sometimes brilliantly. But they're also astonishingly expensive to run at scale, in ways that feel almost wasteful. An agent hunting for the right function in a large codebase can burn through tens of thousands of tokens in what amounts to an elaborate guessing game.
Now, a flurry of startups think they've found the answer. Or at least an answer. Semble's benchmarks from May show 98% fewer tokens than grep+read workflows—a claim that caught attention. It wasn't alone. By the time spring ended, at least half a dozen similar tools had appeared, each promising roughly the same thing—make AI coding agents economically viable, or watch the whole category stall out.
Whether any of them actually deliver remains an open question.
When Retrieval Becomes the Bottleneck
The token problem isn't subtle, and it isn't new. Stanford's Digital Economy Lab published research in April showing that agentic coding tasks consume significantly more tokens than standard code chat sessions—with high variance and unpredictable patterns. Throwing more money at the problem doesn't guarantee better results. Accuracy often peaks somewhere in the middle of the cost curve, not at the top end.
Cast Net Technology modeled what this looks like in practice: a five-person engineering team, each running three agent sessions per day. Their estimate? Around $10,500 monthly in token costs that smarter retrieval could eliminate. That's real money, the kind that compounds fast across quarters. The kind CFOs start asking questions about.
And here's the twist—longer context windows don't fix it. You'd think Claude's 200K token window or GPT-4's expanding capacity would eliminate the need for better search. Not so. A 2024 paper in TACL titled "Lost in the Middle" documented what researchers call positional bias: models struggle to reliably use information buried deep in massive contexts. They skim, essentially. Or fixate on what's at the beginning and end. Quality retrieval still matters, maybe more than ever.
The Grep Trap
Most AI coding agents today lean on grep. It's simple, battle-tested, and easy to implement. Search the codebase with pattern matching, grab entire files, dump them into context. Claude Code, according to its developers posting on Hacker News, primarily uses grep and file tools rather than anything fancier.
The approach works—until it doesn't. Sourcegraph's analysis from May put the breaking point around 400,000 lines of code. Beyond that threshold, grep-based agents enter what developers call "random walks." Tool call after tool call, burning tokens, often failing to converge on the relevant context. It's like watching someone search for a book by checking every shelf in the library twice.
Semble's benchmarks, based on 1,251 queries tested in May, quantify the gap. The average grep-and-read loop consumes 45,692 tokens per query. Semble claims to deliver comparable code coverage at 566 tokens—a 98% reduction. The methodology relies on OpenAI's tiktoken encoder and defines coverage as any retrieved unit overlapping an annotated span. Those numbers are striking, if they hold up in production.
A Suddenly Crowded Field
Semble is hardly the only entrant. The category exploded in early 2026, with new projects appearing almost monthly.
cgrep showed up around February, combining BM25 with tree-sitter parsing. Its Hacker News debut claimed 95.2% fewer tokens than grep on PyTorch scenarios.
Cast Net's "Mnemosyne" published head-to-head benchmarks in April: 5.6 times fewer tokens than grep-based navigation, 2.4 times faster task completion.
ogrep advertises roughly 85% token reduction, integrated with the Model Context Protocol and multiple embedding backends.
Hypergrep claims 87% fewer tokens via structural search and code graph analysis.
The pattern repeats. Everyone's targeting 80-98% token reductions. The architectural choices vary—some emphasize symbol graphs and language server protocol integration, others bet on hybrid retrieval—but the economic motivation is identical. Find the right code faster, or the unit economics crater.
What's Under the Hood

Semble's technical approach reveals something like an emerging consensus in the space. It uses tree-sitter for code-aware chunking, preserving semantic boundaries instead of blindly splitting on line counts. Then it applies hybrid retrieval: static Model2Vec embeddings for semantic search, BM25 for keyword matching. A specialized code embedding model called potion-code-16M, distilled from CodeRankEmbed, handles the semantic layer. Reciprocal rank fusion merges results. Reranking follows, using code-specific signals—definition boosts, identifier stems, file coherence penalties, filters for test and legacy noise.
Everything runs locally on CPU. No external API calls. Semble's benchmarks show average repository indexing in about 250 milliseconds, queries answered in 1.5 milliseconds. Retrieval quality (NDCG@10 of 0.854) lands at roughly 99% of the 137-million-parameter transformer baseline—but with 218 times faster indexing, 11 times faster queries.
The tool plugs into Claude Code, Cursor, Codex, and OpenCode through Anthropic's Model Context Protocol, which has seen widespread adoption for AI tool integration. Command-line and Bash integrations are also available.
Security Concerns Cloud the Landscape
MCP's rapid adoption hit turbulence in April when security researchers flagged architectural vulnerabilities enabling remote code execution. The Cloud Security Alliance and other firms issued advisories covering multiple implementations—LiteLLM, Bisheng, Windsurf among them. TechRadar noted in May that the flaws potentially exposed 150 million downloads and thousands of servers to complete takeover scenarios.
Not exactly the kind of headline enterprises want to see when evaluating AI agent infrastructure.
Security concerns are increasing demand for local-first tools with minimal external dependencies. Semble's CPU-only, self-contained design suddenly looks less like a technical limitation and more like a feature. Enterprises scrutinizing agent toolchains for compliance with the EU AI Act—most rules kick in this August—now have data residency and attack surface high on their procurement checklists.
What the Community Is Really Asking
Semble's Hacker News debut on May 19 gathered 434 points and 72 comments. The discussion was revealing, perhaps more than the company expected.
Multiple commenters pressed on the same point: do the 98% token savings translate to real end-to-end improvements? Fewer agent turns? Faster wall-clock time? Higher success rates? The questions expose a gap between retrieval benchmarks and production reality.
The Semble maintainers—Hacker News users "Bibabomas" and "stephantul," identified as Stéphan Tulkens and Thomas van Dongen from MinishLab—acknowledged the concern. End-to-end evaluations measuring full agent loops across different models are on the roadmap. But community feedback highlighted the need for answers now. Do models actually trust better search results enough to stop exploring? Or do they just issue different tool calls and burn tokens anyway?
That question haunts the entire category. Retrieval quality is necessary but not sufficient. Agent behavior, model calibration, task complexity—they all interact in ways that aren't yet well characterized in public benchmarks. The theory is compelling. The practice remains unproven at scale.
A Market Accelerating Into the Problem

The timing is sharp. Microsoft reported in late April that it has over 20 million paid Copilot users. JetBrains research from April showed AI coding tool usage among professional developers trending upward from 2024 and 2025 levels. Market research firms project compound annual growth around 27% for AI coding assistants through 2030, though definitions vary and projections always should be taken with salt.
As adoption accelerates, cost sensitivity rises with it. Entire.io published a blog post in May titled "How we improved agentic search," complete with agent traces showing reduced search calls per task after optimization. These aren't hypothetical problems. Teams are hitting them in production now.
Token prices continue falling across providers, which helps. But it doesn't solve the fundamental issue. Volume is the killer. When an agent might issue dozens or hundreds of tool calls per task, unit cost becomes less important than the structural efficiency of the workflow itself.
What Comes Next

The competition is just beginning. Most projects that launched early this year are still at version 0.1 or 0.2. None have published comprehensive end-to-end studies proving their value in production agent loops. Semble hit 2,600 GitHub stars in its first few days, which is impressive. Stars don't correlate perfectly with adoption or long-term viability, though.
Several factors will likely determine winners. End-to-end validation matters more than retrieval benchmarks. Agents need to actually complete tasks faster and more reliably, with fewer tokens, measured across diverse codebases and models. Security and compliance posture will drive enterprise decisions, particularly as EU regulations take effect in August. Local-first, minimal-dependency designs have an edge there. Integration quality with popular agents and editors determines distribution—MCP support is table stakes, but depth varies widely.
The architectural consensus seems to be converging: hybrid retrieval, code-aware chunking, lightweight reranking, CPU-friendly inference. Static embeddings, which eliminate transformer forward passes at query time, are gaining traction for speed-quality tradeoffs. The real competition will be in the details. How well does chunking preserve semantic boundaries? How effective are the reranking signals? How cleanly do these tools integrate into existing workflows?
There's also a larger platform question lurking. As Sourcegraph and others point out, code search is only part of the solution. Full code intelligence—symbols, call graphs, type information, dependency analysis—compounds the benefits. The companies that deliver end-to-end code understanding may build more defensible moats than point solutions focused solely on search.
For now, the category remains early and fragmented. But that $1.3 million OpenClaw burn crystallized the problem in a way that's impossible to ignore. Founders building AI-powered products, engineering leaders managing agent infrastructure, investors evaluating developer tools—they're all watching the same numbers. Token efficiency isn't a nice-to-have optimization anymore. It's the constraint determining what's economically feasible.
The tools launching in 2026 aren't solving every problem agents face. They're solving one very specific problem: making it cheaper to find the right code. If they succeed, they'll unlock the next wave of agent capabilities. If they fail, the economics will force a different path. Maybe agents will need to get smarter about when to search. Maybe models will improve their ability to work with less context. Maybe the entire architecture shifts in ways we haven't imagined yet.
What seems certain is this: the race to cut token costs isn't a subplot. It's the story of AI coding in 2026.
