Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
SaaS iconSaaSAugust 28, 2026

Maingen launches AI benchmark for industrial operations

Maingen launches AI benchmark for industrial operations
YcAi Benchmarking+3
SaaS iconSaaSAugust 28, 2026

Enact launches layer pushing robot success rates to 99%

Enact launches layer pushing robot success rates to 99%
YcRobotics+2

Founders Mentioned

Michael Zhou

Codag

saas icon
SaaS

Michael Zhou

Codag

saas icon
SaaS
SaaS iconSaaS
August 28, 2026
YcAi AgentsCost OptimizationAi Infrastructure

Codag launches tool compression to cut AI agent costs 30%

YC S26 startup tackles AI agent 'reading tax' with compression layer that cuts token spending 20-30% while preserving accuracy. Already used by 90+ organizations.

Codag launches tool compression to cut AI agent costs 30%

Codag, a developer-tools company backed by Y Combinator, has launched a compression service that it says can reduce token spending for AI coding agents by 20 to 30 percent. The San Francisco startup claims more than 90 organizations have already adopted the product, with some engineering teams numbering as many as 35 people.

The timing reflects a growing panic in enterprise IT departments. AI agents, which automate coding tasks by calling dozens or even hundreds of tools per workflow, are burning through tokens at a pace that has research firms warning of a coming cost crisis. A study from this year found that agentic tasks can consume roughly 1,000 times more tokens than simple code chat, with input tokens—the data the model reads rather than generates—accounting for nearly all the bloat.

"The cost of running agents is dominated by reading, not reasoning," Codag wrote on its product site.

The numbers suggest the problem is real. One major analytics firm predicted that inference costs per agentic workflow could quintuple within a few years, even as prices per token fall. JetBrains reported in August 2026 that 39 percent of professional developers worldwide were using Claude Code at work during the May to July 2026 period, climbing to 47 percent in the United States. OpenAI said in June 2026 that Codex had surpassed 5 million weekly active users. Atlassian's Rovo MCP server logged more than 5 million daily tool calls in a single month.

How Compression Works

Codag's solution compresses tool-call outputs before the agent reads them. Founder Michael Zhou designed the product as a local proxy that strips large outputs down to the evidence lines an agent actually needs, the company said. Test results, search findings, logs, file trees: all get pared to essentials. Omitted data remains retrievable locally by line range or JSON path, so nothing is lost permanently.

The startup says optimized tool results shrink by 75 percent on average. One example on the product site shows a production-incident triage session dropping from 4.8 MB to 158 KB. Another, involving a repository search, falls from 3.2 MB to 61 KB.

Pricing runs about $0.50 per gigabyte processed on a pay-as-you-go basis, according to the company's website. A $499 monthly team plan includes 12,500 credits, enough to optimize roughly 1.25 terabytes.

Codag is hardly alone. Atlassian open-sourced a tool called mcp-compressor earlier this year, claiming reductions of 70 to 97 percent in tool-definition overhead. Anthropic has discussed an internal technique it calls Tool Search, which the company says preserves 95 percent of the context window and cuts tool-definition tokens by around 85 percent. A handful of other startups—Morph, Tamp, gotcontext.ai among them—offer similar compression proxies, and the open-source ecosystem has produced dozens more options.

Model Providers Fight Back

Digital illustration for article section "Model Providers Fight Back" in "Codag launches tool compression to cut AI agent costs 30%" - A conceptual, minimal illustration of a large, vibrant funnel transforming a chaotic jumble of bulky...

The major AI platforms are trying to solve the token problem from the inside. Anthropic recently launched programmatic tool calling in public beta, which processes tool results in code rather than natural language. Internal results from the company's beta testing show the approach cuts tokens by about 37 percent on complex workflows. OpenAI deprecated its Assistants API and introduced a Responses API designed to reduce cost and latency when multiple tools are involved.

Google rolled out FinOps features for Gemini Enterprise to counter what one news outlet called "AI sticker shock." The company had introduced spend caps and billing visibility in Google AI Studio earlier in the year.

The Cache Problem

Digital illustration for article section "The Cache Problem" in "Codag launches tool compression to cut AI agent costs 30%" - A clean, minimalist conceptual illustration of a simple storage box representing a data cache being ...

Compression isn't a silver bullet. A paper published this summer titled "Token Reduction Is Not Cost Reduction" warned that aggressive compression can corrupt cached prompt prefixes, sometimes pushing costs above baseline. The authors reported one case where tool-output tokens fell 38 percent, yet overall cost rose nearly 7 percent. Another paper on cache-aware prompt compression showed that over-compressing can evict frequently reused blocks from high-speed cache tiers, erasing the benefits.

Anthropic's pricing illustrates why cache economics matter. Base input tokens for Opus 4.7 cost $5 per million, but cache writes run $6.25 to $10 per million depending on time-to-live. Cache hits, by contrast, drop to just $0.50 per million.

Open Questions

Digital illustration for article section "Open Questions" in "Codag launches tool compression to cut AI agent costs 30%" - A minimalist, conceptual illustration of a large, bold question mark resting above a perfectly balan...

Compression layers need to prove they preserve accuracy and play nicely with prompt caching, or risk creating new cost centers instead of eliminating old ones. Codag says its proxy offers byte-exact retrieval and fail-open behavior if the service goes down. Independent benchmarks would help clarify whether those savings hold across different cache strategies and model updates.

Meanwhile, regulatory pressure is building alongside cost pressure. The EU AI Act's transparency obligations for generative systems took effect in early August, imposing logging and traceability requirements that compression middleware must accommodate. A Cloud Security Alliance research note flagged supply-chain risks in the Model Context Protocol ecosystem, which now spans Claude Code, ChatGPT, Cursor, and dozens of other clients.

One research firm recently recommended that product leaders deploy tiered inference, routing, and orchestration to keep costs under control. A consulting group wrote that some bank customer-facing agents cost $20,000 to $30,000 per single-agent workflow at current list prices. Compression is one lever. The real winner, perhaps, will be the team that figures out how to pull all of them at once.

More stories

  • DoD Solution raises $2M for AI drone navigation in war zones
  • DesignVerse raises $5.5M to automate enterprise software
  • Maingen launches AI benchmark for industrial operations
  • Enact launches layer pushing robot success rates to 99%
  • Tenor launches AI workers that track token spend to outcomes
  • Marengo launches AI-native data center design platform
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.