Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Climate / Social Tech iconClimate / Social TechMarch 20, 2026

Beyond Reach Labs' Solar Arrays Expand 100x in Orbit to Power Space

Beyond Reach Labs' Solar Arrays Expand 100x in Orbit to Power Space
YcSpace Tech+3
SaaS iconSaaSMarch 20, 2026

Yann LeCun's AMI Labs Raises Record $1.03B European AI Seed Round

Yann LeCun's AMI Labs Raises Record $1.03B European AI Seed Round
World ModelsStartup Funding+3

Founders Mentioned

Ivan Zakazov

Compresr

saas icon
SaaS

Oussama Gabouj

Compresr

saas icon
SaaS

Kamel Charaf

Compresr

saas icon
SaaS

Berke Argin

Compresr

saas icon
SaaS

Ivan Zakazov

Compresr

saas icon
SaaS

Oussama Gabouj

Compresr

saas icon
SaaS

Kamel Charaf

Compresr

saas icon
SaaS
SaaS iconSaaS
March 20, 2026
YcAi InfrastructureLarge Language ModelsModel OptimizationB2b Saas

YC's Compresr Cuts AI Costs with Drop-In Context Compression API

As LLM context windows explode, YC W26 startup Compresr launches API promising up to 200x compression and 70% token savings—without sacrificing accuracy in agentic workflows.

YC's Compresr Cuts AI Costs with Drop-In Context Compression API

Here's a number that should, in theory, make every AI engineer sleep easier: token prices have cratered more than 280-fold between November 2022 and October 2024, according to Stanford's AI Index 2025 report. And yet, across engineering teams building the new generation of AI agents, the LLM bills keep climbing.

The explanation is deceptively simple. It's not the unit cost—it's the volume. When an agent pulls 106,000 tokens of SEC filing data for a single query, or when an MCP tool catalog consumes 50,000 tokens before any actual reasoning even begins, those plummeting per-token prices don't save you. They just make the underlying problem harder to see.

This tension is now driving a fresh wave of infrastructure startups, each betting that the next bottleneck in AI deployment isn't model capability or accuracy. It's cost management—specifically, learning how to use massive context windows without going broke in the process.

Context windows have exploded to a million tokens and beyond. Agentic systems, eager to exploit every inch of available space, happily fill them. What enterprises are discovering, often the hard way, is that inference has become less a technology problem than a capital allocation one.

When Bigger Isn't Better

The raw capacity exists. Anthropic's Claude models offer million-token windows; the company's Sonnet variant reportedly runs $3 per million input tokens and $15 per million output tokens, while Opus doubles those figures. Google's Gemini lineup reached multi-million token contexts in recent releases. The ceiling on what you can load into a prompt keeps rising.

Using that capacity efficiently? That's another story.

Microsoft's Azure team didn't mince words in a recent blog post, framing enterprise-scale inference as "a capital allocation trade-off between accuracy, latency, and cost." Translation: you can have two of the three, but rarely all at once.

The so-called "lost in the middle" problem—first documented in a 2023 computational linguistics paper and confirmed in subsequent work—persists even with state-of-the-art models. Cramming a million tokens into a single prompt doesn't guarantee the model will actually use them effectively. Position bias remains real. So does context rot. So does the basic challenge of retrieval quality when you're asking a model to sift signal from noise across hundreds of pages of text.

Prompt caching, now standard across major providers including OpenAI and Anthropic, offers discounts of 50–90% on cached input tokens. It helps, but it doesn't shrink your prompts—it just makes repeated static prefixes cheaper. And it's fragile. Change your context even slightly, miss the cache, and you're back to full price. Research published earlier this year under the title "Don't Break the Cache" evaluated caching strategies for long-horizon agent tasks and concluded that maintaining cache hits while adapting to dynamic contexts is, perhaps unsurprisingly, harder than it looks.

Enter the Compression Layer

Which brings us to companies like Compresr.

The four-person startup, fresh out of Y Combinator's Winter 2026 batch, is building what CEO Ivan Zakazov calls "LLM-native context compression." The founding team hails from EPFL: Zakazov earned his PhD there and logged time at Microsoft and Philips Research. CTO Oussama Gabouj comes from EPFL's DLab and a stint at AXA. COO Kamel Charaf brings data science credentials from EPFL and Bell Labs. CAIO Berke Argin rounds out the group with EPFL computer science training and prior experience at UBS.

Their pitch is straightforward enough: compress context, cut costs, improve performance. All without requiring developers to rewire their agent stacks.

The company's homepage advertises compression ratios "up to 200x without quality loss," though the most concrete public example is more modest—and arguably more believable. In a FinanceBench test case, Compresr reportedly compressed roughly 106,000 tokens of SEC filing data down to about 10,500 tokens, a 10x reduction. The result: 76% cheaper inference, with a slight accuracy bump to 74.5% versus a 72.3% baseline.

Compresr offers three models via API. "Latte_v1" provides query-specific, token-level compression. "Espresso_v1" is query-agnostic, also working at the token level. "Coldbrew_v1" handles chunk-level filtering. Developers specify a target compression ratio and plug the API into existing pipelines.

More intriguing, though, is the Context Gateway—an open-source proxy that sits between an agent framework and the LLM API. It monitors token usage, pre-computes summaries when context hits 75% of capacity, and compresses bulky tool outputs. The company claims up to 95% token reduction or 20x compression for large tool responses, with support for frameworks like Claude Code, OpenClaw, Codex, and OpenHands. Deployment involves a curl script or Docker container. The promise: "30–70% token savings" with zero code changes.

That last part matters. If you can drop in infrastructure that cuts your bill by half without touching your codebase, the ROI case is obvious.

A Crowded Field

Digital illustration for article section "A Crowded Field" in "YC's Compresr Cuts AI Costs with Drop-In Context Compression API" - A clean, minimalist, and dreamlike conceptual visualization representing a crowded competitive field...

Compresr isn't working in a vacuum. A cluster of competitors and open-source projects are exploring similar territory.

TokenCompress, updated in early 2026, advertises 87% savings on LLM API costs through real-time code context compression, with compatibility across OpenAI, Claude, DeepSeek, Mistral, and others. Compresso API pitches "advanced prompt compression capabilities designed to reduce token count while preserving meaning," according to its documentation.

Academic work has been prolific. Microsoft's LLMLingua and LongLLMLingua, published in 2023–2024, pioneered much of this space. LLMLingua-2, from March 2024, reported latency speedups of 1.6–2.9x at compression ratios of 2–5x. More recent papers—EFPC, SCOPE, FrugalPrompt, and Active Context Compression, all published in 2025 or early 2026—have pushed compression ratios higher while retaining or occasionally improving accuracy on certain benchmarks.

There's also a parallel research track focused on KV-cache compression at the serving layer rather than the prompt level. Systems like xKV, KVzip, Compactor, and FAEDKV claim 2–6.8x or greater KV compression on models like Llama and Qwen with minimal performance degradation.

And the open-source community is chipping in. IBM released MCP Context Forge as a gateway tool in March 2026. LiteLLM has been a popular proxy for cost management and rate-limiting since 2024. Reddit's LLMDevs and MCP communities are full of developers sharing homegrown compression proxies designed to handle MCP tool bloat or strip PII before API calls.

Real Money, Real Questions

The enterprise appetite for this is undeniable. Microsoft's recent blog post framed inference cost management as a core competency for AI deployment—table stakes, not optional. ARK Invest's Big Ideas 2026 report included "blended cost per million tokens" curves in its investor presentation, a sign that Wall Street is watching these economics closely. The vector database market, estimated at roughly $2–2.6 billion in 2025 with CAGR forecasts in the low twenties through the early 2030s, suggests the infrastructure layer surrounding LLM context is maturing into a real category.

Still, there are open questions.

Most performance claims for compression APIs come from vendors or academic papers using synthetic benchmarks. Independent, third-party evaluations on real-world production workloads remain scarce. Compresr's FinanceBench example is illustrative, but it's a single data point from the company's own marketing. It's hard to know how these systems perform across messier, more varied enterprise use cases.

Then there are the security wrinkles. Multiple papers published in 2025 and early 2026 have demonstrated that prompt caching, if not properly isolated, can leak information across tenants. Timing-based audits detected evidence of global cache sharing at certain providers. A paper titled OptiLeak demonstrated reinforcement-learning-driven prompt reconstruction from cache side channels, highlighting risks that earlier reports may have underestimated. As the EU AI Act phases in through 2026–2028, transparency and data handling obligations will pressure providers—and by extension, compression proxies—to document precisely what they're doing with customer data.

There's also the nagging possibility that providers simply integrate this functionality themselves, rendering third-party compression layers obsolete before they've had a chance to scale. Or that some other approach—progressive loading, smarter retrieval architectures, better tool design—solves the problem first.

Infrastructure in Search of Invisibility

Digital illustration for article section "Infrastructure in Search of Invisibility" in "YC's Compresr Cuts AI Costs with Drop-In Context Compression API" - A conceptual and minimalist visual representation of infrastructure becoming invisible through multi...

But the momentum is clear. Token costs are falling, yes. Agentic workloads and long contexts are growing faster. The industry appears to be converging on multi-layer compression: prompt-level, KV-cache-level, tool-catalog-level. Prompt caching handles repeated static prefixes. Compression tackles dynamic, bloated contexts. Together, they're becoming table stakes for production LLM economics.

Compresr's bet—and the bet of startups like it—is that context compression becomes infrastructure in the truest sense: boring, essential, largely invisible. The kind of thing you don't think about until it breaks. If they're right, compression joins caching, routing, and observability as a standard layer in the AI stack, another thin stratum of middleware that quietly keeps the gears turning.

If they're wrong, it's probably because the problem got solved one layer up—or because the big providers decided to bundle it in for free.

Either way, the problem isn't going away. The bills are still climbing, even as the per-token costs fall. Someone, somewhere, is going to have to pay for all those million-token context windows. The only question is whether it'll be infrastructure startups or the hyperscalers who get there first.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Beyond Reach Labs' Solar Arrays Expand 100x in Orbit to Power Space
  • Yann LeCun's AMI Labs Raises Record $1.03B European AI Seed Round
  • YC-Backed RunAnywhere Launches Platform for On-Device AI at Scale
  • AI Agents Come for Chip Design as Visibl Tackles Coordination Crisis
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.