Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Fintech iconFintechAugust 22, 2026

Arbital launches trading terminal with $1.35B in volume

Arbital launches trading terminal with $1.35B in volume
YcCrypto Trading+3
SaaS iconSaaSAugust 22, 2026

Studio launches AI to predict human response to new ideas

Studio launches AI to predict human response to new ideas
YcAi+3

Founders Mentioned

Hannah Chung

TrustAI

saas icon
SaaS

Medha Venkatapathy

TrustAI

saas icon
SaaS

Hannah Chung

TrustAI

saas icon
SaaS

Medha Venkatapathy

TrustAI

saas icon
SaaS
SaaS iconSaaS
August 22, 2026
YcAi InfrastructureCost OptimizationModel Optimization

TrustAI cuts AI inference costs 20% with cache technology

YC-backed startup uses cross-model KV cache transfer to slash inference costs and latency as AI inference spending overtakes training for the first time.

TrustAI cuts AI inference costs 20% with cache technology

TrustAI wants to change how AI systems burn through compute budgets, and the five-person team from MIT thinks it has found a way to do so without sacrificing quality. The Y Combinator-backed startup announced Tuesday that it claims to cut inference costs by roughly 20 percent and slash time-to-first-token by a factor of 2.5, using what it calls cross-model KV cache transfer. The technique sounds arcane, but the mechanics are straightforward: run the prefill phase on a small model, then hand the work off to a large model for decode.

The timing matters. Inference costs are rising dramatically and projected to overtake training costs, with Gartner data released in August showing that inference costs per agentic workflow will balloon more than fivefold by 2028. Inference has quietly become the cost-of-goods-sold of intelligence. Google disclosed in May that it was processing 3.2 quadrillion tokens per month across its products, a figure cited in a July blog post by a16z partners Raghu Raghuram and Sarah Wang. Gartner pegged total AI spending at $2.52 trillion in 2026, with $1.37 trillion flowing into AI infrastructure, as TechRadar Pro summarized in August. Training happens in bursts; inference runs continuously at scale, and that shift has upended the economics.

Prefill latency has emerged as the bottleneck. In long-context and agentic workloads, time-to-first-token is no longer about decode speed but about how fast you can load the context into memory.

The Fragmented Optimization Landscape

Inference optimization has splintered into competing approaches: quantization, speculative decoding, KV cache reuse. AWS wrote in April that speculative decoding "can accelerate token generation by up to 3× for decode-heavy workloads," but Yahav Biran, a principal architect at AWS, and Truong Pham of Annapurna Labs noted that "TTFT remains effectively unchanged… dominated by the prefill phase." OpenAI introduced prompt caching in October 2024, offering a 50 percent discount on cached input tokens that persist five to ten minutes. Anthropic followed with block-level cache breakpoints and optional one-hour persistence at higher write prices, according to platform documentation updated in 2026.

vLLM added FP8 KV-cache quantization in April, and Together AI reported in March that cache-aware prefill-decode disaggregation delivered "up to 40% higher sustainable throughput" on mixed warm-cold long-context workloads, according to engineers Jiejing Zhang and colleagues.

None of those techniques, however, transfer KV state between model families. That's where TrustAI diverges. The startup's approach, based on an August arXiv paper titled "Cross-Model KV Cache Transfer in LLM Families," uses a closed-form linear mapping to reuse prefill work from a small model for decode on a large one. The company's homepage displays a 100,000-token example: standard prefill plus decode on a 14B parameter model takes 4.3 seconds, while KV cache transfer from Qwen3 4B to 14B decode cuts that to 1.7 seconds. TrustAI claims prefill cost for one million requests drops 61 percent, from $11,800 to $4,600.

The startup declined to disclose revenue, customer names, or funding details beyond its Y Combinator backing.

Infrastructure Under Pressure

Data center electricity demand for AI workloads is forecast to represent 15 to 25 percent of total consumption, according to EPRI's "Powering Intelligence 2026" executive summary. The IEA reported that data center electricity demand grew roughly 17 percent in 2025, with AI coupling driving the increase. FERC ordered regional transmission organizations and independent system operators in June to reform large-load interconnection processes to speed AI data center hookups, underscoring the infrastructure strain.

Agentic workloads amplify the economics in ways traditional inference does not. The vLLM project reported in May that integration with the Mooncake distributed KV cache store improved throughput by 3.8× and reduced median time-to-first-token and end-to-end latency by 46× and 8.6×, respectively, on agent traces, according to authors Yifan Qiao and colleagues. Those traces averaged a 131:1 input-to-output token ratio and 33-turn median session length on SWE-bench-like workloads, profiles where shared prefixes dominate cost if not cached. Cache hit rates jumped from 1.7 percent to 92.2 percent with the distributed pool, the vLLM team wrote, and scaling remained near-linear to 60 GB200 GPUs.

Digital illustration for article section "Content Section 3" in "TrustAI cuts AI inference costs 20% with cache technology" - A minimalist and conceptual digital artwork representing amplified throughput and distributed cache ...

Raghuram and Wang at a16z wrote in July that "inference is the COGS of intelligence," framing specialized inference hardware and system-level optimization as the economic battleground at scale. Gartner's August press release framed rising costs as an "inference paradox." Enterprises realize value via inference, not training, but per-workflow costs are climbing faster than efficiency gains can offset them.

A Team Betting on State Mobility

TrustAI is built by CEO Hannah Chung and CTO Medha Venkatapathy, both MIT engineers, plus three additional MIT researchers, according to the company's About page. The team is advised by Robert Shaw, vLLM lead at Red Hat AI, and a Google Distinguished Engineer in AI Infrastructure. The startup's vision statement reads, "Any model should be able to pick up where another left off," positioning cross-model KV transfer as state mobility across architectures.

Y Combinator's directory page originally described TrustAI as "continuous compliance and governance for agents on sensitive systems," suggesting a pivot toward inference optimization; the current site emphasizes KV-transfer technology.

TrustAI published internal research notes between August 9 and August 17, reporting 85.1 to 91.2 percent retained task quality and 65 percent next-token agreement in lower-parameter mappings, along with divergence reductions and partial recomputation strategies, according to blog entries on the company site. The research has not been independently replicated in peer-reviewed venues as of late August. The company said its approach is anchored in the August arXiv paper on cross-model KV transfer, which details closed-form linear mappings for prefill reuse across model families.

Broader KV cache reuse and disaggregation has moved into production-grade tooling, though much of it stays within model families. vLLM baked in FP8 KV caching, kv_transfer connectors, and distributed KV stores in 2026 releases, according to project documentation. SGLang matured RadixAttention for prefix reuse with sliding-window-aware constraints noted in 2026 GitHub issues. NVIDIA's Dynamo and NIXL documentation describes KV transfer over GPUDirect RDMA with topology-aware routing for disaggregated prefill-decode workers.

AWS showcased SageMaker HyperPod in July using vLLM router plus LMCache over EFA and NIXL for disaggregated prefill, according to a blog post. Together AI's cache-aware prefill-decode system ran on Blackwell B200s and demonstrated 35 to 40 percent sustainable QPS uplift in tests, engineers wrote in March.

Academic papers published mid-year include SmartGen (selective KV cache transfer, July), SpectrumKV (per-token mixed-precision transfer policies, June), and C^2KV (compressed and composable KV cache reuse claiming up to 17× speedup, July), according to arXiv postings. Production adoption status for these techniques remains unclear.

What to Watch

Cross-model KV transfer represents emerging research with caveats around quality retention, positional and rotary alignment, layer-wise mapping, and partial recomputation, according to the August arXiv paper TrustAI cites. Independent benchmarks and customer case studies are not yet publicly available. TrustAI's claimed 20 percent cost reduction and 61 percent prefill savings are vendor-provided figures anchored to internal benchmarks; engineers evaluating the approach should validate retention metrics and model compatibility before production deployment.

The economic tailwind is clear. Companies deploying agentic systems with long contexts and multi-turn sessions face prefill-dominated latency and ballooning token bills. Cross-request cache reuse, whether intra-model via vLLM's distributed stores or cross-model via linear mappings, offers a structural cost lever that speculative decoding and quantization alone cannot provide.

Digital illustration for article section "Content Section 6" in "TrustAI cuts AI inference costs 20% with cache technology" - A minimalist and conceptual representation of economic efficiency and streamlined data caching, feat...

TrustAI's bet is that model families will share enough structure for cheap prefill-to-decode handoffs. If quality retention holds at scale, inference pipelines could run small models for context ingestion and large ones for generation, splitting the bill. The five-person team declined to share named customers or hiring targets, but the company's advisor bench signals credibility in a field where production-grade KV systems are written by hyperscalers and inference platforms, not startups.

What founders should watch: independent validation of cross-model transfer quality, named deployments, and whether vLLM or SGLang merge native support for the technique. Until then, TrustAI's approach remains a promising footnote in a field moving fast enough that yesterday's novelty becomes tomorrow's default.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Arbital launches trading terminal with $1.35B in volume
  • Studio launches AI to predict human response to new ideas
  • Lyon builds private AI models for bank transaction data
  • Traceforce launches on-device security for AI agents
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.