OneTriangle, a San Francisco startup backed by Y Combinator, has opened a hosted inference cloud built around an unusual technical gamble: transferring key-value cache between language models of different sizes to reduce compute costs and speed up response times. The company claims the approach cuts prefill expenses and speeds up first-token delivery compared to conventional methods, though these performance figures are vendor-supplied and have not been independently verified.
The premise is straightforward, if technically intricate. Instead of forcing a large model to process an entire prompt from scratch, OneTriangle runs the input through a smaller, cheaper model first. It then hands off the cached internal state to the bigger model, which skips most of the expensive prefill computation and jumps straight to generating an answer. It's a form of computational arbitrage that depends on something the industry hasn't seen deployed at scale: moving those intermediate tensors between models that weren't trained together.
"OneTriangle transfers the KV cache between models, something that has never been done in production," CEO Hannah Chung wrote in the company's Y Combinator launch announcement. The technical details involve stripping rotary embeddings from the smaller model's key-value tensors, projecting them into the larger model's attention space, then re-embedding so the big model can decode against a cache it never computed. If a model pair fails quality or latency checks, the system falls back to standard prefill.
The company published internal benchmarks as illustrative examples of the approach's potential. One shows a 7.9-fold speedup in time-to-first-token for 8,000-token contexts when transferring cache from Minitron-4B to Llama 3.1 8B. Another example on OneTriangle's homepage illustrates a Llama 3.1 8B-to-70B transfer delivering first tokens 6.3 times faster, cutting response time from 16.9 seconds to 2.7 seconds. The company estimates that for a million requests, prefill costs would drop from $46,900 to $7,500 under that configuration. These figures are vendor-supplied calculations and have not been verified by outside parties.
OneTriangle went live with hosted DeepSeek V4 Flash in late August. Current pricing is $0.15 per million input tokens on cache miss, $0.003 per million input tokens on cache hit, and $0.60 per million output tokens for off-peak delivery, with peak pricing at twice the off-peak rate. A second tier, labeled "Flex / Delayed," costs 40 percent less and promises delivery within ten hours. Other models on the platform include Llama 3.3 70B Fast at $0.024 per million input tokens, Qwen 3.6 27B, and a pair of safety classifiers. OneTriangle says its Llama 70B rates represent "rounded GPU-only breakeven at 100% productive measured throughput," a pricing claim that assumes full utilization.
RuntimeWire, covering the DeepSeek hosting launch, noted that Chung's assertion of "fastest lightweight, cheap inference" had not been established through independent comparison. The startup operates in a crowded inference market where performance claims are common and verification is rare.
A Team With Academic and Trading Roots
Chung studied computer science and economics at MIT and spent time at Virtu Financial, a high-frequency trading firm where microseconds matter and infrastructure efficiency is doctrine. She co-founded OneTriangle with Medha Venkatapathy, who holds degrees in physics and computer science from MIT. The company describes a team of MIT engineers and researchers with stints at Google DeepMind, Jane Street, SpaceX, MIT Lincoln Lab, and CSAIL. Advisers include Robert Shaw, who leads vLLM development at Red Hat AI, and a Google Distinguished Engineer working on AI infrastructure.
Dealroom News reported the company previously operated as TrustAI and raised a $125,000 Y Combinator seed round in June 2024. OneTriangle has not disclosed additional financing.
Cache Management as Competitive Edge

The startup's approach arrives as key-value cache optimization becomes a focal point in inference economics. Inference Radar, an industry newsletter, described KV cache management as "the new battleground," arguing that memory movement and compression now rival raw compute as cost drivers. Academic work published earlier this year by Heo and colleagues explored cross-model KV cache transfer within language model families, reporting speedups ranging from 2.7 to 25 times compared to re-prefill. That research found linear structures enabling closed-form mappings between models, validating the theoretical basis for what OneTriangle has commercialized.
NVIDIA, VAST Data, and Dell have all issued infrastructure guidance on KV offload and transfer in recent months. Open-source inference servers like vLLM and SGLang have pushed optimization work in similar directions, though production deployments of cross-model cache transfer remain sparse.
"No API changes, no quality cliff, and we're upstreaming it into vLLM," Chung wrote, signaling the startup's intention to contribute its cache transfer code to the widely used open-source project. Whether that contribution materializes, and whether it gains traction among developers wary of vendor lock-in, will help determine if OneTriangle's technical bet pays off. For now, the company is offering a service that treats AI inference less like a monolithic compute problem and more like a routing puzzle where the right handoff can save both time and money.
