A San Francisco startup thinks the burgeoning market for artificial intelligence has a Wall Street problem: nobody quite knows what they should pay for computing power months from now.
Touchmark, which emerged from Y Combinator's latest cohort and launched on August 14, 2026, is betting that futures markets—the financial plumbing that lets oil refiners lock in crude prices and airlines hedge jet fuel costs—can bring similar clarity to companies buying inference capacity from AI models. The two-person team has built what appears to be the first trading venue where businesses commit to blocks of tokens one to three months out, often at discounts that the company says can reach 35 percent below list rates.
The premise rests on a familiar arbitrage: providers with idle GPU clusters want guaranteed revenue; buyers planning large deployments want predictable costs. Touchmark sits in the middle, structuring prepaid contracts for fixed token volumes on open-weight models and allowing participants to offload commitments they no longer need.
Co-founder Ilia Bolgov, who spent time in product roles at Revolut's wealth and trading arm, framed the friction plainly in an interview around the launch. "Pricing for reserved capacity is bilateral and buyers have very little visibility into what a fair forward price actually is," he said. His co-founder, Roman Yanushevskyi—a gold medalist at the 2022 International Olympiad in Informatics who interned in quantitative research at Citadel Securities—added that finance had already cracked these problems at scale.
Whether inference tokens will trade like pork bellies remains an open question. But the mechanics Touchmark has assembled suggest the founders are serious about replicating derivatives infrastructure, not just running a discount club.
A buyer selects a model, commits to a token volume and delivery window, then pays upfront. The company displays illustrative examples on its site: 1 billion tokens of a model called GLM 5.2, deliverable between September and October, priced around $3,960 against a roughly $4,400 list figure. The startup meters actual usage as it routes traffic to the provider's endpoint and settles payments through a third-party processor.

More ambitious is a request-for-quote board where buyers post specifications—model family, token count, throughput minimums, delivery month—and let providers compete to undercut the visible best offer. Accepting a quote binds both sides to a contract, a structure that carries echoes of over-the-counter energy trading desks.
Touchmark restricted the platform to commercial producers and end-users, language that tracks commodity forward contracts rather than exchange-traded futures open to speculators. That choice reflects advice from Tölt Strategies, a consultancy led by Dorothy D. DeWitt, who previously directed the Commodity Futures Trading Commission's Division of Market Oversight. In a statement issued days after Touchmark's debut, DeWitt said her firm "played a key role" shaping documentation and trading rules, and noted "extraordinary interest" in GPU infrastructure buildouts.
Initial supply tilts toward Chinese open-weight families—Kimi, GLM, Qwen—delivered through Wafer, another Y Combinator-backed inference provider that the startup named as its first counterparty. Touchmark also teased contracts pegged to "whoever is #1" on a public leaderboard at delivery time, a hedge for buyers chasing frontier performance without committing to a specific unreleased model. The company labeled displayed data "illustrative," a hedge of its own while liquidity remains thin.

Funding details are sparse. Y Combinator and Inception Fund have backed the venture, along with an investor the startup identifies only as "Pareto," though this could not be independently verified. Touchmark declined to share amounts or confirm team size beyond the two named co-founders.
The broader question is whether AI inference—still a young, volatile market—will develop the standardization and price discovery that make commodity futures work. Oil contracts succeed partly because West Texas Intermediate crude is West Texas Intermediate crude, wherever it trades. Token outputs from different model versions, hardware configurations, and throughput guarantees are harder to compare. Bolgov and Yanushevskyi are wagering that enough buyers care more about cost certainty than perfect apples-to-apples benchmarks, at least for now.
In his August statement, Yanushevskyi said the team looked forward to developing "additional novel products" with Tölt's guidance. Translation: if this works, expect more exotic instruments.

