A bank executive somewhere has probably had this nightmare: thousands of AI agents humming away in the cloud, each one burning through tokens at pennies per million, the bills piling up, and absolutely no way to know which ones are actually making money. That anxiety is no longer theoretical. Enterprise spending on generative AI surged from $11.5 billion in 2024 to $37 billion in 2025, according to Menlo Ventures data published in December 2025, and a June 2026 IDC blog found that 42% of organizations report assessing AI and generative AI return on investment is "difficult or impossible."
Tenor, a three-person startup that emerged from Y Combinator, launched this month with a pitch that finance chiefs might find familiar: treat AI agents like employees. Give them job descriptions, spending caps, and measurable goals. Then track whether they earn their keep.
The San Francisco company's founders say they are building what amounts to workforce infrastructure for machines. Each AI agent gets persistent context, guardrails, a budget, and a key performance indicator tied to a business outcome. The platform then attributes token consumption directly to results such as closed deals, processed invoices, or cleaned customer records. The idea is that engineering and finance teams can reallocate budgets based on which agents deliver actual returns, rather than guessing.
"Every piece of AI work starts with an outcome and a budget," Amgad Al-Zamkan, Tenor's co-founder and chief product officer, wrote in the company's YC Launch post on August 11. What happens between those two points has become a black box for most organizations.
The problem Tenor is addressing has drawn attention from the industry's largest players. OpenAI introduced granular enterprise spend controls on June 18, letting administrators set monthly limits per workspace, group, or user. Google Cloud previewed spend caps for its Gemini Enterprise Agent Platform in April and added FinOps features in August. AWS Bedrock published cost-allocation tags and budget patterns in early spring. Datadog brought its LLM and Agent Observability product to general availability in June with per-request cost estimation across more than 800 models. Ramp launched AI Token Spend Management in mid-July.
The rush to build spend controls reflects an awkward reality: per-token prices have collapsed while total spending keeps climbing. Stanford's AI Index 2025 showed that GPT-3.5-level inference fell from $20 to $0.07 per million tokens between November 2022 and October 2024, a 280-fold drop. Yet a FinOps X report from May 2026 found that 93% of organizations exceeded their AI budgets. Gartner predicted in April that the average Fortune 500 company will run more than 150,000 AI agents by 2028.
That trajectory explains why Deloitte's Q2 2026 CFO Signals survey found North American finance chiefs citing AI governance, cost uncertainty, and transparency as key concerns. The August 2 effective date for the EU AI Act's transparency obligations and the FTC's substantiation doctrine have added regulatory pressure. McKinsey warned in an August article that "retrofitting KPIs after launch is where attribution dies."
A handful of enterprises have published cost data that suggests the stakes are substantial. A leading U.S. bank working with Cognizant reported $15 million in savings over six months after modernizing its contact center with AI agents, tracked via handling time, customer satisfaction scores, and platform costs. A Middle East bank using DRUID AI reported a 40% drop in cost-to-serve for routine requests, according to a vendor-published case study. Klarna's investor materials from 2025 showed its AI assistant handled two-thirds of customer service chats within the first month and contributed to a 40% reduction in cost per transaction compared to the first quarter of 2023.
McKinsey published economics in late August showing that customer-facing agents can cost $20,000 to $30,000 to run a single-agent workflow, while multi-agent teams can run $100,000 to $200,000. Andreessen Horowitz wrote three days later that companies delivering "a clear and attributable business result" should price the outcome, not the token.

Tenor's founders bring technical credentials to the space, though the company has not disclosed customers, funding, or revenue. Hamze Al-Zamkan previously served as president of TUM.ai and deployed AI at appliedAI. Amgad Al-Zamkan worked at Amazon and studied computer science at UC Berkeley and the Technical University of Munich. Muhtasham Oblokulov was an early contributor to DeepSeek, co-authored the StarCoder and StarCoder2 papers, and contributed to the CodeClash benchmark with Stanford researchers before joining Munich Re's machine learning team.
The company positions itself not as an observability layer or FinOps dashboard but as something closer to an operating system for AI labor. Agents run until the outcome ships or the money runs out. "That's what we're building Tenor for," the founders wrote, "not just to measure the token economy, but to give companies the machinery to operate inside it."
Whether attribution at the agent-outcome level becomes standard practice or remains a specialist discipline will likely depend on how quickly enterprises professionalize AI operations. OpenAI wrote on June 18 that "as AI becomes part of everyday work, organizations need the ability to manage it with the same rigor they apply to any critical business investment." Tenor is betting that finance teams will demand that rigor sooner rather than later, perhaps well before the 150,000-agent future arrives.

