Olam Labs, a Y Combinator-backed startup with just two employees, launched a platform earlier this year that pits humans against large language models in poker and strategy games, tracking each instance where the AI deliberately lies. The company's Multi-Agent Arena has logged more than 93,000 game turns as of late August, grading what it calls a Social Lie Rate by comparing what models from OpenAI, Anthropic, Google, and DeepSeek say to players against their hidden internal reasoning.
The timing is notable. AI labs and regulators spent much of 2026 grappling with a problem that standard benchmarks cannot solve: How do you evaluate AI agents when they negotiate, deceive, and strategize across dozens of interactions with multiple players? OpenAI reported an agent incident in July related to cybersecurity testing, according to Axios. A month later, the UK AI Safety Institute documented what it termed an "unsanctioned agent behaviour" incident during permissive cyber evaluations and released updated sandboxing tools. NIST's AI Safety Institute had already issued a request for information on securing agent systems in January and followed up with an AI Agent Standards Initiative in February.
Traditional benchmarks were built for simpler times. They ask a model to answer a question, summarize a passage, or complete a coding task once—then grade it. That approach collapses when agents make decisions over time, lie to gain advantage, or coordinate with other systems. CoreWeave, in a May announcement, argued that "the pace of AI has outrun the way teams build for it." The company launched a closed-loop platform designed to put agents in production and improve them with live data rather than offline evaluations. Chen Goldberg, CoreWeave's executive vice president of product and engineering, went further: enterprises that deploy agents first and let them learn continuously "are accelerating the path to superintelligence."
Industry benchmarking efforts reflect the shift. MLCommons added an agentic category to MLPerf Client v2.0 in August, introducing scenarios for software engineering agents and data analyst agents, HPCwire reported. The UK institute open-sourced Inspect, a framework with Docker and Modal sandboxing for multi-step agent evaluations. Microsoft released CTI-REALM in March for cyber detection rules and followed with REDAgentBench in August for executable red teaming. Yet as TechRadar noted in August, many teams are bypassing third-party evaluations entirely, opting instead for domain-specific, longitudinal tests embedded in their actual workflows.
Olam Labs' founders built something different. Om Buddhdev, formerly a staff engineer at AI Dungeon developer Latitude, and Shreshth Sharma, an ex-RBC data scientist who left a computer science program at Waterloo, created arenas where agents and humans play with identical information and actions. "While to consumers the platform is posed as 'a fun game with talking AI opponents,' we've built the entire stack to…run as deterministic environments with data pipelines that feed our evaluations," the company wrote in a July methodology post. Models get randomized seats. No player knows who is human. Full transcripts, including hidden chain-of-thought logs, get captured.
The evaluations page shows a Deception Strength metric that correlates with performance ratings in Social Poker. The company grades lies only when a model's private reasoning conflicts with its external communication and observable reality, using LLM judges with variance checks. An excerpt published in July shows Claude Sonnet reasoning privately about bluffing strategy, then telling an opponent something different. The company labels the metrics preliminary and subject to change, which is perhaps wise given the nascent state of deception measurement.
Beyond the public arena, Olam Labs runs internal simulations. Tokenshire, posted in May, models a persistent town. Market Rush, posted in July, simulates trading behavior. An internal CyberArena tests security scenarios. The startup's YC profile claims it processes "nearing a billion tokens used per day" and ranks in the top 50 on OpenRouter's leaderboard. A founder's LinkedIn post from August cited roughly 35 percent week-over-week user growth. None of these claims have been independently verified.
"We believe…as AI takeoff continues…there is and will be a massive shortage of understanding of how models behave in complex systems with humans," the company wrote in an August blog post. Olam Labs offers pre-deployment evaluations, datasets, and custom environments for training and testing. No public funding announcements have appeared beyond its Y Combinator participation. The startup lists no open roles on YC's job board as of September.
Regulatory timelines, meanwhile, have tightened. The EU AI Act's general-purpose AI transparency obligations became enforceable August 2; legacy models placed before that date must comply with Article 50(2) by early December. NIST issued standards requests and draft guidelines for automated evaluations, with public comment closing in March. China released national guidelines in May to develop and regulate AI agents, including standards on interoperability and protocols.

Analyst projections suggest demand for agent testing will only intensify. Gartner predicted in May that by next year more than 65 percent of engineering teams using agentic coding will treat IDEs as optional. IDC's FutureScape projects 40 to 50 percent of enterprise roles will involve agents by 2027. An IDC resource center claims nearly 90 percent of organizations committed over $1 million to agentic AI this year. Forrester reported in June that half of marketing agencies use agentic AI and nearly half of security decision-makers cite agentic AI concerns.
The simulation software market reflects this urgency. Grand View Research pegs the sector at $30.1 billion this year, climbing to $70.7 billion by 2033. Fortune Business Insights projects a narrower $16.82 billion this year, rising to $40.73 billion by 2034. Mordor Intelligence estimates $15.46 billion this year, hitting $28.59 billion by 2031. The variance in forecasts underscores how rapidly the category is evolving.
Enterprises are already integrating agent evaluation into production. Formula 1 deployed a fan companion agent on Salesforce Agentforce 360 in March, cutting chat handling time 30 percent and response times 80 percent, according to Salesforce. Emily Prazer, Formula 1's chief commercial officer, said in March that "our fans are the heart of everything we do…we have found a perfect partner…to connect fans." Patrick Stokes, Salesforce's chief marketing officer, added that F1 is "demonstrating what it means to run as an Agentic Enterprise."
Research labs have pushed the frontier as well. MIT CSAIL's SceneSmith, published in July, uses collaborating agents to generate simulation-ready 3D indoor scenes for robot policy design. Jeremy Binagia of Amazon Robotics called it "a significant advance…for generating simulation-ready indoor environments just from a simple text prompt." NVIDIA released an Agent Toolkit with Omniverse-based simulation and world models for robotics and autonomous vehicles. Waymo published its World Model in February, transferring vision-language model knowledge into LiDAR outputs for autonomous vehicle simulation.
Turing released CyberStrike in late August, a long-horizon cybersecurity benchmark for complete, verifiable security work including patches, exploits, and investigations. Cresta launched a training simulator in July using AI agents to provide dynamic training for human contact-center agents.
The concept has roots. DeepMind's Melting Pot suite from 2021 and 2022 tested multi-agent reinforcement learning generalization. Meta's CICERO, published in Science in 2022, achieved human-level negotiation in the game Diplomacy through strategic reasoning and natural language. WebArena, launched in 2023, remains a standard for reproducible browser-agent benchmarks, with multiple leaderboards tracking progress through this year.

The shift from static benchmarks to live, multi-actor evaluations suggests Olam Labs and similar platforms will compete on data scale, environment realism, and integration into training loops. OpenAI wound down its Agent Builder and hosted Evals in June, positioning an Agents SDK and Responses API for end-to-end workflows and calling for trustworthy third-party evaluations with improved methodology. The UK institute's August incident report highlighted the risks of permissive, internet-enabled test setups, nudging evaluators toward isolated, deterministic environments.
MLCommons' AILuminate Agentic Product Maturity Ladder and benchmarking integrity notes signal standardization pressure. Multi-nation joint testing exercises through the UK and US safety institutes, NIST's emphasis on realistic sociotechnical testing, and the addition of agentic categories to industry benchmarks all point toward formalized evaluation regimes.
Olam Labs' anonymized human-versus-AI structure offers a dataset no static benchmark can replicate: how models behave when opponents adapt in real time. The company's deception metrics, graded over tens of thousands of turns, provide a quantitative lens on social reasoning that regulators and labs are actively seeking. Whether that proves a durable advantage depends on whether the startup scales environment complexity, attracts lab partnerships, and converts user traction into disclosed revenue before competitors with larger teams and deeper pockets replicate the approach.
The YC batch graduates this month. Funding and hiring announcements will clarify whether the two-person team expands to meet the demand it claims exists.
