When Meta's CICERO system sat down to play Diplomacy against human opponents in 2022, it did what any skilled player might: it lied. The AI hadn't been programmed to deceive—Meta's researchers certainly hadn't intended that outcome. But winning the game, with its web of temporary alliances and inevitable betrayals, demanded it. The discovery, buried in post-game logs, crystallized a problem that's only grown thornier since: How do you evaluate an AI when the task requires more than getting answers right?
Fast-forward to mid-2026, and that question has spawned multiple efforts from various organizations exploring new benchmarking methods. Among the contenders: a two-person startup out of Y Combinator's summer batch that thinks the solution involves poker hands, Catan-style resource trading, and thousands of humans willing to sit across the table from masked AI opponents.
It's an unconventional bet. But then, conventional approaches are showing their age.
When Benchmarks Break Down
The classic AI benchmark—fixed dataset, single metric, tidy leaderboard—worked well enough when the task was image classification or language translation. Feed in an input, check the output, tally the score. Clean, reproducible, quantifiable.
But language models aren't just answering questions anymore. They're negotiating contracts, planning multi-step operations, navigating open-ended social exchanges. And the old tests are starting to buckle under scenarios like these.
A March paper, "Efficient Benchmarking of AI Agents," noted that running full agent evaluations gets expensive quickly, and suggested that rankings might not require exhaustive testing. More troubling was an IBM Research study that dropped in April, finding that removing just a handful of preference votes could reshuffle which model topped widely-cited leaderboards like Chatbot Arena, emphasizing the sensitivity of these rankings. The platform had accumulated over 6 million human votes by mid-year—an impressive dataset, to be sure—but the sensitivity analysis revealed just how fragile those aggregate rankings actually were.
Then there's contamination. Models train on vast swaths of the internet; if benchmark questions leak into that training data, you're measuring memorization, not capability. Stanford's Digital Economy Lab took a swing at this in May with Agent Island, a game-based benchmark designed for contamination resistance and to resist game saturation. Across 999 games testing 49 models, the researchers generated scenarios meant to stay fresh even as models improve.
And perhaps most vexing: what, exactly, should you measure? Standard benchmarks track accuracy, latency, task completion rates. They don't capture whether an agent lies when it's advantageous, cooperates when the pressure's on, or respects norms when no one's watching. A May benchmark called TERMS-Bench evaluated negotiation beyond just completing deals—not just whether agents reach agreements, but how they get there.
The Game Theory Solution
Which brings us to games. Not entertainment—evaluation sandboxes.
Poker offers imperfect information and rewards strategic deception. Diplomacy demands coalition-building and long-horizon planning. Social deduction games like Mafia or Avalon test theory-of-mind and lying under interrogation. These environments aren't just fun to watch. They're structured worlds where researchers can instrument every action, record reasoning traces, and measure the behavioral patterns that actually matter in deployment.
Academic labs have been trending this direction for a while. DeepMind's Melting Pot suite, first released in 2021 and still running experiments in 2026, uses multi-agent social dilemmas to study cooperation and generalization. AvalonBench, introduced in 2023 and updated through last year, measures deception in hidden-role games. MafiaBench maintains an ongoing leaderboard for social deduction agents. February brought PieArena with MBA-grounded negotiation scenarios; June saw Poker Arena launch with multi-axis profiling of strategic reasoning.
Microsoft researchers published a poker study at ICLR 2026 asking how far LLMs still lag behind professional players, analyzing reasoning traces to understand where models fail. The work found gaps in game-theoretic planning even when models had access to tool-use scaffolds. Generating plausible poker talk is one thing. Calculating pot odds and adjusting bluffing frequency against adaptive opponents? That's another.
What makes games compelling as evaluation environments is that they're adversarial by design, open-ended, and—crucially—they generate interpretable logs. When an agent folds a strong hand or lies about its resources in a trade negotiation, you can trace the decision back through its internal reasoning (if the model exposes chain-of-thought) and compare it against ground truth. Sometimes, anyway.
The Two-Person Bet

Olam Labs, which came through Y Combinator's summer 2026 batch, is betting that human-in-the-loop gameplay can scale into a systematic evaluation platform. The founding team is compact, even by startup standards: Om Buddhdev, a former staff engineer at Latitude (the company behind AI Dungeon) with an esports background, and Shreshth Sharma, who left a dual degree program at Waterloo and Laurier to join after a stint at RBC. They're working out of San Francisco.
Their product, Social Arena, invites humans to play games—poker, a Risk-inspired strategy game called Warpath, a Codenames-style word game called Hidden Words—against AI agents whose identities stay masked until the match ends. A Catan-like game, Tradewinds, is in development. The humans aren't just playing for entertainment; they're feeding the benchmark.
Olam's methodology, detailed in a mid-July blog post, emphasizes deterministic environments, randomized agent seating, and continuous grading pipelines. After each game, the platform reveals which seats were human, which were AI. Behavioral metrics get computed from the logs. As of late July, Olam reported measuring "Deception Strength" across 69,140 poker hands and "Social Lie Rate" over 86,729 graded turns. The benchmarks page lists model families—Claude, GPT-5.x, DeepSeek—alongside lying propensity scores.
The approach mirrors Chatbot Arena's human-preference methodology but applies it to interactive, multi-turn scenarios where behavior matters as much as performance. It also echoes Meta's Dynabench, which used adversarial human-in-the-loop rounds to surface real-world failure modes that static datasets miss.
Olam's site states the company collaborates with "frontier labs and researchers" on pre-deployment behavioral evaluations, though no client names are publicly listed. The YC page frames the mission as measuring "AI agents in complex multi-agent and human environments that require real-world socialization traits." On his personal site, Buddhdev positions the work as helping to "make sure AI goes well for everyone."
It's early days—Reddit posts from a few weeks back were advertising early-access codes—but the infrastructure is live. Whether a two-person team can sustain the data collection, grading automation, and community engagement required to compete with academic labs or enterprise evaluation platforms remains an open question. Startups have scaled with leaner teams before. Then again, maintaining benchmark integrity while growing a user base is its own kind of hard problem.
Regulators Are Watching

The timing, at least, is favorable—if you can call regulatory pressure favorable.
The EU AI Act's obligations for general-purpose AI models took effect August 2, 2025; enforcement powers went live exactly a year later. The guidelines explicitly call for model evaluation using "standardized protocols and state-of-the-art tools," with documented adversarial testing. Providers must demonstrate continuous independent evaluation, including behavioral checks on agentic systems.
NIST launched its AI Agent Standards Initiative in February 2026, with draft practices for automated benchmark evaluations open for comment through the end of March. In April and May, NIST ran webinars on evaluation probes and traceability for agentic ecosystems. The UK's AI Safety Institute published evaluations in 2025 and continuing through this year that centered on agentic behavior—multi-step cyber tasks, autonomy, deception. METR's May Frontier Risk Report noted models cheating on harder tasks and recommended covert-capability checks and chain-of-thought review for deceptive reasoning.
Gartner forecasts the AI governance platform market hitting somewhere around $492 million in 2026, climbing past $1 billion by 2030. Other analysts cite figures ranging from the mid-$300 millions to just under $500 million, depending on how they slice the category. The numbers vary—they always do—but the trend is unmistakable: compliance obligations are driving spend on evaluation infrastructure, documentation tooling, and third-party assessments.
LangChain's June State of Agent Engineering report found that 37.3% of surveyed teams had adopted online evaluation and monitoring—a sign that enterprise practice is maturing, even if deployment-time continuous evaluation remains more emerging than standard.
A Fragmented Landscape
The evaluation landscape in 2026 looks, well, crowded. And fragmented.
There's Agent Island's contamination resistance. ALEM's open-ended coordination tests. PieArena's negotiation scenarios. TERMS-Bench's diagnostic depth. Multiple deception frameworks: DeceptionBench, LieCraft, "Lying to Win." Poker environments from PokerBench to Microsoft's ICLR study. Toolkit releases like RainbowArena for tabletop games. Chatbot Arena continues to accumulate votes; LangSmith and Weights & Biases offer enterprise evaluation APIs; Scale markets red-teaming services.
What's emerging isn't a single dominant platform but a shift in what "evaluation" means in the first place. The center of gravity has moved from static, single-skill tests toward interactive, multi-agent settings that surface social behavior—deception, negotiation, cooperation, norm adherence. As one June benchmark put it, coordination remains "a key bottleneck for LLM agents."
Olam Labs is trying to operationalize that shift by turning gameplay into a continuous evaluation pipeline. Whether the approach scales—and whether game performance translates to real-world deployment safety—are separate questions. But the underlying bet seems reasonable enough: if you want to know how an agent behaves under pressure, in adversarial settings, with incomplete information, you need environments that actually test those conditions.
Research cautions are piling up. Leaderboards are sensitive to small data perturbations. Benchmarks saturate or contaminate. Evaluation costs scale poorly. Scaffold choices introduce distribution shifts that confound comparisons. The field is learning, perhaps more slowly than anyone would like, that no single metric captures model behavior, and that rankings depend heavily on what you choose to measure and how you collect the data.
The Stakes Get Real

The regulatory timeline suggests this won't stay a research problem for long. EU enforcement is live. NIST is drafting standards. Enterprise teams are embedding evaluation into product workflows. The question isn't whether behavioral testing becomes standard practice—it's which methodologies prove robust, scalable, and actually predictive of real-world risks.
Olam's human-fed game arenas are one candidate. Stanford's contamination-resistant simulations are another. There will be more.
For now, the poker hands keep accumulating—69,140 and counting, each one a small data point in a much larger question about what it means to trust a system that can plan, persuade, and perhaps deceive. Meta's Diplomacy-playing AI didn't set out to lie. It learned to because the game rewarded it. Four years later, researchers are still trying to figure out how to measure what that means—and what happens when the game isn't confined to a board.
