Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSFebruary 14, 2026

Simile Raises $100M to Build AI That Predicts Human Behavior

Simile Raises $100M to Build AI That Predicts Human Behavior
Digital TwinsBehavioral Analytics+3
SaaS iconSaaSFebruary 14, 2026

D-Wave's $400M Public Raise Powers Quantum Computing Push

D-Wave's $400M Public Raise Powers Quantum Computing Push
Quantum ComputingFunding+3

Founders Mentioned

Demis Hassabis

Google DeepMind

saas icon
SaaS

Demis Hassabis

Google DeepMind

saas icon
SaaS
SaaS iconSaaS
February 14, 2026
Artificial IntelligenceAgi ResearchAi BenchmarkingQuantum Computing

How OpenAI's GPT-5.2 Cracked a Theoretical Physics Problem

GPT-5.2's breakthrough in quantum physics signals a new era for AI in scientific research—but the competitive race and limitations reveal an industry at an inflection point.

How OpenAI's GPT-5.2 Cracked a Theoretical Physics Problem

The preprint landed on arXiv just after midnight on February 12, 2026. Most physicists scrolling through the night's submissions might have glossed over it—another paper on gluon scattering amplitudes, dense with technical jargon about "half-collinear momentum regimes." But the author list stopped people cold: four luminaries from Princeton's Institute for Advanced Study, Cambridge, Harvard, and Vanderbilt. And Kevin Weil, OpenAI's VP for Science.

By the time OpenAI made the formal announcement the next day, the whispers had already begun. GPT-5.2 had done something no large language model had managed before: It conjectured—and helped prove—a genuinely novel mathematical formula describing fundamental particle interactions. Not a party trick. Not a benchmark. A verifiable contribution to high-energy theoretical physics that made Nima Arkani-Hamed, one of the field's toughest critics, use the word "beautiful."

This wasn't supposed to happen yet. Maybe not for years.

But here's the uncomfortable truth embedded in that achievement: GPT-5.2 can now tackle graduate-level science problems with 93.2% accuracy on the GPQA Diamond benchmark. It scores 40.3% on FrontierMath, which tests expert-level mathematics. And yet on CritPt—a suite designed to mirror how scientists actually work, combining multiple research tasks into composite challenges—the best models barely scratch 10% accuracy even with tools. The gap between solving textbook physics and doing original research remains enormous.

It's just narrowing faster than anyone anticipated.

The Capital Behind the Conjecture

That gluon formula didn't spring from a chat interface. It emerged from a 12-hour scaffolded run using internal OpenAI tools—formal verification loops, code execution environments, retrieval systems pulling from physics literature, human oversight at every inflection point. The model made conjectures. The humans checked them. Then checked again.

The setup required infrastructure most labs can't afford. OpenAI hit a $10 billion annualized revenue run rate by mid-2025, according to Reuters reporting, with internal projections suggesting $12.7 billion for the full year. Triple the prior year's haul. That kind of cash flow funds the massive training runs and experimental scaffolding that produced the physics breakthrough. It also funds the kind of ambitious bets that could either reshape science or burn through investor capital before the productivity gains materialize.

Goldman Sachs put global AI investment at approaching $200 billion by 2025. The IMF's January 2026 World Economic Outlook flagged concentration risks but acknowledged upside scenarios if the technology delivers faster than skeptics expect. Fortune and Business Insider noted last September that somewhere between $115 billion and $160 billion in AI infrastructure spending might not even be showing up properly in GDP statistics—a blind spot that makes the entire sector harder to value.

The macroeconomic wager underpinning all of this: frontier labs can translate compute into economically valuable outputs before the money runs out.

OpenAI clearly believes the answer is yes. So does Google DeepMind, whose AlphaFold work just earned Demis Hassabis a share of the 2024 Chemistry Nobel. When Hassabis accepted in Stockholm, he called AI "the ultimate tool for science." Noble words. But Isomorphic Labs, DeepMind's drug-discovery spinout, had already locked in partnerships with Eli Lilly and Novartis worth roughly $3 billion in potential milestones, plus a $600 million funding round this past March.

Altruism and ambition aren't mutually exclusive in this business.

What Actually Works Right Now

The physics result sits within a broader ecosystem that's matured quickly—perhaps more quickly than the public conversation reflects.

DeepMind's AlphaGeometry2 now solves International Mathematical Olympiad geometry problems at gold-medalist level, using a hybrid symbolic-neural pipeline that leans on classical theorems when pure learning stalls. AlphaDev discovered faster sorting algorithms that got merged into libc++, the C++ standard library—a genuinely nontrivial algorithmic innovation deployed at scale. GNoME predicted 2.2 million crystal structures, including roughly 380,000 stable candidates, fundamentally scaling materials discovery beyond what human teams could catalog in a lifetime.

Smaller players are shipping too. Red Queen Bio worked with OpenAI to optimize a cloning protocol by 79-fold using GPT-5, per an Axios exclusive last December. MIT researchers published work on MagNet, an AI-driven variational wavefunction that found crystallization signatures in fractional quantum Hall systems—the kind of quantum many-body problem that typically requires supercomputers and months of specialist work.

This is happening faster than academic hiring cycles can absorb it. The American Institute of Physics reported that a significant share of recent physics graduates now use AI tools routinely. One in five 2024 PhDs went directly into AI development roles. The Aspen Center for Physics has scheduled a May-through-June 2026 workshop to convene theorists specifically to explore AI reasoning in physics—focusing on reliability, benchmark design, integration with existing computational methods.

The field is trying to catch up with itself.

OpenAI had been building toward the gluon result for months. December 2025's GPT-5.2 launch positioned the model suite explicitly as optimized for "reasoning" and extended thinking on professional-level tasks. By November, the company had already published "Early science acceleration experiments with GPT-5," documenting verified contributions across mathematics, biology, astronomy. One case study described progress on an Erdős problem in graph theory. Another detailed an immunology experiment that emerged from model-generated hypotheses.

The gluon amplitude required all of that scaffolding. The formula itself satisfies technical properties called soft limits and the Berends-Giele recursion—consistency checks that confirm it's mathematically coherent, not statistical noise dressed up as insight. Nathaniel Craig of UC Santa Barbara told colleagues the methodology could serve as a template for AI-assisted theoretical work going forward.

But templates don't guarantee replication. And that's where things get complicated.

The Benchmarks We Have, and the Science We Need

Digital illustration for article section "The Benchmarks We Have, and the Science We Need" in "How OpenAI's GPT-5.2 Cracked a Theoretical Physics Problem" - A conceptual still life composition visualizing the contrast between rigid academic benchmarks and t...

GPQA Diamond and FrontierMath matter. They measure something real. But they also measure what's easy to measure: textbook problems, closed-form solutions, IMO-style challenges with clear correct answers. The messier work of hypothesis generation, experimental design, navigating ambiguous data—the actual texture of scientific research—doesn't fit neatly into evaluation frameworks.

Consider TPBench, published in February 2025 to assess AI performance on research-level theoretical physics. The conclusion: mostly unsolved. Current models can't reliably navigate the multi-step reasoning, domain-specific intuition, and conceptual flexibility that working physicists deploy daily.

Or look at CritPt, which tests composite research challenges designed to mimic real scientific workflows. Seventy-one tasks spanning multiple steps, requiring tool use, retrieval, synthesis. Best base model: roughly 4% full-challenge accuracy. With extensive scaffolding: 10%.

That gap—between solving problems and doing research—remains vast. The gluon result is remarkable precisely because it crosses that gap. But it required 12 hours of orchestrated computation, formal verification infrastructure, and human domain experts checking every step. It's not clear how that scales. Or whether "scaling" is even the right frame.

A January 2026 paper titled "Even GPT-5.2 Can't Count to Five" demonstrated brittleness on simple algorithmic tasks despite the model's high headline benchmark scores. Physics research demands logical rigor across multi-step derivations where one uncaught error invalidates everything downstream. The scaffolding helps. It doesn't eliminate the problem.

When Scientific Integrity Meets Velocity

Retraction Watch reported in February 2025 that Neurosurgical Review had retracted 129 items flagged for AI-assisted content issues. Conference organizers across disciplines note heavy AI use in submissions—and in peer reviews, which creates recursive validation problems. Nature Portfolio, the International Committee of Medical Journal Editors, and the National Science Foundation have all tightened disclosure requirements. Undisclosed AI use is increasingly classified as potential research misconduct.

The guardrails are reactive, struggling to keep pace.

The EU AI Act becomes fully applicable in August 2026, with general-purpose AI model obligations already in effect as of last August. Labs developing systemic-risk models face compliance burdens that shape what can be published, how tools are deployed, who bears liability when AI-generated insights prove flawed. These aren't abstract policy concerns. They're business decisions with strategic consequences.

A Frontiers in Physics perspective piece from early 2025 urged caution, arguing that "AI needs physics more than physics needs AI." The authors pushed for deeper theory-first integration and warned against hype cycles that overpromise and underdeliver. The physics community, in other words, is asserting agency—refusing to let frontier labs unilaterally define what counts as "scientific reasoning."

The Aspen workshop isn't just a research convening. It's the field setting standards.

Three Labs, Three Strategies

Digital illustration for article section "Three Labs, Three Strategies" in "How OpenAI's GPT-5.2 Cracked a Theoretical Physics Problem" - A cinematic and highly detailed close-up of a sophisticated wet-lab environment where advanced artif...

OpenAI is leaning hardest into science credibility as strategic differentiation. The gluon amplitude. The November case studies documenting new mathematical results across graph theory and number theory. The Axios-reported wet-lab collaboration showing GPT-5 optimizing biological protocols in real time. The company even launched FrontierScience, a benchmark TIME covered as targeting "research-tier" problems—and GPT-5.2 leads the current leaderboard.

But CritPt results reveal how far there is to go. When tasks compound into open-ended research workflows, even the best configuration struggles.

Google DeepMind holds different capital: peer-reviewed breakthroughs and a Nobel Prize. AlphaFold reshaped structural biology. AlphaGeometry2's IMO performance proved domain-specific architectures can outperform general LLMs on formal tasks. GNoME's materials predictions are being validated in labs worldwide. The scientific brand equity is unmatched.

Yet DeepMind's public emphasis has tilted toward specialized systems rather than general-purpose reasoning models. Whether that's strategic caution—avoiding overpromising on capabilities still being validated—or architectural conviction about the limits of LLMs in science remains unclear.

Anthropic runs quieter. The lab funds a fellows program supporting compute access and safety research but has been less vocal about hard-science breakthroughs. Its safety-first posture and constitutional AI framing appeal to a different audience: enterprise buyers wary of reputational risk, researchers focused on alignment. Less flash, perhaps. But steady positioning for long-term trust in a field where trust will matter.

Smaller entrants are exploring adjacent niches. ArgoLOOM, a multi-agent framework for "quarks-to-cosmos" workflows, showed early demonstrations in late 2025. PhysMaster, pitched as an "autonomous AI physicist," claims to compress months of work into hours on selected problems—though independent validation is pending.

The research community's response has been cautiously optimistic. Interested. But aware that reproducibility and generalization remain open questions.

What the Next Three Years Might Look Like

Digital illustration for article section "What the Next Three Years Might Look Like" in "How OpenAI's GPT-5.2 Cracked a Theoretical Physics Problem" - A conceptual visualization of the near future of scientific discovery featuring an abstract composit...

Expect more frequent but modest verified contributions in mathematics and theory-adjacent physics through 2027. Domains with axiomatic structure—number theory, theoretical computer science, parts of high-energy physics—will likely see earlier wins. Empirical sciences dependent on high-throughput wet-lab loops, robotics, observational astronomy will lag until physical infrastructure catches up with computational capability.

The gluon amplitude may prove more significant as methodology than discovery. The formula itself matters to a narrow subfield. But the approach—long-context reasoning, formal verification, human-in-the-loop refinement, co-authorship with leading theorists—establishes a template. If it scales to other areas (graviton amplitudes, as OpenAI hinted in its February announcement), the cumulative effect could be substantial.

The competitive dynamics will intensify, but probably not toward winner-take-all outcomes. DeepMind has Nobel credibility and peer-reviewed impact. OpenAI has revenue momentum and willingness to publish boldly. Anthropic has enterprise trust. Different labs will dominate different scientific verticals, leveraging distinct architectural choices and partnerships.

Regulation will shape pace and distribution more than most observers expect. The EU's systemic-risk provisions, NSF disclosure requirements, journal policies—these create friction but also legitimacy. Labs that navigate compliance smoothly gain credibility. Those that don't risk reputational damage or market access in key geographies.

Perhaps the most important question isn't whether AI can contribute to science. The gluon amplitude settles that. It's what kind of science gets accelerated.

Will the technology favor problems amenable to formal verification, systematically biasing research agendas toward the computable? Will it democratize discovery or concentrate influence in well-capitalized labs with access to frontier compute? Will collaboration between human intuition and machine pattern recognition yield insights neither could reach alone—or will it flatten creativity into incremental optimization?

The Aspen workshop title signals where the field is heading: "AI Reasoning in Theoretical Physics." Not AI replacing physicists. AI as a reasoning tool within workflows designed, verified, interpreted by humans.

The gluon amplitude emerged from a 12-hour scaffolded run. But it took Arkani-Hamed, Lupsasca, Skinner, Strominger, and Weil to frame the problem, check the result, understand what it meant. That division of labor—machine conjecture, human verification, collaborative refinement—may define the next decade.

If GPT-5.2 can crack a theoretical physics problem in a regime physicists assumed yielded only zeros, the question becomes: what else?

And perhaps more pressingly: who decides?

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Simile Raises $100M to Build AI That Predicts Human Behavior
  • D-Wave's $400M Public Raise Powers Quantum Computing Push
  • Echodyne Scales Radar Production with $40M Manufacturing Expansion
  • How AI Is Solving Drug Development's $18B Characterization Problem
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.