Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
SaaS iconSaaSJuly 2, 2026

Sakana's Fugu Bets Multi-Agent AI Can Beat Single-Model Giants

Sakana's Fugu Bets Multi-Agent AI Can Beat Single-Model Giants
Ai AgentsMulti Agent Systems+2
Climate / Social Tech iconClimate / Social TechJuly 2, 2026

Wakeline Raises €2.1M for Bio-Inspired AI That Learns in Real-Time

Wakeline Raises €2.1M for Bio-Inspired AI That Learns in Real-Time
Continual LearningAi+3

Founders Mentioned

Claire Mao

Instance

saas icon
SaaS

Lucy Cai

Instance

saas icon
SaaS

Claire Mao

Instance

saas icon
SaaS

Lucy Cai

Instance

saas icon
SaaS
SaaS iconSaaS
July 2, 2026
YcVideo GenerationAi BenchmarkingRegulatory ComplianceGenerative Ai

Instance Tackles AI Video's Physics Problem as Compliance Deadline Looms

YC S26 startup launches physics-aware quality benchmark as research shows AI video models fail basic physical laws and EU transparency rules take effect in 2026.

Instance Tackles AI Video's Physics Problem as Compliance Deadline Looms

The day OpenAI pulled the plug on Sora—April 26, 2026—the company's earlier boast that the model represented a step toward "simulating reality" felt, well, optimistic. By then, a steady drumbeat of academic papers had already delivered an unflattering verdict: frontier video models produce gorgeous visuals, sure, but trip over basic physics. Gravity bends the wrong way. Objects phase through walls like ghosts. Momentum evaporates mid-collision, as if Newton never existed.

Then there's the regulatory clock. The European Union's transparency deadline for AI-generated content—originally set for August, later shuffled to December 2 under Omnibus tweaks—suddenly forced a new question on anyone generating synthetic video at scale: Can you prove this clip obeys the laws of physics it claims to depict?

Enter Instance, a two-person startup out of Cambridge, Massachusetts, part of Y Combinator's Summer 2026 batch. Co-founders Claire Mao and Lucy Cai, both MIT-trained, describe what they're building as a "physics-aware quality layer for AI-generated video." Think of it as a benchmarking tool that scores what reads as real versus what doesn't—aimed squarely at synthetic data pipelines and the nascent world of "world models."

Their public site remains sparse: "We're working on world models," and not much else. No detailed whitepaper. No leaderboard. But the timing? That's the interesting part. Instance lands at a moment when the industry's physics problem and the compliance calendar are colliding head-on.

The Compliance Clock Ticks Louder

EU AI Act Article 50 isn't subtle. It imposes transparency obligations on generative AI providers: human-readable and machine-readable labels for synthetic content. The original August 2, 2026 deadline has been carved up under the Digital Omnibus package, with machine-readable marking now likely effective December 2, pending final national implementations. (Admittedly, even reputable sources differ on whether certain obligations kick in August or December; treat December 2 as the working target until final texts clarify otherwise.)

The European Commission published its code of practice for marking and labeling in May and June. On June 29, the Council tightened rules around non-consensual sexual deepfakes and child sexual abuse material. For video platforms and ad-tech pipelines serving EU markets, the message is clear: implement automated QA gates with audit trails before year-end, or face enforcement risk.

YouTube, TikTok, and Meta have all scrambled to expand labeling and detection infrastructure throughout 2025 and into this year. TikTok mandated creator labeling for realistic AI content in November 2025, with partial C2PA support. YouTube rolled out creator disclosure requirements and platform-applied labels for "altered/synthetic" content, updating help center guidance through June. Meta's policies, initially focused on images, now extend to video and audio.

The Coalition for Content Provenance and Authenticity (C2PA) released Content Credentials standard version 2.3 in February, marking a five-year milestone and launching a conformance program in late 2025. Adoption is broadening, but compliance doesn't happen automatically. It requires infrastructure to mark, score, and prove provenance at the moment of generation—not as an afterthought.

When Pretty Isn't Physical

The research record doesn't mince words.

In May 2025, the T2VPhysBench benchmark found that every tested model scored below 0.60 average compliance across categories of physical law, using human evaluation and counterfactual tests. A January 2025 study from DeepMind and INSAIT, titled "Do Generative Video Models Understand Physical Principles?," concluded that physics understanding remains "severely limited"—often unrelated to visual realism. Translation: a clip can look stunning and still violate conservation of momentum.

By the time WACV 2026 proceedings were published in April and May, the pattern was undeniable. Sora, Runway Gen-3, Pika, Lumiere, Stable Video Diffusion—all posted low "Physics-IQ" scores. Violations ranged from conservation laws to collision dynamics to contact forces. NVIDIA's PhyWorldBench, updated through this year, expanded the checklist: object permanence, identity consistency, rigid versus deformable behavior, camera-motion consistency, temporal coherence, long-horizon stability.

CVPR 2026 added more evidence. QuantiPhy, a quantitative physics reasoning benchmark covering 3,300+ video-text instances, found that vision-language models miss quantitative accuracy. PAVAS introduced metrics for audio-physics consistency in video-to-audio synthesis, exposing mismatches between what you see and what you hear. FlatSounds revealed how models rely on captions over visual cues, with temporal alignment gaps that feel off even when the output looks polished. PhyGround, released in May and June, provided a criteria-grounded, large-scale human study with fine-grained, law-specific judgments.

The disconnect matters—perhaps more than the industry expected—because the bet now is on "world models." Not just video generators that produce pretty clips, but systems that simulate plausible physical environments for downstream tasks: robotics training, autonomous systems, interactive applications. OpenAI framed Sora that way before the sunset. Runway CEO Cristóbal Valenzuela told podcasters in late April that his company's trajectory is "moving beyond content generation toward world models." NVIDIA CEO Jensen Huang spent CES, Computex, and GTC this year talking about "physical AI" and synthetic data.

If a model can't maintain momentum through a collision or render believable fluid dynamics, its utility as a world simulator is limited—no matter how photorealistic the output appears on first glance.

A Two-Person Answer from Cambridge

Instance's positioning sits right at that intersection. MIT pedigree. YC S26 batch. Physics-aware benchmarking. The YC company profile, accessed in early July, lists a team size of two and names Mao and Cai as co-founders. Their LinkedIn presence shows Cambridge as the location and a headcount range of 2–10 employees, though the core team remains two. The website offers minimal detail beyond the tagline. No press release or product demo has surfaced publicly yet.

What Instance appears to be building is a scoring layer—an API or service that ingests AI-generated video and returns metrics on physical plausibility. Gravity adherence. Inertia consistency. Object interaction realism. The phrase "what reads as real vs fake" suggests both a qualitative, human-aligned dimension (does this feel right?) and a quantitative physics check (does this obey F=ma?).

The mention of "synthetic data pipelines" points to enterprise use cases where video isn't the end product but a training input for downstream models: robotics simulators, autonomous vehicle testing, industrial digital twins. The "world models" frame aligns with where Runway, NVIDIA, and others are headed.

Whether Instance's benchmark correlates with human perception better than existing frameworks—like VBench, updated through 2025 with VBench-2.0 and VBench-Long leaderboards—or outperforms law-specific audits like T2VPhysBench remains to be seen. The company hasn't published leaderboard results or disclosed which models it's tested.

But the gap in the market is real. Full-reference metrics like Netflix's VMAF (version 3.1.0 released April 2) measure perceptual quality against a pristine reference but don't address generative plausibility or physics without that reference. Large-language-model-assisted evaluators like AIGVQA (ICCV 2025 Workshop) and VF-Eval (May 2025) add multi-dimensional assessment, but targeted physics scoring at API speed is still emerging.

Other players are circling. Hyle Labs exposes a "Physics Consistency Score (PCS)" API with 0–100 scoring for gravity, inertia, and object interaction plausibility (site accessed in 2026). WorldJen offers model and media scoring with flags like "physics implausible," "low dynamics," and "tight prompt adherence." Beamr VISTA handles subjective video quality testing at scale, used in AI-enhancement QA but not AIGV-specific. Netflix VMAF remains the streaming industry's baseline for perceptual QA, though its applicability to generative video without a reference is limited.

The detection and forensics side has its own infrastructure—GenVidBench (AAAI-26, March–April) with 6.7 million videos for cross-generator detection, and a Microsoft–Northwestern–WITNESS benchmark (IEEE Intelligent Systems, March)—but those address "is this fake?" rather than "is this physically coherent?"

The World Model Ambition

Digital illustration for article section "The World Model Ambition" in "Instance Tackles AI Video's Physics Problem as Compliance Deadline Looms" - A clean, minimalist conceptual visualization of a 'world model' encoding spatial reasoning and physi...

"World models" as a term gained traction in late 2024, with a TechCrunch primer published December 14, 2024. By now, it's industry shorthand for video generators that don't just render pixels but encode spatial reasoning, temporal causality, and physical dynamics—systems that can predict what happens next in a scene because they've learned underlying rules, not just surface patterns.

OpenAI's initial Sora pitch leaned into this framing. When the product launched (and then ended less than two months later), the company touted content provenance via C2PA but didn't publish physics benchmarks.

Google's Veo line has moved through iterations—Veo 2.0 endpoints were deprecated with a hard shutdown on June 30, and Google now directs users to Veo 3.1, integrated into Google Vids and Workspace at no cost as of April 2. Kuaishou's Kling 3.0, launched February 5, brought native 4K/60fps (15 seconds) with multi-shot and audio, positioning itself as a high-fidelity creative tool; the company's PR cites 60 million creators since the June 2024 launch.

Runway, widely used in film and TV workflows—Netflix and Disney were both reportedly testing the tools in mid-2025—continues to emphasize Gen-4 and Gen-4.5 in reviews this year. CEO Valenzuela's April 29 podcast appearance made the pivot explicit: AI video is a "prequel," and world models are next.

Adobe's Firefly has added video generation and agentic updates throughout the year: unlimited generations announced February 2, a creative agent on April 15, storyboard-to-video features on June 18. Stability AI maintains an open video line (Stable Video Diffusion, with ongoing model cards on Hugging Face updated through this year). Synthesia, focused on avatar-driven video for enterprise, hit a $4 billion valuation and over $100 million ARR as of January, per TechCrunch.

The market backdrop is pressure and opportunity combined. U.S. digital video ad spend is projected to surpass $80 billion in 2026, up 11 percent year-over-year, according to an IAB report published May 5. The same IAB Digital Video report highlights that two-thirds of buyers are live, testing, or planning agentic AI for campaigns this year.

Wyzowl's "State of Video Marketing 2026" survey—sample size roughly 266, so treat as directional—found that 63 percent of video marketers report using AI tools to help create or edit video. Grand View Research estimates the AI video generator market at approximately $946 million in 2026, reaching roughly $3.44 billion by 2033, though methodologies vary and some pages show a 2025 baseline at $788.5 million. A broader "synthetic media" market projection (2025–2033 CAGR of 18.1 percent) provides context for adjacent categories like avatars and localization.

The Evaluation Arms Race

What began as aesthetic scoring—does this video look good?—has fractured into multiple dimensions.

VBench and its successors handle multi-aspect evaluation: subject consistency, background consistency, temporal flickering, motion smoothness, dynamic degree, aesthetic quality, imaging quality, object class. AIGVQA (November 2025 ICCV Workshop) proposed a unified framework for multi-dimensional quality assessment of AI-generated video. WorldModelBench (March 2025) and WorldScore (ICCV 2025) judge video-generation models as world models, unifying evaluation across application domains and world-generation tasks. The AIGVE toolkit (community framework for AIGV evaluation, updated in 2026) aggregates metrics and provides a testing scaffold.

Physics-specific benchmarks emerged in force over the past year. T2VPhysBench (May 2025) uses human evaluation and counterfactual tests to measure compliance with physical laws. PhyWorldBench (NVIDIA Cosmos Lab, updated through this year) checks adherence to physical laws across multiple model families, with a human evaluation procedure. QuantiPhy (CVPR 2026) tests quantitative physics reasoning for vision-language models on 3,300+ video-text instances and finds that VLMs miss quantitative accuracy. PAVAS (CVPR 2026, oral presentation) introduces VGG-Impact and the APCC metric for audio-physics consistency in video-to-audio synthesis. FlatSounds (CVPR 2026, NVIDIA Cosmos Lab) highlights reliance on captions over visual cues in V2A, exposing temporal alignment gaps. PhyGround (May–June 2026) provides criteria-grounded, large-scale human study data with fine-grained, law-specific judgments to audit physics in video generation.

The research consistently shows that realism and physical faithfulness don't correlate. A clip can be visually stunning—photorealistic lighting, convincing textures—and still violate Newton's laws.

Large human studies in 2025 and this year used targeted, law-specific audits and counterfactual prompts to expose brittle dynamics. Ask a model to render a ball rolling downhill, then ask it to render the same ball rolling uphill, and watch the model fail to adjust forces appropriately. Object permanence breaks down. Contact forces vanish. Fluid dynamics look right at a glance but fall apart under scrutiny.

For companies building synthetic data pipelines—training autonomous systems, generating robotics simulation environments, creating interactive game assets—this gap is existential. NVIDIA's Physical-AI frameworks and PAIBench-G (2026 technical report) use physics-aligned evaluation to predict downstream utility. If a generated video doesn't respect momentum transfer in a collision, using it to train a robot's collision-avoidance system introduces bias. If a simulated environment doesn't model friction correctly, an autonomous vehicle trained in that environment will miscalibrate.

Physics plausibility isn't an aesthetic nicety. It's a functional requirement.

What's at Stake

Digital illustration for article section "What's at Stake" in "Instance Tackles AI Video's Physics Problem as Compliance Deadline Looms" - A clean, minimalist, and conceptual representation of AI video transparency and compliance, featurin...

The short-term catalyst is compliance. EU Article 50 transparency obligations—human-readable and machine-readable labels—apply in 2026, with enforcement dates under the Omnibus adjustments likely landing December 2, pending finalization of exact obligations and dates. Enterprises generating AI video at scale—ad platforms, creative tools, social networks, enterprise video services—need automated QA gates with audit trails. Platform policies from YouTube, TikTok, and Meta require creator disclosure and, in some cases, platform-applied labels. C2PA Content Credentials v2.3 provides a standard, but adoption requires integration at the point of generation, not post-hoc stamping.

The medium-term opportunity is differentiation. Physics plausibility correlates with perceived realism and brand safety, per research trends over the past year and this one. Human-aligned scoring plus physics checks reduce costly human review cycles. Ad buyers moving budget into AI-accelerated video creative—IAB reports 2026 digital video up 11 percent year-over-year, with a majority experimenting with genAI and agentic cycles as of May—need confidence that generated assets won't produce viral embarrassments. Objects floating mid-air. Gravity reversing. Collisions with no impact.

For creative tools embedded in Adobe, Google, and Runway workflows—Google Vids integrated Veo 3.1; Adobe Firefly added agentic and storyboard-to-video features; Runway continues Gen-4 and Gen-4.5 across 2025 and this year—physics-aware QA becomes table stakes, or at least a competitive edge.

The long-term bet is world models. NVIDIA's Jensen Huang has spent events this year (CES, Computex, GTC) emphasizing "physical AI," synthetic data, and world models. Runway's Valenzuela positioned the company's trajectory as moving beyond content generation. If video generators evolve into simulators that robotics labs, autonomous systems teams, and interactive-application builders rely on for training data, evaluation standards will push vendors to publish physics-aligned scores alongside aesthetic metrics.

WorldModelBench, WorldScore, and NVIDIA's PAIBench-G normalize physics-aligned reporting. The expectation will be: show us your conservation-of-momentum score, your collision-dynamics fidelity, your long-horizon stability. Aesthetic realism won't be enough.

Instance's two-person team is betting that the industry arrives at that inflection point sooner than the deprecation cycles—OpenAI Sora's April 26 shutdown, Google Veo 2's June 30 hard stop—suggest. Model lineups change rapidly; APIs come and go. But the underlying problem—video that looks real but behaves impossibly—persists across generations.

Whether Instance's physics-aware quality layer becomes the de facto standard or one tool among many in a crowded evaluation ecosystem remains to be seen. For now, it's a Cambridge-based team of two, MIT-trained, YC-backed, working on world models in a year when the world is finally asking whether the models understand the world.

More stories

  • DoD Solution raises $2M for AI drone navigation in war zones
  • DesignVerse raises $5.5M to automate enterprise software
  • Sakana's Fugu Bets Multi-Agent AI Can Beat Single-Model Giants
  • Wakeline Raises €2.1M for Bio-Inspired AI That Learns in Real-Time
  • AI Agents Dominate Y Combinator's Spring 2026 Batch
  • YC's Mireye Builds Data Layer for Physical-World AI Agents
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.