How a little-noticed Tsinghua research project in 2022 became the architecture powering China's answer to OpenAI's Sora—and why the West is only now catching up
---
There's a peculiar rhythm to how foundational AI breakthroughs become commercial juggernauts. Sometimes it happens overnight, with breathless demo videos and immediate VC pile-ons. Other times, the most consequential work sits quietly in arXiv preprints for months, accumulating citations while the industry fumbles toward the same insight.
U-ViT falls firmly in the second category.
When a team at Tsinghua University published "All Are Worth Words: A ViT Backbone for Diffusion Models" in September 2022, the paper landed with barely a ripple outside academic circles. Most AI labs were still wedded to convolutional neural networks—specifically, the U-Net architectures that had dominated image diffusion models since their breakthrough. U-ViT proposed something conceptually simple but architecturally radical: treat everything—time steps, conditions, noisy image patches—as uniform tokens fed into a vision transformer backbone.
The results were striking. FID scores of 2.29 on ImageNet-256 class-conditional generation. 5.48 on MS-COCO text-to-image tasks. State-of-the-art numbers, quietly achieved.
Three months later, researchers at UC Berkeley and NYU released DiT (Diffusion Transformers). That paper gained wider industry traction, perhaps because of Berkeley's name recognition, perhaps timing. But the chronology matters here. U-ViT came first, and the Tsinghua team's follow-on work—UniDiffuser, which landed at ICML 2023—demonstrated that a single transformer-based diffusion model could jointly learn marginal, conditional, and joint distributions across text and images simultaneously. No specialized output heads required.
TechCrunch would eventually describe what followed as an "industry-wide migration" away from CNN U-Nets toward transformer backbones. That migration is now essentially complete. And the academic lineage that started with U-ViT has become a commercial battleground worth hundreds of millions in venture funding.
---
When Architecture Becomes Destiny
U-ViT's core innovation wasn't a new algorithm. It was perspective.
Traditional diffusion models relied on convolutional encoder-decoder structures with long skip connections—the "U-shaped" information flow that gives U-Nets their name. That architecture worked, but it scaled awkwardly. Adding capacity meant stacking more convolutional layers, which created diminishing returns and made multimodal conditioning feel bolted-on rather than native.
The Tsinghua team flipped the paradigm. What if you treated all inputs as tokens and let attention mechanisms handle the routing? Long skip connections between shallow and deep layers preserved the information flow diffusion models need. But now the backbone could scale more predictably, and multimodal conditioning—text prompts, style references, multiple images—became just more tokens in the sequence.
By March 2023, the team had extended this foundation with UniDiffuser, a unified framework capable of text-to-image, image-to-text, and joint generation. Trained on LAION-5B, released with full code on GitHub. Chinese research communities frequently credit this work with pioneering what they call the "transformer plus diffusion" multimodal paradigm, months before similar approaches surfaced in Western commercial products. That framing may overstate the case—research timelines are messy, and multiple labs often converge on similar ideas—but the chronology is documented.
The broader movement accelerated fast. NVIDIA's DiffiT, published at ECCV 2024, introduced time-dependent multi-head self-attention and hit FID 1.73 on ImageNet-256. PixArt-α and PixArt-Σ demonstrated efficient training at resolutions up to 4K using token compression. Apple's DiT-Air (December 2025) and ongoing research into μP scaling laws showed that diffusion transformers could transfer hyperparameters robustly across model sizes—a practical advantage when you're burning millions per training run.
By mid-2024, the architecture shift was essentially settled. Stability AI's Stable Diffusion 3 adopted MMDiT (Multimodal Diffusion Transformer) and explicitly noted it outperformed U-ViT and DiT baselines on visual fidelity and text alignment. The old guard had been replaced.
---
From Lab Bench to Term Sheet
Commercial translation in Chinese AI follows a well-worn path: tight academic-industry loops, rapid capital formation, aggressive feature velocity. ShengShu Technology hit all the marks.
In March 2023, researchers from Tsinghua's TSAIL lab and RealAI—a university spin-out focused on "trustworthy AI," whatever that means in practice—formed ShengShu. The founding team included Jun Zhu, whose lab page now prominently lists U-ViT as the "first diffusion transformer with strong scalability." Angel funding arrived quickly: approximately RMB 100 million (~$14 million) from Ant Group and Baidu Ventures at a $100 million valuation.
That's the kind of valuation that raises eyebrows if you're coming from a Western VC mindset. No product, barely a company. But in China's AI ecosystem, where Baidu and Ant Group function as both investors and strategic partners with guaranteed distribution channels, those numbers make more sense. You're not just buying equity in a startup. You're buying optionality on infrastructure that might power your own products within 18 months.
ShengShu moved fast. By March 2024, the company had closed a multi-hundred-million RMB round led by Qiming Venture Partners—one of the region's most respected early-stage firms. In February 2026, they announced a Series A+ exceeding RMB 600 million (~$83 million), co-led by Zhongguancun Science City and LINK-X Capital. Strategic investors included Wondershare, Visual China Group, and TRS.
The funding announcement explicitly traced the technical lineage: U-ViT in September 2022, global Vidu launch in July 2024, TurboDiffusion open-source inference stack in December 2025. There's a certain brazenness to that framing, given how many research contributions feed into any commercial product. But in pitch decks and press releases, origin stories matter.
Leadership changes accompanied the capital influx. In 2025, ShengShu appointed Luo Yihang—a former ByteDance and Volcengine AI executive—as CEO, while co-founder Tang Jiayu shifted to president. That's a standard professionalization move: bring in someone who knows how to scale revenue operations, give the technical co-founder a title that keeps them involved but clears decision-making bandwidth.
Whether it works depends on execution, chemistry, and how much the founders are willing to cede control. We'll find out.
---
Vidu's Sprint: Ten Million Users in 100 Days

ShengShu's Vidu platform launched globally in April 2024 with a pitch centered on "long-duration, high-consistency video" generation. The product emphasized Subject-to-Video and Reference-to-Video capabilities—maintaining multi-entity consistency across scenes. That's the technical challenge distinguishing commercial video tools from impressive research demos that collapse the moment you ask for two characters interacting across camera cuts.
Version updates arrived at a clip that would make Western product managers dizzy. Vidu 1.5 (November 2024) introduced multiple-entity consistency and long-context generation. Press coverage positioned it as a "Sora competitor," though OpenAI's Sora remained largely inaccessible to the public, making the comparison somewhat aspirational.
Vidu 2.0 (January 2025) claimed "industry's fastest" generation—producing clips in under 10 seconds at roughly 55 percent below "industry average" pricing. The company reported 10 million users in the first 100 days, a metric that probably includes anyone who signed up and generated a single clip. Still, orders of magnitude matter.
By July 2025, Vidu Q1 extended Reference-to-Video to support up to seven image inputs for multi-reference consistency. Think wedding photo sequences animated with consistent lighting and character features, or product shots that maintain brand visual identity across scenes. Vidu Q2 (October 2025) added cinematic language controls—focal lengths, camera movements, transitions that approximate professional production values.
In interviews, CEO Tang Jiayu and CTO Fan Bao framed multi-image consistency as the core commercial problem. Monetization, they suggested, would come from advertisers needing brand-consistent assets at scale, animation studios prototyping scenes, and viral social media use cases—the animated photo pairs flooding Instagram and Douyin.
Distribution strategy has been dual-track. On the B2B side, ShengShu partnered with Honor to embed image-to-video features in the Honor 400 smartphone series. They signed a strategic cooperation with 360 Group's Nano Search to power AI video campaigns. February 2026 funding materials name-dropped enterprise adoption across ByteDance, Samsung, Alipay, JD, Focus Media, L'Oréal, and Anta—though these references surface in PR contexts rather than verified case studies with usage metrics.
The company claims global reach across 200-plus countries and regions, with 10x growth in users and revenue during 2025. Take those numbers with appropriate skepticism. Fast growth is real in this market, but specifics around revenue composition, margins, and retention remain opaque.
---
TurboDiffusion: The Open-Source Insurance Policy
In December 2025, ShengShu and Tsinghua's TSAIL lab open-sourced TurboDiffusion, an inference optimization stack combining low-bit SageAttention, sparse-linear attention, rectified flow consistency models, and W8A8 quantization. The researchers claimed 100-to-200x speedups over baseline implementations—generating a 5-second video clip in roughly 1.9 seconds on an RTX 5090 GPU at 1080p resolution.
Third-party technical summaries and benchmark posts corroborate the magnitude, if not the exact multiplier. Real-world performance varies by model size, hardware configuration, and whether you're measuring wall-clock time or FLOPs. But the directional claim holds: TurboDiffusion represents a meaningful leap in inference efficiency.
Why open-source it?
The cynical read is strategic positioning. As models scale to higher resolutions and longer video durations, per-token computational costs become the primary barrier to real-time or near-real-time generation. Open-sourcing the inference stack establishes technical credibility, drives ecosystem adoption of your optimization techniques, and creates a talent funnel—developers who implement TurboDiffusion become potential hires or partners.
Meanwhile, you commercialize the proprietary pieces: model weights, fine-tuning approaches, API access with SLA guarantees. The SageAttention component has reportedly seen adoption across multiple vendors, suggesting the optimization techniques have legs beyond ShengShu's own stack.
There's precedent here. Hugging Face built a commercial business by open-sourcing infrastructure. Stability AI initially gained traction through Stable Diffusion's permissive licensing. The playbook is established, even if outcomes vary wildly.
---
The Arena Gets Crowded
ShengShu's commercialization path sits within an increasingly crowded field, and the competitive dynamics are sharpening fast.
Google's Veo launched in May 2024, with Veo 2 following in December 2024 and Veo 3/3.1 rolling out through January 2026. The platform offers 1080p-to-4K generation with SynthID watermarking and integrations across Vertex AI, YouTube, and Gemini. That distribution advantage—embedding video generation directly into Google's enterprise and creator tools—is hard to overstate.
Runway's Gen-3 and Gen-4 emphasize character and scene consistency with strong penetration among professional creators. The company has cultivated a brand position as the "serious" tool for film and advertising workflows, complete with New York offices near production studios and a roster of high-profile projects.
Kuaishou's Kling, which the company bills as the "world's first user-accessible DiT video model," has iterated rapidly through versions 2.0 and 2.5 Turbo. Marketing materials cite Arena rankings—more on that shortly. Luma's Dream Machine, launched June 2024, has gained attention for motion realism and continuous feature updates targeting professional pipelines.
Arena-style benchmarking, such as the Artificial Analysis Video Arena, has become a common reference point in marketing and procurement discussions. These platforms use ELO-type scoring where users vote on paired comparisons. The methodology has strengths—capturing subjective quality and preference at scale—but interpretation remains contested. ELO scores reflect user populations and prompt distributions, which may or may not align with enterprise use cases.
ShengShu's funding announcements reference leaderboard traction. The company's February 2026 PR emphasized Vidu Q3's performance on storytelling and audio-visual synchronization at 16-second clips and 1080p native resolution. Whether leaderboard position translates to revenue is the open question. Early adopters of generative AI tools often churn once they realize prompt engineering and post-production still require skill and time.
The competitive dynamic increasingly centers on three axes. First, multi-entity consistency and reference-guided control—can you maintain character identity and lighting coherence across scene changes? Second, inference latency and cost—can you deliver results in seconds rather than minutes at prices that support freemium conversion? Third, integration with existing creative workflows—does your tool export clean alpha channels, support industry-standard formats, and play nice with Premiere, DaVinci, and Blender?
Companies that deliver on all three while navigating regulatory requirements around synthetic media labeling will likely capture disproportionate market share. None have nailed it yet.
---
Money, Regulation, and the Infrastructure Build-Out

The generative AI infrastructure build-out is happening against a backdrop of accelerating enterprise spending and evolving—some might say fragmenting—regulation.
IDC forecasts enterprise GenAI spending rising from $16 billion in 2023 to $143 billion by 2027, a 73 percent compound annual growth rate. Gartner projects worldwide AI spending at $1.5 trillion in 2025 and $2.5 trillion in 2026. Those numbers include everything from model training to inference infrastructure to application-layer software, so parsing what actually flows to companies like ShengShu requires care.
Hyperscalers are responding. Google has signaled potential 2026 capex reaching $185 billion to meet demand. AWS, Azure, and Google Cloud are all racing to offer inference endpoints with better latency, lower costs, and integrated guardrails. For startups, that creates both opportunity—you can build on top of hyperscaler infrastructure—and risk. Once Google or Amazon decides your use case matters, they have distribution and capital advantages that are hard to compete against.
Regulation is messier.
China's Interim Measures for Generative AI Services took effect in August 2023, establishing obligations for public-facing generative services around content moderation, user verification, and algorithm disclosure. Enforcement has been selective, with high-profile shutdowns for apps that stray into politically sensitive territory, but most commercial tools operate with reasonable latitude.
In the U.S., Executive Order 14110 on AI safety was rescinded in January 2025, creating uncertainty around federal AI governance. Some industry observers view this as a net positive for innovation velocity; others worry it eliminates baseline safety requirements before state-level regulation fragments the market.
The EU AI Act's transparency requirements for general-purpose AI went live in August 2025. Article 50 obligations on deepfake disclosure and watermarking are slated for August 2026. The European Commission is drafting a Code of Practice, with consultations continuing into 2026. For companies with global ambitions, navigating this patchwork becomes a significant operational cost.
Adoption of content provenance standards like C2PA is accelerating. Adobe, Google, OpenAI, Microsoft, and others joined the steering committee. Google's SynthID watermarking in Veo and Imagen represents one implementation path—embedding imperceptible patterns that survive compression and editing. But press investigations continue to identify platform-level gaps in honoring provenance metadata. Enforcement will ultimately determine real-world efficacy, and platforms have mixed incentives when synthetic content drives engagement.
---
What Happens Next—and Why It's Hard to Predict

The diffusion transformer architecture appears settled for the immediate future. Continued research will focus on efficiency: quantization techniques, sparse attention mechanisms, scaling laws that reduce hyperparameter tuning and inference costs. "Real-time-ish" video generation is likely to proliferate, given TurboDiffusion-class accelerations and next-generation GPUs. That enables product experiences we've mostly seen in demos—interactive camera controls, lens grammars, integrated audio generation synchronized to motion.
Whether any of this becomes genuinely transformative for creative work, rather than just faster slop generation, remains unclear.
The competitive landscape will continue to bifurcate. Incumbents with distribution advantages—Google, OpenAI if and when Sora gains broader access, Stability AI assuming they sort out their governance issues—have built-in channels to reach millions of users. Challengers need to build verticalized tools for specific creative workflows where integration and specialization create defensible moats.
China-based companies like Kuaishou, ByteDance's internal video stacks, and ShengShu will compete aggressively on feature velocity and pricing, particularly in markets where regulatory barriers are lower and price sensitivity is higher. The global market may segment more than converge, with regional champions dominating their home turf.
For ShengShu specifically, the path from research milestone to sustainable commercial scale remains unproven. The company has capital—over $100 million raised across multiple rounds. It has technical credibility, rooted in legitimate academic contributions. It has a clear product roadmap and some distribution partnerships.
Whether that translates to durable market position depends on execution across enterprise sales, ecosystem partnerships, and continued model performance improvements. Gross margins in AI tooling are under pressure as compute costs remain high and competitors push pricing toward commoditization. Customer acquisition costs for prosumer creative tools can spiral when you're competing with incumbents' free tiers.
The U-ViT paper may have sparked a revolution in generative AI architecture. Revolutions are exciting to write about. But they don't guarantee commercial outcomes. They just reset the terms of competition, and sometimes that's worse for the revolutionaries than for the incumbents who can afford to wait, watch, and then deploy overwhelming resources once the direction becomes clear.
ShengShu is betting that being first matters more than being biggest. In technology, that bet wins sometimes. Just not usually.
