Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

Media & Entertainment iconMedia & EntertainmentSeptember 28, 2026

Palette launches self-improving AI media platform

Palette launches self-improving AI media platform
YcGenerative Ai+3
Media & Entertainment iconMedia & EntertainmentSeptember 26, 2026

Overlord Labs raises $7.3M for AI device battery chips

Overlord Labs raises $7.3M for AI device battery chips
Ai HardwareSemiconductor Tech+3
Healthtech & Biotech iconHealthtech & BiotechFebruary 8, 2026

The Infrastructure Race Behind In-Silico Medicine's Breakthrough

The Infrastructure Race Behind In-Silico Medicine's Breakthrough
Digital TwinsDrug Discovery+3
Climate / Social Tech iconClimate / Social TechFebruary 8, 2026

Varaha Raises $45M Series B to Scale Carbon Removal from Global South

Varaha Raises $45M Series B to Scale Carbon Removal from Global South
Carbon ManagementClean Tech+2

Founders Mentioned

Jun Zhu

ShengShu Technology

saas icon
SaaS

Jun Zhu

ShengShu Technology

saas icon
SaaS
Media & Entertainment iconMedia & Entertainment
February 8, 2026
Video GenerationArtificial IntelligenceAd TechGaming TechDiffusion Models

How ShengShu's U-ViT Framework Beat DiT to the Diffusion Transformer Race

The Chinese startup's 2022 architecture fused transformers with diffusion models before OpenAI's Sora. Now its Vidu platform is reshaping video generation for film, gaming, and ads.

How ShengShu's U-ViT Framework Beat DiT to the Diffusion Transformer Race

When OpenAI finally pulled back the curtain on Sora in February 2024, the reaction from AI researchers bordered on reverent. Here was a system that could conjure minute-long videos from text alone, complete with camera movements and coherent physics. Revolutionary, the headlines declared.

What those headlines missed: the technical breakthrough powering Sora's magic trick had been sitting in plain sight on arXiv for nearly a year and a half.

On September 25, 2022—long before Sora became a household name in tech circles—a team at Tsinghua University quietly uploaded a paper titled "All Are Worth Words." The work, which came to be known as U-ViT, described a vision-transformer architecture that treated everything—time, conditions, even noisy image patches—as tokens within a unified framework. Three months later, researchers from UC Berkeley and NYU published DiT, a related approach that would eventually become the architectural backbone for systems ranging from Stable Diffusion 3 to Sora itself.

That three-month gap? It turned out to be enough.

The Tsinghua team didn't just publish and move on. By March 2023, barely six months after their paper hit the internet, they'd spun up ShengShu Technology and started building a commercial product. Their timing, as it happens, may prove more valuable than the technical elegance of their skip connections.

When Transformers Met Diffusion

U-ViT arrived at CVPR 2023 with what sounded, at first, like an almost boring proposition: run everything through a vision transformer. No fancy tricks, no baroque architectural flourishes.

The move represented a clean break from how diffusion models had traditionally worked. Previous systems relied on U-Net backbones—convolutional layers stacked with intricate skip connections, the kind of architecture that required real expertise to tune properly. U-ViT collapsed that complexity into transformer blocks with long skip connections linking shallow and deep layers. The results spoke clearly enough: an FID score of 2.29 on ImageNet256 class-conditional generation, 5.48 on MS-COCO text-to-image tasks. And they'd achieved it without raiding massive external datasets.

DiT followed in December 2022, presented at ICCV 2023. Where U-ViT maintained a U-shaped structure—preserving those multi-scale feature hierarchies within the transformer framework—DiT went flatter, with uniform depth throughout. Both approaches worked. Both sparked immediate follow-ons.

By 2024, the industry had spoken. NVIDIA's DiffiT arrived at ECCV. Stability AI built MMDiT for Stable Diffusion 3, using separate transformer weights for text and image modalities with joint attention mechanisms. (Their research blog claimed MMDiT beat both U-ViT and DiT on prompt adherence and typography, though vendor claims require the usual grain of salt.) The field was fracturing into variants even as it unified around the core insight: transformers belonged in diffusion pipelines.

What separated winners from also-rans, then, wasn't the architecture itself. It was who moved first and what they built with it.

The Academic-to-Startup Pipeline, Compressed

ShengShu Technology's origin story reads like someone hit fast-forward on the typical research-to-commercialization timeline.

March 2023: company founded. The roster looked like a Tsinghua AI all-star team. Jun Zhu, a vice president at Tsinghua's Institute for AI Industry Research, signed on as founder and chief scientist. Fan Bao, who'd co-authored both U-ViT and the related UniDiffuser papers, became CTO. By June—three months in—they'd closed an angel round of roughly RMB 100 million (about $14 million), led by Ant Group, with Baidu Ventures and Zhuoyuan Capital joining.

The team's publication record ran deeper than U-ViT alone. UniDiffuser, published at ICML 2023, showed how a single transformer could juggle marginal, conditional, and joint distributions for multimodal image-text generation. Earlier work included Analytic-DPM (ICLR 2022), which improved inference without requiring retraining, and DPM-Solver (NeurIPS 2022), enabling high-quality sampling in just 10 to 20 steps. ProlificDreamer—a NeurIPS 2023 spotlight paper on text-to-3D generation—listed both Tsinghua and ShengShu authors.

They shipped their product, Vidu, on April 27, 2024. The positioning emphasized what enterprise customers might actually pay for: 16-second clips at 1080p with temporal consistency. Long-form generation, in other words, not just flashy demos. The company made a point of calling out U-ViT as its core architecture, predating DiT. Whether that distinction mattered to customers remained to be seen.

By June 2024, ShengShu raised a Pre-A round of "hundreds of millions RMB"—Chinese funding announcements tend toward ranges rather than precise figures—led by Baidu and the Beijing AI Industry Fund. The timing aligned neatly with Vidu's market launch, suggesting investors saw a path from research velocity to revenue.

In March 2025, the executive suite reshuffled. Luo Yihang, a ByteDance and Volcano Engine veteran, stepped in as CEO. Co-founder Tang Jiayu moved to president. The shift signaled what typically happens when startups transition from technology development to market capture: operations and go-to-market strategy start mattering as much as the underlying science.

A Crowded Stage

Digital illustration for article section "A Crowded Stage" in "How ShengShu's U-ViT Framework Beat DiT to the Diffusion Transformer Race" - A modern, conceptual flat illustration depicting the intense competitive landscape of AI video gener...

The competitive landscape Vidu entered was, to put it mildly, intense.

OpenAI positioned Sora as a "world simulator," framing video generation as physics modeling rather than just pixel pushing. Google iterated aggressively, shipping Veo through version 2 and reaching Veo 3 by May 2025, adding native audio generation and vertical video support. Runway launched Gen-3 Alpha in June 2024, with subsequent releases like Gen-4.5 claiming advantages on various benchmarks. Luma's Dream Machine went public the same month.

Perhaps more relevant for ShengShu: Chinese competitors weren't sitting still. Kuaishou's Kling launched in June 2024, reaching version 2.6 with simultaneous audio-visual generation by early 2025. Tencent open-sourced HunyuanVideo in late 2024—a 13-billion-parameter model with FP8 weights and xDiT-based parallel inference. Open-sourcing a competitive model? That's the kind of move that changes negotiating dynamics with enterprise customers.

ShengShu bet on controllability. Vidu 1.5 and 2.0, released in January 2025, emphasized "long-context" features and speed. The Q1 update in March 2025 introduced multi-reference inputs supporting up to seven images for character and asset consistency—critical for VFX shops and anyone trying to maintain narrative coherence across shots. By October 2025, Vidu Q2 expanded the reference-to-video capability with story-focused controls.

In February 2026, the company announced an A+ round of over RMB 600 million (roughly $83 million), co-led by Zhongguancun Science City and LINK-X. Strategic investors included Wondershare, Visual China Group, and TRS Information—a notable roster spanning content creation, stock imagery, and information services.

The press release attached to that funding carried an extensive client list: Tencent Animation & Comics, China Literature, CCTV Animation, iQIYI, Mango TV, ByteDance, Samsung, HONOR, TAL Education, Alipay, JD, Alibaba's 1688 platform, Amazon, L'Oréal, Anta, Lilith, and 37 Interactive. Entertainment, internet platforms, hardware, advertising, gaming—basically every category that might plausibly need AI-generated video.

The company also claimed Vidu Q3 ranked "No.1 in China, No.2 globally" according to Artificial Analysis benchmarks. Standard caveats apply. Vendor-reported rankings carry all the credibility you'd expect when a company cites metrics favorable to its own positioning, particularly in a field where evaluation frameworks evolve monthly.

VBench—introduced in 2023 and updated to VBench-2.0 in 2025—emerged as something approaching a standard for assessing video generation quality. The benchmark measured everything from visual fidelity and temporal consistency to physics modeling and commonsense reasoning. Multiple vendors began citing VBench scores in their marketing materials, and the benchmark's expansion to include "intrinsic faithfulness" metrics suggested the field was maturing beyond surface-level quality comparisons. Whether VBench maintains its centrality as models improve remains an open question.

Converging Technology, Diverging Markets

Digital illustration for article section "Converging Technology, Diverging Markets" in "How ShengShu's U-ViT Framework Beat DiT to the Diffusion Transformer Race" - A sophisticated modern editorial illustration visualizing the convergence of AI model architectures ...

By 2025, the architectural debate was essentially settled. Diffusion models married to transformer backbones became table stakes, whether through U-ViT's skip-connected approach, DiT's uniform depth, or MMDiT's multimodal separation.

The competitive battlefield shifted. Native audio generation. Multi-image reference for character consistency. Real-time or near-real-time generation speeds. These became the features that separated premium offerings from commodity tools.

Market dynamics, though, showed sharper geographic splits than the technology itself.

Chinese deployments operated under the Deep Synthesis Provisions—effective January 2023—and the Interim Measures for Generative AI Services, which kicked in that August. Both required watermarking, identity verification, and security assessments for services with "public-opinion attributes," a phrase that gave regulators wide latitude. European vendors navigated the EU AI Act, adopted in 2024 with initial provisions effective February 2025 and transparency obligations for foundation models phasing in through August.

The United States experienced whiplash. Executive Order 14110, issued in October 2023 under the Biden administration, established safety and testing standards. It was revoked in January 2025 under a new administration's "Removing Barriers to American Leadership in AI" order, which aimed to preempt state-level regulation and shift toward deregulation. Whether that approach proves durable depends on factors well outside the technology itself—litigation, state defiance, congressional action, public incidents that shift opinion.

Content provenance efforts like the Coalition for Content Provenance and Authenticity (C2PA) gained traction among major platforms. Adobe, Google, OpenAI, and Meta all joined the steering committee by 2024. Yet investigative tests by outlets like The Washington Post found inconsistent preservation of metadata across social platforms once content got uploaded and shared. The gap between technical capability and practical implementation remained stubbornly wide.

Hardware trajectories reinforced the split between haves and have-nots. NVIDIA's H200—with 141GB HBM3e memory at 4.8TB/s bandwidth—shipped in 2024. The Blackwell platform (GB200/GB300) arrived explicitly marketed for real-time generative AI and "world models." Analysts projected rising AI accelerator TAM with NVIDIA maintaining leadership, barring geopolitical disruptions or unexpected competitive breakthroughs. The roadmap suggested minute-scale video generation and real-time workflows would become economically viable.

Access, however, remained geographically and economically stratified. Export controls, chip shortages, and capital requirements meant cutting-edge compute flowed to well-funded labs in major tech hubs, not necessarily to the best teams or most promising applications.

What Happens Next

Digital illustration for article section "What Happens Next" in "How ShengShu's U-ViT Framework Beat DiT to the Diffusion Transformer Race" - A clean, modern editorial illustration depicting the explosive growth trajectory of multimodal AI so...

The analyst projections paint a picture of explosive growth, though such forecasts merit skepticism given how often they're revised.

Gartner projected 40% of generative AI solutions would be multimodal by 2027, up from 1% in 2023, and that 80% of enterprise software would be multimodal by 2030. The multimodal AI market was forecast to grow from $1.73 billion in 2024 to $10.89 billion by 2030—a 36.8% compound annual growth rate. McKinsey estimated generative AI could add $2.6 trillion to $4.4 trillion annually to the global economy, with 50% of work activities potentially automatable by the 2030-2060 window, midpoint 2045.

Such numbers always sound impressive in slide decks. The question for ShengShu is narrower and more practical: can they translate academic velocity and early-mover timing into lasting commercial advantage?

In a 2025 interview, founder Jun Zhu described that year as "the commercialization year for AI video," emphasizing differentiation through controllability, longer duration, and narrative strength. The company positioned itself as B2B-first—meeting professional video requirements before scaling to consumers—with a roadmap toward real-time generation.

The February 2026 funding announcement hinted at ambitions beyond digital content, referencing a push toward "world-model-grade" video systems and expansion from digital to physical world applications. Automation. Robotics. The usual founder optimism about expanding TAM.

Whether U-ViT's three-month head start in 2022 translates to durable market position depends less on the elegance of skip connections in transformer blocks and more on execution. Regulatory navigation. Enterprise sales cycles. The company's ability to serve production workflows that existing tools genuinely don't.

The architecture arrived first, that much is documented. Whether that timing mattered—whether September 2022 versus December 2022 made a meaningful difference—the market will eventually decide. ShengShu's bet is that it did. By 2027, perhaps sooner, we'll know if they were right.

More stories

  • Palette launches self-improving AI media platform
  • Overlord Labs raises $7.3M for AI device battery chips
  • The Infrastructure Race Behind In-Silico Medicine's Breakthrough
  • Varaha Raises $45M Series B to Scale Carbon Removal from Global South
  • AI-Designed DNA: The Race to Make Gene Therapy Safer
  • Novoic Raises £2M Seed to Detect Alzheimer's Through Speech AI
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.