Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSJuly 16, 2026

MyDecisive Raises $12M Seed to Slash Observability Costs 90%

MyDecisive Raises $12M Seed to Slash Observability Costs 90%
Devops AutomationAi Observability+3
Climate / Social Tech iconClimate / Social TechJuly 15, 2026

AI Mines 'Failed' Experiments to Accelerate Materials Discovery

AI Mines 'Failed' Experiments to Accelerate Materials Discovery
AiMaterials Science+3
SaaS iconSaaS
July 16, 2026
Large Language ModelsEnterprise AiOpen SourceAi InfrastructureStartup Funding

Thinking Machines Releases Inkling: $12B Startup's First AI Model

Mira Murati's AI lab debuts 975B-parameter open-weights model with 1M token context, positioning as leading U.S. alternative to Chinese models for enterprise customization.

Thinking Machines Releases Inkling: $12B Startup's First AI Model

Mira Murati's first move since leaving OpenAI wasn't to build another walled garden. It was to crack one open.

On July 15, 2026, Thinking Machines Lab released Inkling, a 975-billion-parameter AI model, and did something that's become almost countercultural in the current AI race: they posted the full weights to Hugging Face. No API-only access. No licensing gauntlet. Apache 2.0, with fine-tuning support live from day one.

The timing is pointed. For months now, the most capable open-weights models have been arriving not from Silicon Valley but from Chinese labs—Moonshot, DeepSeek, and others—while American frontier labs have doubled down on closed systems and usage-based billing. Thinking Machines, backed by a record $2B seed round at a $10-12B valuation secured in 2025, is placing a different wager: that enterprises will increasingly want AI they can own, modify, and run on their own infrastructure, rather than rent compute by the API call.

Artificial Analysis, which tracks model performance across standardized benchmarks, ranked Inkling at 41 on its Intelligence Index and called it "the new leading U.S. open-weights model"—narrowly ahead of NVIDIA's Nemotron 3 Ultra. Whether that's enough to shift enterprise purchasing decisions remains an open question, but the philosophical break with the industry's prevailing closed-model orthodoxy is unmistakable.

Routing Through 256 Experts

The architecture underneath Inkling is a mixture-of-experts decoder-only transformer—a design that's become something of a standard approach for labs trying to squeeze maximum capability from limited active parameters. The model has 975 billion total parameters but activates only 41 billion during inference, routing queries through 256 specialized experts plus two shared experts per layer. Top-6 routing determines which experts fire for any given input.

Context window support stretches to one million tokens in the open-weights checkpoint, though served versions on Thinking Machines' Tinker platform offer 64,000- and 256,000-token options. The model was pretrained on 45 trillion tokens spanning text, image, and audio inputs, and outputs text. Native multimodality, in other words, rather than bolted-on vision adapters.

Perhaps more revealing of the lab's priorities: the post-training process leaned heavily on synthetic data bootstrapped from other open-weights models, including Moonshot's Kimi K2.5. Over 30 million reinforcement learning rollouts shaped what Thinking Machines calls "controllable thinking effort"—an inference-time parameter that lets users dial up reasoning depth at the cost of latency and token consumption, or dial it down for faster, cheaper responses.

Hardware requirements are, predictably, steep. The BF16 checkpoint requires at least 2TB of aggregate VRAM—eight NVIDIA B300s or sixteen H200s. A calibrated NVFP4 checkpoint for Blackwell-generation chips cuts that to 600GB. Unsloth has already published GGUF quantized versions for lighter deployments, which speaks to how quickly the open-source community mobilizes around models with accessible weights.

Customization as Competitive Moat

Digital illustration for article section "Customization as Competitive Moat" in "Thinking Machines Releases Inkling: $12B Startup's First AI Model" - A conceptual, minimalist illustration representing a strategic competitive moat through open-source ...

The strategic bet here isn't subtle. Where competitors like Anthropic and OpenAI lean into hosted APIs and closed frontier models, Thinking Machines is releasing weights under an Apache 2.0 license—with a Model Acceptable Use Policy that restricts biometric inference and certain manipulative applications—and emphasizing enterprise customization.

Axios noted earlier this year that enterprises are increasingly frustrated with black-box AI they can't adapt to proprietary workflows. Thinking Machines has answered with day-one fine-tuning on its Tinker platform and full weights for local deployment. The Tinker platform now includes an Inkling Playground for chat and search demos, updated cookbooks, and integrations with partners like Together AI, Fireworks, Modal, Databricks, and Baseten. For those who prefer open-source stacks, the model works with transformers, SGLang, vLLM, TokenSpeed, and llama.cpp. Together AI announced same-day endpoint availability with unified multimodal input.

Pricing on Tinker starts at $1.87 per million input tokens and $4.68 per million output tokens. A price adjustment went live shortly after launch—prefill costs increased roughly 50 percent, training costs about 10 percent—though cached prefill remains discounted 80 percent. The economics still undercut many closed-model APIs, especially for organizations that can amortize infrastructure costs across high-volume workloads.

A Western Counterweight

Digital illustration for article section "A Western Counterweight" in "Thinking Machines Releases Inkling: $12B Startup's First AI Model" - A minimalist flat design illustration of a sleek, modern balancing fulcrum supporting abstract geome...

WIRED framed the launch in explicitly competitive terms: the strongest open-weights releases in recent months have come from China. Inkling claims comparable performance while offering a U.S.-based alternative with native multimodality and a ready-made customization stack.

Safety benchmarks reveal some of the trade-offs. The model card shows a FORTRESS adversarial safety score of 78.0 percent—in line with peers like Nemotron but well below closed leaders like Claude Fable, which scored 96 percent. Thinking Machines recommends defense-in-depth approaches and output moderation tools like Llama-Guard for production deployments. That's standard advice for open-weights models, but it underscores a reality: once the weights are out, downstream safety becomes the implementer's problem.

The company also previewed Inkling-Small, a 276-billion-parameter model with 12 billion active that reportedly matches or approaches the larger model's performance on several tasks with better latency and cost. Full weights will be released after internal testing completes; no date has been set. It's a sensible hedge—most enterprise workloads don't need a trillion-parameter sledgehammer, and a smaller, faster variant could prove more practical for production use.

Thinking Machines has secured infrastructure partnerships that suggest long-term ambition. The lab announced a strategic partnership with NVIDIA in March 2026 to deploy at least 1 GW of next-generation Vera Rubin systems, and expanded its use of Google Cloud AI Hypercomputer in April 2026. The leadership roster includes Chief Scientist John Schulman, an OpenAI cofounder, though some early team members reportedly returned to OpenAI earlier this year—a quiet reminder of how fluid talent remains in this corner of the industry.

Early Enthusiasm, Long-Term Questions

Digital illustration for article section "Early Enthusiasm, Long-Term Questions" in "Thinking Machines Releases Inkling: $12B Startup's First AI Model" - A clean, minimalist conceptual illustration representing a massive burst of community enthusiasm and...

Reddit's r/LocalLLaMA lit up within hours of the launch, with a thread exceeding 1,000 upvotes. Discussion centered on the Apache 2.0 license, the million-token context, and the rapid arrival of GGUF quantizations. Techmeme aggregated the release alongside coverage from Fortune, Reuters, and Artificial Analysis. The message from developers was straightforward: they wanted open weights they could run and modify. Inkling delivered that.

Whether enterprises adopt Inkling at scale is another question entirely. Purchasing decisions in the enterprise AI market are shaped by more than benchmark scores and licensing terms—procurement inertia, vendor relationships, and integration complexity all play roles. But the model's design choices reflect a clear hypothesis: that customization and control will matter more than incremental benchmark gains. With weights available, fine-tuning live, and a smaller variant on the way, Thinking Machines is offering companies a path to build AI that fits their specific requirements.

The alternative—adapting their workflows to fit someone else's black box—is what Murati and her team are betting against.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • MyDecisive Raises $12M Seed to Slash Observability Costs 90%
  • AI Mines 'Failed' Experiments to Accelerate Materials Discovery
  • Screenpipe Launches Privacy-First AI Recorder That Generates Agents
  • YC's Assemble Launches AI Agents to Automate Enterprise IT Development
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.