Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSSeptember 12, 2026

Waddle Labs launches AI platform to control robots via prompts

Waddle Labs launches AI platform to control robots via prompts
YcAi Agents+3
Climate / Social Tech iconClimate / Social TechSeptember 12, 2026

Rosberg Ventures closes €100M fund backing sustainability VCs

Rosberg Ventures closes €100M fund backing sustainability VCs
Venture CapitalClimate Investing+2
SaaS iconSaaS
September 12, 2026
Ai Price WarLarge Language ModelsGenerative AiAi Infrastructure

DeepSeek launches V4.1 Flash at $0.15/M token pricing

Chinese AI startup undercuts OpenAI and Google with new model boasting 1M-token context and cache-hit pricing at $0.003/M, as US agencies raise security concerns.

DeepSeek launches V4.1 Flash at $0.15/M token pricing

DeepSeek launched V4.1 Flash with pricing that undercuts American rivals by orders of magnitude, setting rates at $0.15 per million input tokens for standard queries off-peak ($0.30 during peak) and $0.003 per million when the model reuses cached context off-peak ($0.006 during peak)—a 50-fold savings that arrives days after the NSA, CISA, and FBI accused the Chinese startup and five peers of conducting "aggressive, malicious, and targeted distillation activities" against American AI models at industrial scale.

The September 10 announcement positions DeepSeek's new model well below OpenAI's GPT-5.6 Luna at $0.20 per million tokens off-peak ($1.20 during peak) and beneath Google's Gemini Flash tiers. It also crystallizes a split taking shape across the AI industry: relentless price deflation at the commodity end, geopolitical friction at the top.

Four days earlier, the NSA, CISA, and FBI had named DeepSeek alongside Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI in a joint advisory. The agencies accused the companies of "aggressive, malicious, and targeted distillation activities" conducted "likely with Chinese government awareness." Knowledge distillation, a technique for training smaller models to mimic larger ones, has been used by the firms "since at least late 2024," the September 8 advisory stated.

The timing was not lost on industry observers. DeepSeek's founder, Liang Wenfeng, had told investors in May that his company's API pricing already turned a profit. "We can recover the cost of purchasing a batch of equipment in the market within ten months," Liang said in a May 20 meeting, according to a transcript published by Chinese tech outlet 36Kr and later cited by Axios. He went further: "I could raise the price by half, or double it, and token consumption would barely change."

Yet the company keeps pushing prices lower.

Cache Hits at Three-Tenths of a Cent

The $0.003 rate applies when V4.1 Flash reuses previously processed text—what the industry calls a cache hit. For enterprises running multi-turn agent workflows or analyzing million-token codebases, the savings add up quickly. Moonshot AI's Kimi K3, launched in July, charges $3 per million input tokens with a $0.30 cache rate. OpenAI's tiered structure runs higher still, though the company slashed GPT-5.6 Luna pricing by 80 percent on July 30 and rolled out additional promotions three weeks later.

Google Cloud announced a 50 percent monthly billing credit on Provisioned Throughput for Gemini 3.6–3.8 Flash models running from mid-August through the end of December, according to documentation updated in late August. Model Price Watch, an industry tracker, reported that its "frontier token price index" shows sharp declines since 2023, with the widest gap opening at the low end as Chinese models and open-weights alternatives flood the market.

Cache-hit economics have become central to the price war. DeepSeek's V4.1 Flash activates 8 billion parameters during prefill and 16 billion on decode, pulling from a 552-billion-parameter mixture-of-experts backbone. The asymmetric "Causal Encoder–Decoder" architecture, as the company describes it, cuts KV-cache memory requirements to roughly one-quarter the high-bandwidth memory and one-eighth the SSD footprint of earlier generations. Less memory means lower infrastructure costs, which DeepSeek passes along in the form of sub-penny cache rates.

Output tokens run $0.60 off-peak and $1.20 during peak weekday windows—Monday through Friday, 01:00–04:00 and 06:00–10:00 UTC, according to API documentation updated in September. Peak pricing doubles across the board. The company lists a concurrency limit of 2,500 for V4.1 Flash versus 500 for the outgoing V4 Pro model, which it is phasing out. "Starting at 04:00 UTC on Sept 14, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates," the company wrote September 10, though a subsequent note indicated Pro service might continue beyond that date with unchanged billing. The ambiguity left some developers uncertain about migration timelines.

Industry surveys from March and July highlight mixture-of-experts inference optimization and KV-cache compression as the two dominant strategies for cutting per-token costs while managing large memory footprints. DeepSeek, Moonshot, Alibaba's Qwen, and other Chinese developers have adopted MoE broadly, reducing active compute per token while relying on cache hierarchy design to keep latency manageable.

Adoption Meets Advisory

Digital illustration for article section "Adoption Meets Advisory" in "DeepSeek launches V4.1 Flash at $0.15/M token pricing" - A conceptual and minimalist visual representation of pragmatic decision-making and infrastructure tr...

Lindy, a U.S.-based agent platform, moved all production traffic to DeepSeek V4 earlier this year. The company's CEO framed the decision as pragmatic: open-source Chinese models deliver adequate performance at a fraction of the cost, Axios reported in June, while hosting infrastructure remains in the United States.

In China, adoption has been brisk. China Merchants Bank published a CNCF case study September 8 detailing its use of "a DeepSeek base model" with multi-tenant LoRA sharing on internal infrastructure. Industrial Bank said in April it was among "the first batch to deploy DeepSeek-V4." Beijing Rural Commercial Bank cited local DeepSeek and Qwen deployments across multiple functions in corporate announcements from April and May.

The September 8 intelligence advisory complicates that calculus for Western enterprises. The allegations—systematic distillation against U.S. frontier models, potential government involvement—raise questions about training data lineage, intellectual property, and the legal exposure of organizations routing sensitive workloads through China-origin endpoints.

A NIST evaluation from September 2025 flagged security and censorship concerns on models from the People's Republic of China including DeepSeek, though that assessment is nearly a year old. The EU AI Act's general-purpose AI provider obligations took effect August 2, with content-marking deadlines for systems placed on the market before that date running through early December. China's Cyberspace Administration issued "Interim Measures on Anthropomorphic Interactive AI Services" on April 10, layering onto prior generative AI and content-labeling rules.

DeepSeek raised roughly ¥50 billion (about $7.4 billion) in its first external funding round in mid-June, with Tencent and CATL among the reported investors, according to Reuters updates. The valuation topped ¥350 billion, more than $50 billion. That capital positions the company to compete on infrastructure scale as well as price, though U.S. export controls revised in January subject advanced chips like the H200 and MI325X to case-by-case license review for China.

A Bifurcating Market

Digital illustration for article section "A Bifurcating Market" in "DeepSeek launches V4.1 Flash at $0.15/M token pricing" - A conceptual and minimalist representation of massive economic growth and a bifurcating market, feat...

Gartner forecast worldwide AI spending of $2.59 trillion in 2026, up 47 percent year-over-year, with AI-optimized infrastructure accounting for the largest segment. McKinsey reported in late August that roughly 90 percent of surveyed organizations use AI in at least one function, with 44 percent scaling across the enterprise. Costs and "constrained" productivity gains remain hurdles.

The model market is splitting. Frontier models command premium pricing for the hardest reasoning tasks while commodity intelligent models undercut on cost, pushing enterprises toward mixed-stack routing strategies that reserve expensive calls for edge cases. DeepSeek's willingness to cut rates even as demand remains inelastic—Liang's May assertion that doubling prices would "barely change" consumption—suggests that bifurcation will sharpen.

CTOs evaluating multi-vendor stacks now balance price, performance, regulatory exposure, and data-residency requirements in ways that felt simpler six months ago. The September 8 advisory increases reputational and regulatory risk for China-origin models in U.S. and EU procurement. Expect more vendor "sovereign hosting" arrangements and lengthier procurement questionnaires.

For now, DeepSeek's $0.003 cache-hit pricing sets a new floor. Competitors will need to match or justify the premium. The calculus has shifted, perhaps more quickly than anyone anticipated.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Waddle Labs launches AI platform to control robots via prompts
  • Rosberg Ventures closes €100M fund backing sustainability VCs
  • WonderTx uses extrapolative AI to turn injections into pills
  • OS3 launches $10K humanoid robot for business automation
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.