DeepSeek released its V4 Pro AI model to the public in mid-August 2026, along with a new agent framework, both under MIT licenses that allow unrestricted commercial use, in a direct challenge to the closed-ecosystem strategies of OpenAI and Anthropic.
The Chinese lab put a 1.6-trillion-parameter model with a million-token context window into circulation alongside DeepSeek Harness, a modular framework designed to help teams build AI agents without tying themselves to a single vendor. For developers frustrated by the walled gardens of proprietary models, the timing matters. DeepSeek is betting that openness can win developer loyalty even as the frontier labs tighten their grip on cutting-edge capabilities.
A massive model, freely licensed
V4 Pro uses a Mixture-of-Experts architecture with 1.6 trillion total parameters, though only 49 billion activate during any single forward pass, according to documentation on Hugging Face. That efficiency claim is central to DeepSeek's pitch: the company says the model requires roughly 27 percent of the compute and just 10 percent of the key-value cache that its V3.2 predecessor needed at maximum context length.
Those gains translate into aggressive pricing. DeepSeek lists V4 Pro at $0.435 per million input tokens when the cache misses, and $0.87 per million for output. Cache hits drop to $0.0435 per million tokens. A lighter V4 Flash variant runs at $0.14 and $0.28 per million tokens for input and output. Both models handle context windows stretching to one million tokens, with output capped at 384,000 tokens.
The company's official documentation, published in late July and early August, sets concurrency limits at 500 requests for Pro and 2,500 for Flash. Perhaps more significant than the raw numbers is what the MIT license permits: any developer can download the weights, run the model locally, and build commercial products on top of it without paying DeepSeek a cent in royalties.
Three modes of reasoning
The model ships with a feature set that includes tool calling (with strict JSON-schema enforcement), fill-in-the-middle code completion, and a beta Chat Prefix Completion capability. Developers can toggle between three reasoning modes through API parameters: Non-think, Think-High, and Think-Max. The distinctions between those modes remain somewhat opaque in public documentation, though community testing suggests they govern how much computational effort the model devotes to internal deliberation before generating output.

DeepSeek's integration guides cover third-party frameworks including OpenCode, Hermes, and OpenClaw, positioning V4 Pro as a swap-in replacement for proprietary models already embedded in production workflows. The family first appeared in preview on April 24, with DeepSeek announcing both Pro and Flash variants that mimic the interfaces of OpenAI's Chat Completions and Anthropic's Messages APIs. Switching between models requires changing a single parameter, according to the company's change log.
A general-availability build labeled "DeepSeek-V4-Pro-0813" went live across the company's app, web interface, and API in mid-August, based on OpenRouter model listings and chatter in developer communities on Reddit.
An agent framework built on plugins
DeepSeek released version 0.1 of its Harness framework in the same mid-August window, also under an MIT license. The framework treats everything as a plugin: models, tools, skills, even UI components. Developers can launch a local web interface by running "npx @deepseek-ai/dsh web," which spins up a browser session at localhost:3080, according to posts in r/AIDeveloperNews.
Within days of the announcement, community forks began proliferating. Developers published sample plugins including dsh-anchored-standard and dsh-deepseek-vision, and the Harness reportedly supports the Model Context Protocol, with utilities for integrating with Claude Desktop and Cline, per documentation in those community repositories.
The framework offers two operational modes. "Minimal" closely mirrors DeepSeek's reinforcement-learning training setup with stripped-down prompt scaffolding. "Standard" provides a more fully featured agent surface. Community discussions suggest capabilities vary noticeably between modes when running V4-Pro-0813, though specific performance benchmarks remain sparse.
Betting on ecosystems, not just models
The dual release lands as open-source AI labs fight for developer attention against proprietary providers that guard their models behind paid APIs and licensing agreements. TechCrunch covered the April preview as evidence that the gap between open and closed models was narrowing. AP News highlighted the million-token context as a substantial upgrade from V3.2.

The Harness framework's plugin architecture resembles OpenAI's GPT store strategy, an attempt to cultivate a third-party ecosystem around DeepSeek's infrastructure. Early plugin development activity suggests developers are building discovery tools, vision extensions, and custom integrations, though the pace and quality of that ecosystem will determine whether DeepSeek's openness translates into lasting market presence.
Whether MIT licensing and aggressive pricing can pull developers away from the convenience and polish of frontier models remains an open question. For now, DeepSeek is offering something the major labs will not: complete control, no usage restrictions, and the freedom to run a frontier-class model wherever a developer chooses.
