The typing cursor keeps pace with thought. That's the promise, anyway.
When OpenAI quietly flipped the switch on Spark last month, the company wasn't just launching another AI model. It was making a statement about infrastructure—one that bypassed the silicon giant that has powered nearly every breakthrough in generative AI since ChatGPT's debut. For the first time in a major production deployment, OpenAI is serving a model on something other than Nvidia's chips. The vehicle for this shift? Cerebras, a chipmaker whose dinner-plate-sized processors have long been dismissed by some as a curiosity and championed by others as the future of low-latency AI.
Spark—formally GPT-5.3-Codex-Spark—clears 1,000 tokens per second when developers ask it to generate code. That's roughly 15 times faster than its immediate predecessor and several magnitudes beyond what most GPU-based systems deliver in the wild. To put it plainly: You can watch a function materialize on your screen fast enough that it feels like a conversation, not a chore.
The model debuted February 12, 2026, tucked inside ChatGPT Pro as a research preview. It's available through OpenAI's Codex app, the command line, and a VS Code extension. At launch, it handles text-only input across a 128,000-token context window—respectable but not exceptional. What sets Spark apart isn't capacity or versatility. It's velocity.
A Chip the Size of Dinner Plates, Deployed at Scale
Speed this dramatic doesn't emerge from software optimization alone. Spark runs on Cerebras's Wafer Scale Engine 3, the third generation of a processor that defies conventional chip design. Rather than stitching together dozens of smaller dies—the industry-standard approach Nvidia and others rely on—Cerebras manufactures a single, monolithic wafer packed with cores. The result is a chip roughly the size of an iPad that can move data internally at blistering rates, sidestepping the bottlenecks that plague multi-chip clusters.
"Real-time fast inference will unlock new interaction patterns and use cases," Cerebras CTO Sean Lie said in OpenAI's announcement, a statement that manages to be both visionary and vague. "This preview is just the beginning."
For context, Nvidia-powered deployments of earlier OpenAI models typically clock between 50 and 147 tokens per second, according to third-party benchmarking outfits like Artificial Analysis. Cerebras has demonstrated throughput exceeding 2,100 tokens per second on Meta's Llama 3.1 70B model and topped 3,000 tokens per second on other open-weight architectures—proof, at least in controlled settings, that the hardware has headroom.
But OpenAI didn't just swap chips and call it a day. The company rebuilt portions of its serving infrastructure, defaulting to persistent WebSocket connections that slash client-server round-trip overhead by 80 percent. Per-token overhead drops 30 percent; time-to-first-token shrinks by half. OpenAI says WebSocket will become the default transport layer for all models "soon," a signal that the latency crusade extends beyond Cerebras.
The $10 Billion Hedge
Spark is the first tangible outcome of a multiyear partnership OpenAI and Cerebras outlined in January. Industry reports, including coverage by The Financial Times, peg the deal at roughly $10 billion for around 750 megawatts of inference capacity through 2028. That's a staged rollout, one explicitly designed to reduce OpenAI's reliance on Nvidia's H100 and forthcoming Blackwell GPUs. The narrative from Cerebras's own blog frames the arrangement as both a technical bet and a strategic hedge: Cerebras excels at streaming inference with minimal latency, while Nvidia retains its dominance in training workloads and bulk inference tasks.
It's a notable gamble. Cerebras has carved a niche in ultra-fast inference, but it lacks the ecosystem breadth and battle-tested deployment scale Nvidia commands. OpenAI's public endorsement validates Cerebras's architecture while giving developers a concrete reason to care about hardware diversity. Latency, after all, matters acutely when you're watching code generate line by line in real time—when the difference between 50 tokens per second and 1,000 determines whether the tool feels responsive or sluggish.
There's skepticism, too. Ars Technica noted the absence of independent validation at launch, a fair point given OpenAI's history of aspirational performance claims. And some observers wonder whether Cerebras can scale to meet demand if Spark graduates from research preview to general availability. The company's chips are expensive and complex to manufacture; the infrastructure required to support millions of developers hammering endpoints simultaneously is a different beast than controlled benchmarks.
Built for Back-and-Forth, Not Fire-and-Forget

Spark doesn't just run faster. It behaves differently, a distinction OpenAI emphasizes in its documentation. The model is optimized for rapid iteration—think UI tweaks, logic adjustments, or refactoring a single function—with minimal, targeted edits by default. Developers can interrupt generation mid-stream, a feature explicitly designed for the rhythm of pair programming rather than batch workflows. It won't automatically run tests unless you ask, prioritizing responsiveness over exhaustive validation.
OpenAI describes this as "the first step toward a Codex with two complementary modes: real-time collaboration for rapid iteration and long-running tasks for deeper reasoning and execution." The broader GPT-5.3-Codex, which launched February 5, handles the latter. That model tackles multi-step autonomous workflows, notching state-of-the-art scores on SWE-Bench Pro (56.8 percent) and Terminal-Bench 2.0 (75.1 percent). It also earned a "High" classification under OpenAI's Preparedness Framework for cyber capabilities—a red flag that triggered internal safety reviews before deployment.
Spark, by contrast, does not reach those high-capability thresholds in cybersecurity or biology, according to OpenAI's internal checks. It's a tool for speed, not autonomy. Whether that distinction holds as the model evolves remains to be seen.
Behind the Pro Paywall, For Now
Access is restricted during the research preview. Spark lives exclusively within ChatGPT Pro, and API access remains limited to select design partners. OpenAI promises broader availability "over the coming weeks" but hasn't published per-token pricing or opened a public endpoint. Separate rate limits apply because Spark runs on a different hardware tier than the rest of OpenAI's fleet. The company warns of potential queuing during periods of high demand—a reality check for anyone expecting instant access at scale.
The Codex app itself surpassed one million downloads within a week of its February 2 macOS launch, according to TechRadar, underscoring developer appetite even before Spark's speed boost became available. That's promising for OpenAI, but it also raises the stakes. If Spark falters under real-world load—if latency spikes or availability dips—the backlash will be swift.
What the Competition Is Doing

Anthropic recently introduced a "Fast Mode" for Claude Opus 4.6, reportedly delivering up to 2.5 times standard speed at a steep premium. Public benchmarks don't hit Spark's claimed throughput, though Anthropic hasn't released detailed performance data. GitHub Copilot and Cursor dominate the IDE-native coding assistant space; Cursor claims 200 to 250 tokens per second on its in-house Composer model, still well short of Spark's ceiling.
Nvidia itself has demonstrated over 1,000 tokens per second per user on DGX Blackwell nodes running Llama workloads—but that's lab hardware, not a production developer tool. The gap between controlled demos and real-world performance is wide enough to drive a truck through, and OpenAI's history with Codex includes friction over perceived sluggishness compared to rivals. Spark is both a technical flex and a strategic recalibration, an attempt to reset expectations for what "fast enough" means in AI-assisted coding.
If the speed holds, it will. If it doesn't, this becomes another chapter in the long-running saga of overpromised AI capabilities.
Inference Beyond the GPU Monoculture
OpenAI hasn't detailed how Spark and standard Codex will interoperate long-term, though the company envisions developers toggling between modes or running sub-agents in parallel. The WebSocket optimizations rolling out platform-wide suggest latency improvements aren't exclusive to Cerebras, even if the headline throughput is. And the scale of the Cerebras partnership—750 megawatts of capacity—implies infrastructure for far more than a single research preview model. Perhaps OpenAI is hedging against supply constraints. Perhaps it's positioning for a future where inference diversifies across architectures the way training once diversified across cloud providers.
Or perhaps the company simply got tired of waiting for Nvidia to prioritize latency the way it prioritizes raw compute. Either way, Spark offers a glimpse of inference beyond the GPU monoculture—delivered at a speed that might finally make real-time AI pair programming feel less like waiting for an intern to finish typing. Whether that vision scales to millions of developers remains the open question. For now, it's fast. Very fast. And that, in a market where every millisecond compounds user frustration, might be enough.
