Two hundred fifty dollars a month. That's what Google is charging for access to its most sophisticated reasoning technology—a price tag that outstrips most enterprise software subscriptions and raises a fundamental question about who artificial intelligence is really for.
The product in question, called Gemini 3 Deep Think mode, represents Google's answer to OpenAI's vaunted o-series reasoning models. Available exclusively through the company's AI Ultra subscription tier, Deep Think uses what Google describes as "parallel reasoning" to work through graduate-level mathematics and science problems, exploring multiple solution paths simultaneously before committing to an answer. It's the kind of capability that sounds transformative in a demo and might actually be—if you can afford it.
Google began rolling out Deep Think in early December 2025, then pushed a substantial upgrade this past February 12 that specifically targeted scientific and engineering applications. The timing wasn't accidental. OpenAI's o3 and o4-mini had already set new benchmarks for multimodal reasoning, and the broader AI industry had coalesced around inference-time compute—essentially letting models "think longer" on hard problems—as the next competitive battleground.
But Google's approach differs in a crucial way. While OpenAI has distributed reasoning capabilities across multiple pricing tiers, Google is testing a different hypothesis: that a meaningful slice of users will pay premium prices for premium thinking.
The Mechanics of Expensive Thought
Deep Think isn't actually a standalone model. Think of it more as a specialized mode—a layer of additional processing power draped over Gemini 3 Pro that kicks in when users explicitly request it.
The technical architecture, according to Google's documentation, departs from the standard left-to-right token generation that most language models rely on. Instead of marching sequentially through a response, Deep Think explores multiple hypotheses in parallel before settling on its final answer. It's a computationally expensive approach, which partly explains the pricing.
Inside the Gemini app, accessing Deep Think requires two steps: selecting Gemini 3 Pro as the base model, then toggling "Deep Think" in the prompt bar. The feature carries an "experimental" label—tech industry code for "we're still figuring this out"—and Google has imposed daily usage caps that vary by subscription tier. Ultra subscribers get the most generous limits, though Google has declined to publish exact numbers.
The February upgrade narrowed the focus. Google pointed to Duke University's Wang Lab, which used Deep Think to design crystal-growth processes for thin films exceeding 100 micrometers—work that requires navigating dense materials science constraints. It's the kind of highly specialized application that might justify a $250 monthly expense. Maybe.
What the Benchmarks Actually Show
Google's published performance numbers position Deep Think as competitive with, and occasionally ahead of, existing reasoning models on specific tasks. But benchmarks in AI have become something of an arms race themselves, and the nuances matter.
On Humanity's Last Exam—a 2,500-question gauntlet created by the Center for AI Safety and designed to probe difficult reasoning—Deep Think scored 41.0 percent without access to external tools. Standard Gemini 3 Pro managed 37.5 percent on the same test. A meaningful improvement, though hardly a categorical leap.
The gap widens on abstract reasoning. Deep Think hit 45.1 percent accuracy on ARC-AGI-2, a benchmark verified by the ARC Prize organization, when code execution was enabled. Gemini 3 Pro scored 31.1 percent on the same test. Graduate-level science questions from the GPQA Diamond benchmark showed less drama: 93.8 percent for Deep Think versus 91.9 percent for the base model.
Google also highlighted competitive programming performance—a Codeforces rating of 3455 and gold-medal-level results on International Physics and Chemistry Olympiad written exams. An advanced Deep Think variant reportedly achieved gold-medal performance at the International Mathematical Olympiad last July, though that involved a more specialized setup than what paying subscribers actually get.
There's also a 50.5 percent score on a condensed-matter-theory benchmark, positioning the model as capable of graduate-level physics work. Whether any of this translates to practical research acceleration in real labs remains an open question. Early users are just beginning to find out.
The Two-Hundred-Fifty-Dollar Question

Google introduced the AI Ultra tier at its I/O conference last May, bundling Deep Think with other premium features: Deep Research, priority access to Veo video generation, and the kind of computational headroom that consumer plans simply don't offer. At $249.99 monthly, it costs roughly triple what competing AI services charge for their top tiers.
The pricing reflects a specific bet. Google believes a subset of users—researchers, enterprise developers, teams building AI-powered products—will pay substantially more for capabilities that go beyond chatbot novelty. The usage caps suggest the company is managing compute costs carefully, even at these price levels. Inference-time reasoning isn't cheap to run at scale.
For now, app access remains the primary path to Deep Think. Google opened an early-access program for API and enterprise testing on February 12, allowing researchers and businesses to express interest in programmatic access. API pricing and rate limits haven't been disclosed, though developer documentation indicates Google is preparing infrastructure for a broader rollout—whenever that might be.
Developer Controls, Abstracted
Google's Gemini API documentation reveals some architectural choices worth noting. Gemini 3 models expose a thinkingLevel parameter with settings ranging from low to high for Pro, and minimal to high for the Flash variant. This replaces the thinkingBudget token-count approach used in earlier Gemini 2.5 models.
The shift to level-based controls suggests Google is abstracting away the complexity of inference-time compute allocation. Developers specify desired reasoning depth without managing token budgets directly—a usability improvement, though one that also obscures exactly what's happening under the hood. Vertex AI mirrors these controls for enterprise deployments, but Deep Think as a distinct mode remains separate from general thinking-level adjustments.
Whether API access will carry the same per-query costs as app usage caps isn't yet clear. Google's early-access program language suggests the company is still calibrating pricing for programmatic use, perhaps testing different models with initial research and enterprise partners. Or perhaps they're just figuring out what the market will bear.
Safety Reviews and Strategic Positioning

Google delayed Deep Think's initial release to conduct additional safety evaluations—a decision the company attributed to Gemini 3's expanded capabilities. The model underwent what Google described as its most comprehensive safety testing to date: assessments through the company's Frontier Safety Framework, early access by the UK AI Safety Institute, and independent evaluations by Apollo Research, Vaultis, and Dreadnode.
The competitive landscape has shifted rapidly. OpenAI's o3 and o4-mini models demonstrated multimodal reasoning and tool use, creating a baseline expectation for what premium reasoning capabilities should deliver. Anthropic and other labs have invested in inference-time compute techniques too, though they've taken different approaches to productization and pricing.
Google's February upgrade explicitly targeted science and engineering use cases—a positioning choice that distinguishes Deep Think from general-purpose reasoning tools. Named collaborations with university labs suggest Google is courting academic users who might justify the Ultra subscription for research applications. Broader adoption will depend on whether the performance gains hold up across diverse problem types, not just the cherry-picked benchmarks Google has chosen to highlight.
The architecture remains somewhat opaque. Google describes "parallel reasoning" but hasn't published detailed methodology papers explaining exactly how Deep Think differs from extended chain-of-thought approaches or other inference-time scaling techniques. Third-party verification exists for some results—the ARC-AGI-2 score carries "ARC Prize verified" designation—but independent end-to-end testing of the consumer-facing mode remains limited.
A Test of Premium AI Economics

Google's approach represents a fundamental test of AI business models. By gating its most capable reasoning mode behind a $250 monthly subscription, the company is betting that professional users will pay multiples of consumer AI pricing for capabilities that handle complex technical problems more reliably.
The February API early-access program signals Google intends to expand beyond app-only availability, though timing remains unclear. Enterprise and research use cases may prove the early proving ground for whether Deep Think justifies its pricing—particularly for teams working on scientific computing, advanced code generation, or other domains where reasoning depth matters more than speed.
For developers evaluating reasoning models, Deep Think adds another high-end option in a rapidly evolving landscape. The benchmark numbers suggest meaningful capability gains on specific task types, though the limited availability and steep subscription cost mean most teams will need to weigh those gains against OpenAI's more accessible o-series models and other alternatives.
Google's willingness to invest heavily in inference-time compute indicates the company sees reasoning as a key battleground for AI differentiation. Whether that translates to sustainable competitive advantage—or just a fleeting lead in a benchmark war that never ends—remains to be seen. The price tag suggests Google thinks it might be worth finding out.
