Google's latest salvo in the artificial intelligence wars arrived with a number that turned heads: 77.1%. That's the score Gemini 3.1 Pro notched on ARC-AGI-2, a reasoning benchmark considered something of a crucible for models claiming genuine problem-solving ability. More to the point, it's more than twice what the previous generation managed—the kind of leap that doesn't happen often in machine learning, where improvements tend to arrive in smaller increments.
The model debuted February 19, 2026, in preview form. Timing, as they say, is everything. OpenAI and Anthropic have spent recent months rolling out models that excel at complex, multi-step reasoning—the sort of work enterprises actually pay for. Google's pitch for Gemini 3.1 Pro positions it as "a step forward in core reasoning" for tasks requiring more than pat answers. Which is to say: this is Google's attempt to catch up, or at least stay within striking distance.
Whether the company succeeds depends on more than benchmark scores, though those certainly help with the pitch deck.
When the Numbers Look Good (On Paper)
Beyond the ARC-AGI-2 result—verified by ARC Prize, which lends credibility—Gemini 3.1 Pro posted strong showings across a battery of technical tests. On GPQA Diamond, a graduate-level science reasoning evaluation, it hit 94.3% without external tools. For agentic coding tasks, where the model essentially writes and debugs its own software, it scored 80.6% on SWE-Bench Verified in a single pass. On Terminal-Bench 2.0, using the Terminus-2 harness, it managed 68.5%.
Then there's Humanity's Last Exam, an evaluation deliberately designed to push frontier models to their limits. Gemini 3.1 Pro scored 44.4% without tools. Give it search and code execution, and that jumped to 51.4%. Not exactly acing the test, but it demonstrates something perhaps more useful: how the model performs when it can look things up and run code, much like a human developer might.
Google's internal benchmarks pit 3.1 Pro directly against Claude Sonnet 4.6, Opus 4.6, and various GPT-5.x series models. The subtext is clear—this is a competitive move aimed squarely at the high-end reasoning market, where customers are willing to pay for accuracy and reliability.
Access Points and the Usual Enterprise Hooks
Gemini 3.1 Pro ships with a 1,000,000-token input window and a 64,000-token output limit, according to the official DeepMind model card. Input formats include text, images, audio, video, PDFs, and what Google describes as "entire code repositories." Whether that last bit means truly massive codebases or something more modest remains to be seen in production use.
Developers can reach the model through several channels. The Gemini API in AI Studio is live now in preview, alongside the Gemini CLI and Android Studio integrations. Google is also rolling out access via Google Antigravity, the multi-agent coding platform it launched earlier this year. Antigravity offers what the company calls "mission control" for agent workflows, producing verifiable artifacts—plans, screenshots, browser recordings—that make debugging less of a nightmare.
For enterprises, the model is available through Vertex AI and Gemini Enterprise. Consumer access comes via the Gemini app and NotebookLM, though NotebookLM requires a Google AI Pro or Ultra subscription. That's a deliberate gating strategy: keep the power users paying.
The developer documentation highlights new controls introduced with the Gemini 3 series that carry forward here. A thinking_level parameter lets you trade internal reasoning depth against latency and cost—useful if you're trying to balance performance with AWS bills that keep your CFO awake at night. There's also media_resolution control for multimodal inputs and something called "thought signatures" for more reliable multi-turn function calling.
Pricing: Strategic Continuity
VentureBeat reported that Gemini 3.1 Pro maintains the pricing structure of Gemini 3 Pro, though Google Cloud's official pricing page hadn't updated as of launch day. Based on current Vertex AI Generative AI pricing, developers should expect $2 per million input tokens for contexts up to 200,000 tokens, and $12 per million output tokens in that range. Long-context rates above 200,000 tokens climb higher. Caching runs $0.20 to $0.40 per million tokens. Search grounding includes 5,000 free queries monthly before charges kick in at $14 per 1,000 queries.
Offering what Google describes as a significant performance jump at the same price point is smart business. It gives existing Gemini developers an immediate upgrade path without triggering budget renegotiations or procurement delays—the kind of friction that can kill adoption momentum.
Show, Don't Just Tell: The Demo Reel

Google's launch demos lean into creative and technical synthesis tasks, which feels like deliberate positioning. The model generates animated SVGs from text descriptions, producing website-ready graphics with compact file sizes. It built a live aerospace dashboard using public ISS telemetry data, creating an orbital visualization system from scratch. Another demo showed it coding an interactive 3D starling murmuration—complete with hand-tracking controls and a generative audio score.
These examples emphasize visual and interactive output, perhaps aimed at design and prototyping workflows where Google sees an opening against text-first competitors like OpenAI's offerings. Whether designers and creative technologists will bite remains an open question.
The Messy Reality Beyond Benchmarks

Here's where things get interesting. Ars Technica noted that while Gemini 3.1 Pro's benchmark scores look impressive, the model doesn't automatically dominate preference-based leaderboards like Chatbot Arena, where real users vote on outputs. Tom's Guide ran head-to-head tests against Claude Sonnet 4.6 across seven real-world tasks. Claude won five of seven, though Gemini performed better on tone adaptation and technical explanation—small victories, but victories nonetheless.
The Hacker News thread on the launch drew over 893 comments, a mix of skepticism and pragmatic interest that's typical for developer communities. Some praised the model's performance on data-processing and reasoning tasks. Others noted gaps in practical agentic coding compared to Anthropic's offerings, particularly for certain workflows. TechRadar's roundup highlighted similar mixed reactions, with concerns about "emotional depth and creativity" relative to competitors—a softer metric, admittedly, but one that matters for customer-facing applications.
The model remains in preview, not general availability. Google says the preview phase lets them "validate updates" and refine agentic workflow capabilities before a full GA release "soon." For technical decision-makers, that translates to: available for testing and early integration work, but production deployments might want to wait for the stable release. Nobody wants to explain to their CEO why the AI chatbot went sideways during a product demo.
The Longer Game
Google shipped Gemini 3 in November 2025, followed by Gemini 3 Flash and broader ecosystem tooling like Antigravity. The .1 increment represents a new naming convention—marking iterative improvements within a generation rather than a full version jump. Whether this becomes a sustained pattern or a one-time adjustment, we'll know in a few months.
For CTOs and AI product managers evaluating reasoning models, Gemini 3.1 Pro offers a credible option with strong benchmark performance, genuine multimodal capabilities, and competitive pricing. The question isn't whether it's technically capable—the numbers suggest it is. The question is how it performs on your specific tasks, in your specific context, with your specific users.
Which is precisely why Google is running this as a preview. Benchmarks get you in the door. Real-world performance keeps you there. And in a market where OpenAI and Anthropic are shipping updates at a blistering pace, Google can't afford to ship something that looks good on paper but stumbles in production.
The race isn't over. But Google is at least still running.
