On paper, it makes no sense. A model with 26 million parameters—barely a rounding error by today's standards—shouldn't hold its own against one ten times its size. And yet, according to the team at Cactus Compute, that's precisely what's happening with Needle, a specialized open-weight model released May 12 that claims to distill Google's Gemini 3.1 tool-calling prowess into something compact enough to run on a smartphone at blistering speed.
Within 48 hours, the project had racked up 702 upvotes and 199 comments on Hacker News. The GitHub repository collected around 1,600 stars—respectable metrics for a niche technical release. But the real significance, perhaps, isn't the model itself. It's what Needle represents: tangible evidence that the agent stack is fracturing into specialized layers, with function calling emerging as its own distinct category, separate from the heavy reasoning engines that have dominated the conversation.
That shift matters more than any individual release, because it arrives at a moment when enterprises are scrambling to deploy agentic systems at scale. Gartner projected that 40% of enterprise applications will feature task-specific AI agents by the end of 2026, up from under 5% in 2025—a projection that feels both aggressive and entirely plausible given the current pace. Yet infrastructure teams deploying these systems are discovering an uncomfortable truth: routing every tool dispatch through GPT-4 or Claude burns budget fast. Very fast.
Needle joins a growing cluster of function-calling specialists—Google's FunctionGemma at 270 million parameters arrived in December, Liquid AI's LFM 2.5 at 350 million came in March, along with IBM's Granite 4.1 and Mistral Small 4—and a body of research suggesting that 60 to 80% of certain workloads can run on sub-20-billion-parameter local models if you route them correctly. What was once a monolithic call to a frontier model is splintering into purpose-built components. The question isn't whether this unbundling will happen. It's how fast, and who gets left behind.
The Architecture of Parsimony
Tool calling—some vendors prefer "function calling," but the distinction is mostly branding—sits at the heart of every agent system. An LLM reads a user request ("set a timer for 20 minutes"), matches it to a tool schema in its training, and spits out structured JSON arguments an application can execute. Most production stacks route these requests through general-purpose models: OpenAI's GPT-4, Anthropic's Claude, Google's Gemini. It works, certainly. But it's expensive, and it adds latency when a cloud round-trip isn't strictly necessary.
Needle's design takes a different approach entirely. It's an encoder-decoder "Simple Attention Network" with 12 encoder layers and 8 decoder layers—all self-attention and cross-attention, with no feed-forward networks. Co-founder and CTO Henry Ndubuaku laid out the reasoning on Hacker News: tool calling is fundamentally retrieval-and-assembly. Map an utterance to a tool name, fill the JSON slots. For that task, cross-attention is "the right primitive," and FFN parameters are "wasted at this scale."
The model was pretrained on 200 billion tokens over 27 hours using 16 TPU v6e chips, then post-trained on 2 billion tokens of single-shot tool-calling data synthesized from Gemini 3.1 in 45 minutes. The result: 26 million parameters shipped under an MIT license. Cactus Compute claims 6,000 tokens per second prefill and 1,200 tokens per second decode on consumer hardware—speeds that make local tool dispatch viable for voice assistants, smart-home commands, calendar operations. All without touching a server.
Cactus Compute, incorporated in the UK on February 21, 2025, builds an on-device inference runtime optimized for ARM chips in smartphones, wearables, and laptops. Ndubuaku frames the project as "an agent core for phones/watches/glasses," positioning it squarely in the privacy-first, offline-capable space that's gained traction since Microsoft's Windows Recall privacy debacle in 2024 and Google's rollout of Gemini Nano as a default assistant on Pixel 9 devices. The model card lists 15 tool categories it was trained on, from timers to navigation.
Needle doesn't operate in isolation, of course. Google released FunctionGemma—a 270-million-parameter open model targeting edge function calling—in December 2025, four months before Needle appeared. Liquid AI launched LFM 2.5 (350 million parameters) on March 31, emphasizing tool use and on-device structured outputs. Mistral Small 4, a mixture-of-experts model with strong tool-calling reliability according to community tests, shipped March 16. IBM's Granite 4.1, out in April, highlights enterprise function calling. The Needle team claims their model beats all these baselines on single-shot function calling, though they acknowledge the larger models have broader conversational range.
Needle's performance has not yet been independently benchmarked on standardized suites like Berkeley's BFCL v4 by neutral parties. Current performance claims come from the project authors and community discussion—enthusiastic, to be sure, but not exactly disinterested.
Three Forces Converging
Why 2026? Three trends converged to make this the year of function-calling specialists.
First, cost pressure. TrafficBench, a research project presented at ICLR in February, found that 80.7% of LLM workloads could run on small local models under 20 billion parameters with proper routing, yielding 60% cost savings, 77% energy reduction, and 67% compute savings versus cloud-only deployments. OptyxStack, a vendor case study from the same month, reported 25 to 60% inference cost cuts in production pipelines by routing simple queries to small models and reserving frontier models for complex reasoning. These aren't peer-reviewed academic studies, but the directional signal is unmistakable: developers are hunting for cheaper ways to handle high-frequency, low-complexity tasks. And finding them.
Second, on-device momentum. The edge AI market was estimated at roughly $30.3 billion in 2026 by one recent report, with mobile silicon advancing quickly—Apple's A18 Pro, new NPU designs, LPDDR6 spec arriving in devices by mid-2026. Google made Gemini Nano the default assistant on Pixel 9 phones, expanding offline features. Privacy-sensitive use cases, from health data to financial transactions, increasingly demand local processing. A 26-million-parameter model that fits in RAM and runs at 6,000 tokens per second can handle tool dispatch entirely offline, sidestepping latency, connectivity issues, and the specter of data leakage.
Third, agent adoption is accelerating—but cautiously. Gartner placed agentic AI at the "Peak of Inflated Expectations" in an April Hype Cycle report, noting only 17% of organizations have deployed agents to date, though over 60% plan to within two years. A separate Gartner press release from August 2025 projected that 40% of enterprise apps would feature task-specific AI agents by the end of 2026. Surveys from late 2025 and early 2026 showed a surge in planned investments, but also cited gaps in reliability, trust, and governance as reasons many pilots stall. Function calling is one of the brittlest parts of the agent stack—Berkeley's BFCL leaderboard and related research show even frontier models fail formatting constraints or compound errors in multi-turn scenarios.
Needle's architectural thesis is that tool calling doesn't need the general knowledge stored in feed-forward layers. It needs schema-conditioned cross-attention to map user intent to JSON. Whether that holds beyond single-shot scenarios remains to be tested—multi-step, stateful tasks or tools with complex interdependencies might expose the limits of a pure-attention design. But for the narrow problem Needle targets, the approach seems defensible. Maybe even elegant.
Separation of Concerns

The pattern gaining traction across the industry is separation of concerns. Instead of one monolithic model handling reasoning, tool dispatch, and execution monitoring, production systems are layering specialized components. A small router picks the tool. A small dispatcher fills arguments. Only if absolutely necessary does a large model step in for ambiguity resolution or complex planning.
This architecture mirrors the evolution of other infrastructure categories. Databases split into storage and query engines. Networks separated control planes from data planes. In AI, the equivalent is emerging as developers realize that not every token needs frontier-model compute. Needle positions itself as the routing layer. Cactus describes it as "an agent core," but agent periphery might be more accurate—it handles the mechanical step of matching intent to tool, leaving semantic reasoning to something else. Or nothing, if the task is simple enough.
The Model Context Protocol (MCP), donated by Anthropic to the Linux Foundation's Agentic AI Foundation on December 9, 2025, provides a standardization layer for tool orchestration. MCP has been integrated across major agent stacks and IDE tools through 2025 and 2026, though an RCE-class vulnerability disclosed April 22 underscored the ecosystem's relative immaturity. Still, the protocol's adoption signals that multi-server tool coordination is becoming a solved problem at the plumbing level. Once tools are discoverable and invocable via a standard interface, a tiny local model can handle the dispatch without needing awareness of every possible API.
Patch Window, an industry brief published May 14, discussed the cost-reduction potential of using Needle over larger models: replace the expensive API call for tool selection in GPT-4 or Claude pipelines with a local 26-million-parameter specialist. Call the frontier model only when genuinely needed. AgentConn's review the same day positioned Needle as "an agent routing/tool-emission layer, separate from heavy reasoning." The architectural split is becoming a pattern, not an experiment.
The Thorny Bits
Distillation from proprietary models raises legal and competitive questions that Cactus hasn't fully addressed publicly. Needle trained on data synthesized from Gemini 3.1, and Google's API terms explicitly state, "You may not use the Services to develop models that compete with the Services (e.g., Gemini API or Google AI Studio)," according to a September 2025 archive of the Gemini API Additional Terms. OpenAI's terms (updated January 1, 2026) and Anthropic's support documentation (March 2026) carry similar prohibitions on using outputs to train competing models.
Cactus hasn't publicly detailed how it navigated these restrictions—whether it holds a special agreement, interprets Needle as non-competing, or relied on open Gemini releases rather than API outputs. The ambiguity matters. In February, Ars Technica reported attempts to clone Gemini via over 100,000 automated prompts, illustrating why providers guard against output-based training. The legal lines remain blurry, and that's a problem for any enterprise considering deployment.
The EU AI Act, with general-purpose AI obligations that became effective on August 2, 2025, adds another compliance layer. Open-weight models with publicly available architecture may be exempt from some transparency obligations, but still require a training-data summary and copyright-compliance policy under the General-Purpose AI Code of Practice published July 10, 2025. Needle's MIT license and open weights simplify some of that. But any enterprise deploying a distilled model needs to verify the provenance chain, and that's easier said than done.
Performance claims also await independent validation. The project asserts Needle beats FunctionGemma-270M, Qwen-0.6B, Granite-350M, and LFM2.5-350M on single-shot function calling, but no neutral benchmarks on standardized suites have been published yet. Berkeley's BFCL v4, released in 2026, increased multi-turn and tool-difficulty challenges compared to earlier versions. Single-turn success doesn't guarantee multi-step reliability. IFEval-FC (2025) tests instruction-format adherence in tool calls, and DispatchQA (EMNLP 2025) evaluates small function-calling models specifically. Running Needle through these would provide a clearer picture—one the community is waiting for.
The architecture's trade-offs matter, too. Removing FFNs means the model stores less world knowledge. It can't fall back on parametric memory if a tool schema is vague or a user request is ambiguous. That's fine if the task is deterministic (turn on the lights). Less fine if the task requires interpretation (book a restaurant "somewhere nice"). Multi-step agentic workflows, where tool calls depend on previous outputs and state tracking is essential, might expose brittleness. Needle is optimized for a specific wedge. How far that wedge extends remains an open question.
What Comes Next
If Needle's speed and accuracy claims hold under independent testing—and that's still an if—the model represents a proof point for ultra-specialization. A 26-million-parameter specialist that runs faster and cheaper than a 270-million-parameter generalist suggests there's headroom to push even further. Perhaps a 10-million-parameter model for a narrower tool set. Or a 50-million-parameter model that handles multi-turn with state. The trend toward purpose-built components is likely to accelerate, especially as more devices ship with capable NPUs and enterprises hunt for cost-efficient agent architectures.
The competitive landscape will clarify over the next few months. Google, Liquid AI, IBM, and Mistral all have function-calling offerings in market, and none are standing still. If FunctionGemma or LFM 2.5 receive updates that close the gap on speed or efficiency, Needle's advantage narrows. Conversely, if Cactus or another team demonstrates a 10x parameter reduction with equivalent accuracy, the entire category shifts. The race is on, and the finish line keeps moving.
Standards and tooling will matter as much as the models themselves. MCP's move to the Linux Foundation and the Agentic AI Foundation, despite the recent security scare, signals industry alignment on interoperability. If tool discovery and invocation become genuinely plug-and-play, the barrier to swapping in a new dispatcher drops to near zero. Developers will pick the fastest, cheapest, most reliable option that fits their latency and privacy constraints. That could be a 26-million-parameter local model as easily as a cloud API.
For builders, the takeaway is straightforward: the agent stack is unbundling. Tool dispatch is becoming a separable layer, and running it locally is now viable—not theoretical, not aspirational, but actually viable. Whether Needle itself becomes the standard or simply proves the concept, the architectural pattern it represents looks increasingly like the shape of production systems to come.
The 26-million-parameter challenge to Google's 270-million-parameter spec may be less about who wins a benchmark and more about proving that the game has changed. In infrastructure, as in most things, once you see the pattern, you can't unsee it. And the pattern here—small, specialized, on-device routing layers feeding into larger reasoning engines when needed—is starting to look inevitable.
