Cactus Compute on Thursday released Needle 3, a family of AI models small enough to run on phones, smartwatches, and even microcontrollers — yet capable, the YC-backed startup claims, of matching DeepSeek's V4 Flash on tool-calling tasks after a single round of fine-tuning. The entire package spans 8 to 29 megabytes, a footprint hundreds of times smaller than what the industry typically calls "small."
The pitch here is architectural efficiency over raw scale. Needle 3 ships as a single binary containing stackable subnetworks ranging from 2 to 20 layers (29 million to 121 million parameters), compressed to roughly 2.125 bits per weight. Developers select the layer count that fits their hardware. The smallest configuration, a 2-layer variant weighing about 9 megabytes, targets wearables. The 20-layer version, at 29 megabytes, can process 400 to 4,000 tokens per second on a Raspberry Pi 5, according to the company.
A Different Kind of Transformer
The architecture departs from standard transformer design in a notable way. Cactus calls it a Laddered Simple Attention Network, and the key move is replacing traditional feedforward blocks with a Hadamard-initialized mixer, as detailed in a September blog post from the team. That decision cuts per-layer parameters from 4.7 million down to 25,600, based on the same post. Co-founder Henry Ndubuaku acknowledged the internal resistance in an August LinkedIn post: "When I suggested getting rid of FFNs in transformers, the team looked at me with contempt haha, but it worked."
Every layer functions as a standalone subnetwork, Cactus explained in documentation released with the model. The company trained Needle 3 on 360 billion tokens of what it describes as proprietary structured data. The result handles three tasks in one pass: exact JSON tool calls (including multi-step sequences), typed structured extraction, and text embeddings. It also outputs calibrated confidence scores, letting applications decide whether to execute locally or escalate to a cloud API.
Performance figures are striking, if preliminary, based on benchmarks provided by Cactus Compute. On Mobile Actions, a Google benchmark of 961 Android intent invocations, the 20-layer Needle 3 scored 86.0 versus 82.4 for Liquid AI's LFM2.5 1.2B and 76.0 for Qwen3.5 0.8B, both at full 16-bit precision. Apple's on-device model managed 57.6, Ndubuaku wrote in a Hacker News thread. The company's own charts show the 4-layer subnetwork (around 29 million parameters, 29MB on disk) matching DeepSeek V4 Flash on the DroidCall benchmark after one fine-tuning epoch. DeepSeek's model, for context, is a 13-billion-active-parameter mixture-of-experts architecture with 1-million-token context.
Hardware Tailwinds and Memory Headwinds

The timing may be deliberate. Industry forecasts suggest a shift in where AI computation happens. Counterpoint Research projected in mid-2024 that 45 percent of global smartphone shipments in 2026 will carry GenAI capabilities, while ABI Research forecast in June that TinyML AI chipset shipments will exceed 4.1 billion by 2031, propelled by a 90 percent compound annual growth rate in neural processing units. Gartner estimated the AI processor market for edge endpoints could surpass $90 billion by 2030.
Yet memory costs are tightening hardware economics. Tom's Hardware reported in September that DRAM and NAND had climbed to 60 percent of materials costs for some manufacturers, forcing smaller OEMs to redesign products and test for counterfeit chips. IDC data from August showed smartphone shipments fell 7.4 percent year-over-year in the second quarter, with demand concentrating in premium tiers as average prices rose.
Cactus believes that extreme compression could enable a category of automation that can run entirely offline: calendar additions, smart-home commands, Android intent routing. The company distributed Needle 3 under a source-available license. The engine runtime weighs under 1 megabyte and targets iOS, Android, macOS, Windows, WebAssembly, and microcontroller platforms including ARMv7, RISC-V, and MIPS, per documentation updated last week.
Crowded Territory

DeepSeek released V4.1 Flash on September 12, refining what it calls its "small-tier" offering with improved key-value cache compression. The model remains a reference point for open-weights efficiency at scale, though at 13 billion active parameters it's roughly 450 times the size of Needle's 4-layer variant. Google's FunctionGemma 270M, released late last year for edge function calling, occupies adjacent space but ships as a base model meant for downstream fine-tuning. Liquid AI's LFM2.5 family (230M and 350M configurations) launched over the summer, targeting high local throughput in MLX and GGUF formats.
A newer entrant is Jev, launched September 15 by TypeSafe AI. It isn't a language model at all — it returns typed choices with confidence scores for tool routing, designed to replace LLM-based agent selection loops. Hacker News commenters flagged it as a direct alternative for the narrow automation tasks Cactus addresses.
Cactus appears to be a small team of between 2 and 10 employees in San Francisco, co-founded by Roman Shemet and Ndubuaku. The company disclosed in August via LinkedIn posts that it accepted a strategic investment from an unnamed large technology company. The Needle repository on GitHub had accumulated 11,400 stars and 727 forks as of Thursday. The Hacker News launch post drew 198 upvotes and 89 comments by early evening UTC.
Shemet framed the project plainly in a recent LinkedIn post: "How can we make mobile AI powerful, personal, and private?" The EU AI Act's enforcement timeline adds some regulatory weight to the privacy claim. The general application rules took effect in August 2026, with high-risk embedded systems covered by August 2028.
Whether developers adopt the architecture or stick with quantized versions of larger models remains the open question. If Cactus' benchmarks hold beyond the datasets it published, the approach could meaningfully shift where certain automation runs. For now, it's a compelling technical demonstration with real deployment uncertainty ahead.
