The timing was, if nothing else, impeccable.
Just days after Anthropic found itself scrambling to contain a rather spectacular leak—approximately 1,900 TypeScript files from its Claude Code product, accidentally exposed to the internet—a new open-source repository appeared on GitHub. Its tagline: "The best Claude Code that $200 can buy."
On April 5, developer Salman Mohammadi quietly released nanocode, a pure JAX library designed to train a 1.3 billion parameter AI coding agent in roughly nine hours on Google Cloud TPUs. The total cost? About what you'd pay for a decent pair of running shoes, or perhaps a somewhat extravagant night out.
Within 18 hours, the project had racked up 185 points on Hacker News. The developer community, still processing the implications of Anthropic's security lapse and the uncomfortable congressional questions that followed, suddenly had a working alternative to contemplate. Alternatives, it turned out, matter quite a bit when proprietary infrastructure proves less than watertight.
A Leak, Then a Flood
Anthropic's troubles began unraveling on March 31, when security researchers discovered that the company had inadvertently published nearly 1,900 TypeScript files. The leak—covered by outlets including TechRadar Pro and PCGamer—contained internal orchestration logic, unreleased features, and architectural decisions the company had presumably hoped to keep under wraps.
By early April, Rep. Josh Gottheimer had dispatched a letter to Anthropic. It was the company's second source-related incident in roughly a year, escalating what had been murmurs of concern into something closer to alarm about AI vendor security practices.
But the leak did something else, too. Within 24 to 72 hours, multiple developer communities had posted forks, plugins, and "source builds" attempting to reconstruct Claude Code-like functionality. The legal and security implications remain murky—those efforts exist in a gray zone that no one seems eager to define—but the appetite was unmistakable.
Nanocode arrived into this environment not as a pirated rebuild, but as something arguably more pointed: an educational pipeline for training coding agents from scratch, using transparent tooling, public datasets, and spot TPU access. Mohammadi's README made the case plainly. You don't need leaked proprietary code to build capable AI coding assistants. What you need is JAX, a few hundred dollars, and most of a workday.
The Dinner-Tab Economics of AI Training

Nanocode's central pitch is refreshingly specific. The project's d24 configuration—a 1.3 billion parameter model—trains on a single TPU v6e-8 node in approximately nine hours for an estimated $200. A smaller variant, the d20 model at 477 million parameters, finishes in 1.5 hours for $34. Both figures assume Google's Trillium (v6e) TPU pricing, which third-party sources estimate at roughly $2.70 per chip-hour, though official documentation remains fragmented and varies by region.
That's not a research grant. Not a Series A runway expense. It's the cost of a decent dinner, the sort of amount a solo developer might charge to a credit card while evaluating whether TPU-based training makes sense for their use case.
The economic case sharpens further when you factor in Google's TPU Research Cloud program, which offers free preemptible TPU access for a month, and promotional credits that are sometimes bundled with new Google Cloud accounts. Mohammadi's README points explicitly to these avenues, positioning the project as accessible to students, hobbyists, and small teams who wouldn't otherwise dream of touching large-scale model training.
It's a compelling pitch. Whether it holds up in practice—across different regions, with varying availability, and under real-world debugging conditions—remains an open question.
JAX All the Way Down

Nanocode is written entirely in JAX, Google's functional programming framework for high-performance numerical computing. That decision carries weight. JAX's compilation model, which funnels Python code through 149 compiler passes before execution on TPU hardware, offers extreme optimization potential. It also demands careful kernel design and a tolerance for sharp edges.
Mohammadi's code leans on JAX Pallas, a low-level kernel language, to implement "SplaSh attention"—a TPU-optimized attention mechanism that replaces standard eager-mode operations with hand-tuned block sizes for v6e hardware. The model architecture itself follows familiar transformer conventions: rotary position embeddings with a base theta of 10,000, RMSNorm for normalization, grouped key-value heads, logit softcapping at 15. There's AdamW and Muon optimizers, zarr-based checkpointing, a two-phase prefill-decode inference loop with KV caching.
Not exactly cutting-edge research. But production-adjacent, the kind of stack you'd expect from an engineer who's debugged JAX's idiosyncrasies more than once.
The repository draws inspiration from MaxText and Seqax, both established JAX codebases for large-scale TPU training. What distinguishes nanocode is the assembly: it packages these components into a turnkey pipeline. Tokenizer training on FineWeb-Edu and The Stack v2-dedup datasets, pretraining, synthetic data generation for tool-use scenarios, supervised fine-tuning for agentic behavior, direct preference optimization for constitutional alignment. The speedrun scripts automate the entire flow, from provisioning a v6e-8 node to generating a final evaluation report.
It's surprisingly polished for a project that appeared days after a major industry incident.
Synthetic Rollouts and Public Datasets

Nanocode's training data highlights another dimension of contemporary open-source AI development: synthetic rollout generation. The project combines existing instruction datasets—tulu-3, self-oss-instruct, evol-codealpaca—into roughly 134,000 tool-call rollouts, published on Hugging Face as smohammadi/nanocode-tulu-selfoss-evol. Long-context multi-turn scenarios are synthesized using Gemini with what the project calls "constitutional critique," creating preference pairs for alignment training.
The datasets lack detailed documentation cards. But the pipeline itself is transparent: download, restyle, critique, filter, train. No black boxes, no licensing negotiations.
This reflects a broader shift in how small teams and individual developers access capable AI models. Rather than waiting for frontier labs to release artifacts or navigating enterprise licensing agreements, they're assembling pipelines from public datasets, open-source frameworks, and cloud credits. The March Claude Code leak accelerated this trend—perhaps more than Anthropic expected—by surfacing orchestration patterns that had been opaque. Internal MCP hooks, multi-agent buses, tool-use abstractions. Suddenly, the "how" of building coding agents felt less proprietary, more reverse-engineerable.
Nanocode doesn't replicate Claude Code's leaked architecture. It doesn't need to. The project demonstrates that the core capability—a model that can write code, use tools, and respond to multi-turn instructions—is achievable through well-understood training techniques and accessible compute. The "SOUL.md" gist, created on March 29, outlines the agent persona and CLI interaction model, framing nanocode as a conversational coding assistant rather than a drop-in replacement.
What It Signals
The confluence of events in late March and early April marks a potential inflection point, though the industry has seen enough "inflection points" to warrant skepticism. Still. Proprietary AI vendors face heightened scrutiny after security lapses. Open-source alternatives are emerging with cost structures that make them viable for small teams. JAX and TPU adoption, once confined to Google-affiliated research labs, is spreading through the developer community via projects like nanocode, vLLM's TPU backend, and SGLang-JAX.
Perhaps more significant is the psychological shift. When a $200 training run can produce a functional coding agent, the economic moat around proprietary assistants narrows. Enterprises evaluating vendor lock-in may look harder at internally trained models for tool-use, especially in regulated environments where auditability matters. Startups building AI applications have a clearer path to prototyping without burning runway on API costs.
The technical landscape remains volatile, though. JAX's Pallas attention kernels on v6e hardware are still experimental—community reports from March mention scheduler crashes and instability. TPU pricing remains opaque and regionally variable. The legal status of projects inspired by or reconstructed from leaked code is uncertain at best. Nanocode sidesteps these issues by building from first principles, but the broader ecosystem includes forks of unclear provenance and questionable legality.
What's harder to dispute is the momentum. Within days of the Claude Code leak, repositories sprouted across GitHub, Reddit threads filled with "you can now build" posts, and Hacker News discussions churned through the implications for AI security and competition. Nanocode arrived into that conversation not as a reaction, exactly, but as evidence that the infrastructure for distributed, low-cost AI development is maturing faster than some anticipated.
For ML engineers, AI researchers, and technical leaders evaluating infrastructure alternatives, the question has shifted. It's no longer whether open-source coding agents are feasible. It's whether the economics and transparency advantages outweigh the integration friction and the comfort of vendor-managed solutions.
Nanocode's answer is in the README: nine hours, $200, pure JAX.
The rest, as they say, is up to you.
