A 15-megabyte program that fits on a floppy disk. No cloud connection required. Zero API keys to manage.
Those specs might sound like throwback computing, but they describe Ante, a new coding assistant from Antigma Labs that the company says runs entirely on a developer's own machine. The tool, which shipped March 31, 2026, as a single Rust binary, targets an audience that has remained wary of AI coding helpers: security teams at companies reluctant to send proprietary code through external servers.
"Today Anthropic 'open-sourced' Claude Code — and it's the perfect day to introduce Ante," the company wrote in a blog post timed to Anthropic's announcement. The jab was pointed. While Anthropic's tool requires cloud connectivity, Ante runs local GGUF models through a bundled llama.cpp engine. Developers can shut off telemetry with a single environment variable.
Antigma Labs claims the tool uses a fraction of the resources that cloud-based alternatives demand. Internal benchmarks comparing Ante to Claude Code across parallel Docker tasks showed roughly nine times lower average CPU consumption and five times lower disk I/O, according to documentation the company published. Peak memory usage, the team said, dropped by roughly seven times. The company did not provide independent verification of those figures, and the benchmark data includes timestamps that appear inconsistent with the product's recent launch.
The installation process involves a single curl command. No runtime dependencies. The alpha version supports macOS and Linux, with Ante automatically detecting GGUF models stored in standard directories like Hugging Face cache or its own ~/.ante/models folder. On Apple Silicon, it installs Metal-optimized llama.cpp binaries; on Linux, it pulls CUDA or Vulkan support as needed.
Among the models Ante supports is Qwen3.6 27B, a 17-gigabyte file with a 256,000-token context window. The company lists performance metrics for various models on Terminal-Bench 2.1, a coding evaluation suite. Results vary. One test with DeepSeek V4 Flash achieved 82.7% accuracy across several hundred trials, though these benchmarks remain unaudited by outside researchers.

Developers who prefer not to run models locally can point Ante at commercial APIs from Anthropic, OpenAI, Google, xAI, or OpenRouter. Custom endpoints work too, provided they match OpenAI's specification. Authentication happens through OAuth for subscription services or environment variables for direct API keys.
The agent operates in four modes. There's an interactive terminal interface for conversational debugging. A headless mode handles one-shot execution without user input. Server mode shares a single local model over WebSocket and JSONL, useful for teams running internal instances. Gateway mode hooks into Slack or Discord for collaborative workflows.
Permission management spans three presets that the company labels Strict, Auto, and YOLO. Developers cycle through approval workflows with a keyboard shortcut. The tool also supports multi-agent orchestration and integrates with the Model Context Protocol, a spec for sharing context between AI systems.

Antigma Labs keeps the core binary closed during the alpha period, though it released the documentation, protocol definitions, SDK, and evaluation pipeline under an Apache 2.0 license on GitHub. The company says it runs Terminal-Bench continuously and publishes live, continuously updated results with pinned builds for reproducibility.
Behind Ante sits a small team with roots in FAANG engineering orgs, quantitative trading shops, and competitive programming. Monk Zero leads as founder, with Zhaochen She listed as partner. The company's other product, Antix, functions as an LLM proxy and governance layer for enterprises navigating compliance requirements around AI use.
Ante enters a crowded field. Anthropic's Claude Code made waves with its cloud-first approach. Cursor offers autonomous agents with IDE integration. Rig.ai and Octo-agent have pitched local-first alternatives, while containerized solutions like herm and OpenJet provide offline coding assistance through different architectures.

Between mid-June and early July, Antigma Labs pushed a dozen preview updates. Features added during that stretch include goal-driven sessions, LLM error recovery, OpenRouter web search integration, and refined permission controls. The pace suggests active development, though the alpha label carries the usual caveats about stability and breaking changes.
Whether privacy-conscious enterprises will adopt Ante at scale remains an open question. The pitch is straightforward: keep your code on your own hardware, run whatever model you trust, pay nothing for inference if you stick to local execution. For companies that have resisted cloud-based AI tools on principle, that proposition might prove compelling enough to warrant a 15-megabyte download.
