The problem sounds almost philosophical until you're actually building it: How do you let an AI agent write its own Python code without also handing it the keys to your entire system?
Pydantic's answer just landed on GitHub, and it's not what you might expect. No containers. No full-featured sandboxes. Just a stripped-down Python interpreter called Monty, written in Rust, that boots faster than you can blink—literally, in about 0.06 milliseconds—and assumes nothing about trust.
The release comes marked "experimental" and licensed under MIT, which in startup parlance usually means "we're shipping this now and we'll see what breaks." But the timing matters. Large language models have gotten disturbingly good at writing code, and agent frameworks increasingly treat code generation as a first-class capability rather than a parlor trick. The calculus shifts fast when your AI assistant starts executing Python it wrote three seconds ago.
Speed Versus Safety, Reconsidered
Container-based isolation remains the gold standard for running untrusted code. Spin up a Docker instance, execute whatever the model generated, tear it down. Simple enough. But containers boot in seconds—maybe two, maybe five—and that delay shatters the conversational rhythm agents need. Nobody wants their AI assistant pausing mid-thought while infrastructure provisions.
Monty positions itself in the gap between those extremes. According to benchmark scripts in the project's repository, it starts in microseconds, not seconds. Runtime performance lands somewhere between five times faster and five times slower than standard CPython, which Pydantic evidently considers acceptable. Perhaps more telling is what the interpreter doesn't do: touch the filesystem, read environment variables, make network calls. Not by default, anyway.
Every interaction with the host system flows through what the documentation calls "external functions"—essentially developer-controlled callbacks. An agent can't accidentally (or maliciously) delete your files because file operations simply aren't exposed unless you explicitly grant them. A pull request merged February 3 added pseudo-filesystem support, but even that routes through explicit trait mechanisms. The mediation layer stays intact.
Resource controls let developers set hard ceilings on memory, stack depth, execution time. Code that exceeds those boundaries gets cancelled. The interpreter captures stdout and stderr, bridges async and sync execution, supports deterministic pause and resume operations. Standard stuff for sandboxing, but implemented with an eye toward speed.
The Agent-First Architecture

Type checking ships as core rather than optional. Monty bundles Astral's "ty" type-checker directly into a single binary—modern Python type hints work out of the box. State snapshotting lets developers dump and load execution state to bytes, which becomes important when agent systems need to persist work across sessions. Or when they crash and need to pick up where they left off.
Pydantic plans to use Monty for "code mode" in Pydantic AI, their agent framework. The concept draws from recent thinking at Cloudflare and Anthropic around programmatic tool calling: agents write code against generated APIs of tools rather than invoking tools directly through JSON-RPC style calls. Early February 2026 discussions on Hacker News included comments from Pydantic AI's lead linking to a pull request adding code mode support with Monty and pluggable runtime options. The pull request exists. Whether code mode ships as promised remains to be seen.
The most revealing technical decision might be what Monty excludes. The standard library shrinks to sys, typing, and asyncio. Dataclasses and json are marked "soon." No third-party packages. No class definitions yet—though those are apparently coming. No match statements. The README states this explicitly: Monty is "designed specifically for 'run code written by agents.'" Not a general-purpose interpreter. Not trying to be.
That narrow focus creates obvious constraints. An agent that writes code using NumPy or pandas or requests will hit a wall. Developers either provide equivalent functionality through external functions or accept that their agents work within Monty's deliberately constrained universe.
Embedding Across Languages
Monty can be called from Rust, Python, or JavaScript/TypeScript, with no CPython dependency required. For Python users, installation follows familiar patterns: pip install pydantic-monty or uv add pydantic-monty. A conda-forge package exists for multi-arch installations, updated as recently as February 7.
The JavaScript story appears more fluid. The README references npm install @pydantic/monty with a roughly 4.5 MB download, though the package wasn't visible on npmjs.com when various sources documented the project. Developer Simon Willison compiled Monty to WebAssembly and produced both a direct JavaScript module and a Pyodide-loadable wheel, with working browser demos posted to his blog February 6. Whether that represents official support or enthusiast experimentation isn't entirely clear.
A basic async Python example from the repository shows the pattern:
from pydantic_monty import Monty
async def run_code():
monty = Monty()
result = await monty.run_monty_async(
"print('Hello from Monty')"
)
return resultExternal function callbacks let the host application expose specific capabilities to the sandboxed code. Strict control over what agent-generated Python can actually do.
How It Stacks Against Alternatives

Monty's comparison table pits it against Docker, Pyodide, Starlark-rust, and remote sandboxing services like Daytona, E2B, and Modal. Each represents a different set of compromises.
Docker offers complete language support and strong isolation but brings startup latency measured in seconds and operational complexity that scales badly. Pyodide, which compiles CPython to WebAssembly, provides full Python compatibility but suffers from slow cold starts—the documentation includes security caveats for server-side use. Remote execution services handle sandboxing infrastructure but introduce network round-trips, authentication flows, cost at scale.
Starlark-rust offers hermetic execution but uses a different language entirely. RustPython, another Rust-based Python interpreter, takes the opposite approach from Monty: full language compatibility at the cost of more complexity and weight. Monty deliberately stays minimal, betting that agent-written code doesn't need Python's full feature set.
That bet might not hold. Or it might turn out agents rarely need more than basic computation and orchestration. Nobody knows yet.
What's Missing, What's Coming
The "experimental" label carries weight here. The tiny standard library rules out most typical Python workflows. Third-party packages—the entire ecosystem that makes Python useful—are off the table. If an agent tries to write code using those tools, execution fails.
The repository notes that class definitions and match statements are "coming soon," suggesting active development toward broader language coverage. But it's unclear how far Pydantic intends to push that expansion. Too much Python support might undermine the minimal, secure-by-design philosophy that makes Monty interesting to begin with. There's a tension there.
Kinnaird McQuade, commenting on LinkedIn about the launch, emphasized the value for "server-side agent loops without external interpreters" and highlighted the "no host FS access" model as key for reducing latency and complexity compared to AWS Bedrock AgentCore or containerized approaches. That focus on latency and integration friction appears central to Monty's design thesis.
Pydantic already ships a different solution for Python execution called mcp-run-python, which runs full CPython via Pyodide inside Deno for isolation. That server moved from npm to JSR for better sandboxing but represents a heavier approach than Monty. The two tools target different points in the security/convenience/completeness tradeoff space. Why maintain both? Presumably because different use cases demand different tradeoffs.
The Broader Context

Monty arrives as AI agent development hits an inflection point. Models keep getting better at writing code. Agent frameworks increasingly treat code generation as a capability rather than an edge case. But execution safety remains genuinely hard—letting an LLM write Python that touches your filesystem or makes network calls is a non-starter for production systems. Building container infrastructure for every code execution adds operational burden and latency that many applications can't tolerate.
For teams building agent systems that need programmatic tool calling or code-based reasoning, Monty offers a middle path: fast, isolated Python execution with explicit control over host access. The tradeoff is accepting a restricted language subset and building external function interfaces for any capabilities agents need beyond basic computation. Whether that tradeoff makes sense depends entirely on your agent's requirements.
If your LLM needs to process data, perform calculations, or orchestrate tool calls through code rather than direct function invocation, Monty's microsecond startup and strict isolation model might fit. If your agents need full Python with rich libraries, you're still looking at Pyodide or container solutions. Or you accept the network latency of remote execution services.
The project remains experimental. Pydantic hasn't published a formal announcement beyond GitHub and social posts. But conda-forge packages are available, WASM builds work, and code mode integration is planned for Pydantic AI. Monty appears to be moving beyond proof-of-concept toward something developers can actually use, though exactly when it crosses that threshold remains uncertain.
Whether it becomes a standard tool for agent execution or remains a niche solution for specific use cases will likely depend on how well the restricted Python subset serves real-world agent workflows. And whether the constraints it imposes—no third-party packages, no full standard library—prove liberating or limiting. Right now, three months after release, it's too early to know. But the fact that Pydantic is betting on it for their own agent framework suggests they believe the minimal approach has legs.
Or maybe they just really wanted something that boots in 0.06 milliseconds. Sometimes speed is reason enough.
