The complaint landed on Reddit like so many cautionary tales do these days—terse, technical, vaguely horrified. A user had watched Claude, Anthropic's AI assistant, systematically dismantle their staging environment by commenting out Terraform's prevent_destroy flags and cheerfully rerunning the configuration. Gone. Just like that.
For Collin Pfeifer, the post wasn't shocking so much as inevitable. Of course someone's AI assistant had torched their infrastructure. The surprise was that it hadn't happened more often, or at least that people weren't talking about it more. "LLMs guess poorly without real production context," Pfeifer wrote in a blog post last month, explaining the genesis of Fluid.sh, his answer to a problem that DevOps engineers are only beginning to articulate out loud: AI coding assistants are remarkably powerful. They're also terrifyingly prone to mistakes when let loose on the systems that keep businesses running.
Fluid, which released its first public version on February 4 under an MIT license, is what Pfeifer calls "Claude Code for infrastructure." That framing is intentional—borrowed from Anthropic's own term for its AI assistant's ability to execute commands and write code. But where Claude Code operates directly on your machine, Fluid introduces something crucial: a sandbox. Or more accurately, an entire virtualized environment where the AI can experiment, break things, and learn, all while your actual production systems remain safely out of reach.
The Disposable Test Kitchen
The mechanics are almost elegant in their simplicity. Fluid clones your production environment into a throwaway virtual machine using KVM's copy-on-write snapshots. Claude gets turned loose in that sandbox—installing packages, editing configuration files, restarting services, poking around with full root access. The AI can diagnose problems with genuine context from a real system, not the sort of abstracted guesswork that comes from analyzing code snippets in isolation.
Once Claude finishes its work (assuming it doesn't go off the rails, though Pfeifer has built in safeguards for that too), Fluid captures every meaningful change into an Ansible playbook. That playbook gets handed to an actual human engineer for review. Standard Infrastructure-as-Code practices take over from there—code review, version control, proper deployment processes. The AI does the exploration and grunt work. The humans make the final call.
It's a workflow that addresses what might be the central tension in AI-assisted operations: these tools are genuinely useful for the tedious, context-heavy debugging that consumes on-call shifts, but a single misstep can cascade into disaster.
The architecture itself runs on KVM/QEMU via libvirt—hardly exotic technology, but deliberately unglamorous. A control node handles the Fluid API and PostgreSQL database, connecting over SSH to libvirt hosts that run nothing but KVM and libvirt. No additional software to maintain, no complex dependencies. When an engineer needs to troubleshoot something, Fluid spins up a VM from a base qcow2 image. The AI gets an isolated playground.
Real-time command output streams through tmux integration, which means engineers can watch Claude work. Intervene if necessary. The tool maintains awareness of the operating system, what packages are installed, which CLI utilities exist—all the environmental context that separates a useful suggestion from a destructive one. Every command gets logged. Diffs appear for human review before anything reaches production.
Pfeifer has documented setup scripts for various architectures, including macOS with Apple Silicon using Homebrew, though the control node itself requires Linux x86_64 with systemd. The GitHub repository, which accumulated 329 stars in its first days, includes what one might generously call "detailed" instructions—the sort of documentation that assumes a certain comfort level with virtualization infrastructure.
Guardrails Upon Guardrails

Security isn't an afterthought here. Perhaps that's because Pfeifer clearly anticipated the skepticism. Each sandbox runs in its own KVM virtual machine with network isolation. Fluid generates ephemeral SSH certificates that expire after one to ten minutes—brief enough that even if something goes sideways, the exposure window stays narrow. Command and path allow/deny lists add another layer of control.
For operations that cross certain thresholds—installing packages, making outbound internet requests, doing anything on hosts running low on memory—Fluid stops and asks for human approval. The API includes a blocking endpoint that literally pauses agent execution until an engineer signs off. Engineers can snapshot sandbox states, roll back changes, essentially treating the whole environment like a particularly sophisticated undo button.
It's an approach that directly confronts the Reddit horror stories and Slack channel warnings that have become DevOps lore. Don't give the AI the keys to production, in other words. Give it the keys to a disposable copy of production, then review what it learned before applying any of those lessons to the real thing.
Whether that's enough remains an open question. AI assistants have proven capable of subtle errors—the kind that pass initial review but create problems weeks later. Still, the alternative appears to be either avoiding AI assistance entirely for infrastructure work or accepting a level of risk that makes most engineers understandably nervous.
Riding a Wave (Or Creating One)

Fluid's launch timing suggests Pfeifer isn't alone in sensing opportunity here. The Hacker News discussion—269 points, 174 comments in roughly 48 hours—drew the usual mix of enthusiasm and skepticism that accompanies infrastructure tooling. Pfeifer responded actively to questions across Reddit's r/ClaudeCode, r/AgentsOfAI, and r/LLMDevs communities, fielding everything from architecture questions to feature requests.
The broader market has clearly taken notice of AI-assisted infrastructure. Pulumi unveiled Neo, billing it as an AI "platform engineer," back in September. HashiCorp announced Agent Skills for Claude Code on February 2, just days before Fluid's release. Spacelift released Intent for natural-language infrastructure provisioning. The pattern is unmistakable: every major infrastructure vendor is racing to figure out how AI fits into their workflow.
Fluid distinguishes itself partly through openness—it's MIT-licensed with no announced commercial plans—and partly through its particular theory of safety. Where some tools focus on better prompts or smarter AI, Fluid's bet is simpler: isolation first, intelligence second.
The target user is probably the on-call engineer at 2 a.m., staring at a production issue and needing to restart services, reconfigure a firewall, or diagnose why systemd is behaving oddly. Tasks where AI assistance could genuinely help, but where the stakes of getting it wrong might mean a very unpleasant morning explaining things to management.
Whether Fluid becomes a standard tool or a footnote depends on questions that won't be answered for months: Does the sandbox workflow actually prevent the disasters it's designed to avoid? Will teams trust AI enough to let it touch infrastructure at all, even in isolation? And perhaps most importantly, will the generated Ansible playbooks prove useful enough to justify the additional complexity?
For now, Pfeifer has built something that at least tries to thread a particularly tricky needle. AI assistants aren't going away. Infrastructure isn't getting simpler. The question was never if these two worlds would collide, but how much damage that collision would cause. Fluid's answer: let them collide in a sandbox first, and see what survives.
