Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSMarch 9, 2026

Palo Alto Networks Founder Raises $45M for Sovereign Security Startup

Palo Alto Networks Founder Raises $45M for Sovereign Security Startup
CybersecurityEnterprise Security+3
Healthtech & Biotech iconHealthtech & BiotechMarch 8, 2026

ROXFIT Raises £1.9M to Scale AI Training Platform for Hybrid Athletes

ROXFIT Raises £1.9M to Scale AI Training Platform for Hybrid Athletes
Fit TechAi+3
SaaS iconSaaS
March 9, 2026

Agent Safehouse Brings Native Sandboxing to macOS AI Agents

Open-source tool leverages Apple's Seatbelt kernel to lock down local AI coding agents with minimal overhead, addressing security concerns as agent adoption accelerates.

Agent Safehouse Brings Native Sandboxing to macOS AI Agents

Eugene didn't set out to launch a product. The New York-based developer just wanted his AI coding agents to stop having full run of his laptop.

On March 9, he posted Agent Safehouse to Hacker News with a pitch that could've been a tweet: "macOS-native sandboxing for local agents. Move fast, break nothing." Within hours, it had climbed to 529 points. The comments piled up—127 in total—and most struck a similar chord. Finally.

Anyone running AI coding tools locally knows the uneasy bargain. Cursor, Aider, Claude Code—they're stunningly useful. They write code, execute terminal commands, parse your entire project structure. They also run with your user permissions, which means a single garbled prompt could theoretically expose SSH keys, browser session cookies, or that side project you're not ready to show anyone.

Agent Safehouse tries to fix that, at least on macOS, by wrapping agents in Apple's Seatbelt kernel—the same sandboxing primitive the operating system uses under the hood for system apps. Not a VM. Not Docker. Just a thin policy layer that intercepts filesystem requests at the kernel level before anything dangerous happens. Agents can still modify your project files. They just can't peek into ~/.ssh/id_ed25519 or wander into unrelated directories unless you explicitly allow it.

Installation That Doesn't Feel Like Installation

One curl command. That's it. You pull down a self-contained Bash script, drop it in ~/.local/bin, and then prefix your agent commands with safehouse. The tool auto-detects your git root as the working directory, applies a deny-all baseline, and layers on permissions for common toolchains.

Running Cursor's agent in a sandbox? safehouse cursor-agent. Need read-only access to a config directory? Tack on --add-dirs-ro ~/.config/myapp.

Eugene—who declined to share his last name—spent what he describes as "way too many hours" profiling a dozen popular agents to identify their minimum required permissions. Aider, Auggie, Claude Code, Cline, Codex, Cursor, Goose. The investigations are published as detailed docs on the project site, each noting architecture quirks, config paths, and the occasional sandboxing gotcha.

There's a Policy Builder, too: a point-and-click web app where you select an agent, pick toolchains (Node, Python, Docker), toggle optional capabilities like clipboard access or process control, and download a ready-to-run .sb policy file. Or, if you'd rather, paste some instructions into Claude or GPT and let it scan your dotfiles, ask one question about writeable directories, then emit a least-privilege profile. The whole thing feels almost too simple, which is probably the point.

What It Actually Stops (And Where It Shrugs)

The threat model is refreshingly candid. Safehouse aims to reduce the "blast radius" of an LLM agent gone sideways. It won't protect you from a determined adversary exploiting a kernel vulnerability. It's not a hypervisor. It won't prevent network exfiltration of data the agent can already read. The docs state it plainly: "usability prioritized over paranoid lockdown."

By default, the sandbox denies writes outside your project directory. It blocks access to SSH private keys, browser profiles, shell startup files, raw /dev device access, and setuid binaries. Network access stays open—these agents need to install packages and call APIs, after all. Optional toggles exist for Docker socket access, Kubernetes configs, browser native messaging, LLDB, and a "wide-read" mode for agents that need broader filesystem visibility.

The homepage includes live proof examples. Try to cat ~/.ssh/id_ed25519 inside a sandboxed shell and the kernel stops you cold. Same for listing files in an unrelated project directory. Operations inside the granted workdir? Those proceed normally.

Whether that's enough depends on your paranoia level. For developers who just want basic guardrails without rearchitecting their entire workflow, it's probably sufficient. For teams worried about sophisticated supply chain attacks or insider threats, Eugene's advice is explicit: layer Safehouse inside a VM.

The Timing Is No Accident

Agent Safehouse lands as the industry converges—belatedly, some would say—on OS-native sandboxing for AI tools.

Digital illustration for article section "The Timing Is No Accident" in "Agent Safehouse Brings Native Sandboxing to macOS AI Agents"

Cursor published an engineering post a couple weeks back detailing their own cross-platform approach: Seatbelt on macOS, Landlock and seccomp on Linux. They claimed sandboxed agents "stop 40% less often" because clearer error boundaries help developers troubleshoot faster. Anthropic released sandbox-runtime, an experimental open-source library that also leans on sandbox-exec and dynamically-generated Seatbelt profiles on macOS, with bubblewrap handling Linux.

Other projects take different angles. Nono emphasizes kernel-enforced capability models with attestation. Microsandbox and Agent Harbor lean on VM-level isolation. DevCage and AgentSphere target multi-platform or cloud deployments. Kilntainers gives each agent an ephemeral Linux sandbox via containers or microVMs.

Safehouse distinguishes itself by staying host-native, macOS-only, and ruthlessly lightweight. The docs include an isolation model comparison table: VMs offer the strongest boundary but add overhead; containers sit somewhere in the middle; Safehouse trades some isolation strength for "very low" performance impact and seamless local integration.

It's a narrow niche. But perhaps a growing one.

The Work Continues

Recent commits show active refinement. Around the time of the Hacker News launch, Eugene added GitHub Copilot CLI support, split process-control from LLDB toggles (developers wanted debugging without full process control), integrated agent-browser for headless automation, and improved the LLM generation instructions. Earlier work included broadening allowances for the macOS open command to keep UX smooth. A license shift to Apache-2.0 signals an eye toward adoption-friendly terms.

Testing infrastructure is thorough: sectioned tests, end-to-end TUI simulation via tmux to verify expected deny/allow behavior, live agent checks with parallelization options. You can run --explain to see a human-readable summary of effective grants, or --stdout to inspect the raw SBPL policy.

Eugene suggests wrapping the tool in shell functions for daily use. Alias claude to safehouse claude, and the sandbox becomes your default. Need to bypass it? Run command claude to invoke the unwrapped binary.

In the Hacker News thread, Eugene engaged directly with commenters. One asked about ~/.gitignore access; Eugene added read-only grant logic. Another raised concerns about headless Chrome brittleness; Eugene acknowledged the complexity and pointed to recent browser integration work. The tone throughout felt pragmatic—this is a tool built by someone running local agents daily, not a venture-backed security startup hunting for enterprise contracts.

Where This Goes Next

With no formal company, no funding, and no pricing (Agent Safehouse is pure open source), adoption will track through GitHub stars, pull requests, and downstream integrations. The Hacker News launch generated genuine interest, though the usual caveats apply. One commenter noted the irony of using an LLM to generate sandboxing policies for other LLMs. Another asked about Windows support—sandbox-exec is macOS-specific, though Eugene's detailed investigations could inform analogous tooling elsewhere.

Digital illustration for article section "Where This Goes Next" in "Agent Safehouse Brings Native Sandboxing to macOS AI Agents"

The project positions itself within an emerging category rather than claiming to invent anything. The docs reference Anthropic's sandbox-runtime, restrictive policy experiments, and an overview of sandboxing patterns in the field. That's probably the right posture. As agent adoption accelerates and more developers run autonomous code on their personal machines, the question isn't whether sandboxing becomes standard—it's which implementation wins on ease of use.

For now, Agent Safehouse offers macOS developers a zero-dependency option that takes five minutes to install and doesn't require rethinking your stack. Whether it gains traction beyond early adopters depends on how well the composable policy model scales to edge cases, how actively Eugene (or potential contributors) maintain agent-specific profiles as tools evolve, and whether the single-script simplicity holds up under production complexity.

The bones seem solid. Seatbelt is battle-tested kernel infrastructure. The deny-first model maps cleanly to how developers already think about least privilege. And the LLM-assisted policy generation is a clever bootstrap for customization, even if it does feel slightly recursive.

If you're running agents on macOS and haven't locked them down yet, it's worth a curl. Or at least a few minutes of healthy paranoia about what exactly those helpful coding assistants can see.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Palo Alto Networks Founder Raises $45M for Sovereign Security Startup
  • ROXFIT Raises £1.9M to Scale AI Training Platform for Hybrid Athletes
  • EU Project Unveils Portable AI Recycling Plant for Remote Areas
  • Chime vs Nubank: Two Paths to Neobank Profitability in 2026
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.