Two founders, a library of 10,000 jailbreaks, and more than $100,000 advertised in weekly rewards. That's the wager Fabraix, a scrappy San Francisco outfit from Y Combinator's latest cohort, is making as it tries to carve out space in the crowded, occasionally chaotic world of AI security.
The bounty program went live this summer—specifically July 13—and it's less publicity stunt than stress test. Anyone who can successfully compromise one of the AI agents running in Fabraix's open-source Playground stands to collect. Human reviewers still vet the breaks before money changes hands, which is perhaps the smartest hedge the company has built into the whole arrangement.
But step back from the bounty theater for a moment. The real story here isn't the weekly prize pool. It's Nyx, the autonomous testing system that Fabraix has built underneath, and whether it actually delivers on what the founders claim: a way to probe AI agents the same way an attacker would, at scale, without ever seeing the source code.
An Arms Race With No Referees
Fabraix bills itself as "the world's frontier hacker for AI agents," which is the kind of claim that sounds either visionary or absurd depending on how the next twelve months unfold. The company says Nyx has already identified vulnerabilities in AI systems at Fortune 500 companies—though it hasn't publicly disclosed customer names or provided independent verification of these claims—but the suggestion is that corporations are paying for this kind of scrutiny, even if quietly.
The technology itself operates in what the company calls "blackbox, multi-turn" mode. Translation: Nyx interacts with AI agents purely through their external interfaces, adapting its attack strategy in real time based on how the target responds. No privileged access, no inside knowledge. Just relentless, methodical probing until something cracks or the budget runs dry.
Teams submit a target, set an objective, allocate resources. Then Nyx goes to work, drawing on what Fabraix describes as the largest library of jailbreaks it knows of—north of 10,000 exploits and counting. Whether that's genuinely the biggest repository out there is hard to verify independently, but it's a number meant to signal both ambition and seriousness.
The scope is wide: chatbots, autonomous agents, multi-agent systems, browser-based tools, voice assistants, even reinforcement learning setups. A recent LinkedIn post from the company showed Nyx testing voice agents with audio samples, hinting that the system can pivot across modalities without missing a beat.
Numbers, Claims, and the Absence of Outside Verification
On paper, the performance looks sharp. Fabraix says Nyx posted a 78% attack success rate on AgentHarm, a widely cited adversarial benchmark published last fall—though as of mid-July, this figure comes from the company's own testing without independent replication. That edges out what the company reports as a 67% success rate for GPT-5.6 Sol. Adaptive, multi-turn attacks—Nyx's signature move—succeed roughly 20 times more often than static or replayed attacks, according to the company's internal testing.
These figures come directly from Fabraix. As of mid-July, there's no independent third-party validation in the public record, which is worth keeping in mind. Vendor-supplied benchmarks have a way of aging poorly once scrutiny arrives.
The company has also rolled out its own benchmark called ACE (Adversarial Cost to Exploit), which quantifies how expensive it is to break various AI systems. In a blog post from earlier this year, Fabraix reported a 44-fold variance in adversarial cost across models, with mean costs ranging from $0.23 to $10.21. It's an interesting lens—framing security through the economics of exploitation—but the methodology has not yet been externally validated, and it's self-reported data awaiting broader verification.
Who's Building This

Behind Fabraix are two co-founders who logged time in fraud detection, compiler infrastructure, and AI debugging before deciding to tackle agent security.
Ahmed (Zach) Aly was the first data scientist at Two, a Sequoia-backed fraud detection platform. The company claims to move more than $1 billion in annual B2B transactions and credits its system with preventing $50 million in losses, though such figures are always tricky to verify after the fact. Aly published research during his time at University College London before leaving a PhD program to pursue the startup full-time.
Ibrahim Abdu spent a stint at Meta building AI agents designed to catch and fix production errors—unglamorous work, but the kind that teaches you where systems actually break under load. Before Meta, he was an early engineer at Two and built compiler and database tools at TradingHub. He graduated near the top of his cohort at Oxford with a degree in Philosophy, Politics and Economics, which is either relevant or a biographical footnote depending on how much stock you put in academic pedigree.
Together, they're betting that the chaos emerging around AI agents—the race to deploy, the opacity of behavior, the surface area for exploitation—creates an opening for a company willing to automate the red-teaming work most organizations still handle manually, if they handle it at all.
The Playground as Proof of Concept
The bounty program runs through what Fabraix calls its Playground, an open-source environment where security researchers can take live swings at AI agents. It debuted on Product Hunt in mid-July and landed at #5 for the day. As of then, the repository had collected 65 stars on GitHub—a modest following, though the company says an "agent runtime" feature is in the works.
Fabraix actually introduced an earlier version of the Playground this past March, drawing some coverage from niche outlets and a "Show HN" post on Hacker News. The reception was warm but not overwhelming. Building developer tools is a slow burn. You win by being useful, repeatedly, until people start recommending you without prompting.
A Market That's Already Crowded, and Getting More So

The challenge for Fabraix—and this is not a small one—is that the AI red-teaming space is filling up fast.
Promptfoo has an open-source suite with agent testing modules and documentation that's been updated within the past week. Giskard offers continuous red-teaming services. Robust Intelligence and Knostic both provide platform-level defenses with adversarial testing baked in.
On the research side, the activity is even more frenetic. Anthropic published ExploitBench earlier this year, a capability-ladder benchmark for LLM cybersecurity agents. Wiz introduced an AI Cyber Model Arena in February for testing offensive security tasks. Microsoft Research released AgentRx in March, a framework for systematic agent debugging.
Fabraix's counter is straightforward: Nyx is fully autonomous, adapting attacks without human intervention. That's a meaningful distinction if it holds up under real-world pressure. Most red-teaming still requires human judgment at critical junctures. Automating that step is valuable—if the automation is actually reliable, which is always the question with systems that learn and adapt on the fly.
Business Model and What Comes Next

The company offers a free Research tier, though you need to apply for access. Larger teams pay custom pricing—details undisclosed—and there's a complementary product called Arx that provides runtime defense and pre-execution validation using insights from Nyx. Whether organizations will pay for both remains to be seen, but the logic is sound: find the holes, then help close them.
Fabraix emerged from Y Combinator's Summer 2026 batch. No other funding rounds have been announced publicly, which means the company is either bootstrapping off early revenue or preparing a seed round that hasn't landed yet.
It's early. The technology is promising but unproven at scale. The market is getting louder by the month. And the founders are making a very public bet—literally, given the bounty program—that their approach to AI security testing can stand out in a field where everyone claims to have the sharpest tools.
Whether Fabraix becomes the standard or just another footnote in the agent security wars will depend less on clever benchmarks and more on whether enterprises actually trust Nyx to find what human red teams might miss. That's a harder proof point than a Product Hunt ranking or a GitHub star count.
But for now, the company has captured attention—and put real money behind the claim that its system works. In a sector full of promises and whitepapers, that counts for something.
