The three Israeli entrepreneurs had already cashed out once—$95 million when Fiverr acquired their talent marketplace Stoke in late 2021. They could have retired. Instead, Shahar Erez, Hilik Paz, and Tal Salmona spent two years inside Fiverr building out enterprise software, watching a different problem metastasize in real time.
Companies were racing to deploy AI agents. Customer service bots, coding assistants, financial advisors powered by large language models. And almost no one, the founders noticed, had a reliable way to know whether these systems would behave when actual humans—confused, impatient, or outright hostile—started using them.
So they left. In March 2024, they launched Arato.ai out of Tel Aviv. In June 2026, the startup disclosed it raised $10 million in seed funding led by TLV Partners, with backing from Jibe Ventures and notable angels including Raghu Raghuram, VMware's former CEO now at Andreessen Horowitz, and Marianna Tessel, who spent years as CTO at Intuit.
Their pitch? Simulate thousands of user interactions before launch. Stress-test AI systems against not just ideal users but the full spectrum: people who misunderstand prompts, adversaries probing for vulnerabilities, bad actors trying to extract sensitive data or manipulate outputs.
Beyond Traditional Testing
Arato's approach diverges from conventional software QA. Traditional testing tools check for bugs in code; model evaluation frameworks measure accuracy or bias in training data. Arato sits somewhere else entirely—what the company describes as a "behavioral assurance layer" designed specifically for autonomous AI systems.
The platform generates what it calls a "simulation context and scenario matrix." In practice, that means spinning up synthetic personas—each with different intents, communication styles, and levels of technical sophistication—and running them through multi-turn conversations across text, voice, image, and data interfaces. The system logs where things fracture: misunderstood queries, hallucinated responses, security breaches, inappropriate escalations.
Then it prioritizes. Not every failure matters equally. Arato's analysis weights breakdowns by potential business impact, surfacing which risks demand immediate fixes and which can wait.
"Not a testing tool. Not a model eval framework," the company stated in a recent LinkedIn post. It's an assertion that reveals how crowded—and fragmented—this emerging market has become.
A Pattern Emerges

Arato's funding arrives amid a concentrated burst of capital into AI reliability startups. General Analysis pulled in $10 million in late April for adversarial testing and red-teaming. ZeroDrift closed a $10 million seed round in early June for what it calls a compliance firewall. Tenet raised $6 million around the same time, focusing on runtime protection for agents accessing sensitive enterprise data.
The clustering isn't accidental. Investors are betting that a widening gap exists between AI labs shipping powerful models and enterprises actually deploying those models at scale without catastrophic failures. Each startup is staking out different territory—runtime monitoring, compliance guardrails, pre-deployment simulation—but the underlying thesis is the same: blind faith isn't a release strategy.
Arato's founders are narrowing in on that pre-launch moment. They're not promising to catch every misbehavior in production or guarantee regulatory compliance in perpetuity. Instead, they're selling peace of mind before the launch button gets pressed.
The Go-To-Market Play

With capital secured, Arato has been building out its commercial operation. The company recently posted for a founding account executive role targeting U.S. and European buyers—a signal that the founders see the problem extending well beyond Israel's tech ecosystem.
They've also been showing up where enterprise QA leaders gather. StarEAST, a software testing conference held in Orlando in 2026, featured Arato among its exhibitors. That's deliberate positioning: speaking the language of quality assurance, even as the underlying challenge—testing non-deterministic AI behavior—demands entirely new playbooks.
The team has grown to somewhere between 11 and 50 employees, according to recent company disclosures. And Arato recently shipped connectors for Claude and OpenAI's Model Context Protocol, allowing engineering teams to query simulation results directly from chat interfaces. It's a meta move, perhaps—using AI to interrogate AI.
The Bet

Whether Arato becomes the standard or just another entrant in an increasingly noisy space remains unclear. The founders have one exit under their belts, which buys credibility. The investor roster—particularly Raghuram and Tessel, both veterans of scaling enterprise infrastructure—suggests serious people believe the problem is real.
But the window might be narrower than it appears. Enterprises are already deploying AI agents, often without rigorous testing. Waiting for perfect tooling could mean missing the market entirely. Arato is wagering that enough decision-makers will slow down, just long enough to simulate what happens when things inevitably go sideways.
Simulation, not faith. That's the pitch. Whether it lands depends on how many companies learn the hard way first.
