Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
Fintech iconFintechMay 6, 2026

AI Takes the Wheel: Inside the First Fully Autonomous Hedge Funds

AI Takes the Wheel: Inside the First Fully Autonomous Hedge Funds
YcAi+3
Healthtech & Biotech iconHealthtech & BiotechMay 6, 2026

The Race to Build Factories in Space for Pharma and Semiconductors

The Race to Build Factories in Space for Pharma and Semiconductors
Space TechAerospace+3

Founders Mentioned

Aum Upadhyay

Silmaril

saas icon
SaaS

Eduardo Velasco

Silmaril

saas icon
SaaS

Aum Upadhyay

Silmaril

saas icon
SaaS

Eduardo Velasco

Silmaril

saas icon
SaaS
SaaS iconSaaS
May 6, 2026
YcAi AgentsCybersecurityEnterprise SecurityB2b Saas

YC's Silmaril Launches Self-Healing AI Security as Agent Attacks Surge

As AI agents proliferate, a YC startup unveils security that auto-updates within an hour. Claims 2x better blocking with 10x lower latency amid wave of high-profile breaches.

YC's Silmaril Launches Self-Healing AI Security as Agent Attacks Surge

The email sat in the inbox, unassuming. No link to click, no attachment to download. And yet, within moments, Microsoft 365 Copilot began leaking data—quietly, obediently, to an attacker who never had to persuade a single human being to do anything foolish. EchoLeak, as the vulnerability came to be known, earned a CVSS score of 9.3 when researchers disclosed it in June 2025. Zero-click prompt injection, they called it. The AI assistant, in essence, had been turned into an unwitting courier.

It wasn't an outlier, either. By this spring, the roster of AI agent vulnerabilities read like a hacker's wish list: GitHub Copilot's "YOLO mode" that researchers weaponized into remote code execution, disclosed in late June 2025 and patched that August. A Perplexity browser agent manipulated to exfiltrate 1Password vaults. Framework-level flaws in LangChain that exposed enterprise data through half a dozen different attack vectors.

Into this mounting chaos stepped Silmaril, a two-person San Francisco outfit that emerged from Y Combinator's Spring 2026 batch. Their pitch? A self-healing defense system that blocks twice as many attacks as existing guardrails, with one-tenth the latency. Former AWS security architect Aum Upadhyay and ex-Amazon ML engineer Eduardo Velasco claim they can ship updated classifier weights within an hour of discovering new threats.

Whether that's fast enough to matter when the UK's National Cyber Security Centre warns that prompt injection "might never be properly mitigated" is another question entirely. But the timing, at least, couldn't be sharper.

When the Attack Surface Became Infinite

The technical write-up of EchoLeak, published last September, detailed a vulnerability chain elegant in its cruelty. Attackers could sidestep Microsoft's XPIA classifier, evade link redaction, and route stolen data through Teams—all set in motion by an email that required nothing from the victim except the bad luck of receiving it. Microsoft patched the server-side flaw in May, but didn't disclose it publicly until a month later. That lag matters more than it used to, now that AI agents are proliferating faster than security teams can catalog them.

The GitHub Copilot exploit, disclosed in August, went further still. Researchers showed how a malicious prompt could modify Visual Studio Code settings to auto-approve tool calls, which then escalated to full remote code execution. The attack was almost embarrassingly simple: inject instructions telling the agent to trust everything. Then ride that trust all the way down to the operating system.

Perhaps most unnerving was the Perplexity Comet browser agent attack that Zenity Labs detailed in March. An indirect injection, hidden inside a calendar invite, could exfiltrate 1Password vault contents and escalate to full account takeover. Vendors got the disclosure back in November, but the public write-up didn't surface for four months—a reminder, if anyone needed it, that the vulnerability pipeline is always fuller than we're told.

By April, LangChain and LangGraph had accumulated their own catalog of critical issues. Path traversal bugs. Deserialization flaws. Secret leaks. SQL injection vectors. Anthropic's Model Context Protocol, released with some fanfare as a way to extend Claude's capabilities, introduced a new class of chained exploits that could lead to remote code execution and file tampering. Vendors scrambled. Some issues remained unresolved well into spring.

The Trilemma

Security researchers have started framing the problem in starker terms: it's architectural, not just tactical. A paper published in April described a "defense trilemma" facing prompt injection wrappers—they can't simultaneously maintain security, utility, and performance. Static rule-based filters generate false positives that break legitimate workflows. Simple classifiers trained on known attacks miss the novel variants that inevitably follow. And LLM-based guardrails that attempt to reason about inputs introduce latency that renders real-time applications unusable.

OWASP's 2025 Top 10 for Large Language Model Applications, updated this past April, put prompt injection at number one. NIST formalized the taxonomy in its AI 100-2e2025 guidance, distinguishing between direct and indirect attacks. But formalizing the problem and solving it turn out to be very different endeavors.

OpenAI, in a December blog post on hardening Atlas against prompt injection, acknowledged it as "a long-term AI security challenge." The post emphasized continuous loops and attack-trace analysis rather than any silver bullet. Microsoft's internal research, published in February, demonstrated that a single unlabeled prompt could unalign multiple models—a safety failure that compounds when agents start interacting with each other. Their Spotlighting technique, which separates trusted instructions from untrusted content, helps. But it doesn't eliminate the fundamental issue: LLMs lack a hard boundary between data and instructions.

Gartner predicted last August that 40 percent of enterprise applications would feature task-specific AI agents by the end of this year, up from less than 5 percent in 2025. An April CIO survey, though, found that only 17 percent of organizations had actually deployed them. That gap—between prediction and reality—is the gap between ambition and risk tolerance.

The Self-Healing Bet

Digital illustration for article section "The Self-Healing Bet" in "YC's Silmaril Launches Self-Healing AI Security as Agent Attacks Surge" - A modern, minimalist editorial illustration representing a three-layered self-healing architecture t...

Silmaril's architecture, detailed in its YC launch post, layers three components. Autonomous threat-hunting agents probe for vulnerabilities. A runtime classifier firewall, built on a modernBERT variant with FlashAttention, sits in the middle. Then there's the continuous retraining loop. When the threat hunters discover a new exploit, verified attacks feed back into the model. The company says it ships updated weights within an hour.

That's an order of magnitude faster than traditional security update cycles, assuming it works as advertised.

The performance numbers—all company-provided, it's worth noting—are striking. Silmaril claims a 96 percent block rate on real-world contextual and emerging threats, compared to 61 percent for what it characterizes as "leading guardrails." Latency hovers around 20 milliseconds at the 90th percentile, which the founders insist is fast enough for production environments. Integration happens through five lines of code in frameworks like LangChain, LangGraph, CrewAI, and Vercel AI SDK.

The team backs its claims with case studies citing $68 million and $20 million in damages averted for customers. Those figures are unverified by independent parties. Silmaril also reports disclosing 15 critical vulnerabilities to OpenAI, Anthropic, Google, and Microsoft—evidence, if accurate, that the founders understand the attack surface from multiple angles.

Upadhyay's background includes leading a security and privacy framework at AWS that reportedly prevented $1.8 billion in damages. Velasco has publicly demonstrated chaining a prompt injection into root access inside ChatGPT. It's a pedigree suggesting familiarity with both defensive infrastructure and offensive techniques. But Silmaril remains a two-person, bootstrapped YC company with no announced funding beyond its batch participation.

The broader market for "LLM firewalls" is nascent—IT-Harvest estimated roughly $30 million in revenue for the category in 2025, with expectations for 100 percent growth this year, according to a March TechTarget feature. That's a rounding error in a global security spend that IDC projects will reach $308 billion in 2026.

A Crowded Field, Moving Fast

Silmaril isn't operating in a vacuum. Microsoft's Entra Global Secure Access, updated in April, offers prompt injection protection at the network edge. Google Cloud's Model Armor provides scanning for prompt injection, jailbreaks, and data exfiltration. OpenAI released GPT-OSS-Safeguard last October, an open-source classifier that developers can deploy inline. Anthropic published browser-specific prompt injection defenses in November.

Commercial offerings are proliferating, too. Lakera Guard, acquired by Check Point, maintains widely cited datasets like Gandalf and PINT. Superagent, a YC Winter 2024 company, positions its product as an "AI Firewall" sitting between applications and LLMs. NVIDIA's NeMo Guardrails and Meta's Llama Prompt-Guard models are open-source options that integrate into orchestration layers.

Academic research is racing to catch up, though perhaps racing is the wrong word—scrambling might be more accurate. AgentDyn, published in February, introduced a dynamic benchmark showing that many defenses either over-block legitimate inputs or under-secure against novel attacks. ClawGuard, released in April, proposed runtime rule enforcement for tool-augmented agents. PCFI, also from March, applied control-flow integrity concepts to prompt execution.

The challenge they all face is the one Silmaril is betting it can solve: speed. Threats evolve daily. Static defenses decay. A classifier trained on January's attacks may miss February's variants by a wide margin. The self-healing loop—if it works as advertised—offers a different paradigm. It also introduces its own questions about false positives, model drift, and the operational overhead of continuous retraining.

The Regulatory Clock Is Ticking

Digital illustration for article section "The Regulatory Clock Is Ticking" in "YC's Silmaril Launches Self-Healing AI Security as Agent Attacks Surge" - A conceptual, modern editorial illustration featuring a large, abstract stylized clock face where th...

The regulatory environment is tightening in parallel with the technical arms race. The EU AI Act's general-purpose AI obligations entered force last August, with full enforcement for GPAI providers coming this August. ISO/IEC 42001, the first certifiable AI management systems standard, saw its first accreditation granted by UKAS in January. Enterprises evaluating AI agents increasingly face procurement requirements demanding evidence of continuous security controls.

KPMG's March Global AI Pulse survey found that 32 percent of organizations are deploying or scaling agents, with 27 percent orchestrating multiple agents. Gartner forecasts $2.52 trillion in worldwide AI spending this year, a 44 percent jump year-over-year. Security controls that can't keep pace with deployment velocity become deployment blockers. Simple as that.

The UK NCSC's December warning about prompt injection perhaps never being "properly mitigated" wasn't defeatist, exactly. It was a call for defense in depth. No single layer will suffice. Runtime firewalls like Silmaril's, if they prove effective, become one control among many: input validation, output sanitization, privilege minimization, network segmentation, audit logging. The full suite.

Whether a two-person startup with hour-scale update cycles can outmaneuver the collective ingenuity of attackers targeting trillion-dollar AI deployments remains unproven. The benchmarks are company-claimed. The case studies are unaudited. The founders are talented, sure, but the problem they're tackling has humbled much larger organizations.

Still, the underlying bet is sound. In a world where agents move faster than human security teams, only automated defenses that learn in real time have much of a chance. Silmaril is wagering that self-healing isn't just a feature—it's the only architecture that scales. Time will tell if they're right, or if they're just the latest in a long line of security startups that discovered the gap between laboratory conditions and the messy, chaotic reality of production systems is wider than anyone wants to admit.

More stories

  • DoD Solution raises $2M for AI drone navigation in war zones
  • DesignVerse raises $5.5M to automate enterprise software
  • AI Takes the Wheel: Inside the First Fully Autonomous Hedge Funds
  • The Race to Build Factories in Space for Pharma and Semiconductors
  • AI Research Goes Meta: How Labs Use Agents to Discover Better AI
  • Y Combinator's First All-Crypto Investment: $500K in USDC on Solana
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.