Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSJune 7, 2026

AI Agents Replace Consultants: Ontora Maps Workflows in Hours, Not Months

AI Agents Replace Consultants: Ontora Maps Workflows in Hours, Not Months
YcAi Agents+3
SaaS iconSaaSJune 7, 2026

Interfaze Merges DNNs with LLMs for Deterministic AI Tasks

Interfaze Merges DNNs with LLMs for Deterministic AI Tasks
YcAi Infrastructure+3

Founders Mentioned

Aum Upadhyay

Silmaril

saas icon
SaaS

Eduardo Velasco

Silmaril

saas icon
SaaS

Aum Upadhyay

Silmaril

saas icon
SaaS

Eduardo Velasco

Silmaril

saas icon
SaaS
SaaS iconSaaS
June 7, 2026
YcAi AgentsCybersecurityEnterprise SecurityThreat Intelligence

Silmaril's Self-Healing Defense Blocks AI Prompt Injection Attacks

YC-backed security startup claims 2x better threat detection with 10x lower latency as AI agents face rising attacks. Autonomous retraining deploys updates within an hour.

Silmaril's Self-Healing Defense Blocks AI Prompt Injection Attacks

The attack arrives disguised as innocuous web content. A malicious instruction, buried in a product description or embedded in a support document, hijacks an AI agent mid-conversation and tricks it into exfiltrating sensitive data. By the time anyone notices, the damage is done.

That scenario played out in Microsoft's Copilot Studio this January, when researchers disclosed CVE-2026-21520—a vulnerability that let attackers manipulate the tool through carefully crafted public forms. Microsoft patched the flaw on January 15, though the full scope of the vulnerability's impact across different configurations continued to be assessed. Salesforce dealt with something similar around the same time, an incident dubbed "PipeLeak" that followed an earlier breach in September 2025.

Welcome to the world of prompt injection, a threat that's gone from academic curiosity to production nightmare in the span of eighteen months. And into that chaos steps Silmaril, a Y Combinator-backed startup founded by two former Amazon engineers who think they've cracked a problem that the UK's National Cyber Security Centre calls "the most persistent and difficult-to-fix threat" in AI security.

Maybe. The venture-backed promise of silver-bullet solutions rarely survives contact with determined adversaries, particularly when those adversaries are learning just as fast as the defenses evolving to stop them.

When the Threat Went Mainstream

The numbers tell a story of acceleration that's caught most enterprises flat-footed. Google's security team documented a 32% surge in malicious injection attempts scraped from public web content between November 2025 and February 2026. By April, Forcepoint X-Labs had identified ten verified injection payloads actively deployed on live websites. Not proof-of-concept demonstrations tucked away in academic papers. Actual attacks, running in the wild.

The timing couldn't be worse for organizations racing to deploy AI agents. Stanford's AI Index 2026, published in May drawing on 2025 McKinsey survey data, found that 88% of organizations now use AI in at least one business function, with 79% deploying generative AI regularly. Gartner's May Hype Cycle revealed an even more revealing gap: while only 17% of organizations currently run AI agents in production, more than 60% plan to do so within two years.

That gulf between intent and execution has created what security professionals describe, with uncharacteristic bluntness, as a feeding frenzy for attackers.

OWASP—the nonprofit that maintains the canonical list of web application vulnerabilities—ranks prompt injection as LLM01, the highest-severity risk category for large language model applications. It's a designation that carries weight precisely because OWASP doesn't traffic in hyperbole. The group that spent two decades cataloging SQL injection attacks knows what a fundamental architectural flaw looks like.

Two Ex-Amazonians Walk Into a Security Crisis

Aum Upadhyay and Eduardo Velasco launched Silmaril in April 2026 with the kind of credentials that make venture capitalists reach for their checkbooks. Upadhyay says he built AWS security frameworks that prevented $1.8 billion in damages during his tenure. Velasco's claim to fame is more baroque: he reportedly chained a prompt injection into root access inside ChatGPT, a hack that reads like either brilliant security research or the plot of a techno-thriller depending on your perspective.

Their pitch centers on what they call a "self-healing" defense system. Vendor-reported benchmarks show 96% threat blocking compared to 61% for leading guardrails, with ten times lower latency and a p90 overhead of just 20 milliseconds. Those figures, if they hold up under independent testing, would represent a meaningful improvement. The guardrails market has been plagued by a painful tradeoff: comprehensive protection that introduces latency measured in hundreds of milliseconds, or lightweight filters that miss sophisticated attacks.

Silmaril claims to thread that needle by reasoning over what it calls "application snapshots"—a combination of user intent, execution context, and state that's designed to catch multi-turn attacks building incrementally across conversations. More unusually, the company says it operates a continuous threat-hunting loop, feeding synthetic training data back into its classifier and redeploying updated weights within an hour of discovering new attack patterns.

That autonomous retraining cycle is the "self-healing" part. Whether it works as advertised remains an open question. The company cites prevented damages of $68 million and $20 million in separate incidents, though without independent verification those numbers land somewhere between impressive and impossible to assess. Silmaril also says it's reported vulnerabilities to Microsoft, OpenAI, Anthropic, and Google that resulted in significant patches, claims that await independent verification. If true, it suggests the founders have been busy in their first months.

The company integrates with LangGraph, LangChain, and similar agent frameworks via what it describes as five-line SDK implementations. Whether that simplicity masks underlying complexity or genuinely reflects elegant engineering is the sort of question that only becomes clear when enterprises start running it in production.

Why This Problem is Different (And Harder)

Digital illustration for article section "Why This Problem is Different (And Harder)" in "Silmaril's Self-Healing Defense Blocks AI Prompt Injection Attacks" - A beautifully crafted, vintage wooden sorting box resting on a clean, uncluttered surface, serving a...

Security veterans keep reaching for the SQL injection analogy, and it's not a terrible comparison. Two decades ago, web applications routinely mixed user input with database commands, creating vulnerabilities that let attackers manipulate queries and steal data. The solution—parameterized queries that strictly separated code from data—eventually became table stakes for web development.

Prompt injection exploits a similar architectural flaw, but one that's proving far more stubborn to fix. Large language models don't distinguish between instructions and data the way databases do. Everything is text. An agent reading a web page can't reliably tell the difference between legitimate content and malicious instructions embedded by an attacker. The model processes it all through the same mechanism.

That fundamental characteristic means defenses can't rely on the kind of clean separation that solved SQL injection. Microsoft's Learn documentation, updated March 24, acknowledges this reality with its recommendation for layered controls: input filtering, output validation, execution-time policy checks, least privilege, allowlists, and human-in-the-loop oversight for risky actions. Notice the plural—controls, not control.

Security experts have argued for monitoring actual agent actions rather than relying solely on intent detection. The logic makes sense. If you can't perfectly filter malicious instructions going in, watch what the agent actually does and block unauthorized actions before they cause damage. It's defense-in-depth translated into the language of AI security.

Perplexity's experience illustrates why single-layer defenses keep failing. The company published a methodology paper in late 2025 describing BrowseSafe, their approach to protecting browser agents. By January 19, security researchers had documented significant bypasses. Research published on arXiv in late April found that adaptation and attacker evolution consistently broke model-led defenses, particularly when those defenses relied on a single layer of input filtering.

OpenAI acknowledged as much in design guidance published March 11, noting that prompt injection may never be "fully solved" but can be mitigated through rapid response loops. It's the kind of measured language that doesn't fit neatly into startup pitch decks but probably reflects reality more accurately than claims of 96% effectiveness.

The Economics of Moving Fast and (Hopefully Not) Breaking Things

The market opportunity here is small but explosive. IT-Harvest, an analyst firm tracking the AI security space, pegged the "AI firewall" market at roughly $30 million in March 2026, with expectations for 100% growth over the year. That's driven M&A activity that suggests larger players recognize both the opportunity and the threat.

Check Point acquired Lakera in a deal that closed January 27, with the acquisition amount widely reported to be around $300 million based on industry sources. Palo Alto Networks reportedly paid north of $500 million for Protect AI in 2025, according to industry estimates. Those valuations reflect something beyond typical acqui-hire dynamics. Established cybersecurity vendors see enterprises deploying AI agents at scale and recognize that their existing products don't address the new attack surface.

The regulatory pressure compounds the urgency. The EU AI Act entered general application on August 2, 2026, with staggered high-risk system timelines already in force. A White House executive order issued June 2 directed DHS and CISA to publish binding operational directives within 30 days. NIST released a request for information in January specifically addressing AI agent security and indirect prompt injection.

For security leaders and engineering teams, this creates an uncomfortable dynamic. Your competitors are deploying AI agents to accelerate legal workflows, customer service, productivity tools—the Stanford data shows organizations using generative AI regularly are doing so across multiple functions. The temptation to prioritize speed over security becomes nearly irresistible when every quarter of delay potentially cedes ground to faster-moving rivals.

Perhaps that explains why 60% of organizations plan agent deployments within two years despite only 17% currently running them in production. The gap between those numbers represents thousands of architectural decisions being made right now, often by teams without deep expertise in adversarial machine learning.

What the NSA Thinks About Your AI Agent

Digital illustration for article section "What the NSA Thinks About Your AI Agent" in "Silmaril's Self-Healing Defense Blocks AI Prompt Injection Attacks" - A minimalist, conceptual illustration of a stylized, luminous key hovering just outside the keyhole ...

An NSA Cybersecurity Advisory released May 20 warned about security design considerations for AI-driven automation leveraging the Model Context Protocol and similar frameworks. The timing wasn't coincidental. MCP and comparable tools have made it trivial to give AI agents access to enterprise systems and data—connect a few APIs, write some configuration, and suddenly your agent can read Salesforce, manipulate spreadsheets, and query internal databases.

The advisory highlighted tool-response and memory poisoning vectors that create new attack surfaces. An attacker who can manipulate what an agent remembers from previous interactions, or poison the results returned by tools the agent calls, can effectively control the agent's behavior across subsequent conversations. Multi-turn attacks that build incrementally across sessions are particularly insidious because they don't trigger defenses looking for obvious malicious payloads in individual requests.

Google's April 23 security blog noted that current in-the-wild attacks remain mostly low-sophistication, but both scale and complexity are expected to climb. That assessment tracks with historical patterns—early-stage attacks tend to be crude because they don't need to be sophisticated. As defenses improve, attackers invest in more complex techniques.

Academic research published in May went darker. A paper titled "AI Agents May Always Fall for Prompt Injections" argued that data-instruction separation cannot serve as a complete solution, given how language models fundamentally operate. If that assessment holds up, it implies organizations will be managing prompt injection risk indefinitely rather than solving it and moving on.

The Arms Race Ahead

Digital illustration for article section "The Arms Race Ahead" in "Silmaril's Self-Healing Defense Blocks AI Prompt Injection Attacks" - A minimalist, conceptual illustration of a sturdy, abstract protective shield deflecting swift, styl...

Silmaril isn't operating in a vacuum. Arcjet announced edge-based prompt injection protection on March 18. Armorer Guard claims 3.4-millisecond inline defense as of May 21. The competitive dynamics will likely consolidate quickly, with larger security vendors either acquiring promising startups or building competing solutions in-house. The Check Point and Palo Alto acquisitions suggest that playbook is already in motion.

Latency matters more in this market than in traditional security tools because overhead compounds across multi-turn agent interactions. If every agent action adds 200 milliseconds of security checking, a complex workflow involving dozens of tool calls becomes noticeably sluggish. Silmaril's claimed 20-millisecond p90 latency falls between ultrafast solutions advertising single-digit milliseconds and heavier guardrail systems that prioritize comprehensiveness over speed.

Whether that positioning proves optimal depends on questions that won't be answered for months, maybe years. Do organizations prioritize lower false negatives (catching more attacks) or lower false positives (avoiding legitimate actions getting blocked)? How much latency can they tolerate? Does the self-healing approach actually adapt faster than attackers evolve their techniques?

The UK NCSC's joint guidance with Five Eyes partners in May emphasized starting with low-risk tasks, implementing strong governance, and layering mitigations. Sensible advice, assuming organizations have time for a measured approach. The Gartner data suggests many don't feel they have that luxury.

What Founders Should Actually Do

The most important takeaway—and this gets lost in vendor pitches about detection rates and benchmark scores—is that prompt injection defense cannot be bolted on as an afterthought. The attacks exploit fundamental aspects of how language models process instructions and data. Defenses need to be embedded at the agent framework level, with runtime monitoring of what actions agents actually take, not just what inputs they receive.

Organizations waiting for a silver-bullet solution that eliminates the problem entirely will find themselves vulnerable while competitors deploy layered defenses and ship securely, if imperfectly. The market will reward those who move fast but carefully, not those who move fastest.

Vendor benchmarks are only the starting point. Security teams need to pressure providers for transparent testing against evolving attack datasets, not static benchmarks that attackers will quickly learn to evade. Silmaril's self-healing concept is compelling precisely because it acknowledges this is an arms race, not a one-time engineering problem. But arms races reward participants who test their defenses against real adversaries, not marketing claims.

NIST's April 7 concept note on an AI Risk Management Framework profile for critical infrastructure signals that sector-specific requirements are coming. The White House executive order's 30-day directive to CISA means federal guidance on AI vulnerability detection and mitigation will arrive soon—likely before many organizations have thought seriously about agent security.

By the time agents handling financial transactions, legal documents, and enterprise workflows become ubiquitous, the organizations with mature prompt injection defenses already in place will have accumulated significant advantages. Those scrambling to retrofit security after a breach will face both remediation costs and reputation damage that compounds in an environment where AI competence is increasingly seen as competitive differentiation.

The window to prepare is open. Whether Silmaril's approach proves durable or becomes another cautionary tale in a market littered with overpromised security solutions remains to be seen. But the threat itself isn't going anywhere. And that, more than any startup's benchmark claims, is the story worth watching.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • AI Agents Replace Consultants: Ontora Maps Workflows in Hours, Not Months
  • Interfaze Merges DNNs with LLMs for Deterministic AI Tasks
  • AI Discovering AI: Inside Aster Lab's Autonomous Research Platform
  • YC-Backed Imperfect Launches AI Coach That Adapts Training to Real Life
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.