On the afternoon of June 18, at Y Combinator's Demo Day, a two-person team named Silmaril made a pitch that would feel audacious if the problem weren't so widely acknowledged: they believe they've cracked prompt injection, the vulnerability that security researchers and intelligence agencies have deemed perhaps the hardest challenge in artificial intelligence defense.
TechCrunch listed them among the batch's standout companies. The pitch? A defense system that learns from attackers—getting stronger, not weaker, as adversaries evolve their tactics.
It's the kind of claim that raises eyebrows. Plenty of vendors promise adaptive security. Few deliver. But the timing of Silmaril's emergence isn't random. Enterprises are betting heavily on AI agents—autonomous systems that can book travel, analyze spreadsheets, or orchestrate complex workflows without constant human oversight. Gartner estimated in May that global AI spending would surge to $2.59 trillion in 2026, a 47% jump from the prior year. Yet according to Forrester's June 2026 data, nearly half of security decision-makers flag these same autonomous systems as a top concern, even as three-quarters of enterprise leaders say they're already deploying them.
That gap—between ambition and safeguards—is where startups like Silmaril see an opening.
The Vulnerability That Won't Go Away
OWASP, the nonprofit that tracks web application security risks, ranked prompt injection as LLM01 in its Top 10 for Large Language Model Applications in its 2025 edition. Number one. The UK's National Cyber Security Centre was blunter still when it issued guidance late last year: prompt injection "may never be totally mitigated." The reason is architectural. Unlike SQL injection, which can be contained through proper input sanitization, large language models lack an inherent boundary between instructions and data. Tell an LLM to "ignore previous instructions and do this instead," and you're exploiting a design characteristic, not a bug.
The threat landscape backs up the alarm. In April, researchers disclosed "GrafanaGhost," an indirect prompt injection attack against Grafana's AI features that exfiltrated sensitive data without triggering conventional security logs, CyberScoop reported on April 7. Months earlier, academic work documented in September 2025 revealed "EchoLeak" (CVE-2025-32711)—a zero-click exploit path in Microsoft 365 Copilot, triggered by specially crafted emails. Google's security team noted in spring reports that they're seeing indirect prompt injection payloads in the wild, targeting AI agents that browse the web or ingest external documents.
When CISA and the Five Eyes intelligence alliance released guidance on agentic AI adoption at the end of April, they didn't mince words. Prompt injection remains an unresolved threat. Their advice: defense-in-depth controls, least privilege access, human-in-the-loop checkpoints, and comprehensive audit trails. In other words, layer every safeguard you can think of and hope it's enough.
Regulation Moves Faster Than Solutions

Europe's AI Act obligations for general-purpose AI providers took effect on August 2, 2025, with phased compliance requirements running through mid-2027. IDC data from March shows European chief information security officers are already earmarking nearly 17% of AI budgets specifically for agent security and governance—a nontrivial chunk of spend driven by regulatory pressure as much as genuine risk mitigation.
OWASP responded with a dedicated "Top 10 for Agentic Applications" framework in December, cataloging risks unique to autonomous systems: tool misuse, memory poisoning, vulnerabilities in how agents talk to each other. MITRE's ATLAS threat matrix now tracks prompt injection (AML.T0051) with technique filters tailored to agent scenarios. On May 20, the NSA's AI Security Center issued design considerations for systems using the Model Context Protocol, warning that MCP adoption in development environments and agents expands attack surfaces through tool poisoning and cross-origin hijacking.
The infrastructure giants have taken notice. Amazon announced on June 16 that Bedrock Guardrails added new APIs targeting agentic workflows—standalone checks for jailbreak attempts, prompt injection, and prompt leakage. Google detailed what it called a "continuous mitigation program" against indirect prompt injection in early April, emphasizing layered defenses backed by automated red teaming and bug bounty programs.
But there's a catch. Research published on arXiv in June suggests that detection-only filters degrade when attackers paraphrase or slightly evolve their techniques. Pattern matching, in other words, has limits against adaptive adversaries. The implication: you need system-level architectural controls, not just smarter regex.
Self-Healing, Or So They Say

Silmaril describes itself as "the world's first self-improving prompt injection defense," according to its company website. The architecture, as the founders explain it, has three layers. First, autonomous threat-hunting agents that generate verified exploits—essentially, internal red teams running continuously. Second, a low-latency firewall classifier operating at around 20 milliseconds at the 90th percentile. Third, a retraining loop that deploys updated defenses in under an hour from the moment a new exploit is discovered.
The company's Y Combinator profile claims the system blocks twice as many threats as current state-of-the-art solutions at ten times lower latency, though this claim lacks third-party verification. On its website, Silmaril reports benchmark accuracy of 95.6%, compared to Lakera Guard at 74.3%, Perplexity BrowseSafe at 79.2%, GPT Safeguard at 64.8%, and ModelArmor at 61.4%. Worth noting: these are vendor-run tests. No independent verification yet.
Integration, at least in theory, is straightforward. The company says it plugs into agent frameworks like LangGraph with five lines of code, available as both SaaS and self-hosted deployments. The differentiator, according to Silmaril, is context- and outcome-aware filtering—intercepting harmful executions like unauthorized tool calls or data exfiltration paths, not just scanning input text for suspicious patterns.
The founding team brings security credentials, if their bios are accurate. CEO Aum Upadhyay claims to have built security and privacy frameworks at AWS that prevented over $1.8 billion in damages. CTO Eduardo Velasco, formerly a tech lead at Amazon, says he "chained a prompt injection into root access inside ChatGPT." Neither claim appears verified in the independent press coverage available during research for this piece.
Case studies on Silmaril's site are anonymized, which is common but limits verification. One references preventing $68 million in damages for the "#1 AI-native productivity app." Another notes "Microsoft patched" in relation to a Copilot SSRF data exfiltration vulnerability via email-based prompt injection. The YC profile asserts the company has "stopped $28 million of damages for customers"—again, vendor-reported figures without third-party corroboration. Perhaps that's inevitable for a startup this young, but it makes independent assessment difficult.
A Crowded Field, but Gaps Remain

Silmaril isn't entering a vacuum. Lakera Guard, with documentation updated earlier this year, offers prompt-attack detection aimed at SOC2-compliant operations with use-case tuning. HiddenLayer provides runtime LLM security aligned to OWASP and ATLAS frameworks, with research and product updates as recent as late January. Microsoft's Prompt Shields, launched in 2024 for jailbreak and indirect prompt injection defense in Azure AI Content Safety, remains deployed alongside PyRIT, an open-source red-teaming framework. Superagent, a Y Combinator Winter 2024 company, markets an "AI Firewall" and agent red-teaming service.
What these solutions share—and what Silmaril claims to transcend—is reliance on pattern recognition or rule-based filtering without continuous adversarial feedback loops. The concept of autonomous threat-hunting feeding directly into production defenses represents a shift, assuming the execution matches the pitch.
The broader security market has room for multiple approaches. IDC projects global security spending will exceed $300 billion this year and reach $430 billion by 2029, per a March 20 report highlighting consolidation toward AI-driven platforms. Forrester data from late last year noted that a quarter of planned AI spend was being deferred into 2027 due to governance and security pressures—a dynamic that should favor vendors who can demonstrate measurable risk reduction, not just sell on fear.
Open Questions
As of late June, Silmaril has disclosed no public funding beyond Y Combinator participation. The team size remains two, according to the YC company page. Whether the startup can scale operationally while maintaining the sub-hour retraining cadence it advertises will be an early stress test. Hiring engineers who understand both adversarial machine learning and production reliability isn't easy, and two people can only stretch so far.
The industry consensus—from the NCSC, OpenAI disclosures late last year, and Google's security blogs this year—is that prompt injection won't be solved at the model layer. Defense-in-depth plus architectural governance is table stakes. The value proposition for a product like Silmaril hinges on whether context-aware, continuously adapting filters can shrink the window between detection and deployment faster than attackers evolve their payloads, all while fitting inside the latency budgets of real-time agent workflows. That's not a trivial engineering challenge.
OWASP's frameworks have become procurement checklists. The Five Eyes guidance from late spring is accelerating enterprise requirements for OWASP and ATLAS mappings, audit trails, and least-privilege tooling. Vendors that publish formal control mappings and demonstrate compliance with these emerging standards will likely have an edge in vendor risk reviews through the rest of this year and into next.
Silmaril's appearance on TechCrunch's Demo Day standout list signals investor appetite for AI security infrastructure. There's clearly demand. The question is execution. Autonomous red-teaming and self-healing defenses sound compelling in a pitch deck—maybe even inevitable, given the trajectory of AI deployment. But enterprises will demand proof. Independent benchmarks. Customer references that aren't anonymized. Evidence that the system performs under the kind of adaptive, multi-turn attacks that academic research shows can evade static filters.
For now, the startup has carved out a narrative in a market that desperately needs better answers. Whether it can deliver them—well, that's the story to watch.
