← Back to the latest

Research resultSep 14, 2026Entry № 417

CAITLYN: a self-evolving defense middleware against prompt injection for LLM agents


CAITLYN, a self-evolving defense middleware for LLM agents, has been released with its code and paper. It targets prompt injection attacks arriving through web text, search results, local files and tool returns. The system pairs a fast runtime layer (script, signature and heuristic filters plus a Tier-1 LLM classifier) with a second, counterexample-driven layer that turns missed attacks into new defense skills and writes them back to a reusable skill library.

On AgentDojo-S250, ASPI-S, SafeClawBench-S240 and an Emerging setting focused on adapting to new attacks, the CAITLYN-evolved system cut the success rate of emerging attacks by roughly 40 percentage points while keeping false positives low.

Original sources (Chinese)

网页藏毒,劫持Agent?你需要会「追凶」的自进化女警系统aiera