Code and data used in the paper: AI Agents May Always Fall for Prompt Injections
-
Updated
May 19, 2026 - Python
Code and data used in the paper: AI Agents May Always Fall for Prompt Injections
Leakage-verified contextual-integrity benchmark: does an agent disclose a sensitive attribute to a recipient the context forbids. Defensive, synthetic-only. Sibling of leakgauge.
Generative model of Contextual Integrity (CI) in multi-actor LLM systems: scenarios from CI taxonomies, generation pipeline, multi-agent simulation, and a two-sided appropriateness metric. No actor-level lever closes the protection/utility gap.
Research notes and surveys on privacy-preserving memory for AI agents: machine unlearning, contextual integrity, differential privacy, and memory-architecture defenses.
Mechanistic study of contextual-integrity post-training in Qwen2.5-7B, testing whether improved privacy behavior comes from new mechanisms or better use of machinery already present in the base model.
How much private information leaks incidentally between LLM agents as a function of communication topology — an evaluation built on Terrarium.
To associate your repository with the contextual-integrity topic, visit your repo's landing page and select "manage topics."