When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
New attack poisons AI agent skill libraries by injecting malicious trajectories that get promoted to trusted instructions — a supply-chain threat for agentic AI deployments.
Summary written by editorial AI · Source link below
arXiv:2608.05563v1 Announce Type: new Abstract: Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion process. Our skill-visible black-box attacker can inspect a target skill and contribute bounded evidence, but cannot observe private pools or evolution logic or edit the skill bank. Artifact poisoning requires Inclusion, Evol
Editorial Analysis
Enterprises deploying self-learning AI agents face a novel integrity risk: adversaries can corrupt the experience-to-skill pipeline, turning learned behaviours into persistent backdoors.
Audit any agentic AI system that promotes operational experience into reusable skills; enforce cryptographic integrity and human review gates before skill adoption.
Self-learning AI agents can be compromised through poisoned experience data, creating persistent backdoors that evade traditional security controls.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the AI Security Desk
- OpenAI admits it didn't disclose rogue AI wiki hijacking incident2d
- Thousands of OpenAI Agents Quietly Turned an Abandoned Wiki Into Their Coordination Channel3d
- Using a VM to Contain an AI Agent3d
- Companies Have 6 Months to Prepare for Automated Attacks3d
- [NEU] [mittel] Ollama: Schwachstelle ermöglicht Offenlegung von Informationen3d