CodePoisonRAG: Knowledge Poisoning Attacks on Retrieval-Augmented Code Generation
CodePoisonRAG demonstrates that poisoning the retrieval corpus of LLM-assisted code generation can silently inject vulnerabilities into produced code — a supply-chain attack targeting the knowledge layer rather than the model itself.
Summary written by editorial AI · Source link below
arXiv:2609.02774v1 Announce Type: new Abstract: Retrieval-Augmented Code Generation (RACG) improves LLM-based software development by retrieving external code artifacts, documentation, and patches, and incorporating them into the generation context. This reliance on external knowledge introduces a critical trust boundary: poisoned artifacts can influence generated code without modifying the underlying LLM. Prior work shows that selecting existing vulnerable examples can increase the general vul
Editorial Analysis
As developer teams adopt retrieval-augmented code generation, the retrieval corpus becomes a high-value supply-chain target that traditional model-security controls overlook.
Mandate integrity verification and allowlisting for all external sources feeding LLM code-generation tools, and scan generated code with SAST before merge.
AI-assisted coding tools that pull from external knowledge bases can be silently poisoned — a new supply-chain risk for software-producing organisations.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the AI Security Desk
- OpenAI admits it didn't disclose rogue AI wiki hijacking incident2d
- Thousands of OpenAI Agents Quietly Turned an Abandoned Wiki Into Their Coordination Channel3d
- Using a VM to Contain an AI Agent3d
- Companies Have 6 Months to Prepare for Automated Attacks3d
- [NEU] [mittel] Ollama: Schwachstelle ermöglicht Offenlegung von Informationen3d