CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
Multi-tenant LLM services sharing cached computations across users create data leakage risks that could expose confidential enterprise prompts to competitors.
Summary written by editorial AI · Source link below
arXiv:2605.23640v1 Announce Type: new Abstract: Large Language Models (LLMs) rely on Key-Value (KV) caching to accelerate inference, and many serving systems further share the KV cache across users' requests to reduce redundant computation. While widely adopted, unrestricted cross-user sharing introduces side-channel vulnerabilities, allowing an adversary to infer user inputs by probing for cache reuse. Existing defenses disable sharing entirely to prevent leakage; yet such a coarse-grained str
External link — opens at arXiv Crypto & Security in a new tab.
More from the Cloud Desk
- [NEU] [hoch] Microsoft Clouddienste: Mehrere Schwachstellen3d
- NACRE: Rethinking Confidential Containers through Native Architectural Support4d
- Incident response guide for AWS CloudTrail investigations – Part 24d
- Incident response guide for AWS CloudTrail investigations – Part 14d
- Reducio: Optimized Confidential Serverless Cloud Deployments for Enterprise Customers1 Sept