CyberRota Analysis
AI-GeneratedvLLM versions up to 0.29.0 are vulnerable due to improper cleanup of decode-side metadata for rejected inference requests, allowing remote attackers to exploit this flaw by submitting requests with max_tokens set to 0. This can lead to memory exhaustion of decode workers, potentially causing service disruptions until the worker is restarted. Organizations utilizing vLLM in their deployments should prioritize patching this vulnerability to mitigate the risk of denial-of-service attacks.
Public Exploit Signal
A public exploit, PoC, GitHub repository or Metasploit reference was detected for this CVE.
Note: these links are listed for security research and verification purposes only.
Original NVD Description
vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.
Related CVEs
Other vulnerabilities affecting the same vendor(s)