SOURCE-LINKED INTELLIGENCE
GHSA-8pw2-6jv3-mj5j: vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
## Summary Current vLLM `main` lets an inference request choose the PyNvVideoCodec GPU video decoder through `media_io_kwargs.video.video_backend`, but engine GPU memory reservation is computed only from static startup configuration and `VLLM_VIDEO_LOADER_BACKEND`. If the server starts with the default OpenCV/software backend and no `--mm-ipc-gpu-memory-gb` budget, a client can still route a video request into the PyNvVideoCodec path after startup, causing frontend CUDA-context, decoder-surface, and decoded-frame GPU allocations that were not carved out of the engine KV-cache budget. ## Techni
Read original source ↗ Open in workspace
- recordType
- vulnerability
- status
- active
- evidenceStatus
- reported
- region
- Global
Evidence & attribution
- OSV AI package advisories · 2026-09-17T17:17:39.000Z
First collected: 2026-09-20T22:31:48.298Z. This is not the publication date.