SOURCE-LINKED INTELLIGENCE
GHSA-8737-qx52-hjff: vLLM: Derender endpoints decode caller-supplied GenerateResponse token IDs without output bounds
## Summary The `/v1/completions/derender` and `/v1/chat/completions/derender` endpoints accept caller-supplied `GenerateResponse` objects and postprocess every nested `choices[*].token_ids` list directly. Unlike the normal render/generate path, derender does not enforce model context length, resolved `max_tokens`, `max_num_seqs`, choice-count, or response-size bounds before detokenizing and returning the supplied token IDs. An authenticated API client can therefore make the CPU-only render frontend, or any server exposing these `/v1` derender routes, spend CPU and memory proportional to attack
Read original source ↗ Open in workspace
- recordType
- vulnerability
- status
- active
- evidenceStatus
- reported
- region
- Global
Evidence & attribution
- OSV AI package advisories · 2026-09-04T21:32:07.000Z
- OSV AI package advisories · 2026-09-10T09:45:00.060Z
First collected: 2026-09-20T22:31:48.298Z. This is not the publication date.