vLLM 0.31 Adds Draft-Model Speculative Decoding to Model Runner V2, a Preload CLI for Fast Restarts, and Gates Per-Request Multimodal Kwargs
vLLM 0.31.0 brings draft-model speculative decoding to Model Runner V2, adds a vllm preload CLI and a max-num-active-seqs cap, and removes tokenizer_mode="slow".