vLLM 0.30 Adds a Persistent GPU Weight Cache for Faster Engine Restarts and Removes GPTQ Activation Ordering
vLLM 0.30.0 adds a GPU-resident weight cache for faster engine restarts and removes GPTQ g_idx; llm-d v0.10.0 then moved its base image to 0.30.0.
Overview
The vLLM project published version 0.30.0 of its open-source LLM inference engine on September 22, 2026, according to the vLLM GitHub release page. The release notes say it contains “762 commits from 315 contributors (104 new).” Its headline infrastructure feature, called Fast Start, keeps model weights resident in GPU memory so that restarting an engine does not require reloading them from disk. A week later, the llm-d project moved its bundled vLLM base image to 0.30.0 in its own v0.10.0 release.
What Fast Start Does
Per the vLLM release notes, Fast Start is a persistent per-GPU weight-cache daemon that holds post-quantized, TP-sharded weights in GPU memory. Restarting engines map those weights over CUDA IPC using the --load-format ipc_cache option instead of reloading from disk. The notes say the mechanism now covers FP4 checkpoints and multi-node TP.
A separate change targets engine startup time. The release notes state that freezing garbage collection during CUDA graph capture cuts capture from 12 seconds to 2 seconds, and engine initialization from 28.9 seconds to 8.2 seconds on an H200 GPU.
Other Serving Changes
- HiSparse: a host-resident tier for sparse-MLA decode that, according to the release notes, spills KV pages to pinned host memory under GPU pressure and serves top-k misses from a per-request GPU hot buffer. It is enabled through
HiSparseConnector. - Watermarking: the release adds Gumbel-max watermarked generation and detection with a keyed PRF, per-request opt-out and an example detection endpoint. A dual-key variant is described as compatible with speculative decoding.
- Models: new model support listed in the release notes includes DeepSeek-V4.1-Flash, GLM-5.3-Flash, K2-Horizon and Cohere Compass.
- Quantization: the notes list AutoRound support for 2/3/5/6/7-bit on CUDA.
Breaking Changes
The vLLM release notes list several changes operators may need to act on:
- Scale-out endpoints are opt-in on plain
vllm servethrough--enable-scale-out, replacing theVLLM_ENABLE_SCALE_OUT_ENDPOINTSsetting. - GPTQ activation ordering (
g_idx) was removed. - The
allMamba cache mode was deprecated. python -m vllm.entrypoints.grpc_serverwas deprecated in favor ofvllm serve --grpc.- YaRN was aligned with Transformers, so vendor YaRN aliases no longer re-scale
max_model_len.
Downstream: llm-d v0.10.0
The llm-d project, which The Machine Herald previously reported was donated to the CNCF, published v0.10.0 on September 29, 2026. Its release notes name two themes, the first being “Operational hardening” aimed at making the production path “safe and boring: rollouts, HA, and failure behavior that operators can trust.” A component table in the notes lists the vLLM base image moving from v0.26.0 to v0.30.0.
The same notes show llm-d shifting maintenance of inference images upstream. They state that ghcr.io/llm-d/llm-d-cuda image builds are deprecated in favor of docker.io/vllm/vllm-openai, and that ghcr.io/llm-d/llm-d-aws builds are deprecated in favor of public.ecr.aws/deep-learning-containers/vllm. The llmd-fs-connector was integrated upstream into vLLM under the OffloadingConnector and is deprecated as a separate component.
The release also changes its container registry arrangement. Images built from code merged to main or from official releases go to ghcr.io and are signed with cosign, while images built from open pull requests go to quay.io, where, per the notes, “images will intentionally NOT BE SIGNED.” The notes add that the repository llm-d/llm-d-kv-cache has been migrated to llm-d/llm-d-router and that llm-d/llm-d-workload-variant-autoscaler was renamed llm-d/llm-d-autoscaling.
What We Don’t Know
The release notes for both projects do not publish benchmarks for Fast Start’s restart improvement beyond the H200 engine-initialization figures quoted above, so real-world gains on other hardware and model sizes are not documented in the sources reviewed. Neither set of notes states how many production deployments use the new features.