## vllm-project/vllm — v0.31.0…v0.31.1rc0

_466+ commits._

### Features
- **[Model] Support EmbeddingGemma2 multimodal pooling architecture (#60254)** (02b8391)
- **[Docs] Add model recipes skill (#57786)** (1a86733)
- **[vllm-bench, feature] Added support for Mooncake-style, timed-traces replay. (#55937)** (3faf214)
- **[ROCM][Perf] Add ROCM_SEGMENTED_ATTN attention-backend works across RDNA3/3.5/4 (#59132)** (edfccc9)
- **[Profiler] Add per-session profiling controls to Python and Rust frontends (#57875)** (6fddc9b)
- **[ZenCPU][Test] Add zentorch/torch version pin consistency test (#52888)** (533d81b)
- **[Docs] Add reviewer area of interest for TheEpicDolphin (#60092)** (88f3c60)
- **[Model] Add Nemotron 3.5 ASR transcription support (#59827)** (64cb683)
- **[CI] Add CODEOWNERS for HiSparse (#60060)** (9e931a0)
- **[Model][GLM-5.3-Flash] Support DCP for the kpool sparse indexer (#59211)** (30d4032)
- **[Attention][MLA] Support fp8_ds_mla KV cache for NoPE-512 models on SM90 (#59246)** (ff53f32)
- **[Performance] Add SM121 TP=2 skinny-GEMM plans (#59632)** (c04c79e)
- **feat: Add video support for the Transformers backend (#57441)** (34051ad)
- **[Bugfix][Attention] Support int4_per_token_head on non-power-of-two head sizes (#56198)** (e8a53a5)
- **[Kernel][LoRA] Add deterministic split-K=8 LoRA shrink for batch invariance (#59377)** (1388100)
- **[Frontend] Add output_mode to /inference/v1/generate (RFC #56851 Phase 1) (#58588)** (bc21cba)
- **[ROCm][Kimi-K3] Add opt-in gfx942 MXFP4-to-int4 conversion (#51274)** (b0e21b3)
- **[ROCm][CI][AiterExperts] Add test coverage for hidden/intermediate padding correctness (#59333)** (faa9860)
- **[EPLB] Add contention-aware expert migration batching (#52641)** (7dfe333)
- **[Docs] Add ERNIE 4.5 to batch invariance tested models (#58476)** (6e517b1)
- **[Perf][MoE] Support fp8 combine in FlashInfer one-sided MoE all2all (#57995)** (8459395)

### Fixes
- **[Bugfix] Persist FlashInfer autotune cache per rank to fix the rank-0-only cache-hit deadlock (#57635)** (3403e0f)
- **Fix grammar in `BlockPool.reset_prefix_cache` docstring: "to invalid" → "to invalidate" (#59886)** (68088ed)
- **fix: exclude flashinfer-jit-cache from wheel install_requires (#60052)** (6ad2e06)
- **[CI/Build] Fix incorrect CLI arg validation tests in test_cli_args.py (#46779)** (1e5d0ea)
- **[Bugfix][Model] Fix Qwen3-Omni interleaved M-RoPE boundary positions (#59842)** (8c417a2)
- **[Bugfix][KVConnector][NIXL] Fix TP handshake test after per-region FA desc flags (#60157)** (48eb44b)
- **[Bugfix] Fix Qwen4Exp FP8 PLE pinned lookup kernel failing to compile below SM89 (#60142)** (b0b23a4)
- **[TEST][XPU][CI] Fix ci uva kernel test instability (#60041)** (4e198ce)
- **[CI] Fix test_mixed_warmup_gate after #57053 (#60116)** (8542c3e)
- **[Bugfix] Fix Mamba page size AssertionError with spec decoding on GraniteMoeHybrid, FalconH1 and Zamba2 (#59975)** (d0d6e5f)
- **[Bugfix] Fix Qwen4Exp PLE embedding rejecting INC (AutoRound) checkpoints (#59990)** (4f52fa3)
- **[CI][Docs] Fix decode/prefill consistency test prefix construction; add GLM-4-9B and Phi-4 to batch-invariant tested models (#53692)** (89439db)
- **[ROCm][Bugfix] Fix ROCM_ATTN sliding-window boundary (#59550)** (f98d1fc)
- **[Bugfix] Fix FlashInfer all_reduce backend selection (#56891)** (ce49174)
- **[ROCm][RDNA3] Fix W4A16 split-K accuracy and determinism (#54706)** (4ac0d0e)
- **[Bugfix][Model] Fix M-RoPE offset double-count in Qwen3-Omni (#58890)** (44198f5)

### Backend
- **[Metrics] Expose cached prompt tokens by cache tier (#56318)** (e37e51d)
- **[Bugfix][Spec Decode] Keep heterogeneous-vocab draft models on Model Runner V1 (#59541)** (ad0f67a)
- **[Frontend] Integer token IDs for generate output logprobs (GenerateLogProbs) (#58181)** (7b665f7)
- **[Perf][HiSparse] Cache per-request residency instead of rescanning every page each step (#60083)** (e82b800)
- **[GLM5.3 Perf] Switch to fp8 kv cache by default for GLM, 2.3%~5.5% E2E Throughput Improvement (#60140)** (fba31f3)
- **[Quantization] Layer re-quantization for linear layers through online quantization API (MXFP8 -> FP8 PTPC showcase) (#55684)** (f8d92ca)
- **[ROCm] Fuse MLA dual RMSNorm + FP8 group quant for DeepSeek-R1 (#54857)** (29f955c)
- **[CI/Build] Warm the DP engines before measuring load balance in test_load (#60241)** (4516896)
- **[Spec Decode] Remove eager metadata rebuild during MTP fused multi-step decode (#58463)** (fcf53d2)
- **[Bugfix] Warn only on per-request do_normalize/do_rescale overrides (#60240)** (4ea0c28)
- **Upgrade tpu-inference to v0.31.0 (#60190)** (aad4579)
- **[Build] Migrate vendored DeepGEMM from pybind to TORCH_LIBRARY (abi3) (#48962)** (cc73cca)
- **[NIXL] Bump NIXL version to 1.5.0 (#56907)** (cbe1f97)
- **[CI][Rust Frontend] Pass `--locked` to cargo binstall (#60182)** (0719975)
- **[Docs] Clarify that vLLM does not isolate tenants sharing a server (#59811)** (8bd7737)
- **[Perf][MoE] Eliminate staging from NCCL symmetric reduce-scatter (#49194)** (4a99ca4)
- **[Docs] Run pr-checklist before opening a PR (#60235)** (bf049dd)
- **[ROCm][Perf] W4A16: pad gfx11 weight and activation strides (#56301)** (b267fe4)
- **[Bugfix][Structured Output] Flag multi-branch allOf as unsupported for xgrammar (#59061)** (049507a)
- **[Bugfix][KV Offload] Honor speculative cacheability in SimpleCPU (#60071)** (2538510)
- **[Deprecation] Deprecate sonet dataset as scheduled (#60129)** (3e18218)
- **[Bugfix][DiffusionGemma] Read the causal mask as int32 and compile the sample step (#59992)** (3313825)
- **[Bugfix][Frontend] Apply model-default reasoning parser in GPU-less render server (#54835)** (09e0ce3)
- **[CPU][Recipes] Detect model head constraints for automatic tensor parallel selection in vLLM Recipes Tool (#60170)** (888074b)
- **[Rust Frontend] Pass tool defer_loading through to chat templates (#60203)** (6bbad6a)
- **[Model] LongCat-Flash: scale MLA norms while loading and drop the post-load sweep (#60080)** (31e2443)
- **[Feat][Mamba2] Enable internal prefill checkpoints (#57329)** (e11962f)
- **[Bugfix][EPLB] Use group-local ranks for torch P2P transfers (#55804)** (15c1cc1)
- **[Kernel] Enable OAI Triton MXFP4 MoE on SM12x (#58877)** (40c86ae)
- **[Rust Frontend] Remove the PyO3 tool-parser bridge (#59744)** (b3082d3)
- **[Frontend] Port the MiniMax M3 tool parser to the parser engine (#59743)** (5145e35)
- **[Model][DeepSeek-V4.1] Permute mega-attention wq_b / wo_a while loading (#60064)** (9a1f6fd)
- **[Perf][DSv4.1] Run decoder replay layers in CUDA graphs and trim in PIECEWISE graphs (#59532)** (eb2901b)
- **[Attention][Model] FP8 KV cache for Triton DiffKV and MiMo-V2.6-Flash (#58128)** (21d93d0)
- **[Bugfix][Attention] Restrict FlashInfer sparse MLA FULL graphs to decode (#59751)** (30ab428)
- **[Bugfix][Parser] Mistral pre-v11: don't fail requests on unexpected tool call JSON (#54844)** (a3e0243)
- **[Bugfix][Core] Retire stale Mamba spec blocks in prefill checkpoint steps (#59759)** (6fa67af)
- **[Bugfix][Core] Drop stale block hashes when a session truncates tokens (#59103)** (2b32d9a)
- **[Bugfix][NIXL] Recover locally invalidated pull peer metadata (#55471)** (9465e1c)
- **[Bugfix][Frontend] Malformed `$defs` in a tool schema should be a 400, not a 500 (#54850)** (9bf4e51)
- **[ROCm] Disable AITER attention query quantization on RDNA3 (#57750)** (efd0017)
- **[Bugfix][KVConnector][NIXL] Size per-region replicate flags by each region's block count (#60107)** (2537629)
- **[Bugfix][KVConnector][NIXL] Return NumPy descriptors from mixed-memory local registration (#60108)** (0e8f8ec)
- **[ROCm][DSv4.1][Perf] One-pass dequantize and a gather-sized grid for the compressed K cache (#56720)** (97bd8c7)
- **[Kimi-K3] Use the fused K/V pack kernel on the decode-context-parallel prefill path (#56881)** (b2863c0)
- **[Bugfix][Structured Output] Cap JSON schema nesting to prevent 500s and API-server stalls (#60036)** (5cd8a26)
- **[CI][ROCm] Wait for engine teardown between LM Eval models (#60100)** (fa423b6)
- **[Bugfix][MoE] Allow MiniMax2 routing in TRTLLM BF16 monolithic backend (#59571)** (6dd813e)
- **[Bugfix] Route every speculation-capable row through the speculative path so recurrent state stays consistent (#56531)** (0eac152)
- **[Bugfix][Frontend] Preserve logprobs when top_logprobs is null (#47838)** (4a705c7)
- **[CI] Read each profiler round's trace as it stops in test_gpu_profiler (#59769)** (89a0099)
- **[Bugfix][Frontend] Build tool-call grammars from the tools the prompt renders (#59879)** (2e06395)
- **[Frontend] Extract parse and assembly out of batch chat derender (#60073)** (3fb1cd7)
- **[HiSparse] Log steady-state max concurrency and expose host-tier utilization gauges (#58949)** (2e3154a)
- **[BugFix][Frontend] Pass reasoning_ended through render → generate (#60059) (#60062)** (18c8a65)
- **[Bugfix][MRV2][Spec Decode] Honor dynamic K in autoregressive speculators, 1.2~1.3x kernel perf improvement (#57053)** (74c5cbc)
- **[watermark] compatibility validation (#56801)** (c32495d)
- **[Bugfix] Give each data-parallel engine its own global RNG streams (#59788)** (d09ff77)
- **[Bugfix][Spec Decode] Reserve the bonus KV slot for fill-in DSpark (#59105)** (877ddcc)
- **[ROCm][Perf] Use a stacked hipBLASLt GEMM for the bf16x3 router  (#52668)** (1822472)
- **[Bugfix][HiSparse] Reject cudagraph_mode=FULL at startup (#59688)** (54d93af)
- **[CI/Build] Declare the video modality on the CohereCompass reference model (#60063)** (87954c5)
- **[ROCm][Perf] Reach the fused QSA pre-indexer from the AMD path (#57947)** (d23af12)
- **[Bugfix][Core] Retain encoder cache references for repeated multimodal inputs (#59942)** (528772a)
- **Remove redundant dependency requirement for TPU (#59977)** (edde9d2)
- **[CI/Build] Run the prefill token scoring test under batch invariance (#60056)** (f2d8fbf)
- **[Bugfix][Structured Output] Preserve literal values in Guidance disable_additional_properties (#58709)** (60932a6)
- **[LoRA] Remove tensorizer (#60024)** (f9c9e8a)
- **[Test] Re-enable Voxtral HF reference test on Transformers v5 (#59771)** (d3547f9)
- **[Model] Extend device-side mm normalization to Kimi K2.5 / K3 (#59278)** (0e468ad)
- **[Bugfix] Load stacked expert weights for non-gated MoE (#59031)** (51eeb0c)
- **[Bugfix][Qwen4Exp] Honor --kv-cache-dtype-skip-layers in QSA attention (#60023)** (4ff028d)
- **[Perf][Qwen4Exp] Keep the HC up projection on the skinny GEMM path (#60027)** (55b8022)
- **[Security] Restrict Pillow image formats on untrusted media paths (#60022)** (710ac56)
- **[Bugfix][Qwen4Exp] Load PLE tables unquantized under Quark checkpoints (#59443)** (edca360)
- **[Bugfix][NIXL] Count heartbeats as remote engine activity (#59873)** (b1401e0)
- **[Bugfix][GLM-5.3-Flash] Keep batch x heads out of gridDim.z in the GLM-5.3-Flash fused recurrent KDA kernel (#56974)** (ae53b06)
- **[Model] Declare SupportsEagle3 on Qwen3ASRForConditionalGeneration (#52824)** (b1f229f)
- **[LoRA] Code cleanup (#60017)** (e3af5bf)
- **[Model][DeepSeek-V4] Make MegaMoE shared-expert finalize independent of linear post-load order (#59927)** (4a30c4c)
- **[Perf][Qwen4Exp] Merge QSA QKVG and indexer QK projections (#59533)** (042ab03)
- **[Bugfix][ROCm] Preserve config during GPU memory profiling (#58014)** (9204424)
- **[ROCm] Bump AITER to 0.1.24.post1 (#59794)** (ac8c213)
- **[Bugfix][KV Offload] Reset SimpleCPU eager-store placement state on MRV2 resume (#57816)** (0c16eee)
- **[Bugfix][DSv4.1] Keep the compressor ring out of the null block (#58560)** (49e0f47)
- **[Frontend] Name the served models in the model-not-found 404 (#59889)** (7867d6c)
- **[Bugfix][Determinism] Preserve output dtype in batch-invariant mean (#59106)** (155488d)
- **[Bugfix][Frontend] Return streaming errors before the first token (#40986)** (b0eb87f)
- **[Misc][Platform] Check aligned KV block sizes against every attention backend (#58457)** (bd42276)
- **[CI] Stop requiring both API servers to record weight sync metrics (#59521)** (bdd31c3)
- **[CI] Auto-label pooling PRs and issues (#59913)** (d64f6cd)
- **[CI] Bump peft to satisfy transformers 5.18 minimum (#59932)** (d61081d)
- **[Bugfix][XPU] Make DeepSeek V4 FP8 sparse decode graph-capturable (#59159)** (155d23c)
- **[Bugfix][Multimodal] Normalize single-channel audio to 1D (#56691)** (0872ddf)
- **[Bugfix][KV Offload] Release pending CPU lookup pins on cache reset (#59862)** (18f8f96)
- **[Bugfix][Responses API] Reuse streamed item ids in final harmony response (#59859)** (6f75cc7)
- **[Bugfix][Frontend] Preserve abort finish_reason for scale-out token streams (#47933)** (50b404e)
- **[Feature] Release the CUDA graph pool on sleep (#59160)** (5e56e9a)
- **[CI] Bump Transformers version to 5.18.0 (#59621)** (9e0b6fc)
- **[DCP] Enable TokenSpeed MLA with block-interleaved DCP (#59462)** (84bcbc6)
- **[Bugfix][MoRIIO] Keep discovery heartbeats running while workers hold the GIL (#59441)** (f03026a)
- **[Feature] Per-row candidate IDs for prefill token scoring (M2 of #56860) (#56984)** (e319f86)
- **[Bugfix][KV Connector][Mooncake] Suppress completion for empty pulls (#59347)** (5f30fc7)
- **[Bugfix][Core] Exempt exactly the blocks an async KV load writes from zeroing (#59504)** (ca0df1f)
- **[CI] Drop duplicate bf16 skinny GEMM test from Kimi K3 B200 job (#59850)** (1a001d5)
- **[ROCm][CI] Extend AMD coverage for distributed, model, and eval tests (#56679)** (c28834d)
- **[Bugfix] Log CRIU failure details before snapshot cleanup (#59661)** (28556f8)
- **[Bugfix] Avoid InfiniBand state in TP1 snapshots (#59699)** (9fdb174)
- **[TEST][XPU][CI] disable xpu tests for nonexistent input norm kernels (#59568)** (8c3b720)
- **[CI/Build][NVIDIA] Build Rubin images on the public nvidia/cuda base image (#59288)** (8e44387)
- **[GLM 5.3 Perf] Enable fused multi-step decode, 13.3% E2E throughput improvement for concurrency 1 (#57443)** (8bdd8d8)
- **[feat] FlashInfer CuteDSL MegaMoE integration  (#54049)** (111f71a)
- **[Bugfix][Frontend] Avoid generation for empty streaming input (#59015)** (c4973d9)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/vllm-project/vllm?utm_source=github-action)._