## vllm-project/vllm — v0.31.0rc2…v0.31.0rc3

_16 commits._

### Features
- **[Model Runner V2] Support randomized dummy inputs (#58411)** (f426292)
- **[CPU][Zen] Add DA8W4 (W4A8) int4 support for dense and MoE layers (#54024)** (443ad8e)
- **[CPU][Whisper] Support W4A16 quantized Whisper on the CPU WNA16 kernel (#58268)** (3700ae3)

### Fixes
- **[Bugfix][Core] Fix mamba prefill checkpoint block reservation and prompt-end eviction in align mode (#59175)** (2e3e990)

### Backend
- **[Perf][PP] Skip sampled-token broadcasts whose requests leave the engine (#58542)** (024c957)
- **[Core] Include the LoRA path in prefix-cache block hashes (#59335)** (41b8365)
- **[Bugfix][HiSparse] Stop the host pool feeding device KV cache residency metrics (#58725)** (1ea61d3)
- **[Bugfix][HiSparse] Adopt GPU prefix copies after the hit's allocation (#59282)** (f9f8c13)
- **[Bugfix][HiSparse] Preserve host prefix publication after request completion (#59007)** (65182ef)
- **[KV-Offloading][TP] : Expand replicated_layout detection to multi-group MLA  (#57652)** (a5738bf)
- **[Core] Bound UniProc EngineCore startup threads to available CPUs (#58946)** (cc28ed1)
- **[Bugfix][HiSparse] Never allocate GPU pages without host backing (#59036)** (197712e)
- **[Bugfix][HiSparse] Resolve MTP verification rows with a union residency kernel (#59235)** (25a029c)
- **[Dependency] Upgrade FlashInfer to 0.7.0.post1 (#59323)** (bffd162)
- **[Bugfix][Frontend] Reject LoRA adapters named after a served model (#59286)** (1d17b13)
- **[Core][BugFix] Tag prefix-cache extra keys by source (#51899)** (7f99308)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/vllm-project/vllm?utm_source=github-action)._