A high-throughput and memory-efficient inference and serving engine for LLMs
This week focused on stability and correctness across multiple fronts. Several bugfixes addressed determinism in batch-invariant operations, KV cache offloading state management, and attention kernel behavior across ROCm and XPU backends. The team also shipped frontend improvements including better…
Get this in your inbox every Monday →A deterministic 0–100 hygiene score — README, license, CI, tests, docs, and freshness.
Who ships this repo — author concentration and the bus factor across the last 300 mainline commits.
How welcoming this repo is to contributors — issue throughput, close time, responsiveness, and good-first-issue count.
What this project is built on — dependency count by ecosystem, the license mix, and anything worth a legal look before you adopt it.
Whether this project's CI can be trusted — pass rate, run times, flaky runs, and which workflow is the weak link.
Grounded in vllm's README, structure, and recent commits — answers won't invent code they haven't seen.
A Monday email with what shipped, in plain English — no account needed.
Showing raw commit titles for the newest commits. Sign in to generate AI summaries.
[Misc] Report a missing DeepSelect extension only when it is requested (#60242)
[Bugfix][MoRIIO] Exclude synchronous READ destinations from KV zeroing (#59164)
[CI][RL] Consolidate RL entrypoint tests under tests/entrypoints/rl (#59948)
[Rust Frontend] Share `SchemaRoot` between argument coercion and grammars (#59408)
[CI/Build] Widen the ragged prefill scoring tolerance, drop batch invariance (#60192)
[Bugfix][Scheduler] Let pooling chunked prefill use full context (#48039)
[Docs][RL] Move pause/resume and batch invariance examples to examples/rl (#59949)
[Bugfix][Frontend] Return 4xx for client errors in /cohere/v2/chat (#60309)
A floor, not a guess: counts only commits whose author, co-author trailer, or message explicitly credits an AI tool (Claude, Copilot, Cursor, aider, Codex…). Based on 30 mainline commits. Unattributed AI code isn't counted here — the full audit estimates that separately.