TokenSpeed is a speed-of-light LLM inference engine.
This week brought substantial work on query-context-parallel (QCP) prefill and inference, including prompt logprobs support and attention head tensor parallelism across query shards. The team also shipped online expert rebalancing for mixture-of-experts models, static expert placement from recorded…
Get this in your inbox every Monday →A deterministic 0–100 hygiene score — README, license, CI, tests, docs, and freshness.
Who ships this repo — author concentration and the bus factor across the last 300 mainline commits.
How welcoming this repo is to contributors — issue throughput, close time, responsiveness, and good-first-issue count.
What this project is built on — dependency count by ecosystem, the license mix, and anything worth a legal look before you adopt it.
Whether this project's CI can be trusted — pass rate, run times, flaky runs, and which workflow is the weak link.
Grounded in tokenspeed's README, structure, and recent commits — answers won't invent code they haven't seen.
A Monday email with what shipped, in plain English — no account needed.
Showing raw commit titles for the newest commits. Sign in to generate AI summaries.
feat(moe): run TRT-LLM BF16 MoE at 64-aligned intermediate sizes (#1878)
perf(spec): load tree causal_conv1d weights and ancestors ahead of the PDL wait (#2006)
fix(sampling): keep the whole support at top_p >= 1 in fused top-k/top-p (#1836)
A floor, not a guess: counts only commits whose author, co-author trailer, or message explicitly credits an AI tool (Claude, Copilot, Cursor, aider, Codex…). Based on 30 mainline commits. Unattributed AI code isn't counted here — the full audit estimates that separately.