## lightseekorg/tokenspeed — v0.1.0…v0.1.1

_961+ commits._

### Features
- **feat(kernel): add an optional per-head bias to sigmoid_mul (#1890)** (9422ed2)
- **feat(layernorm): round the residual sum to BF16, scale residual inputs and enable PDL in the Triton RMSNorm (#1930)** (afb4c7b)
- **feat(rl): advertise the control endpoint; 501 for unsupported refit sources; NCCL split guard; flat sidecar server info (#1736)** (e60f386)
- **feat(spec): draft trees on Nemotron-H Mamba2 targets (#1968)** (7389f2a)
- **feat(dsa): bf16 index-K history gather under KVP pages (#1985)** (e035f28)
- **feat(qcp): attention head TP over the query shards (#1986)** (6e9f992)
- **feat(qcp): prompt logprobs on query shards (#1987)** (329a2c3)
- **feat(qcp): query-context-parallel prefill (#1979)** (1f76736)
- **feat(mla): head-TP, LM-head TP and TP-batch-invariant layouts for attention-DP decode (#1976)** (0a3d139)
- **feat(numerics): rl-bitwise reproduces the trainer's operation order; in-switch all-reduce (#1978)** (00b613a)
- **feat(moe): online expert rebalancing through the same-round gate (#1977)** (911fa25)
- **feat(logprobs): SGLang-dialect prompt (input) logprobs (#1958)** (a704cea)
- **feat(moe): static redundant-expert placement from recorded load (#1975)** (f6f0eac)
- **feat(pd): KV-page sharding (KVP) on the prefill role, routed by page ownership (#1974)** (af4da68)
- **feat(scheduler): per-request cap on the admission prefix probe (#1957)** (0c56476)
- **feat(rl): attention-DP weight updates and a Mooncake weight-store update route (#1955)** (e6c208e)
- **feat(spec): draft-prob rejection sampling for chain speculative decoding (#1954)** (3aaf452)
- **feat(pp): MTP speculation on a pipeline prefill server (#1956)** (29a11d6)
- **feat(drafter): multi-depth MTP under attention DP, layerwise PD and DSA k-row draft metadata (#1953)** (22c217e)
- **feat(spec): draft-tree speculative decoding (#1916)** (338f6df)
- **feat(moe): add a Triton block-FP8 MoE fallback for GLM-5.3 (#1936)** (54e5a18)
- **Add stateful kernel class support (#1914)** (a1c88af)
- **feat(model): support nemotron-3-super nvfp4 (#1918)** (52f9151)
- **feat(gemm): add Gluon gfx950 decode GEMM for Kimi K3 (#1913)** (932e186)
- **feat(amd): serve AMD Quark MXFP4 Kimi K3 with per-layer FP8 attention (#1898)** (2225168)
- **Add gfx1201 support through portable AMD kernels (#1922)** (5b9b86c)
- **feat(spec): ReplaySSM on by default for Qwen GDN targets (#1910)** (b403cf2)
- **feat(mla): add Gluon gfx950 MLA extend kernel for Kimi K3 (#1859)** (29060ee)
- **feat(DCP): Add Kimi K3 DCP support to the CuTe MLA backend (#1853)** (5905f20)
- **Add daily TokenSpeed nightly publication (#1888)** (6a90ad1)
- **Add daily tokenspeed-kernel nightly publication (#1886)** (4523807)
- **feat(ci): Add kernel benchmark filtering and Proton profiling (#1882)** (e565a05)
- **feat(amd): Add K3 support in Gluon MegaMoE (#1855)** (67aff99)
- **feat(scheduler): export pages_to_zero as zero-copy int32 arrays (#1872)** (7fa8acb)
- **Add opt-in startup phase timing (#1851)** (87f8e2c)
- **feat(amd): Integrate Petit Gluon MegaMoE (#1755)** (4248122)
- **feat(dsv41): support vit batching (#1844)** (b954417)

### Fixes
- **fix(epd): leave the receive pool out of the KV cache budget (#1998)** (4112f54)
- **fix(sampling): give each n>1 replica its own seed (#1964)** (1ddb046)
- **fix(test): restore loading dtype on failure and update V4.1 runtime tests (#1972)** (0cd4363)
- **perf(runtime): capture Mamba2 layers in the prefill graph (#1993)** (043cd4b)
- **fix(kernel): keep the n-gram history request chunk within the default stack (#1995)** (6fa1084)
- **Fix exact logprob arithmetic test inputs (#1994)** (1e577a7)
- **fix(ci): drop expert-placement flags from Kimi-K3 AMD job (#1992)** (4cbb277)
- **fix(router): forward the DSA query-shard surface to the sole leaf (#1991)** (ebb2f9c)
- **fix(ci): GQA prologue scalar v_ptr index; stale test stubs (#1984)** (8854e2a)
- **fix(moe): hand the flashinfer CUTLASS MoE one persistent zeroed workspace (#1989)** (31c47f2)
- **fix(ci): reconcile today's merges (unit tests, native libraries) (#1990)** (35dc7c2)
- **fix(dsa): slot_order=sorted is ascending position order, emitted by the top-k leaf (#1983)** (3919874)
- **fix(gdn): copy a row-strided QKV split input that spans past int32 offsets (#1967)** (3ca4ee3)
- **fix(longcat): count the identity zero-expert residual once under MoE TP/EP (#1973)** (33c6467)
- **fix(longcat): keep the attention branches in the dense row layout under attention DP (#1952)** (179b482)
- **perf(nemotron): trim launches and host time from Nemotron-3 Super serving (#1965)** (45db70f)
- **perf(amd): Fuse K3 latent up-projection add3 for decode batches (#1966)** (eb04745)
- **perf(kernel): pick llvm scheduling for gfx950 BF16 DSA attention (#1963)** (bf46f90)
- **perf(kernel): loop over gfx950 AttnRes block snapshots at runtime (#1962)** (dc37884)
- **perf(amd): route Kimi K3 prefill projections to Gluon large-M GEMM (#1946)** (7da7d6c)
- **perf(kernel): speed up gfx950 KDA prefill state scan (#1951)** (254fc6b)
- **perf(kernel): bound rel-MHA prefill KV loop and cut SGPR pressure (#1949)** (a842aeb)
- **perf(amd): route long BF16 GQA through packed prefill (#1939)** (e88f555)
- **perf(amd): unify k3 attention all reduce selection (#1923)** (b395a8c)
- **fix: stabilize FlashInfer rc2 startup and mixed draft batches (#1932)** (34eecae)
- **fix(kernel): key MoE tuning policy by target profile (#1935)** (d3e0338)
- **perf(amd): optimize gfx1250 GQA prefill with packed query rows (#1924)** (fc37d80)
- **perf(kda): fuse the `f_b` decay projection into gfx950 fused decode (#1928)** (bdd0326)
- **fix(kernels): 64-bit row offsets for verify sampling and GDN replay commit (#1919)** (fdab688)
- **fix(kernel): chain the Triton backend probe's underlying error (#1875)** (6acbe95)
- **perf(amd): reuse expert weights across K3 decode routes (#1817)** (af58b36)
- **perf(multimodal): move prefilled encodings to host memory (#1907)** (dbaca84)
- **perf(amd): pipeline and specialize gfx1250 MHA prefill (#1883)** (a4178a9)
- **fix(gdn): launch replay commit with the layer-batch-head product on grid x (#1915)** (d36f9bc)
- **fix(gemm): accept 0-dim per-tensor FP8 scales in mm (#1908)** (2f5e1b2)
- **perf(autotune): tune exact decode shapes and refresh GEMM routes (#1421)** (d3454db)
- **Fix KimiLinearMoE native layer state for attention DP (#1912)** (8bf3ef5)
- **perf(kimi-k3): shard prefill attention reduction and attnres (#1790)** (9e003ae)
- **fix(mla): avoid mask kwargs for flattened full-attention drafts (#1894)** (12d7cc3)
- **perf(loader): parallelize safetensors checkpoint prefetch (#1893)** (c8bf59e)
- **perf(amd): reuse MLA KV across causal verification queries (#1892)** (6cda9da)
- **perf(kimi-k3): split post moe all reduce in prefill and shard the moe tail (#1771)** (f46a87e)
- **perf(dsv41): skip decoder work for incomplete prefill chunks (#1863)** (cb127fd)
- **perf(pd): expand cache transfer SGEs per field, not per page (#1871)** (5a390b2)
- **perf(kimi3): prefetch KV tiles in FP8 MLA decode on gfx1250 (#1879)** (370268e)
- **perf(kimi3): speed up FP8 MLA prefill on gfx1250 (#1869)** (76b3f41)
- **perf(kimi3): retile decode GEMMs and pad large-M LDS rows on gfx1250  (#1865)** (0212562)
- **perf(cache): expand page x field zero ranges on the device (#1870)** (dbbe05c)
- **fix(kernel): stop dsv41 address kernels recompiling on page-table geometry (#1868)** (f69da88)
- **fix(ci): skip local branch guard during lint (#1866)** (bd5cbe3)
- **perf(prefill-graph): share output and handoff buffers across prefill graphs (#1816)** (16c951d)
- **fix(dcp): support FP8 query gathers and DeepGEMM V4 indexing (#1841)** (e0c2f51)
- **fix(cache): allow uneven ordinary cache groups** (3c3de31)
- **perf(amd): stream split partials in the MHA extend reduce kernel** (2199c66)
- **fix(kda): clamp gfx1250 prefill scan addresses into the sequence (#1846)** (7bb3582)
- **perf(moe): optimize TP MXFP4 tiles and fuse GEMM2 combine on gfx950 (#1848)** (0bbc7ce)
- **perf(kimi3): skip the C16 MoE cat on gfx1250 (#1796)** (a6962f5)
- **fix(kernel): pass the remaining per-batch kernel parameters at runtime (#1842)** (1070c65)
- **perf(amd): avoid spills in MXFP4 expert output reduction (#1838)** (8721be8)

### Backend
- **Show parallel ranks in scheduler process names (#1927)** (485701a)
- **Upgrade TokenSpeed Scheduler to 0.1.24 (#1906)** (7767d94)
- **Update TokenSpeed Scheduler to 0.1.24 (#1905)** (3cb09cd)
- **Update tokenspeed-kernel MLA to 0.2.16 (#1904)** (5d29434)
- **Update tokenspeed-mla to 0.2.16 (#1903)** (a12182c)
- **Bound startup health-probe reconnect delays (#1856)** (ae8bcfe)
- **Update tokenspeed-kernel MLA to 0.2.15 (#1850)** (701a62b)
- **Update tokenspeed-mla to 0.2.15 (#1849)** (49225ce)
- **Allow wider TVM-FFI versions in TokenSpeed MLA (#1847)** (95dd8e8)

### Tests
- **test(amd-kernel): trim duplicate and low-value AMD kernel tests (#2002)** (cb91188)
- **test(kernel-amd): Fill gfx1250 gaps in MI450 simulation coverage (#1997)** (340b009)
- **test(logprob): the admission handler stub declares the layout fields (#1988)** (741e719)
- **test(drafter): Mtp finalizes layerwise, like Eagle (#1981)** (bcbb065)
- **test(iris): barrier before the MoE tail in the attention-mix test (#1970)** (6bb3259)

### Docs
- **docs: wrap the AMD ops README to 80 columns (#1843)** (90eea27)

### Chore
- **build: release tokenspeed 0.1.1 (#2009)** (3d22d47)
- **build: release tokenspeed-kernel 0.1.4 (#2007)** (608a6f7)
- **build: release tokenspeed-kernel-amd 0.1.4 (#2005)** (12b41b6)
- **ci: orchestrate weekly releases across PyPI, wheelhouse, and Docker (#2001)** (9f3138c)
- **ci: validate pull request commit trailers (#1996)** (60faf07)
- **chore(deps): require tokenspeed-scheduler>=0.1.25 (#1982)** (45f2216)
- **chore(scheduler): bump tokenspeed-scheduler to 0.1.25 (#1980)** (5cfc8ce)
- **refactor(linear): select V4.1 FP8 methods during construction (#1971)** (408f926)
- **ci(amd-kernel): Retune gfx950 kernel benchmark regression budgets (#1969)** (7e2f1de)
- **ci(amd): Run kernel benchmarks alongside model tests (#1961)** (729fec0)
- **ci(amd): Run all Inkling AIME25 problems concurrently (#1960)** (0697a71)
- **chore(runtime): remove stale weight loader method names (#1959)** (aee3618)
- **ci(amd): Raise Kimi K3 EAGLE3 50K/500 perf reference (#1948)** (aeee7be)
- **ci(amd-kernel): Benchmark Kimi K3 AttnRes prefill mixing (#1947)** (ab14b32)
- **ci(amd-kernel): Benchmark Kimi K3 MLA cached extend (#1944)** (b2dd427)
- **ci(amd-kernel): Update Kimi K3 decode GEMM cases (#1943)** (7bdeb03)
- **chore(server-args): default gpu_memory_utilization to 0.95 at every world size (#1885)** (2f45e96)
- **ci: make glm-5.2 nvfp4 aime26 eval manual only (#1942)** (8066037)
- **ci(amd-kernel): Benchmark Kimi K3 MLA verify on the query axis (#1937)** (3616aa0)
- **ci(amd): use 1gpu-bench runners for gfx950 benchmarks (#1934)** (c623fb8)
- **ci(amd-kernel): Benchmark Kimi K3 fused KDA decode, verify and replay (#1896)** (9ec647a)
- **ci: use hardware-based names for GPU test workflows (#1926)** (a0254c6)
- **ci(amd): run gfx950 benchmarks on managed MI350 runners (#1925)** (bbedc7a)
- **ci: allow explicit TokenSpeed nightly replacement (#1909)** (018c0c0)
- **ci: unblock CUDA wheel builds with a local repository (#1902)** (76fc28f)
- **ci: include H100 kernels in published CUDA wheels (#1901)** (e8ff04e)
- **refactor(attention): one prologue entry for QK norm, RoPE, KV quantize and the cache write (#1772)** (6930215)
- **ci: publish dated tokenspeed-kernel-amd nightlies (#1900)** (4464c43)
- **ci: move ROCm nightly pip index under nightly/rocm7.2 (#1899)** (6e3cfe2)
- **ci: publish ROCm kernel and TokenSpeed nightly wheels (#1897)** (9a87835)
- **refactor(kimi-k3): refactor and optimize latent MoE tail (#1845)** (2a8dbc6)
- **ci(amd-kernel): Add Kimi K3 KDA prefill benchmark cases (#1884)** (0fbc6fa)
- **ci(kernel-benchmark): Publish kernel benchmark comments for fork PRs and merged PRs (#1873)** (aafc8a2)
- **chore: clean up moe benchmark (#1881)** (b3a18b8)
- **chore(pd): drop per-request Mooncake INFO logs, log requests by default (#1867)** (54ac713)
- **ci: run lint on main pushes (#1864)** (cf1bc65)
- **refactor(attention): validate MLA DCP backend and cache capabilities (#1862)** (2cda9d9)
- **ci(amd-kernel): Add Kimi K3 decode projection benchmark cases** (fa74a28)
- **refactor(kimi3): select packed sigmoid top-k from the registry (#1768)** (066edaf)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/lightseekorg/tokenspeed?utm_source=github-action)._