Ongoing research training transformer models at scale
The week focused on stability and correctness across inference and training. Notable fixes include handling edge cases in multimodal models (Qwen3 tool names, LLaVa hybrid architecture, RADIO temporal grouping) and resolving state management issues in Mamba's prefill caching. On the training side,…
Get this in your inbox every Monday →A deterministic 0–100 hygiene score — README, license, CI, tests, docs, and freshness.
Who ships this repo — author concentration and the bus factor across the last 300 mainline commits.
How welcoming this repo is to contributors — issue throughput, close time, responsiveness, and good-first-issue count.
What this project is built on — dependency count by ecosystem, the license mix, and anything worth a legal look before you adopt it.
Whether this project's CI can be trusted — pass rate, run times, flaky runs, and which workflow is the weak link.
Grounded in Megatron-LM's README, structure, and recent commits — answers won't invent code they haven't seen.
A Monday email with what shipped, in plain English — no account needed.
Showing raw commit titles for the newest commits. Sign in to generate AI summaries.
fix(inference): re-salt unprefilled requests when resuming after refit (#7891)
A floor, not a guess: counts only commits whose author, co-author trailer, or message explicitly credits an AI tool (Claude, Copilot, Cursor, aider, Codex…). Based on 30 mainline commits. Unattributed AI code isn't counted here — the full audit estimates that separately.