## pytorch/torchtitan — v0.2.2…v0.3.0

_842+ commits._

### Features
- **Add GPU release candidate validation (#4290)** (7e83924)
- **add release.md (#4290)** (086e36e)
- **[ci] Run fake-PG for functionality tests, add numerics guard for models, enable GPU unittests (#4127)** (d328155)
- **Add TestPyPI release workflow (#4219)** (e4a581b)
- **[DistMuon] Support oversharded QKV heads (#4078)** (858b9e2)
- **Drop Python 3.10 support, require Python 3.11+ (#4180)** (233c22f)
- **[RFC][Config] Add torchtitan_configs for full training configurations (#4114)** (c82675e)
- **Add torch_checkpointing save config plumbing (#4058)** (7c6f9f7)
- **Add a common checkpoint manager bridge (#4141)** (6bc2108)
- **Add remat as an optional dependency (#4163)** (0d2438f)
- **[Reland] Add dist GEMM attention and FFN backend (#4162)** (9a71152)
- **Add CUDA graph capture to the core Trainer (#3559)** (f84224a)
- **[Kimi K2.7] Add DistributedMuon with all-to-all/compute overlap (#4060)** (bb90ff9)
- **Muse Glimmer: add varlen (FA3) attention + vLLM generator support (#4126)** (c39b844)
- **Support single-visible-device accelerator launches (#4103)** (90251b3)
- **Add MuseGlimmer GraphTrainer support (#4113)** (ef923dd)
- **Support context, tensor, and pipeline parallelism with MinimalAsyncEP (#4089)** (15db18b)
- **Add Muse Glimmer (30B) model and integration tests (#4106) (#4106)** (5de7d11)
- **Experimenting MoE new sharding (#3996)** (45ea2f0)
- **[MTP] Add MTP module for deepseek_v3 model (#3392)** (f59c215)
- **Add an NVFP4 quantization converter (#3914)** (bcc0929)
- **[kimi k2_7] add kimi k2_7 (#3532)** (85c549b)
- **[rl] add window FIFO scheduling (#3927)** (fd27765)
- **Add Google Cloud TPU to get_peak_flops (#3900)** (363709f)
- **Add single-node Qwen3 DAPO math example (#3951)** (e00b27d)
- **Add MoE support to the transformers modeling backend (#2679)** (af35cfb)
- **add a debug config for deepseek + mxfp8 (#3934)** (7a3dab7)
- **[trainer] add model config to Trainer to_dict (#3919)** (7e3f2eb)
- **Support remote fsspec checkpoint paths (e.g. gs://) via filesystem helpers (#3887)** (fec3e19)
- **Add local install command to README (#3902)** (bee12f1)

### Fixes
- **Fix legacy GPU release validation setup** (5926201)
- **Fix GPU validation artifact permissions** (ca02f47)
- **Fix release GPU validation compatibility** (d885754)
- **fix integration test cmd to include batch dim** (f0916d0)
- **Fix quoting of legacy integration test overrides** (62b6ca3)
- **[ci] fix transformers_backend routed_experts MoE tests (#4249)** (d0881ff)
- **[ci] fix transformers backend test arguments (#4309)** (82e98c8)
- **[rl] Fix SPMD HF state-dict loading (#4236)** (4f9e9a4)
- **Fix TorchFTCheckpointManager's contract with BaseCheckpointManager (#4183)** (b838036)
- **[spmd_types] Qwen3.5 spmd_typecheck error fix (#4179)** (66f8e0d)
- **Fix YaRN zero correction boundary to pass Helion RoPE tests  (#4215)** (48cb7ae)
- **[qwen3.5] Fix GatedDeltaNet A_log initialization range (#4178)** (e0549d1)
- **Fix final checkpoint export dtype conversion (#4166)** (575dd2f)
- **[spmd_types] fix MoE routing bias TP reduce (#4154)** (f3660a4)
- **[spmd_types] fix fake PG + AC interaction (#4150)** (07eaa22)
- **Fix CUDA graph capture with FSDP always reshard (#4143)** (04326fc)
- **Muse-Glimmer - text-config torchvision fix + HuggingFace StateDictAdapter (#4118)** (ad4a0e4)
- **fix RL spmd_types state dict mesh resolution (#4128)** (3ed5001)
- **fix async tp and move it into compile (#4045)** (7375947)
- **[spmd_types] fix spmd->DTensor translation on partial mesh (#3913)** (801fe17)
- **[MoE] fix: pass the computed `in_grad_placements` without EP too — TP gradients below the experts lose their reduction (#4054)** (ecae62f)
- **Fix float8 filter_fqns to exclude lm_head, not output (#4008)** (b3cf840)
- **Fix AMD 8-GPU-feature CI (#3896)** (1c40dd2)
- **[MinimalAsyncEP] Fix int32 address overflow in top-k kernels (#3969)** (b3a13ee)
- **[rl] fix rl tests (#3981)** (c30a800)
- **fix: support transformers 5.9.0 and hub 1.24.0 for tokenizer download (#3962)** (d93d7d0)
- **[rl] Fix CI: Synchronize bitwise parity test teardown (#3959)** (b36e4b6)
- **Patch readline-import issue which causes VLLMGenerator hang (#3950)** (cadbf37)
- **fix RL ci (#3930)** (fbceec0)
- **[CI] Fix lychee error (#3917)** (fd16c70)
- **[rl] Revert generator initialization race condition fix (#3809) (#3871)** (e622688)
- **[qwen3.5] fix DTensor TP vision position caches (#3899)** (491596e)
- **Fix:  Dataloader restarting on second resume (#3908)** (6ad8c3d)
- **[ci][models] block qwen3.5 spmd_types, fix dsv3 compile (#3901)** (0d7f4ab)

### Backend
- **Prepare TorchTitan 0.3.0** (086bf6c)
- **Prepare TorchTitan 0.3.0rc4** (6f56096)
- **Revert #4127 and restore legacy release validation** (baceb10)
- **Remove unstable Qwen3.5 golden numerics check** (b57b076)
- **Update release SFT A10G numerics** (b41f658)
- **Make GPU release validation container-safe** (c5f3270)
- **Make config validation tests Tyro-version agnostic** (c02a44e)
- **Prepare TorchTitan 0.3.0rc3** (b7ed933)
- **Update stale reference links** (1338c31)
- **use torch checkpointing stable API in linter** (ba61cd9)
- **update to rc2** (c68364c)
- **[spmd_types] enable transformers modeling backend** (ad4ad69)
- **[Config] Move the integration tests to full configurations (#4177)** (1e7672e)
- **[ci] pin rl+hf tests to partial_dtensor (#4228)** (49e8485)
- **[graph_trainer][ci] Disable CUDA graphs for blocking EP tests (#4246)** (f844314)
- **[ci] disable-cuda-graphs for GraphTrainer PP tests (#4229)** (c9aaefb)
- **Initialize SPMD backend during graph precompile (#4227)** (38c91ef)
- **Set version to 0.3.0rc1 (#4216)** (1fdf90d)
- **[Forge] Normalize PP loss before backward (#4194)** (f4a575f)
- **[Flux] Normalize validation loss by latent elements (#4195)** (6bdd6f4)
- **[rl] Exclude non-finite old-policy tokens from the global loss denominator (#4174)** (9a99528)
- **Fused MLA override (#4134)** (db1ce7f)
- **Use concurrent_endpoint instead of endpoint (#4204)** (ea1db70)
- **pin vllm page size = 256 for FA2 compatibility (#4181)** (6508194)
- **spmd_types as default backend (#4085)** (5ab3a0f)
- **Move the checkpointing-disabled guard into BaseCheckpointManager (#4173)** (862f966)
- **Remove the legacy torchtitan.components.lr_scheduler import path (#4172)** (203f9a8)
- **[spmd_types] Muse Glimmer enablement  (#4161)** (f0c5aad)
- **[validation] Aggregate loss by valid token count (#4167)** (cbe4bfa)
- **Group the optimizer components into a package (#4140)** (970fbdd)
- **[DistMuon] workaround `missing_for_each_optimizer` linter as it's false (#4169)** (a6df53f)
- **[rl] Cycle least-loaded routing between tied candidates (#4109)** (2f10a25)
- **Do not count MoE tokens during evaluation (#4148)** (c4dff13)
- **[Kimi 2.7] rename DistributedMuon to DistMuon (#4155)** (5b70899)
- **[spmd_types] Kimi 2.7 enablement (#4079)** (624c312)
- **In-place loss accumulation (#4146)** (2807d3f)
- **[spmd_types] qwen3.5 enablement (#3895)** (126e26c)
- **Make torch_checkpointing an optional dependency (#4123)** (41c36ae)
- **[muse_glimmer] Change vision encoder forward internally (#4130)** (2d8b2ec)
- **[qwen3_5] Reset GatedDeltaNet state at document boundaries under flex attention (#3984)** (f4e7818)
- **Run InterGeneratorRouter inside a monarch actor (#4072)** (4b71e5f)
- **qwen3: align initialization with Qwen defaults (#4104)** (031b84c)
- **[RL] Test RL integration with spmd_types (#4112)** (65fa556)
- **Apply YaRN scaling independently of cache length (#4117)** (23bcf4e)
- **Keep global valid token counts on device (#4099)** (1f3ae09)
- **Mask loss at document boundaries (#4075)** (3f71477)
- **move padding to inference (#4100)** (d94ecb9)
- **[MoE] Unify expert parallel token dispatcher API (#3970)** (4a93ee4)
- **inference moe expert sp padding (#4080)** (f4f7cf7)
- **[CP] Move context parallel code into a context_parallel package (#3977)** (547b0b4)
- **Allow fake process groups to simulate any rank (#4018)** (5c0b804)
- **Decouple memory snapshot frequency from profiler (#4092)** (e061028)
- **Update TitanRL readme (#4086)** (95007d2)
- **Exclude fake-backed axes from get_all_one_dimensional_meshes (#4068)** (96276d8)
- **enable varlen full cudagraph (#3893)** (a6948f5)
- **[kimi2.7] pass correct fsdp mesh for backend=spmd_types (#4070)** (57cfb27)
- **Bump pypa/gh-action-pypi-publish from 1.14.1 to 1.14.2 in the github-actions group (#4066)** (d905f73)
- **Skip GraphTrainer H100 EP-overlap MoE failures (#4053)** (b175497)
- **enable graph trainer + mxfp8 composability (#3558)** (95e4226)
- **Skip GraphTrainer H100 FSDP toposort failure (#4048)** (681fd4b)
- **Relax GraphTrainer stack trace node count test (#4046)** (d84e54e)
- **Always Pre-Split Microbatches for PP (#3856)** (9228564)
- **Raise 8 GPU model test timeout to 60 minutes (#4036)** (df51ae9)
- **Bump DeepEP HybridEP pin (#4033)** (c91448d)
- **Temporarily disable failing ROCm CI jobs (#4002)** (20f12e3)
- **Update CUDA llama3 loss golden (#4024)** (b5eb9d9)
- **Select FA4 varlen attention on Blackwell (#4012)** (4bed502)
- **Qwen 3.5 Varlen Attention (#3801)** (d24ae40)
- **[cp] Allow PTRR load balancer to select a mask from a dict (#3972)** (4c64811)
- **[rl] simple rename changes (#3926)** (725b995)
- **Gate deepseek_v3_671b float8 converters on hardware capability (#3945)** (cc286a6)
- **Bump the github-actions group with 2 updates (#3965)** (31acbb3)
- **[Checkpoint] clarify initial load and resume behavior (#3732)** (7463c16)
- **[override] Allow passing kwargs in override from both CLI and config (#3894)** (0fb45d4)
- **[spmd_types] VarlenAttention: remove hardcoded spmd_types (#3937)** (5059f32)
- **mxfp8 moe: enable ep=1 (#3935)** (c95b211)
- **Make AutoParallel tests GPU-agnostic using device_type (#3805)** (ab46124)
- **[MoE] sibling token_dispatcher + grouped_experts for composable override (#3859)** (3101b42)
- **[spmd_types] FlexAttention: remove hardcoded spmd_types of FlexAttention._complex_flex_attn output (#3924)** (fc5a729)
- **[common] Move Linear and ScaledBiasRowwiseLinear into common/linear.py (#3923)** (659a0d2)
- **[gpt_oss] Build flex masks via the shared decoder helper (#3814)** (ec8c906)
- **batch invariant configs should have max off policy = 0 (#3910)** (ea3562e)
- **Unify CosSinRoPE and ComplexRoPE YaRN computation (#3787)** (51c197c)
- **[rl] Remove RL H100 CI to A10G and drop oversized sample before packing (#3855)** (4b041be)

### Chore
- **ci: use writable HF cache in H100 integration tests (#4006)** (1ac4653)
- **ci: run one backend per model test (#3961)** (16e3fac)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/pytorch/torchtitan?utm_source=github-action)._