## deepspeedai/DeepSpeed — v0.19.6…v0.19.7

_80 commits._

### Features
- **Add DCO sign-off to release commits (#8527)** (7075430)
- **feat(rollout): add continuous batching profiling (#8494)** (8950dc3)
- **[Phase 2] Add NEON SIMD path for CPU Adam on AArch64 (#8453)** (c102bbd)
- **feat(rollout): add continuous batching generation prototype (#8368)** (e5887a3)
- **Add an opt-in ZeRO-1/2 gradient norm fast path (#8331)** (f992dc2)
- **Skip fp16-config tests on accelerators without fp16 support (#8398)** (493dafa)
- **Add nccl_version to the source-checkout torch_info fallback (#8383)** (56de570)
- **Add an opt-in fused weighted restore for AutoEP (#8326)** (80f19f3)
- **Add an opt-in DeepEP transport for the AutoEP expert all-to-all (#8213)** (534dc0e)
- **Add macOS (MPS) CI workflow and a torch floor check for the MPS accelerator (#8335)** (ba3246d)
- **feat(cpu-adam): add ARM SVE update kernel (#8365)** (020893e)
- **Add pin_empty helper for empty scratch destinations (#8332)** (725811f)
- **Add configurable dtype for ZeRO checkpoint export (#8318)** (92843ad)

### Fixes
- **Fix sequence overlap backward gradient permutation (#8342)** (2ce20d6)
- **Fix ZeRO parameter alignment for grouped_mm (#8277)** (080eb6b)
- **Fix universal checkpoint resume across AutoTP sizes (#8474)** (e76a280)
- **Fix comms logger KeyError when log_name is omitted (#8267)** (bd6e968)
- **[AutoEP]Fix optimizer and replaced MOE parameter mismatch (#8377)** (5daeffe)
- **fix(lr_schedules): make --lr_range_test_staircase an opt-in flag (#8337)** (8d4074c)
- **Fix AutoTP + deep compile collectives silently drop when AC is on (#8355)** (92269f4)
- **Replace cuda_graph assertion with explicit ValueError; fix custom_op type annotations (#8336)** (26b2d99)
- **Fix flops profiler counts for transposed convolutions (#8323)** (7fabe47)
- **fix: Only bind device id when needed,  Fixes #8248 (#8269)** (3bdabae)
- **perf(rollout): profile prefill and decode forwards (#8350)** (1c5cc6b)
- **Fix Triton NFS detection crash when df wraps long device names (#8259)** (c389bae)
- **Fix the seq-first Ulysses all2all output layout (#8317)** (cd5206a)

### Backend
- **[bugfix] _DimZeroAllToAll silently fails to send grads under torch.compile (#8491)** (f5af15c)
- **Cache the Modal sandbox install chain in image layers (#8536)** (83dd543)
- **Emit affine maps from AutoTP layers (#8519)** (69797b0)
- **Update version.txt after 0.19.6 release (#8333)** (78db78d)
- **Give Muon's momentum the dtype of the gradient it is combined with (#8483)** (431f8ac)
- **Read rope_theta from rope_parameters across Inference V2 (#8345)** (1b892a7)
- **Partition AutoEP expert parameters per layer under ZeRO-3 (#8424)** (6d5c56f)
- **Gate the offload-state memory deltas on allocator-backed stats (#8409)** (e6d2d40)
- **Probe the device module for train_cifar's fork_rng device entries (#8407)** (0306300)
- **Deprecate sparse attention (#8493)** (a3f9d4a)
- **Make wait() idempotent on AllGatherHandle and NoGatherHandle (#8487)** (10e358b)
- **Honor --include/--exclude in the SLURM launcher (#8304)** (8bdd8f1)
- **Offload a saved view when it is the last value holding its storage (#8388)** (5663128)
- **Validate positive inference and HybridEngine max output tokens (#8343)** (7c1b0f2)
- **Muon silently discards the param groups it is given (#8440)** (113c675)
- **Remove triton compatibility check for fp_quantizer (#8492)** (cc05b81)
- **[muon] Keep the momentum out of steps the loss scaler discards (#8435)** (b5e000c)
- **Deprecate unused DeepSpeed features (#8490)** (71d316d)
- **[muon] Reconcile the momentum dtype when a checkpoint is restored (#8433)** (b726f4e)
- **Carry the affine scale on the replicated map, not the split (#8477)** (bce0adf)
- **Describe universal checkpoint shards as affine maps (#8385)** (1190946)
- **Split DeepCompile ZeRO-3 memory scheduler (#8233)** (3856146)
- **[MPS] Update C++ Standard in CPUAdamBuilder  (#8466)** (bb7ad0d)
- **Avoid collective token preparation for AutoEP DeepEP (#8423)** (29d0abb)
- **Muon is silently disabled under ZeRO-3 when the model is built with zero.Init (#8438)** (0166cc8)
- **Filter --include against the real slots, not against itself (#8239)** (3b1c14a)
- **Muon runs no Newton-Schulz at ZeRO stage 0, the default: run it (#8442)** (cbd303e)
- **Stop the debug name maps from pinning the model they snapshot (#8356)** (e9680c7)
- **Read rope_theta from rope_parameters in the Llama injection policy (#8341)** (6474bc5)
- **[Workflow] run modal GPU workflows only from the merge queue (#8412)** (1e90aa9)
- **Split the modal CI budget into acquisition and test phases (#8404)** (666720b)
- **Reshape instead of view in TiledFusedLogitsLoss (#8362)** (7119936)
- **op_builder: use C++20 for nvcc on CUDA 13+ (#8422)** (abf3dfa)
- **Preserve AutoEP score correction bias buffers (#8369)** (32110a9)
- **[Workflow] Raise modal CI timeouts to absorb slower sandbox provisioning (#8403)** (05daf05)
- **[Ulysses] Carry the KV head count per DistributedAttention (#8316)** (d4ed1f1)
- **Use device names, not rank ids, for device placement in test helpers (#8397)** (4fd3c82)
- **Do not pin DDP device_ids for CPU reference models (#8399)** (d3a1b68)
- **Preallocate the static KV cache with config.head_dim (#8389)** (e74c707)
- **Stop the curriculum schedule starting below min_difficulty (#8334)** (8e09ed2)
- **Keep the elasticity batch overrides out of the caller's config dict (#8329)** (c7eed15)
- **Broadcast elementwise flops from the trailing dimension (#8324)** (b8bfd26)
- **Forward barrier device_ids to communication backends (#8312)** (177c5aa)
- **Forward free_data through partition() instead of hardcoding True (#8305)** (9666fe5)
- **Recognize Qwen3.5's RMSNorm variants in AutoTP module loading (#8306)** (8e64a09)
- **Make the WarmupCosineLR ratio flags reach the config (#8268)** (cdd206f)
- **DeepCompile: stabilize ZeRO-3 parameter guards (#8328)** (4189091)
- **remove the dead non-Triton attention path from TritonSelfAttention (#8349)** (5b6da2b)
- **[AutoTP] Replace tp_shard process-wide globals with per-model AutoTPMeta (#8241)** (183c7f9)
- **[tiled mlp] reshape instead of view (#8348)** (87d9ecd)
- **Fallback for unsupported Hybrid Engine policies (#8265)** (a36f78e)
- **Count each module object once when aggregating flops profiler totals (#8320)** (f4ee787)
- **Raise the per-tensor norms to norm_type when combining them (#8313)** (37cf8f2)

### Docs
- **docs: add test discipline rules to agent guidelines (#8372)** (f15c990)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/deepspeedai/DeepSpeed?utm_source=github-action)._