## deepspeedai/DeepSpeed — v0.19.5…v0.19.6

_55 commits._

### Features
- **feat(opsd): share prompt prefill across rollout samples (#8296)** (a9505fd)
- **[Apple Silicon Support Phase 1] Add Metal FusedAdam kernel and CPU Adam build for Apple Silicon (#8300)** (715965e)
- **Enable DeepSpeed support on Apple Silicon (MPS) with ZeRO Stage 1-3 (#8293)** (64fcec6)
- **[typo] Add raise for RuntimeError (#8281)** (11b518a)
- **Add  fused triton kernel for swiglu (#8244)** (24a41af)
- **[AutoTP] Complete uneven sharding and universal checkpoint support (#8185)** (aa3914d)
- **Add CUDA graph support for HybridEngine generation (#8271)** (80c8e5b)
- **Add compile.offload_activation_pin_memory for DeepCompile activation offload (#8258)** (c952d92)
- **Add native (DeepNVMe) host-memory pinning backend for accelerators (#8211)** (aa0e91b)
- **Add configurable sum gradient reduction (#8232)** (1cae492)

### Fixes
- **[AutoTP] Fix training lm_head routing (#8302)** (1b956e6)
- **Fix ZeRO-3 synchronization during OPSD rollout (#8264)** (e40c681)
- **Fix shared loss gradient accumulation (#8245)** (4ffb4e5)
- **Fix AutoTP metadata updates for unsharded modules (#8299)** (7154a00)
- **Fix device mismatch in test_gate_up_partition_covers_the_whole_weight (#8308)** (84fd92a)
- **perf(rollout): add HybridEngine rollout profiling (#8295)** (da3ca68)
- **Fix zero reduce bucket size validation (#8266)** (e70b903)
- **fix(engine): resolve inference workspace attribute lookup (#8288)** (fb57b81)
- **Fix repeated FlopsProfiler metric accumulation (#8246)** (ebb75d7)
- **fix: stop DeepSpeedConfig writing max_grad_norm back into the caller's config dict (#8289)** (f567175)
- **Fix ZeroDivisionError in compute_elastic_config return_microbatch on non-0.2 elasticity (#8286)** (ad1c516)
- **Fix ZeRO-1/2 with zero-sized parameters (#8280)** (edaa722)
- **Fix docstring Args entries that name a parameter the function does not take (#8223)** (32d51a1)
- **Fix ZeRO-3 crash in AutoTP universal-checkpoint metadata (#8270)** (313ce47)
- **Fix int32 overflow in Triton grouped-GEMM expert offset (#8261)** (c5331bc)
- **Fix DeepCompile last use for unconsumed waits (#8254)** (9bd89f9)
- **[DeepCompile] Fix staticmethod handling on Python 3.9 (#8240)** (1d580d6)
- **[DeepCompile] Fix KeyError on frozen parameters in ZeRO-3 (#8214)** (04b4b0c)
- **Fix checkpoint rank selection for Ulysses sequence parallelism (#8226)** (b39e07a)

### Backend
- **Honor adam_w_mode in the CPU multi_tensor_adam binding (#8307)** (c7cc64a)
- **Route ZeRO/SuperOffload pin sites through accelerator pin_memory (#8256)** (eebe24a)
- **Isolate DeepCompile list-schedule test ops (#8319)** (28ed612)
- **stop allocating per-element temporaries in the overflow check (#8325)** (32e301f)
- **Enable optimized Adam backend for MuonWithAuxAdam optimizer (#8278)** (d47e4f4)
- **Avoid redundant copies in MPS P2P staging (#8314)** (eba5d27)
- **Reject invalid ZenFlow ratio and update interval boundaries (#8274)** (6e3bd08)
- **Import print_dist in auto_tp (#8311)** (b753aec)
- **Route isend/irecv to nonblocking backend methods and stage them as async on MPS (#8303)** (9311fd5)
- **Skip zero-sized parameters in HP fragment mapping (#8298)** (573ce48)
- **Drop the documented grad_hooks ZeRO option, which does not exist (#8242)** (f76ab88)
- **Register native pinned host memory with CUDA for GPU DMA (#8283)** (cb26080)
- **[AutoTP] Convert HF embedding_rowwise tp_plan entries to SKIP specs (#8294)** (20ea454)
- **Enable activation offloading (#8255)** (9a297cf)
- **Non-reentrant activation checkpoint CPU offload (#8282)** (858e91e)
- **Make pass contracts cover the passes DeepCompile actually schedules; Renaming files (#8251)** (35c8b03)
- **Route FPDT and checkpoint writer pins through accelerator pin_memory (#8257)** (310eb8c)
- **Keep parameter dtype through ZeRO-3 weight quantization (#8215)** (6098f79)
- **NCCL and other backend PG timeout 30min -> 10min (#8253)** (24402c3)
- **(1/2) Implementing Compiler Pass for AutoTP (#8204)** (7904603)
- **Upgrade pip in Python install-smoke jobs (#8247)** (22969fa)
- **Keep the elastic batch size within max_train_batch_size (#8237)** (2cfebbd)
- **Implement the documented per-param-group lists in OneCycle (#8201)** (8cae9d2)
- **Update version.txt after 0.19.5 release (#8243)** (66f9a36)
- **Return a copy from OnebitLamb.get_lamb_coeffs (#8227)** (5dceeea)

### Docs
- **docs: note async cpu_checkpointing perf and expandable_segments (#8287)** (96ffef2)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/deepspeedai/DeepSpeed?utm_source=github-action)._