## NVIDIA-NeMo/Automodel — v0.4.0…v0.5.0

_531+ commits._

### Features
- **cp: `feat(diffusion): support Hugging Face datasets (2816)` into `r0.5.0` (#2831)** (aa589f1)
- **cp: `ci: add gb200 cluster specification for nemotron_ultra recipe (2733)` into `r0.5.0` (#2734)** (b8413d5)
- **cp: `ci: add time budgets for 12 new timeout failures (2707)` into `r0.5.0` (#2724)** (e030120)
- **cp: `ci: add cluster_tag to gb200 benchmarks (2714)` into `r0.5.0` (#2715)** (7200554)
- **cp: `perf(distributed): add retrieval tuning knobs (2452)` into `r0.5.0` (#2607)** (142eafe)
- **cp: `feat(examples): add Nemotron-3-Ultra-550B benchmark and full-SFT recipes (2539)` into `r0.5.0` (#2550)** (1ab1bfe)
- **feat: Add query functionality of Model Capability Registry  (#2423)** (d2fcf72)
- **feat(moe): enable MXFP8 MoE training on GB200 (TransformerEngine + torchao) (#2394)** (77f1123)
- **feat(diffusion): add Wan2.2 T2V-A14B two-stage finetuning support (#2284)** (6ff1c8b)
- **feat(diffusion): improve qwen image finetuning configs (#2442)** (c19cd94)
- **feat(distributed): add selective activation checkpointing for FSDP2 (#2389)** (02c0077)
- **feat(speculative): add EAGLE-3 sequence packing and reasoning-mode control (#2444)** (f69851e)
- **feat: add use_memory_efficient_lora knob (#2239)** (9924ae4)
- **feat(models): add Qwen3.5 MTP support (#2417)** (76bb18f)
- **feat(bagel): add multimodal Bagel training support (#2275)** (809270d)
- **feat(examples): add Falcon H1 fine-tuning recipes (#2334)** (e971fdf)
- **feat(speculative): add target_attn_implementation knob for EAGLE-3 target (#2415)** (1af381d)
- **feat(models): support fused linear cross-entropy across custom models (#2397)** (e2b1227)
- **feat: mixture of mutiple ASR datasets training recipe (#2414)** (b42152d)
- **feat(dllm): add Qwen3-4B dflash recipe, surface FSDP2 prefetch knobs (#2412)** (0bc85e2)
- **feat(speculative): log EAGLE training metrics to Weights & Biases (#2408)** (e7634b1)
- **feat(speculative): add P-EAGLE sequence partitioning for long-context training (#2409)** (ef2f1b3)
- **feat(speculative): add DFlash draft-model training recipe (#2406)** (8252101)
- **feat(speculative): log the draft model summary in EAGLE training (#2407)** (62d5177)
- **feat(speculative): add P-EAGLE parallel-drafting training for EAGLE-3 drafts (#2376)** (16983de)
- **feat: Enable cycling through all positive documents in biencoder training #907 (#933)** (f2ef677)
- **feat(speculative): add step-based and final checkpointing to EAGLE recipes (#2405)** (0fb503e)
- **feat: staging (#2411)** (8c262d9)
- **feat(speculative): add remote target serving for EAGLE-3 training (#2398)** (6aa5e5b)
- **feat(speculative): add gpt-oss EAGLE-3 draft model (#2399)** (f56e92b)

### Fixes
- **cp: `fix(ci): stabilize failed benchmark recipes (2817)` into `r0.5.0` (#2818)** (0f282dd)
- **cp: `fix(optim): align Dion mesh with FSDP sharding (2808)` into `r0.5.0` (#2812)** (13296ba)
- **cp: `fix(ci): drop base-image uv/wandb copies flagged for CVEs (2800)` into `r0.5.0` (#2803)** (d5081c2)
- **cp: `fix: qwen3.5 and 3.6 mtp expert checkpoint layout (2778)` into `r0.5.0` (#2795)** (d67da2d)
- **cp: `fix(vlm): keep Qwen3.5 media tokens aligned (2772)` into `r0.5.0` (#2793)** (42f017f)
- **cp: `fix(ci): HybridEP bench + LoRA OOM fixes (2789)` into `r0.5.0` (#2794)** (956ff1e)
- **cp: `fix(ci): stabilize diffusion finetune smoke tests (2788)` into `r0.5.0` (#2791)** (1653a29)
- **cp: `fix(ci): address go-git/go-billy and rustls-webpki CVEs (2780)` into `r0.5.0` (#2787)** (63d6802)
- **cp: `fix: skip fused LoRA MLP install for meta weights (2775)` into `r0.5.0` (#2781)** (7c3db40)
- **cp: `fix: Qwen3.5 MedPix EP32 NCCL timeout (2777)` into `r0.5.0` (#2785)** (a0564b3)
- **cp: `fix(benchmark): skip unsupported MTP flops (2767)` into `r0.5.0` (#2784)** (3b3f52d)
- **cp: `fix: Remove dali from container (2770)` into `r0.5.0` (#2771)** (4f2b883)
- **cp: `fix(distributed): use flattened CP FSDP mesh (2768)` into `r0.5.0` (#2769)** (3160010)
- **cp: `fix(sdpa): apply resolved backend constraints to custom models (2761)` into `r0.5.0` (#2765)** (9617f8f)
- **cp: `fix(mistral3): preserve medium VLM checkpoint layout (2758)` into `r0.5.0` (#2762)** (24ecc7e)
- **cp: `fix(diffusion): reuse warm HF cache instead of re-downloading models (2747)` into `r0.5.0` (#2754)** (8751227)
- **cp: `fix(diffusion): raise qwen-image dist timeout for checkpoint consolidation (2748)` into `r0.5.0` (#2756)** (79b9b31)
- **cp: `fix(fsdp2): guard uninitialized accumulated grads (2744)` into `r0.5.0` (#2752)** (84e8579)
- **cp: `fix(qwen3_moe): step-0 NaN in MXFP8 packed finetune — expert unload + fused RoPE (2722)` into `r0.5.0` (#2753)** (b5aa132)
- **cp: `fix(ci): reduce mixtral release smoke batch (2728)` into `r0.5.0` (#2751)** (76614e7)
- **cp: `fix(deepseek-v4): avoid bf16 -inf overflow in additive attention mask (2658)` into `r0.5.0` (#2746)** (4b262fd)
- **cp: `fix(deepseek-v4): restore batch axis for packed-sequence (THD) forward (2651)` into `r0.5.0` (#2745)** (3651b94)
- **cp: `fix(ci): pin qwen3_moe_30b mxfp8 finetune to gb200 (2735)` into `r0.5.0` (#2737)** (de509da)
- **cp: `fix(qwen3_5): handle packed MTP attention (2727)` into `r0.5.0` (#2729)** (7dd335b)
- **fix(gemma4): avoid DynamicCache OOM on dense E2B/E4B via kv-share holder (#2725)** (d90462e)
- **cp: `fix(models): use bool sparse masks for sdpa (2624)` into `r0.5.0` (#2721)** (3016d91)
- **cp: `fix(checkpoint): super-49B consolidated reload and vllm_deploy (2626)` into `r0.5.0` (#2718)** (ca46223)
- **cp: `fix(moe): handle non-EP expert weight DTensors (2697)` into `r0.5.0` (#2712)** (7bec0d7)
- **cp: `ci: fix 26.06 release cves (2705)` into `r0.5.0` (#2711)** (86d70de)
- **cp: `fix(gemma4_moe): re-tie lm_head to active embed_tokens on MoE path (2601)` into `r0.5.0` (#2709)** (c6601c9)
- **cp: `fix(loss): reuse LM head gather for MTP loss (2694)` into `r0.5.0` (#2708)** (1cfbabc)
- **cp: `fix(mistral3): remap FP8 VLM checkpoint prefixes (2692)` into `r0.5.0` (#2702)** (ff9bffd)
- **cp: fix(llama3_3) (#2673) to r0.5.0 (#2684)** (9da75ee)
- **cp: `fix(vlm): bump mistral3p5_128b_medpix max_length 1024->2048 (2689)` into `r0.5.0` (#2693)** (ba8123d)
- **cp: fix(qwen3_moe) (#2687) to r0.5.0 (#2688)** (b708e90)
- **cp: fix(checkpoint) (#2682) to r0.5.0 (#2685)** (7dedcd5)
- **cp: fix(training) (#2672) to r0.5.0 (#2679)** (8115f09)
- **cp: `fix(docker): bump DeepEP to `42144303` to pad HybridEP token capacity (2678)` into `r0.5.0` (#2680)** (7835dcd)
- **cp: `fix(recipe): disable fused RoPE for MLA packed-sequence MoE recipes (2675)` into `r0.5.0` (#2677)** (7239598)
- **cp: fix(model) (#2657) to r0.5.0 (#2671)** (d94e911)
- **cp: `fix(moe): preserve fp32 A_log in Qwen3.5-{MoE,Next GatedDeltaNet} (2484)` into `r0.5.0` (#2664)** (293ae9f)
- **cp: fix(models) (#2652) to r0.5.0 (#2669)** (af6f203)
- **cp: fix(distributed) (#2655) to r0.5.0 (#2668)** (adbc657)
- **cp: fix(datasets) (#2649) to r0.5.0 (#2665)** (fd4453c)
- **cp: fix(merge_lora) (#2653) to r0.5.0 (#2667)** (35cb928)
- **cp: `feat(moe): MTP FLOPs accounting fix (2486)` into `r0.5.0` (#2660)** (92da994)
- **cp: `fix(examples): enable ac for phi_4_squad (2634)` into `r0.5.0` (#2661)** (631a49f)
- **cp: `fix(devstral2,ministral3): load FP8 checkpoints via custom mistral3_vlm path (drop HF FineGrainedFP8) (2654)` into `r0.5.0` (#2666)** (61cdd05)
- **cp: `fix(datasets): decode MedPix images on demand instead of up front (2645)` into `r0.5.0` (#2650)** (5e4bcdc)
- **cp: `fix(parallelizer): resolve NemotronH decoder blocks for Nemotron-V3 (2638)` into `r0.5.0` (#2642)** (302b3ed)
- **cp: `fix(moe): default ignore_router_for_ac=True for activation checkpointing (2635)` into `r0.5.0` (#2636)** (9f01cfd)
- **cp: `fix(config): validate pp_size against distributed.pipeline (2616)` into `r0.5.0` (#2631)** (7228549)
- **cp: `fix(qwen3_moe): keep native forward under PP so CP+THD works (2625)` into `r0.5.0` (#2628)** (f05c15c)
- **cp: `fix(docker): build DeepEP against the NVSHMEM wheel matching the apt runtime (2614)` into `r0.5.0` (#2629)** (954adfb)
- **cp: `fix(transformers): keep gemma3n KV sharing working under FSDP2 (AM-454) (2594)` into `r0.5.0` (#2619)** (815309c)
- **cp: `fix(bagel): distributed setup init (2608)` into `r0.5.0` (#2613)** (263b4e6)
- **cp: `fix(qwen35moe):convert MTP experts as grouped(AM-442)(2595)` into r0.5.0 (#2618)** (b187d4b)
- **cp: `fix(moe): weight GroupedExpertsTE down-projection bias by routing probability (2591)` into `r0.5.0` (#2610)** (dfd7ed0)
- **cp: `fix(oom): use FusedLinearCrossEntropy in qwen3 tulu3 configs to avoid OOM (2609)` into `r0.5.0` (#2612)** (7a4f3c3)
- **cp: `fix: use TE attention for gpt_oss packed-sequence recipe (AM-438) (2587)` into `r0.5.0` (#2611)** (bfc3548)
- **cp: `fix(models): keep RoPE frequency buffers fp32 under bf16 model cast (2549)` into `r0.5.0` (#2606)** (dae417f)
- **cp: `fix(distributed): register Falcon-H1 TP plan to fix 34B PEFT OOM (2589)` into `r0.5.0` (#2605)** (e4196bd)
- **cp: `fix(vlm): use FusedLinearCrossEntropy for qwen3_5_9b to avoid logits OOM (2603)` into `r0.5.0` (#2604)** (9baa42a)
- **cp: `fix(gemma4): FSDP2-safe kv-sharing + skip frozen audio tower on grad-accum (2566)` into `r0.5.0` (#2599)** (dffc89f)
- **cp: `fix(vlm): enable activation checkpointing for 35B Qwen3.5/3.6 VLM recipes (2600)` into `r0.5.0` (#2602)** (e588588)
- **cp: `fix(peft): LoRA MLP QLoRA/PP/gemma3n fixes (AM-435, AM-447, AM-453) (2584)` into `r0.5.0` (#2597)** (fa3a1f1)
- **cp: `fix(test): load checkpoint-robustness HF reference via device_map (2582)` into `r0.5.0` (#2588)** (cfb034c)
- **cp: `fix(recipe): reshard MoE experts after forward in nemotron_nano_v3_cp_test (2577)` into `r0.5.0` (#2580)** (1d5529b)
- **cp: `fix(ci): set node counts for multi-node VLM finetune recipes (2574)` into `r0.5.0` (#2579)** (9799a9f)
- **cp: `fix(checkpoint): preserve tied lm_head on resume (2511)` into `r0.5.0` (#2573)** (94269ea)
- **cp: `test: fix all 5 vllm_deploy tests (token drift, nemotron OOM + mamba merge) (2559)` into `r0.5.0` (#2576)** (2aa3ec8)
- **cp: `fix(ci): bump ling_1t_lora_pp local_batch_size to satisfy PP assert (2575)` into `r0.5.0` (#2578)** (d76d109)
- **cp: `fix(diffusion): resolve flux nightly CI failures (2529)` into `r0.5.0` (#2567)** (c267b55)
- **cp: `fix(qwen3_5): make dense VLM pipeline-parallel safe (2524)` into `r0.5.0` (#2554)** (7b4d66a)
- **cp: `fix: unwrap ModelOutput to extract logits (2523)` into `r0.5.0` (#2530)** (8196ce6)
- **cp: `fix(gemma4): cast dense params without casting buffers (2359)` into `r0.5.0` (#2525)** (1389b2a)
- **cp: `fix(config): glm4.7 yaml (2527)` into `r0.5.0` (#2528)** (360a70e)
- **fix(moe): include MTP modules in FSDP sync traversal (#2441)** (05f8707)
- **fix(speculative): embed d2t/t2d vocab remap in EAGLE-3 draft checkpoint (#2447)** (08bab25)
- **fix(checkpoint): exclude TE _extra_state keys from load-time mismatch warning (#2247)** (99187d0)
- **fix(precision): dtype contract bug fixes for FSDP2 mixed-dtype loads (#2419)** (65691cb)
- **fix(deepseek_v3): initialize weights in fp32 and default router to fp32 (#2450)** (df3f42a)
- **fix: fp32 master weights for custom MoE models under FSDP2 (#1896)** (80f212e)
- **fix(tokenizer): make NeMoAutoTokenizerWithBosEosEnforced picklable (#2439)** (7af6e0e)
- **fix(speculative): guard PEAGLE flex attention compile (#2443)** (b84db2f)
- **perf(diffusion): optimize Wan2.1 finetuning recipes (#2403)** (0230cce)
- **perf(speculative): add fused Triton soft cross-entropy kernel for EAGLE-3 (#2428)** (ccf537a)

### Backend
- **cp: docs: document opt-in media extras (vlm-media/diffusion-media) (#2799) (#2848)** (d02f49c)
- **cp: `build(deps): move ffmpeg/opencv deps to opt-in media extra (2743)` into `r0.5.0` (#2792)** (d233564)
- **cp: `ci: run dsv32_lora, kimi_k2 and qwen3_moe_235b deepep benchmarks online (2773)` into `r0.5.0` (#2774)** (693d7ef)
- **cp: `ci: bump benchmark glm_4.7_flash_te_deepep time (2757)` into `r0.5.0` (#2759)** (56fe9a6)
- **cp: `build: install TileKernels for DeepSeek V4 (2740)` into `r0.5.0` (#2750)** (2a5f9ac)
- **cp: `ci: address thrift cve bump to 0.23.0 (2736)` into `r0.5.0` (#2742)** (ffe28f6)
- **cp: `ci: address diffusers cve (2706)` into `r0.5.0` (#2738)** (61dfa4d)
- **cp: DeciLM Nemotron TP plan (#2703) to r0.5.0 (#2726)** (f2cdc78)
- **cp: `build: install tilelang + tile_kernels for DeepSeek-V4 recipes (2683)` into `r0.5.0` (#2717)** (62b7390)
- **cp: `perf(checkpoint): mmap HF DCP read_data to avoid host-RAM OOM on large loads (2690)` into `r0.5.0` (#2701)** (3104ed7)
- **cp: `test(models): speed up qwen3.5 moe/vl-moe from_pretrained unit tests (2698)` into `r0.5.0` (#2699)** (879baed)
- **cp: `feat(config): disable W&B in example configs (2643)` into `r0.5.0` (#2662)** (bb6e3f5)
- **cp: `feat(qwen3_5): port dense Qwen3.5 to a custom-model(2557)` into `r0.5.0` (#2663)** (bf7f98b)
- **cp: `refactor(moe): remove enable_deepep, switch failing ep recipes to hybridep (2630)` into `r0.5.0` (#2644)** (2494607)
- **cp: `ci: cap MAX_STEPS to 10 for slow vlm_finetune recipes (2639)` into `r0.5.0` (#2641)** (5143e1b)
- **cp: `ci: raise ci.time for slow finetune recipes hitting 10-min default (2637)` into `r0.5.0` (#2640)** (4be5f0c)
- **cp: `ci: Enable activation checkpointing for gemma_2_9b_it_squad (AM-464) (2585)` into `r0.5.0` (#2586)** (0f667e5)
- **cp: `ci: use digits for spark recipes (2581)` into `r0.5.0` (#2583)** (e8af076)
- **cp: `ci: schedule ep-parallel finetune recipes at documented node counts (2546)` into `r0.5.0` (#2558)** (3da763f)
- **cp: `feat(model): flux2 (2145)` into `r0.5.0` (#2489)** (c45d323)
- **cp: `feat(vlm): enable Qwen3.5 MoE VLM CP (2432)` into `r0.5.0` (#2483)** (6a436d9)
- **cp: `feat: make mesh accept meshcontext (2266)` into `r0.5.0` (#2474)** (7390a64)
- **cp: `ci: update package version to 0.5.0 (2472)` into `r0.5.0` (#2473)** (4b5fba0)

### Docs
- **docs(speculative): add subsystem README, fold in regeneration guide (#2448)** (3822f9e)
- **docs(fern): add Nemotron-3-Ultra-550B fine-tuning guide (#2420)** (5d1a7cb)

### Chore
- **refactor(speculative): reuse shared dflash mask and loss in the trainer (#2433)** (3675f86)
- **ci: add nemo-run, split qwen-vl-utils from decord for arm (#2456)** (5bb1e61)
- **chore(ci): update codeowners to use NVIDIA-NeMo/core-am (#2453)** (a20d9c5)
- **refactor(datasets): unify reasoning_content coercion in agent chat (#2440)** (873fe27)
- **chore(skills): refresh distributed training signature (#2438)** (355f526)
- **refactor: expose shared recipe builders from components (#2190)** (1f4f83e)
- **build: add managed = true to [tool.uv] (#2434)** (e2c8895)
- **refactor(speculative): decouple P-EAGLE from the EAGLE-3 code path (#2429)** (cbddd0b)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/NVIDIA-NeMo/Automodel?utm_source=github-action)._