## NVIDIA-NeMo/Automodel — v0.5.0…v0.6.0

_498+ commits._

### Features
- **cp: `feat(gemma4): add E-series tensor parallelism (3512)` into `r0.6.0` (#3552)** (10b08ee)
- **cp: `feat: Add {% generation %} chat template for DiffusionGemma SFT/LoRA examples (3353)` into `r0.6.0` (#3417)** (592a7f7)
- **cp: `perf(moe): add scoped partial CUDA graphs (2917)` into `r0.6.0` (#3314)** (1e3e20b)
- **cp: `ci: add Tulu-3 convergence + eval flow (3171)` into `r0.6.0` (#3312)** (2ca0893)
- **feat: add ViSpec VLM draft training (#3173)** (eec041f)
- **feat(checkpoint): defer distributed async consolidation (#3125)** (d60eb60)
- **feat(vlm): integrate packed CP and vision sharding for Qwen3.5-MoE (#3186)** (fb0a7ed)
- **feat: add Kimi K3 model support (#3259)** (8859d3c)
- **feat: support pre-extracted frame sequence of video sample (#3211)** (a3fe65f)
- **feat(diffusion): add LTX-2.3 video+audio finetuning (training + prepr… (#3165)** (5bf1a7a)
- **feat(distributed): frame-level context-parallel vision-tower sharding (#2990)** (69899bb)
- **feat(diffusion): add Qwen-Image-Edit-2511 training support (#3217)** (1d60a72)
- **feat(speculative): add the multimodal speculative decoding (MSD) core (#3166)** (dab67fc)
- **feat(diffusion): context parallelism via diffusers ContextParallelConfig (#3157)** (24a1e2d)
- **feat(retrieval): add normalized Arrow dataset tooling (#2596)** (788c1cf)
- **feat(glm_moe_dsa): expose update_moe_gate_bias on GLM MoE DSA models (#3207)** (622cc24)
- **feat(dllm): add DiffusionGemma generation via the built-in HF sampler (#3161)** (380423c)
- **feat(kernels): add QuACK backend (#3115)** (4027eb1)
- **feat(checkpoint): preserve intrinsic fp32 during offline consolidation (#3193)** (d412564)

### Fixes
- **fix(release): set version to 0.6.0 (#3692)** (89c248a)
- **cp: `fix(peft): support mixed-dtype memory-efficient LoRA backward (3675)` into `r0.6.0` (#3682)** (e20ca02)
- **cp: `fix(deps): resolve 26.08 rc10 container CVEs (3647)` into `r0.6.0` (#3658)** (922e83e)
- **cp: `fix(deps): resolve 26.08 rc9 container CVEs (3607)` into `r0.6.0` (#3611)** (94387e4)
- **cp: `fix(deps): resolve msgpack, wandb, and mistune container CVEs (3563)` into `r0.6.0` (#3566)** (8bf69b4)
- **cp: `fix(fsdp): uniform reduce dtype and EP-local expert gradients (3540)` into `r0.6.0` (#3561)** (448bf50)
- **cp: `fix(docker): build bitsandbytes for SM121 (3553)` into `r0.6.0` (#3557)** (c14ec3b)
- **cp: `fix(retrieval): support canonical Sentence Transformers metadata (3546)` into `r0.6.0` (#3551)** (85dfd54)
- **fix(minimax): declare packed-sequence and CP attention ownership (#3548)** (fa81d57)
- **cp: `fix(deps): resolve Starlette and GitPython CVEs (3523)` into `r0.6.0` (#3524)** (b2eb9c3)
- **cp: `fix(checkpoint): preserve FSDP2 mixed precision during recompute (3513)` into `r0.6.0` (#3515)** (d7f43c2)
- **cp: `fix(deepseek_v4): avoid TileLang boolx8 backward codegen (3467)` into `r0.6.0` (#3506)** (1a35152)
- **cp: `ci(convergence): fix gemma4 eval setup and re-baseline Qwen3-MoE (3493)` into `r0.6.0` (#3505)** (2eb0edc)
- **cp: `fix(checkpoint): validate PEFT adapter-only state (3501)` into `r0.6.0` (#3503)** (6aa8cd7)
- **cp: `fix(models): preserve lm-head dtype boundaries (3491)` into `r0.6.0` (#3502)** (3948d66)
- **perf(recipes): run GPT-OSS 120B with EP64 and no activation checkpointing (#3492)** (fedc908)
- **cp: `fix(ci): register flux2/wan2.2/qwen-image-edit diffusion recipes in CI (3485)` into `r0.6.0` (#3487)** (9ae8148)
- **cp: `fix(recipe): shard GPT-OSS 120B across 64 experts (3483)` into `r0.6.0` (#3489)** (94d6b8b)
- **cp: `fix(training): all-reduce grad-norm scalars on the mesh device (3461)` into `r0.6.0` (#3481)** (b58e52c)
- **cp: `fix(peft): align frozen LoRA tensors with compute dtype (3470)` into `r0.6.0` (#3480)** (b2fdcf8)
- **cp: `fix(checkpoint): save ranked RNG state per global rank (3437)` into `r0.6.0` (#3471)** (d4b9e8d)
- **cp: `fix(peft): enable v5 expert adapters for Nemotron and MiniMax (3439)` into `r0.6.0` (#3468)** (6a26a13)
- **cp: `fix(bagel): type SFT backend before AutoModel init (3449)` into `r0.6.0` (#3469)** (a4dae84)
- **cp: `fix: normalize MoE auxiliary loss during gradient accumulation (3359)` into `r0.6.0` (#3459)** (684c42a)
- **cp: `fix(gemma4): exact router scalar and fp32 reference routing for MoE parity (3456)` into `r0.6.0` (#3466)** (f859c91)
- **cp: `fix(retrieval): repair Ministral3 training recipe (3441)` into `r0.6.0` (#3454)** (efea0a4)
- **cp: `fix: Gemma4-31B CoderForge CP8/64K data processing and recipe (3446)` into `r0.6.0` (#3452)** (14a85a1)
- **cp: `fix(kimi_k25_vl): keep PEFT expert LoRA keys under the base_model.model. prefix (3431)` into `r0.6.0` (#3451)** (205e038)
- **cp: `fix(distributed): avoid duplicate FSDP2 prefetch all-gathers (3411)` into `r0.6.0` (#3445)** (edea1ca)
- **cp: `fix(retrieval): use stock Ministral embedding backbone (3103)` into `r0.6.0` (#3433)** (e5bede8)
- **cp: `fix(transformers): default missing THD capability (3406)` into `r0.6.0` (#3429)** (8204ef7)
- **cp: `fix(ci): run Kimi K3 HellaSwag on GB200 (3420)` into `r0.6.0` (#3421)** (056cb05)
- **cp: `fix(docs): AUT-1324 qualify LTX model coverage slug (3418)` into `r0.6.0` (#3419)** (0beb4a4)
- **cp: `fix(gemma4): enable expandable_segments on the CP tulu3 recipes (3408)` into `r0.6.0` (#3410)** (6756fd8)
- **cp: `fix(retrieval): guard optional W&B imports (3380)` into `r0.6.0` (#3402)** (89ee771)
- **cp: `fix(checkpoint): harden PEFT PP checkpointing and Step-3.7 coverage (3316)` into `r0.6.0` (#3403)** (9c21948)
- **cp: `fix(moe): support EP-free DTensor state-dict conversion (3397)` into `r0.6.0` (#3409)** (e3b5efa)
- **cp: `fix(deps): resolve OSS CVE findings (3398)` into `r0.6.0` (#3412)** (e4376a6)
- **cp: `fix(checkpoint): survive an interrupted save instead of hanging all ranks (3261)` into `r0.6.0` (#3394)** (9d4713d)
- **cp: `fix(model): correct Mistral4 attention and distributed MoE routing (3348)` into `r0.6.0` (#3393)** (3ab8c21)
- **cp: `fix(ci): enable LTX-2.3 diffusion finetuning (3372)` into `r0.6.0` (#3391)** (f5369a6)
- **cp: `fix(fsdp): resolve fp32 master-weight compute dtype per parameter (3328)` into `r0.6.0` (#3384)** (a0daec9)
- **cp: `fix(ci): use a single container cache donor (3377)` into `r0.6.0` (#3381)** (61cf7ee)
- **cp: `fix(distributed): stabilize SAC replay with TE and FSDP (3330)` into `r0.6.0` (#3376)** (f517ee2)
- **cp: `fix(tokenizer): make special token insertion opt-in (3337)` into `r0.6.0` (#3375)** (e3fa641)
- **cp: `fix(diffusion): use spawn start method for GPU preprocessing pools (3339)` into `r0.6.0` (#3366)** (80386d0)
- **cp: `fix(kimi_k25_vl): don't int4 quantize LoRA adapter keys on save (3295)` into `r0.6.0` (#3362)** (30f53c8)
- **cp: `fix(pp): route Qwen3.5 MoE pre-embedded inputs (3294)` into `r0.6.0` (#3364)** (e544802)
- **cp: `fix(deps): resolve 26.08 rc2 container CVEs (3346)` into `r0.6.0` (#3363)** (47413a2)
- **cp: `fix(pp): preserve VLM media cursor with static metadata (3344)` into `r0.6.0` (#3361)** (dc4dddb)
- **cp: `fix(docs): make Fern autodoc metadata MDX-safe (3355)` into `r0.6.0` (#3357)** (0e35479)
- **cp: `fix(training): prewarm Mamba SSD autotune kernels (3296)` into `r0.6.0` (#3338)** (5f2d4d6)
- **cp: `fix(distributed): reuse default group for world-sized meshes (3319)` into `r0.6.0` (#3347)** (015e5ae)
- **cp: `fix(nemotron-parse): sync RADIO preprocessing (3331)` into `r0.6.0` (#3341)** (956b535)
- **cp: `fix(distributed): restore Nemotron Flash TP2 training (3345)` into `r0.6.0` (#3349)** (cdf67ad)
- **cp: `fix(kd): use mesh-safe gradient clipping (3302)` into `r0.6.0` (#3334)** (9b80ea3)
- **cp: `fix(kd): preserve tensor-valued hidden states (3324)` into `r0.6.0` (#3332)** (00c9701)
- **cp: `fix(config): repair GLM-5.2 LoRA recipe (3313)` into `r0.6.0` (#3323)** (0189fb5)
- **cp: `fix(docs): avoid literal ampersand in Kimi autodoc (3317)` into `r0.6.0` (#3320)** (77e86ad)
- **cp: `fix(distributed): preserve ERNIE router and dense Qwen3.5 SSM precision (3255)` into `r0.6.0` (#3310)** (89842b6)
- **cp: `fix(pp): use static metadata with PyTorch 2.13 (3290)` into `r0.6.0` (#3305)** (7e9493d)
- **cp: `fix(vlm): fused linear CE in gemma4 31B FFPA 8k recipe (AMINT-203) (3273)` into `r0.6.0` (#3307)** (21b2050)
- **cp: `fix(docs): hide Kimi tokenizer regex from autodoc (3288)` into `r0.6.0` (#3304)** (755ecde)
- **fix(KD): Fix variable-length teacher PP batches on separate meshes (#3280)** (d724c2a)
- **fix(inkling): support HF 5.14 attention fields and embed norm FSDP (#3281)** (e111add)
- **fix(docker): cap uv install concurrency to avoid ARM FD exhaustion (#3275)** (a664ea4)
- **fix(ci): install CUDA torchvision with CUDA torch (#3279)** (acd8060)
- **fix(checkpoint): save all PEFT EP optimizer parts (#3250)** (3e795ab)
- **fix(checkpoint): harden PP safetensors consolidation (#3248)** (332b2a1)
- **fix: harden checkpoint torch loads (#3240)** (6462b78)
- **fix(vlm): preserve mRoPE axes in PP chunking (#3208)** (d3f3b2b)
- **fix(ci): narrow CUDA wheelhouse cache keys (#3269)** (6e77221)
- **fix(moe): preserve gate load across activation recompute (#3247)** (15d1ac2)
- **perf(benchmark): tune Qwen3-VL LoRA pipeline batches (#3230)** (34a4d0c)
- **fix(gpt-oss): route packed attention through THD (#3226)** (422ac0d)
- **fix(moe): preserve top-k routing under full AC (#3140)** (4e49fb0)
- **fix(nemotron-v3): recompute deterministic LoRA router (#3258)** (5a350eb)
- **fix(thd): derive packed padding mask from the pack layout, not token ids (#3223)** (57a43e2)
- **fix(deepseek-v4): preserve fp32 model dtype (#3227)** (2d9397f)
- **fix(deepseek-v4): accept cu_seqlens for THD packing (#3231)** (e9ef66b)
- **fix(ci): omit token type IDs from vLLM parity prompts (#3198)** (75e2e8d)
- **fix(ci): preserve main wheelhouse cache builds (#3253)** (987af05)
- **fix(distributed): control frozen multimodal FSDP sharding (#2763)** (a8e9ce7)
- **fix(checkpoint): infer Qwen MTP layout from checkpoint keys (#3229)** (6230da4)
- **fix(speculative): validate mask_token_id when resuming DSpark (#3238)** (f59b242)
- **fix(speculative): align the decode-eval cadence on resume (#3237)** (2dafc38)
- **fix(speculative): honor --dflash-causal on the vLLM remapping export (#3235)** (daabc6c)
- **fix(speculative): average DFlash train metrics over the log window (#3234)** (ab7fbf7)
- **fix(glm): own the GLM MoE DSA config so qk_rope_head_dim survives (#3222)** (a3aa09b)
- **fix(distributed): checkpoint linear attention with compile (#3213)** (bce392b)
- **fix(transformers): custom configs override builtin ones by default (#3221)** (a3570d4)
- **perf(bagel): grouped MoT routing + fused SwiGLU/RoPE (#3214)** (9e0ebe0)
- **fix(dllm): recipe defaults, corruption seeding, sampler and docs fixes (#3162)** (0ff7020)
- **fix(distributed): support packed CP for Llama, Qwen2, and Qwen3 (#2999)** (3ebd99a)
- **fix(distributed): activation-checkpoint Qwen3-Next linear_attn layers (#3192)** (d8c354b)
- **fix(glm): size GLM5.2 release CI jobs (#3210)** (bd592f8)
- **fix(ci): shard Nemotron Super vLLM deploy across 8 GPUs (#3061)** (35b1fca)
- **fix(vlm): size qwen3.6 medpix CI configs (#3209)** (37298f4)
- **fix(dist): keep profiler record-function ops out of SAC replay accounting (#3133)** (dfd1401)

### Backend
- **beep boop 🤖: Bumping NeMo-Automodel to v0.5.1 [skip ci]** (c93ffca)
- **cp: backport packed THD VLM context parallelism to r0.6.0 (#3558)** (d1ccd9d)
- **cp: `ci: route DGX Spark recipes to GB10 (3538)` into `r0.6.0` (#3556)** (859b68e)
- **cp: feat(models): make Inkling standalone (#3358) into r0.6.0 (#3498)** (b12ec20)
- **cp: `chore: Bump gitpython to >= 3.1.59 (3482)` into `r0.6.0` (#3488)** (a7a0543)
- **cp: `perf(benchmarks): recompute deterministic MoE routers under AC (3474)` into `r0.6.0` (#3490)** (2d1ab88)
- **cp: backport VLM CP gradient and vision sharding fixes to r0.6.0 (#3479)** (74de94a)
- **cp: backport Sentence Transformers metadata export to r0.6.0 (#3464)** (c299a3d)
- **cp: `test(checkpoint): stabilize MoE checkpoint parity gates (3440)` into `r0.6.0` (#3460)** (ef3794b)
- **cp: `test(checkpoint): redesign resume robustness around a shared trajectory (3427)` into `r0.6.0` (#3457)** (503d5de)
- **cp: backport PEFT v5 adapter output fixes to r0.6.0 (#3458)** (3781b1a)
- **cp: `ci(vlm): raise Slurm wall time for MiniMax-M3 and Gemma4 recipes (3448)` into `r0.6.0` (#3453)** (b809ebc)
- **cp: `ci(minimax): extend M2.7 LoRA timeout (3428)` into `r0.6.0` (#3430)** (9b78650)
- **cp: backport MoE checkpoint reload parity fixes to r0.6.0 (#3438)** (d89695e)
- **cp: `ci(convergence): (3416)` into `r0.6.0` (#3432)** (edf3a0f)
- **cp: `docs(training): mark Mamba prewarm sections for review (3340)` into `r0.6.0` (#3399)** (ba157a2)
- **cp: `perf: remove Python overhead from model hot paths (3374)` into `r0.6.0` (#3401)** (e9b3c7d)
- **cp: `ci: tune Nemotron single-GPU model load threads (3370)` into `r0.6.0` (#3395)** (5d5f70c)
- **cp: `perf(checkpoint): reduce distributed save overhead (3369)` into `r0.6.0` (#3389)** (8e6acff)
- **cp: `refactor(docker): build torchao & FlashAttention as isolated wheel stages (3329)` into `r0.6.0` (#3343)** (73e8dc4)
- **cp: `docs(distributed): review frozen multimodal FSDP guidance (3272)` into `r0.6.0` (#3306)** (2062c42)

### Tests
- **test(retrieval): add Nemotron VL checkpoint coverage (#3276)** (2751329)
- **test(vlm): stabilize Qwen3.5 checkpoint robustness (#3233)** (16b801e)

### Docs
- **docs(models): add Kimi K3 coverage (#3283)** (758f04c)
- **docs: fix remaining broken link sources (#3249)** (30bbb5b)
- **docs(fern): cut v0.5.0 version train (#2970)** (07a0edf)
- **docs: explain embedding row repair (#3175)** (f5e94f0)

### Chore
- **chore(ci): AUT-1135 pin GitHub Actions to commit SHAs (#3277)** (cdf6b63)
- **build: add MagiAttention optional dependency (#3070)** (5282a5f)
- **build: cache CUDA extension source builds (#3204)** (8124b58)
- **build(deps): bump ffpa-attn to 0.2.2 (#3225)** (1edffb8)
- **refactor(diffusion): migrate recipe onto typed RecipeConfig build() path (#3122)** (fc57de7)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/NVIDIA-NeMo/Automodel?utm_source=github-action)._