## NVIDIA-NeMo/Automodel — v0.3.0…v0.4.0

_374+ commits._

### Features
- **cp: ci: add known_issue_id / allow_failure keys + triage (#2028) to r0.4.0 (#2033)** (6c61d3f)
- **cp: `ci: add --tb=short to pytest invocations in CI test scripts (2018)` into `r0.4.0` (#2019)** (90edb21)
- **cp: `ci: Support per-recipe env_vars in CI config (1999)` into `r0.4.0` (#2000)** (4a23c9d)
- **cp: `chore: add @zyzhou5 and @athitten to codeowners (1968)` into `r0.4.0` (#1969)** (f9b19e5)
- **cp: `docs: Add container version to docs version picker (1965)` into `r0.4.0` (#1966)** (c6d2d1e)
- **cp: `ci: Add test_recipes for custom test scope (1915)` into `r0.4.0` (#1954)** (0748d6e)
- **cp: `ci: Add Dockerfile.deploy for deploy test environment (1804)` into `r0.4.0` (#1902)** (04eee44)
- **cp: `ci: add NMP customizer contract test configs (1712)` into `r0.4.0` (#1858)** (1baf6c2)
- **cp: `docs: Add nightly CI test summary for LLM and VLM finetune configs (1791)` into `r0.4.0` (#1792)** (95e3139)
- **cp: `ci: add missing recipe owners (1775)` into `r0.4.0` (#1776)** (7fce82d)
- **cp: `ci: Update test timeout and add ci_tests readme (1752)` into `r0.4.0` (#1759)** (914fdc2)
- **cp: `build: drop rc0 pre-release tag and add dynamic git versioning (1729)` into `r0.4.0` (#1756)** (d06f1bb)
- **cp: `test: add vLLM deployment tests ` into `r0.4.0` (#1745)** (83062e2)
- **cp: `feat: Add lora recipes for gemma4 (1731)` into `r0.4.0` (#1741)** (d81f478)
- **feat: enable packed sequences for Qwen3.5-MoE with EP+PP (#1685)** (263fcb1)
- **feat: implement NEFTune noisy embeddings for instruction fine-tuning (#1686)** (1c3944a)
- **feat: Add dllm generation support (#1692)** (15fb121)
- **feat: Allow to conditionally skip malformed jsonl lines when loading dataset (#1694)** (3eb5bc0)
- **feat: Add Llada SFT support (#1672)** (34097e4)
- **feat: add reasoning_content and tool-calling support to ChatDataset (#1644)** (e9bee96)
- **feat: add cp2 convergence configs and eval fixes (#1602)** (90f625f)
- **feat: context-parallel with nemotron v3 (#1441)** (4ebde17)
- **feat: Add discrete diffusion LLM (dLLM) supervised fine-tuning support (#1665)** (10a3b03)
- **feat: add HybridEP example config for Qwen3-30B-A3B (#1666)** (cefa53f)
- **feat: add gemma 4  (#1660)** (6606a0f)
- **feat: add gemma4 configs (#1658)** (c1b78f1)
- **feat: GPT-OSS 20B and Moonlight 16B convergence results (#1577)** (48d18d8)
- **feat: enable TE Linear layers for PEFT/LoRA (#1626)** (980f23d)
- **feat: add UCCL-EP as alternative dispatcher for expert parallelism (#1635)** (67b9bca)
- **feat: add missing recipe in yaml (#1642)** (fedc9d2)

### Fixes
- **cp: `fix: lora checkpointing (2037)` into `r0.4.0` (#2046)** (89bb1e7)
- **cp: `fix: gradient clip with torch_mm + EP (gpt-oss 120b recipe) (2012)` into `r0.4.0` (#2035)** (08cc896)
- **cp: `fix: regression in tokenizer+auto_map with transformers 5.5.0 (2025)` into `r0.4.0` (#2029)** (f25e374)
- **cp: `fix: add discover pp seq len (2024)` into `r0.4.0` (#2026)** (115f85c)
- **cp: `fix: switch from match_all_linear to target_modules (2022)` into `r0.4.0` (#2023)** (e1349bf)
- **cp: `fix: transformers v5.5.0 validation (2010)` into `r0.4.0` (#2013)** (689e408)
- **cp: `fix: change drop_long_samples to True by default (2009)` into `r0.4.0` (#2014)** (0671a08)
- **cp: `fix: batch Flash 1B + Super-49B PEFT + qwen2.5-7B ckpt-robustness (1984)` into `r0.4.0` (#2008)** (3a5143a)
- **cp: `fix: Address pillow CVE (1994)` into `r0.4.0` (#1998)** (670b269)
- **cp: `fix: Address ci timeout test from rc8 (1991)` into `r0.4.0` (#1993)** (5f392d8)
- **cp: `fix: Move benchmark recipe out of llm_finetune nightly (1989)` into `r0.4.0` (#1990)** (bbc231b)
- **cp: `fix: vllm deploy test should fail if vllm is not present (1987)` into `r0.4.0` (#1988)** (16c6fec)
- **cp: `fix(devstral): point 24B Squad recipes at official FP8 model (1980)` into `r0.4.0` (#1982)** (8d626d5)
- **cp: `fix: nemotron flash (1973)` into `r0.4.0` (#1978)** (0e881a4)
- **cp: `fix: batch ckpt-robustness fixes for pipeline 48953745 (supersedes 9 PRs) (1971)` into `r0.4.0` (#1979)** (0364ea8)
- **cp: `fix(vlm): qwen3_5_4b_neat_packing OOM - reduce seqlen to 4096 (1975)` into `r0.4.0` (#1977)** (81310d2)
- **fix: ministral tp plan (#1963) (#1974)** (303aa77)
- **cp: `fix: qlora ckpt loading (1549)` into `r0.4.0` (#1920)** (8f08318)
- **cp: `fix: Update gemm4 26b ci timeout (1962)` into `r0.4.0` (#1967)** (dfc9972)
- **cp: `fix: Patch wandb-core Go CVEs: bump otel SDK, add go-jose (1957)` into `r0.4.0` (#1964)** (555af67)
- **fix: Step-3.5-Flash layer_types mismatch and related recipe fixes (#1… (#1936)** (5f24def)
- **cp: `fix: AC silently skipped on all registered VLMs — flatten ModuleList  (1941)` into `r0.4.0` (#1958)** (f23c168)
- **cp: `fix: Update recipe test time based on release test run (1955)` into `r0.4.0` (#1956)** (7d1fb87)
- **cp: `fix: disable packed sequences for nemotron_nano_4b_squad (1929)` into `r0.4.0` (#1935)** (ada885b)
- **cp: `fix: make _get_logits pp aware in ckpt robustness (1923)` into `r0.4.0` (#1934)** (401c8e5)
- **cp: `fix: chat dataset (1921)` into `r0.4.0` (#1931)** (d85acd5)
- **cp: `fix: update defer_fsdp_grad_sync in recipes (1919)` into `r0.4.0` (#1930)** (d0d9105)
- **cp: `fix: Update recipe_owner for gemma4 (1925)` into `r0.4.0` (#1926)** (39fcc3b)
- **cp: `fix: Coerce plain-dict backend to BackendConfig in model init (1784)` into `r0.4.0` (#1803)** (38da59e)
- **cp:  `fix Qwen3.5+Phi4MM CI after transformers v5.5 update(1906)` into `r0.4.0` (#1908)** (97a500f)
- **cp: `fix: baichuan dynamic cache (1865)` into `r0.4.0` (#1909)** (67f0021)
- **cp: `fix: gradient checkpointing broken for MoE models on single GPU (ep_size=1) (1873)` into `r0.4.0` (#1879)** (285262f)
- **cp: `fix(gemma4_moe): vision-aware mask when use_bidirectional_attention==vision (1905)` into `r0.4.0` (#1907)** (7011ed4)
- **cp: `fix: pass unnormalized residual to MoE gate in Gemma4 decoder layer (1895)` into `r0.4.0` (#1903)** (5dd1393)
- **cp: `fix: Fix bug in diffusion generation (1850)` into `r0.4.0` (#1900)** (ab87c7f)
- **cp: `fix: relax KL thresholds and remove invalid kwargs in Qwen3Next linear attn (1867)` into `r0.4.0` (#1899)** (feb7d0e)
- **cp: `fix: Create diffusion_kernels group to fix HF_HUB_OFFLINE compatibility (1842)` into `r0.4.0` (#1886)** (7857667)
- **cp: `fix: handle transformers.FineGrainedFP8Config quantization config (1864)` into `r0.4.0` (#1888)** (cc78874)
- **cp: `fix: Allow use_cache w/ activation_checkpointing (1726)` into `r0.4.0` (#1760)** (7d794aa)
- **cp: `fix: Skip snapshot_download when HF_HUB_OFFLINE=1 (1834)` into `r0.4.0` (#1880)** (1006012)
- **cp: fix: Setup vllm testing with uv --no-config (#1875) (#1881)** (6c78230)
- **cp: `fix: gpt oss ci (1877)` into `r0.4.0` (#1878)** (7e8a323)
- **cp: `fix: pre-cache HF dynamic modules to prevent filesystem race in robustness test (1840)` into `r0.4.0` (#1856)** (c182044)
- **cp: `fix: trust_remote_code guard in robustness test (1845)` into `r0.4.0` (#1857)** (8c0e124)
- **cp: `fix: relax checkpoint robustness HF KL threshold for nemotron_nano_8b_v1 (1839)` into `r0.4.0` (#1855)** (b17b6ac)
- **fix: FSDP2 meta-device crash for Qwen3.5 GatedDeltaNet fp32 params (#1813)** (65d5526)
- **cp: `fix: rotary embeddings for v4 (1821)` into `r0.4.0` (#1851)** (0c1e059)
- **cp: `fix: install ffmpeg and rebuild torchcodec for phi4mm audio decoding (1826)` into `r0.4.0` (#1849)** (f955996)
- **cp: `fix: Re-apply PyTorch dependency overrides after full COPY in Dockerfile (1847)` into `r0.4.0` (#1848)** (835c06b)
- **cp: `fix: stop resolve_yaml_env_vars from scanning runtime data in instantiate() (1827)` into `r0.4.0` (#1836)** (f314a3d)
- **cp: `fix: gpt_oss_20b_single_gpu_peft CI crash with nproc_per_node override (1835)` into `r0.4.0` (#1844)** (48a8229)
- **cp: `fix: Align benchmark TEST_LEVEL check with generate_ci_tests scope (1831)` into `r0.4.0` (#1832)** (bd6f6f1)
- **cp: `fix: tie weights outside _init_model (1817)` into `r0.4.0` (#1829)** (caedd87)
- **cp: `fix: enable dequantization for ministral3 and dataset limit  (1807)` into `r0.4.0` (#1820)** (932a090)
- **cp: `fix: meta init with force_hf=True (1810)` into `r0.4.0` (#1822)** (5941d5a)
- **cp: `fix: Restrict auto-discovery scopes in generate_ci_tests.py (1805)` into `r0.4.0` (#1812)** (0d72395)
- **cp: `fix: NotImplementedError: aten::equal on meta tensors during multi-GPU init (1769)` into `r0.4.0` (#1797)** (e3517d8)
- **cp: `fix: Add per-tensor conv. in gemma4 sd adapter (1764)` into `r0.4.0` (#1795)** (b8c8a9c)
- **cp: `fix: update yamls for vllm_deploy (1780)` into `r0.4.0` (#1781)** (deb38a4)
- **cp: `fix: skip embedding[padding_idx] = 0 with TP (1675)` into `r0.4.0` (#1771)** (8cee6ad)
- **cp: `fix: launcher option from being a config override. (1766)` into `r0.4.0` (#1772)** (75a0770)
- **cp: `fix: Update lora configs for gemma4 (1748)` into `r0.4.0` (#1749)** (f304ecc)
- **cp: `fix: Baichuan2 checkpoint robustness test CI failures (1727)` into `r0.4.0` (#1754)** (4bdb284)
- **cp: `fix: Qwen3.5 dense CP support (1710)` into `r0.4.0` (#1737)** (58af018)
- **cp: `fix: fixing the pooling error  (1645)` into `r0.4.0` (#1736)** (9cd5299)
- **cp: `fix: handle dict-typed chat_template (1696)` into `r0.4.0` (#1716)** (9f2e48a)
- **cp: `fix: mute warning spam (1721)` into `r0.4.0` (#1724)** (721b7a9)
- **fix: freeze dead KV-sharing params to fix checkpoint resume (#1698)** (3fadac9)
- **fix: swap DTensor shard placements after transpose in Step3p5 state dict adapter (#1691)** (5018039)
- **fix: add best_metric_key field to CheckpointingConfig dataclass (#1641)** (4ecba76)
- **fix: move .claude/skills to skills (#1673)** (86c686b)
- **fix: add tp plan for phi2 (#1674)** (6f2643a)
- **fix: update gemma4 configs and doc with correct model IDs (#1670)** (7d9b3f7)
- **fix: Finetune DeepSeek V3 (issue #1496) (#1654)** (ca68ef7)
- **fix: Mistral4 FP8 dequant on multi-dim mesh (#1594)** (5792a9b)
- **fix: link in readme. (#1664)** (a3a59d7)
- **fix: move skills to .claude/skills (#1662)** (bd9e79c)
- **fix: Float32RMSNorm torch.compile crash on PyTorch 2.11+ (#1650)** (ec2f724)
- **fix: skip initialize_weights for Phi3ForCausalLM with TP sharding (#1648)** (784b1f8)

### Backend
- **cp: `docs: Bump docs version (2073)` into `r0.4.0` (#2074)** (ee69711)
- **cp: `ci: Update base container pillow version for cve (2065)` into `r0.4.0` (#2066)** (69cacee)
- **cp: `ci: triage vllm_deploy rc9 failures (2047)` into `r0.4.0` (#2050)** (9606fc7)
- **cp: `ci: triage rc9 finetune failures (2043)` into `r0.4.0` (#2049)** (2a8049e)
- **cp: `ci: triage pipeline benchmark failures (2040)` into `r0.4.0` (#2044)** (aae650d)
- **Cherry-pick #1728 to r0.4.0 (Qwen refs removed) (#2031)** (79674c9)
- **cp: `ci: Update test recipe list (2001)` into `r0.4.0` (#2002)** (3dd89d2)
- **cp: `resolve VLM CI failures for PP recipes and collate_fn(1799)` into `r0.4.0` (#1889)** (2d0fed2)
- **cp: `chore: move recipes to have perf CI/CD coverage (1885)` into `r0.4.0` (#1897)** (38d102f)
- **cp: `ci: Reduce default finetune step count from 100 to 50 (1874)` into `r0.4.0` (#1876)** (e540964)
- **cp: `ci: Update to transformers v5.5 (1734)` into `r0.4.0` (#1854)** (302b096)
- **cp: `chore: Update GPT-OSS and Qwen3 recipe configs (1811)` into `r0.4.0` (#1815)** (8d774e1)
- **cp: `ci: Increase benchmark timeout for GLM and Qwen3.5 MoE LoRA recipes (1818)` into `r0.4.0` (#1819)** (3bb7521)
- **cp: `ci: RC6 timeout fixes for release test recipes (1801)` into `r0.4.0` (#1816)** (e7e90eb)
- **cp: `feat: Enable CI benchmark with {llm,vlm}_benchmark (1793)` into `r0.4.0` (#1794)** (aab1504)
- **cp: FSDP2 w weight prefetching and async TP optimization (#1711) (#1779)** (60172f3)
- **cp: `ci: Resolve cve and remove uv cache (1774)` into `r0.4.0` (#1778)** (22823b8)
- **cp: `ci: Address container and source code cve (1753)` into `r0.4.0` (#1758)** (08d8454)
- **cp: `test: Checkpoint robustness skips atexit-registered destroy_process_group() (1730)` into `r0.4.0` (#1739)** (5c10e96)
- **cp: `ci: Address timeout is ci tests (1733)` into `r0.4.0` (#1738)** (79e3594)
- **cp: `feat: MoE model benchmarks, LoRA configs (1676)` into `r0.4.0` (#1707)** (4ade51e)
- **cp: `feat: adding lora to diffusion (1653)` into `r0.4.0` (#1708)** (6e65bf1)
- **cp: `(#1696)` into `r0.4.0` (#1732)** (84a8b30)
- **cp: `feat: integrate NeMo-Run launcher (1668)` into `r0.4.0` (#1706)** (df6b281)
- **cp: feat: VLM pretokenized data pipeline with neat packing (#1618)** (6264a70)

### Tests
- **test: add checkpoint robustness functional tests (#1606)** (cadbba7)

### Docs
- **docs: Update docs version to 0.4.0 (#2075)** (b651aa8)
- **docs: add per-model pages (#1683)** (a3cac3d)
- **docs: update the finetune guide (#1678)** (0dea775)
- **docs: add gemma4 tutorial (#1657)** (a8bdc4f)

### Chore
- **ci: cherry-pick #2048 (LoRA nightly tests) to r0.4.0 (#2051)** (9687b04)
- **ci: Update version to 0.4.0 (#1703)** (b9a2154)
- **ci: Remove duplicate ci config (#1702)** (e68cbe1)
- **ci: Associate recipe owners (#1690)** (62d2f8d)
- **ci: Add code freeze workflow (#1688)** (f7afa1f)
- **ci: Set target version for ruff (#1636)** (91c6e41)
- **chore(beep boop 🤖): bump FW-CI-templates workflow pins to v0.88.0 (#1669)** (cccf771)
- **ci: Add recipe golden values (#1647)** (f7b98a2)
- **ci: Update mistral4 medpix ci run time (#1646)** (9a0e0df)
- **ci: Update run time for nemotron super ci (#1614)** (dbe0f65)
- **refactor: CLI app and launching (#1406)** (35b5ed9)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/NVIDIA-NeMo/Automodel?utm_source=github-action)._