## NVIDIA-NeMo/Automodel — v0.2.0…v0.3.0rc4

_211+ commits._

### Features
- **feat: add test for biencoder + inline dataset (#1152)** (aceaf81)
- **feat: update checkpoint auto-loading with explicit restore_from (#1009)** (579cb30)
- **feat: v5 refactor for device mesh only model init (#1181)** (8a4f070)
- **feat: tp plan for ministral (#1190)** (6e45b57)
- **feat: Add GroupedExpertsTE backend (pre-requisite for MoE fp8) (#1109)** (037d2a0)
- **feat: Add Minimax M2 model implementation (#1191)** (5f63eb4)
- **feat: Implement DoRA (#1150)** (a6a9d2e)
- **feat: add dion optimizer (#1035)** (c9ae0b9)
- **feat: support Qwen3 VL 235b  (#1160)** (73cd446)
- **feat: streaming safetensors writer (#1164)** (2da5f18)
- **feat: add deepseek 3.2 support  (#1143)** (f258802)
- **feat: Add Step3p5ForCausalLM (#1168)** (1a3ef23)
- **feat: change default to sdpa if fa unavailable (#1165)** (e769c13)
- **feat: inline text dataset format (#1115)** (8c8ed48)
- **feat: remove checkpointer from Automodel class (#1147)** (5ffc16d)
- **feat: add support for kimi K2.5 VL (#1132)** (5b25e73)
- **feat: support kimi-vl model (#1103)** (f942ec3)
- **feat: transformers v4 API (#1116)** (6fa2b47)
- **feat: nano-v3 custom model (#1091)** (d86b960)
- **feat: add PP to vlm (#1080)** (5fd6482)
- **feat: Add faster fp8 dequant kernels and fix dtensor dequantization for dsv3 (#1110)** (5d7710b)
- **feat: databricks deltalake dataset support (#920)** (8c1994f)
- **feat: dereference env vars in yaml (#1112)** (8401c54)
- **feat: add norm fusion and rope cache for dense model (#1120)** (822effa)
- **feat: Add GLM 4.7 config (#1119)** (5809056)
- **feat: improve import time (#1076)** (89cbcff)
- **feat: Support LoRA in Biencoder (#1098)** (ab37350)
- **feat: refine codeowners (#1100)** (10cf77b)
- **feat: update moe parallelizer - remove unused arg, support ignoring router in ac (#1089)** (f85e420)
- **feat: Implement hard negative mining (#1069)** (a50fd05)
- **feat: add consolidation for Databricks (#1066)** (846b2ae)
- **feat: Add validation support for pipeline parallelism (#1061)** (aea3732)
- **feat: add configurable remote logging frequency via `step_scheduler` (#1008)** (50253d1)
- **feat: nemotron flash configs (#1034)** (3c316df)
- **feat: support nemotron_parse vlm (#999)** (ff022b2)
- **feat: Add TE rope fusion for custom MoE models (#897)** (3e23904)
- **feat: Add unroll pos docs script for biencoder training data (#1018)** (7711f7b)
- **feat: Support LoRA for custom MoEs (#1010)** (2a20947)
- **feat: backport devstral to v4 (#977)** (6bb6db7)
- **feat: add support for parquet files (#919)** (4e9d32b)
- **feat: allow passing model-id to from_config (#984)** (274fc3f)
- **feat: add functiongemma yaml (#985)** (e713fbe)
- **feat: add xlam toolcall dataset (#975)** (ca7f4e8)
- **feat: simplify from_pretrained/from_config (#967)** (c1cc88d)
- **feat: add nano-v3 to README (#978)** (6d51619)
- **feat: add nsys model layer name scope and benchmark support (with nsys) in app (#951)** (ab56f2f)
- **feat: add more PEFT lora recipes (#959)** (507bfb7)
- **feat: nano v3 configs and FSDP fix (#964)** (1645e61)
- **feat: Support for Llama-Embed-Nemotron-8B Training Pipeline (#963)** (1efd3e8)
- **feat: improve yaml logging to stdout (#882)** (99d214e)

### Fixes
- **fix: config fixes (#1207)** (4ab28b4)
- **fix: temp walk around VocabParallelEmbedding to address OOM (#1197)** (c444569)
- **fix: raise exception in tests (#1196)** (99e8db0)
- **fix: add force_hf for vanilla hf llama in hf recipe (#1177)** (7fbfd24)
- **fix: add nemotron parse custom loss (#1189)** (a3a8c68)
- **fix: fix pp batch issue in step3p5 (#1171)** (1c5c02d)
- **fix: re-enable megatronfsdp tests (#1134)** (b0ca574)
- **fix: surface trust_remote_code (#1139)** (801f63e)
- **fix: file io based safetensors index creation (#1137)** (46e6253)
- **fix: leave num_epochs unset if max_steps is specified (#1107)** (c1b5705)
- **fix: Update container workdir (#1124)** (83db658)
- **fix: output_hidden_states (#1121)** (8896d62)
- **fix: Update Nemotron Nano v3 Configs (#1122)** (8421d6b)
- **fix: checkpoint consolidation for custom MoEs (#1117)** (992a8e0)
- **fix: fix ckpt loading missing key in llama/qwen PEFT run (#1104)** (3af9075)
- **fix: make TEParallelCrossEntropy DTensor-aware (#1087)** (9cd12a7)
- **fix: remove skipping lm_head (#1074)** (58f0fae)
- **fix: gpt2 weight init with meta device (#1086)** (581c1df)
- **fix: save_pretrained with NeMoAutoTokenizer (#1073)** (5dd540e)
- **fix: fix custom llama and qwen state dict adapter concat and split (#1053)** (eab757a)
- **fix: make model accept args again (#1079)** (ffbda88)
- **fix: QLoRA checkpoint saving (#1054)** (314913b)
- **fix: update nemotron parse model id in config (#1070)** (27dbef3)
- **fix: seq_lens in batch (#1050)** (c0458f9)
- **fix: Make nemo_automodel dependency optional in biencoder model checkpoint (#1048)** (1033f36)
- **fix: Update checkpointing to support Databricks unity catalog (#1023)** (a363149)
- **fix: missing hidden_states return (#1036)** (1697e1e)
- **fix: Fix checkpoint saving for biencoder models with empty FQN mapping (#1025)** (4c364c2)
- **fix: gemma-2 tokenizer (#993)** (3bb8789)
- **fix: update attn_implementation in configs, add FA fallback + quantization fix (#1019)** (bfecb5d)
- **fix: custom llama (#1012)** (b7031bb)
- **fix: add pooler weights to biencoder state dict adapter (#998)** (4d319d4)
- **fix: expose llama and qwen2 to model registery (#997)** (272c9d7)
- **fix: respect trust_remote_code when building AutoConfig (#1007)** (caa5177)
- **fix: fix qwen3omni model registration (#1003)** (0a18faf)
- **fix: fix qwen3vlmoe state dict adapter (#1002)** (672806f)
- **fix: handle nonfloat dtype during ckpt consolidation (#996)** (a8d9ca3)
- **fix: Add DeepEP fallback logic and tests (#1000)** (1939fc4)
- **fix: resolving errors in the hf decorator function (#983)** (1d42deb)
- **fix: add nvtx config (#974)** (6b8367d)
- **fix: prevent hang in in download_model_weights (#991)** (1841c3c)
- **fix: misplaced parenthesis; Thanks @jbross-ibm-research (#973)** (47ec481)
- **fix: move print_trainable_parameters calculation to device (#966)** (1f8d645)
- **fix: Biencoder consolidated checkpoint and transformers issue (#936)** (5f27227)

### Docs
- **docs: update paths (#1194)** (619d540)
- **docs: update paths (#1194)** (f0122ef)
- **docs: update new models (#1188)** (75d8bae)
- **docs: update coverage tables (#1180)** (1411078)
- **docs: explain patch_inner_model and patch_causal_lm_model more (#1133)** (2d23be7)
- **docs: 🤗 Hugging Face Transformers API compatibility (#1146)** (7536edf)
- **docs: update readme (#1162)** (6ec24b7)
- **docs: Update docs/guides/dataset-overview.md (#1145)** (25f5c25)
- **docs: add all copywrite changes from #1120 (#1123)** (8473ee6)
- **docs: Fix code highlighting typo in databricks doc (#1114)** (9f81989)
- **docs: Updates to Databricks docs (#1102)** (7261e73)
- **docs: bump docker to 25.11.00 (#1096)** (e13515c)
- **docs: Documentation update for release 0.2.0 (#1041)** (62547bd)
- **docs: Created guide for quantization aware training (#1088)** (70ab1ef)
- **docs: Added ml flow guide  (#1045)** (54f23a9)
- **docs: Add Databricks example (#1043)** (5273467)
- **docs: update feature roadmap for 26.02 (#1056)** (819f2a0)
- **docs: Fixing version switcher (#1049)** (c3e48cf)
- **docs: Add documentation for the new ChatDataset class (#990)** (aea43ef)
- **docs: update pretraining docs (#981)** (b32dc1f)
- **docs: update readme with FunctionGemma (#988)** (ed04d5d)
- **docs: functiongemma docs (#986)** (e1569ed)
- **docs: Update LLM coverage table (#982)** (87411d7)
- **docs: Update news section for nano-v3 in README.md (#969)** (47191a3)
- **docs: update vlm coverage (#961)** (4de5d0f)
- **docs: Fix images not rendering in docs (#954)** (5d8cf31)

### Chore
- **ci: Update Automodel CI to use AWS infra (#1067)** (28bd5c6)
- **ci: Upgrade gnupg to address cve (#1202)** (31e3e1e)
- **ci: lower-bound megatronfsdp version (#1192)** (d9bceaa)
- **chore: remove _original_strings from repr(ConfigNode) (#1185)** (bbc0ed2)
- **ci: add docs render date (#1186)** (691da05)
- **ci: Fix release docs syntax error (#1183)** (4ed266e)
- **chore: refactor common te module (#1172)** (22dffd7)
- **ci: Address CVEs (#1175)** (ba91bfb)
- **ci: add duration time logging (top-32) for tests  (#1173)** (e2a9268)
- **ci: Release docs set default values for when push on main (#1158)** (423bb65)
- **ci: Update transformers upperbound to v5 (#911)** (07b0700)
- **ci: Add detect secrets (#1144)** (ee89252)
- **ci: Add doc-strings to update_pyproject_pytorch.sh (#1155)** (5234f4c)
- **ci: Update contrib and developer workflow (#1140)** (c008250)
- **ci: Auto publish on main if docs is updated (#1135)** (8f21f34)
- **ci: Reorganize automodel dep management (#1127)** (bbad52e)
- **ci: Introduce remote caching and optimize docker (#1113)** (00b822e)
- **ci: Add release docs workflow (#1105)** (4c16d73)
- **ci: consolidate functional tests (#1058)** (8312568)
- **ci: Update transformers to latest version 4.57.5 (#1051)** (92eb14d)
- **chore: Include formats in ``CheckpointingConfig`` validation error message (#1042)** (519823f)
- **ci: Bump version to 0.3.0 (#1044)** (cdd15cf)
- **ci: Revert codecov to use patch (#1037)** (00dc51a)
- **ci: Update TE install to v2.11 (#1031)** (51b5f4e)
- **ci: Bump deepep & nvshmem (#1030)** (0d80ef8)
- **ci: Debug test coverage (#1021)** (6053625)
- **build: Bump to pytorch 25.11 (#1004)** (d395f5a)
- **ci: add QLoRA test (#1017)** (4fe74cd)
- **ci: update codeowners 2 (#992)** (6960091)
- **ci: update owners (#958)** (ce09c48)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/NVIDIA-NeMo/Automodel?utm_source=github-action)._