## NVIDIA-NeMo/Automodel — v0.1.2…v0.2.0

_190+ commits._

### Features
- **cp: `feat: add glm 4.5 air finetuning config (873)` into `r0.2.0` (#879)** (d4cf179)
- **cp: `ci: Add additional dep for model support (861)` into `r0.2.0` (#866)** (f406dd2)
- **cp: `feat: sft qat support (704)` into `r0.2.0` (#854)** (6b390dc)
- **cp: `build: Add OSS NOTICES.txt file to docker build (838)` into `r0.2.0` (#842)** (2bb2835)
- **cp: Add internvl recipe #823 #810 (#835)** (78ffedf)
- **cp: `feat: add answer only masking in ColumnMappedDataset (832)` into `r0.2.0` (#834)** (acbb932)
- **cp: `ci: Add mamba-ssm and causal-conv1d dep (811)` into `r0.2.0` (#820)** (3069ca4)
- **cp: `feat: Add NeMo Biencoder (745)` into `r0.2.0` (#819)** (e2b0d07)
- **cp: `perf: add qwen2.5 32b lora perf (802)` into `r0.2.0` (#816)** (4bfc496)
- **feat: removing split across pack (#792)** (ba32d63)
- **feat: Support clip_grad_norm for all parallelisms (#791)** (9c5854c)
- **feat: vlm ep and qwen omni custom implementation (#742)** (8134b0c)
- **feat: add mfu estimation for lora (#786)** (91fd03c)
- **feat: add fully_shard_by_dtype (#614)** (a5f0652)
- **feat: Review docs (#749)** (dcb1bee)
- **feat: mlflow integration (#715)** (11c8f59)
- **feat: Add wandb to benchmark recipe + qwen3 next and glm 4.5 air configs (#613)** (8df18da)
- **feat: introduce ScopedModuleOffloading in KD to reduce memory usage (#774)** (6eca242)
- **feat: add streaming ds (#778)** (9995e4a)
- **feat: update news (#785)** (43b75ca)
- **feat: Sharding Optimization for Sequence Parallelism/LoRA (#733)** (875fe69)
- **feat: Support PEFT with PP (#740)** (c4afaf6)
- **feat: adding mlflow (#776)** (aa7cd25)
- **feat: fp32 lm_head and fp32 apply_rope options for MoE (#769)** (8316227)
- **feat: add oot parallelism decorator (#736)** (e913574)
- **feat: configurable router precision for moes (#767)** (39b1038)
- **feat: add qwen3 omni ootb recipe (#739)** (a2d91b6)
- **feat: support lora with te (#766)** (27d70ef)
- **feat: Add convert_single_tensor_to_hf API for state dict adapter (#759)** (d195771)
- **feat: adding rank0 download for custom models (#754)** (d8b3778)
- **feat: adding SIGTERM handling (#755)** (666d391)
- **feat: add num_nodes to benchmark configs (#758)** (5aa7965)
- **feat: sequence classification finetune (#688)** (8ccdb2f)
- **feat: introduce NEMO_ENABLE_USER_MODULES (#747)** (55aff83)
- **feat: support multiple validation datasets and per-dataset logging (#537)** (f6ff70d)
- **feat: Support TE attention for gptoss (#732)** (2b7c02c)
- **feat: surface truncating & padding options (#719)** (38b330c)
- **feat: force hf flag for custom models (#708)** (b27761b)
- **feat: Add GLM 4/4.5/4.6 MoE (#705)** (824408f)
- **feat: add offline consolidation script (#711)** (a4b22fb)
- **feat: load dataset subset (#706)** (62e5181)
- **feat: Add Qwen3 Next  (#672)** (897d6f2)
- **feat: support hf metadata checkpointing for offline consolidation (#709)** (45c803e)
- **feat: consolidate HF safetensors backport (#698)** (99f4663)
- **feat: add configs for qwen3-vl 4B & 8B (#700)** (1fbc65c)
- **feat: support for blend json file via config (#699)** (f8a8f8b)
- **feat: add Qwen3ForSequenceClassification to registry (#617)** (d26e182)
- **feat: multiturn chat dataset (#680)** (71844e8)
- **feat: Add packed sequence + context parallel support for custom MoEs via TE (#633)** (88bff26)
- **feat: add warnings around CP/SP/sdpa (#642)** (3ee57d6)
- **feat: add Metric logger with jsonl output (#429)** (df2959d)
- **feat: GPT-OSS for DGX Spark (#657)** (37e513e)
- **feat: move transformers/diffusers out of components (#644)** (a311e7b)
- **feat: add phi4 tp plan (#641)** (8245c23)
- **new changelog-build (#660)** (e50ce9f)
- **feat: checkpoint refactor + async checkpointing (#635)** (8e15338)
- **add Qwen3-8B SFT recipe for DGX Spark (#645)** (3eb3d08)
- **feat: move mask creating to data pipelining for better perf (#615)** (24ff3ed)
- **feat: parallel diffusers generate (#573)** (3de81c6)
- **feat: Update model coverage for llm.md and vlm.md (#631)** (2b79e93)
- **feat: Update vlm.md (#632)** (4d6bfb8)
- **feat: add tool call dataset and recipe (#588)** (bfd9fd2)
- **feat: Add fsdp optimizations for custom MoE models  (#574)** (2c236e0)
- **feat: update docs (#590)** (56b2b13)
- **feat: custom validation step for KD (#575)** (ef39176)
- **feat: Add performance summary document (#562)** (89e3a9f)
- **feat: add multinode config and update readme (#552)** (c1805ee)
- **add trust remote option (#580)** (78a1caa)
- **feat: Add benchmarking recipe and configs (#548)** (79d0ccf)
- **feat: Update README.md (#555)** (31200de)
- **feat: Model Registry (#553)** (6aadbae)

### Fixes
- **cp: `fix: NeMoAutoTokenizer (878)` into `r0.2.0` (#880)** (0504d0e)
- **cp: `fix: no meta init when `force_hf` (874)` into `r0.2.0` (#876)** (adf5845)
- **cp: `fix: test process launcher error propagation (871)` into `r0.2.0` (#872)** (358e20b)
- **cp: `fix: remove validation for packed seq moe configs (867)` into `r0.2.0` (#868)** (2063302)
- **cp: `fix: update moe finetuning configs to use from_pretrained (863)` (#864)** (91ccc82)
- **cp: `fix: revert recipe change for memory fragmentation OOM in Llama3 70B (818)` into `r0.2.0` (#862)** (78d509e)
- **cp: `fix: deepseek v3 pretrain config parallelizer (851)` into `r0.2.0` (#852)** (6ac97d6)
- **cp: `fix: Cast norm to fp32 in clip_grad_norm (825)` into `r0.2.0` (#830)** (bd1f952)
- **cp: `fix: ep shard state dict conversion (815)` into `r0.2.0` (#829)** (5bacb0c)
- **cp: `fix: torchrun single proc (814)` into `r0.2.0` (#821)** (7845427)
- **cp: `fix: Include megatron ``Makefile`` in package data (798)` into `r0.2.0` (#813)** (5ad44ab)
- **fix: Ensure target device for clip_grad_norm (#793)** (17055ab)
- **perf: Update lora tflops explanation and best LoRA datapoint and recipe (#790)** (b3c51c3)
- **fix: Update is_hf_model to exclude custom models  (#789)** (a579f04)
- **fix: peft bug fix inside sequence classification recipe (#781)** (a198c9d)
- **fix: fixing local TPS for CP (#775)** (9512b25)
- **fix: custom model fixes (#773)** (19484e8)
- **fix: remove stale param visualization_font_size_offset (#772)** (c8b3c8a)
- **fix: update truncate logic (#757)** (5e995e9)
- **fix: LR scheduler state loading fix (#764)** (4fae7f6)
- **fix: explicitly remove lm head from fqn key mapping (#728)** (9b85cf9)
- **fix: relative symlinking (#720)** (bf4517b)
- **fix: enable checkpoint saving by default (#661)** (dc45ba4)
- **fix: some bug when running finetune script (#654)** (a6d88dd)
- **fix: set max_steps inside constructor (#650)** (2889cfc)
- **fix: checkpointing unnecessary arg (#655)** (c96c27a)
- **fix: misc minor perf issues (#634)** (4e0c9d9)
- **fix: step scheduler switch to zero based indexing (#627)** (ebd03b0)
- **fix: merge sft.md and peft.md (#618)** (8457d9d)
- **fix: Update peft.md (#619)** (f345168)
- **fix: Add query proj to hidden size ratio in Qwen3 flops calculator (#611)** (8d9bfb9)
- **fix: 0 token routed expert bug fix & update moe bias (#532)** (f06e8b5)
- **fix: cp fix (#591)** (3774006)
- **fix: Skip grad clipping when tp>1 and pp>1 (#576)** (0a7a495)
- **fix: update readme (#452)** (411e990)

### Backend
- **cp: `feat: update change log for r.0.2.0 (921)` into `r0.2.0` (#930)** (0be83ba)
- **cp: `ci: Bump to 0.2.0 (927)` into `r0.2.0` (#928)** (9b9a634)
- **cp: `ci: ci: Update changelog for Automodel 0.2.0 (894)` into `r0.2.0` (#898)** (f207bdd)
- **cp: `feat: adding flags for special tokens & chat template in column mapped dataset (844)` into `r0.2.0` (#865)** (9a29d45)
- **cp: `feat: combine projection refactor (804)` into `r0.2.0` (#856)** (5f02a33)
- **cp: `docs: Update version and contrib (849)` into `r0.2.0` (#850)** (9f0704f)
- **cp: `feat: ckpt val loss + run val at ckpt + symlink best ckpt (828)` (#847)** (a24530c)
- **cp: `ci: Build bitsandbytes from source (837)` into `r0.2.0` (#846)** (822577a)
- **cp: `feat: (809)` into `r0.2.0` (#822)** (2be8d2f)
- **cp: `feat: enable meta device by default (797)` into `r0.2.0` (#812)** (67c3478)
- **cp: `feat: auto detect base weights dequant (796)` into `r0.2.0` (#801)** (33da51a)
- **cp: `(803)` into `r0.2.0` (#806)** (75eb85f)
- **Limit dataset samples for Column Mapped datasets (#521)** (a5eb299)

### Docs
- **docs: Update README to reflect PyTorch Distributed instead of DTensor (#752)** (6a4aa51)
- **docs: split job launching guide (#724)** (87d06f8)
- **docs: add link to recipes (#707)** (3918fdc)
- **docs: add model coverage overview (#692)** (6dae71a)
- **docs: update pretraining guide (#686)** (1d6ed7d)
- **docs: Update README with NeMo Framework link and image (#684)** (b116615)
- **docs: Added repo overview image (#685)** (e41e67b)
- **docs: Update README.md (#602)** (2112f2b)

### Chore
- **ci: Address CVE and remove duplicate (#765)** (8b246ad)
- **ci: Bump torch upperbound (#649)** (1e2632c)
- **ci: Update DeepEP commit (#738)** (34f73e9)
- **ci: Pip install deepep (#729)** (3e6952d)
- **chore: bump pytorch index from cuda 12.8 to 12.9 (#695)** (3ebf595)
- **ci: Remove coverage combine in functional test script (#722)** (91cfae4)
- **ci: Use env var for setting test data dir path (#702)** (979ed85)
- **ci: Update pytorch uv lock (#716)** (60752e0)
- **ci: Add 0.1.1 changelog (#710)** (e131f2e)
- **ci: Update codeowners (#697)** (f442296)
- **ci: Update testing to use pytorch container (#696)** (4587b99)
- **ci: Update docker and TE install (#694)** (9dfe28b)
- **ci: Fix pytorch uv lock logic (#671)** (6d96f5d)
- **ci: Cancel outdated pipelines (#653)** (7177d45)
- **ci: Add max-parallel (#624)** (5ccab49)
- **ci: Add uv lock to pre commit hook (#626)** (693bf2d)
- **refactor: simplify  `_extract_model_layers` (#402)** (00dc3fe)
- **ci: Update dockerfile (#592)** (ccc728c)
- **ci: Update rest of preflight (#601)** (4b4489d)
- **ci: Bump preflight version (#595)** (8b23507)
- **ci: Update changelog for Automodel 0.1.0 (#560)** (27b414d)
- **refactor: TP parallel style with LoRA (#520)** (1fa9894)
- **ci(fix): pre-flight (#550)** (f8358ee)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/NVIDIA-NeMo/Automodel?utm_source=github-action)._