## huggingface/transformers — v5.16.1…v5.17.0

_152+ commits._

### Features
- **[Quantizaiton]support 5/6/7 bits in AutoRound (#48481)** (e8bcd79)
- **Add Fun-ASR-Nano model (#46180)** (fc50134)
- **Add `supports_context_parallel` to `PreTrainedModel` (#48442)** (8eaf75f)
- **[docs] Add a LiteRT page under community integrations (#48540)** (26600f6)
- **Add PR comment CI for AMD (MI300) (#48065)** (e898a10)
- **Add h4 (#48473)** (8914bc8)
- **QA: Add noisy comment checker (#48484)** (03a4d7b)
- **Support nested FLA kernel imports for fla-core (#48221)** (c119ec3)
- **add xpu expectations for hunyuan_vl model tests (#48504)** (904b555)
- **Add support for NeuCodec (#47143)** (9c8478e)
- **Support per-layer MTP configuration (#48264)** (57ddd70)
- **model: Add NVIDIA Canary-1B-v2 to Transformers (#46825)** (0cc72b4)
- **Add offload to gradient checkpointing (#48444)** (69a7fb1)
- **Add NeoMME and NeoMME-Retriever (#47992)** (e2b3550)
- **Bump transformers-mlinter to 0.1.5 and clear the new findings (#48259)** (ba356d2)
- **[Glm 5.3 Flash] GLM 5.3 Flash Support (#48342)** (eb4d9e2)
- **Add Qwen4Exp model (#48337)** (fc5c5bd)

### Fixes
- **[fix] Update stale expected strings in HunYuanVL integration tests (#48646)** (50bbcc6)
- **[fix] Update stale golden values and fix expected_logits shape in FlavaForPreTraining integration tests (#48639)** (3283d5f)
- **[tests] Fix integration test golden values broken by fast image processor default (PR #41388) (#48637)** (5f47b5a)
- **Fix YOLOS device mismatch with device_map="auto" (#46886)** (d9fe823)
- **[KimiLinear] Fix test_cpu_offload: set num_local_experts=4 in model tester (#48624)** (338b33d)
- **Fix expert-parallel training: NaN gradients and missing gradient contributions (#48205)** (2ca0705)
- **Fix GPU memory teardown in CLI serve tests (#48618)** (0a959de)
- **Fix `generate_flags` parsing in `transformers chat` (#48597)** (0514b65)
- **[fix] Fix how we read package versions - triggered by torch 2.14+ (#48615)** (361d009)
- **Fix AttributeError in gradient_checkpointing_enable(offload=True) (#48590)** (2dedac3)
- **[GLM 5.3 Flash] Fix NaN gradients in chunked KDA (#48455)** (0b179b2)
- **[docs] Fix [[autodoc]] directives in ALBERT model documentation (#48593)** (f67ec2b)
- **[Fix] Fix A10 expectations for a test (#48454)** (6622f6f)
- **[serge] Fix 2 integration tests for model `glm4_moe` failing with `OOM` (other (2)) (#48551)** (be8cf9d)
- **[serge] Fix 2 integration tests for model `nemotron` failing with `import_or_config` (other (2)) (#48582)** (5bea5af)
- **[serge] Fix 2 integration tests regressed by commit 83d46aa2a2c4 (PR #47625) (#48580)** (299c538)
- **[serge] Fix 2 integration tests for model `kosmos2` failing with `import_or_config` (other (2)) (#48552)** (99e19a9)
- **Fix failing tests for cohere_compass (#48005)** (e4052f5)
- **[Fix] Use dedicated helpers for DeepGEMM and SonicMoE tests (#48523)** (b8a1979)
- **fix failed test cases for glm5_next (#48497)** (3197876)
- **vibevoice: fix bug for quant cache (#48487)** (ec657b1)
- **Fix sliding-window mask `layer_idx` in Gemma3/Gemma4 `create_masks_for_vision_model` (#48482)** (c676202)
- **[`Qwen 3.5 Moe`] Fix decorators (#48436)** (92732e3)
- **[serge] Fix 2 integration tests for model `fsmt` failing with `output_mismatch` (tensor values differ (2)) (#48496)** (8ce59c1)
- **Fix some tests by removing the deprecation cycle (#48503)** (ea8f1f7)
- **Fix Pix2StructTextAttention init using hidden_size instead of d_kv (#47558)** (c66311b)
- **Fix pre patch release utility (#48499)** (23fbf2b)
- **Fix Inkling inputs_embeds and add more tests (#47827)** (774a675)
- **Simplify and fix qwen4 tests (#48340)** (8f54202)
- **doc: fix syntax error and typos in VibeVoice documentation (#48489)** (dd35e07)
- **[`DSA`] Fix vanillas SDPA with padding (#48363)** (fc236b6)
- **[serge] Fix 1 integration tests regressed by commit bd9509355c8a (PR #47493) (#48426)** (e84aa52)
- **Fix VoxtralRealtime rejecting non-static cache implementations (#48082)** (58d5a2b)
- **Fix rotary embedding regression (#48477)** (e15d467)
- **[docs] Fix failing audio doctest (#48380)** (ac32445)
- **[docs] Fix code snippets (#47772)** (05e078a)
- **[Fix] Sparse TikToken tokenizers silently fail (#48446)** (c155593)
- **[serge] Fix 2 integration tests for model `cwm` failing with `import_or_config` (other (2)) (#48414)** (49995e1)
- **fix: Add DEIMv2 attribution (#48448)** (1c05ce2)
- **Fix MTP generation test regex gate for escaped layer ignore keys (#48003) (#48262)** (8e35c98)
- **fix qwen4exp-fp8 ple embedding (#48368)** (2686521)
- **[fix] inkling: mps + cuda mel spec extraction (#47432)** (dce3921)
- **fix: decode() batch path respects self.clean_up_tokenization_spaces (#47793)** (4ff98bf)
- **Add TP-plan for muse_glimmer and fix dflash for sharded embedding (#47913)** (899f55f)
- **[serge] Fix 2 integration tests for model `hyperclovax` failing with `other` (other (2)) (#48440)** (4da0548)
- **fix(generation): Enforce the auto-compile cache check for encoder-decoder models (#48364)** (e37b269)
- **fix(models): Drop the position-indexed token type lookup in RoPE encoders (#48407)** (a3f3da8)
- **[serge] Fix 6 integration tests for model `seamless_m4t_v2` failing with `other` (other (6)) (#48425)** (dea7065)
- **[serge] Fix 4 integration tests for model `generation` failing with `output_mismatch` (list output differs (4)) (#48133)** (ccba41e)
- **Fix incorrect tuple return annotations on forward methods returning a Tensor (#48359)** (5f8ab9b)
- **Fix interval merge invariant in _find_disjoint (#47860)** (0b80c4a)
- **fix some failure in xpu (#48252)** (d4dc3f2)
- **[CB] Fix wrong device scoping (#48370)** (a8e8d14)
- **fix: flash-attn fallback failing on torch2.13 (#48388)** (281dd53)
- **[LongcatFlash] Fix test_longcat_generation_cpu: use device_map="cpu" to avoid MoE disk offload issue (#48377)** (83d024e)
- **Fix `safe_open` mmap memory exhaustion on Windows by using `pread` backend (#48341)** (8631167)
- **Fix Zamba2 construction for num_mem_blocks > 1 checkpoints (#48325)** (323f243)
- **[conftest] Use get_cpu_ram_total_gib for psutil patch (cgroup-aware) (#48290)** (35a66fa)
- **Fix flaky test_training_gradient_checkpointing for BigBirdPegasus (fp noise filter) (#48332)** (9b4f0a0)
- **Fix missing FP8 TP layer overrides (#48343)** (d65f37c)
- **[docs] Fix links and remove TokenizerFast (#47748)** (762fae6)
- **Fix incorrect token classification prefix for ESMC (#48348)** (4e5538a)
- **Fix kernel commit and repo paths for ESMFold2 (#48186)** (ba5cb08)

### Backend
- **v5.17.0** (856157a)
- **MRoPE continued (#48594)** (5b7dcb0)
- **[`Generate`] Avoid unconditionally downloading remote hub file (#48620)** (cbc1651)
- **Honor `shift_labels` in decoder-only LLM/VLM losses (#48493)** (bd05a4b)
- **[docs] Per-layer config (#48601)** (bcd59bc)
- **[Docker] Upgrade CPU torch to <=2.14.0, torchcodec to <=0.16.0 (#48614)** (95294e7)
- **Another day fixing CI (#48591)** (ada0a91)
- **esmfold2: keep `distogram_head` in fp32 as well (#48488)** (0df4ef3)
- **[docs] mlinter reference (#48460)** (0a14f2f)
- **extend some case to xpu as well (#48502)** (f16442e)
- **Compress the agent conventions file and document two modular pitfalls (#48586)** (254e62f)
- **Guard against a None video processor class when the backend is unavailable (#48557)** (d82653f)
- **[nit] use requires_backends (#47576)** (15a2691)
- **Kimi linear (#48250)** (c93057d)
- **Refactor GGUF to speed-up inference (#47779)** (f62dc9b)
- **[generate] stop synchronizing the accelerator on every decode step (#47975)** (4b24a19)
- **Pass kwargs to the Mamba2 mixer in Nemotron-H, Falcon-H1 and Mamba2 (#48490)** (86a6c97)
- **Retire test_multi_gpu_data_parallel_forward (#48508)** (37ed014)
- **Processing tests [part 2] (#47922)** (976e7a4)
- **:rotating_light: Vision (2d/3d) rotary embeddings  (#48105)** (4177486)
- **Allow nested rope params for tiny models (#48435)** (d3f494b)
- **Infinite loop in dependency search (#48393)** (cfa309c)
- **Remove deprecation (#48500)** (a8d5f2c)
- **[AfMoE] Standardize past_key_values argument naming across forward and generate (#48430)** (2f70c41)
- **Update dev version (#48498)** (e530c28)
- **Warn once when a hub-kernel function falls back to its reference PyTorch path (#48185)** (78fa17c)
- **[docs] Partial checkpointing and group_by_length (#48463)** (64d071c)
- **[docs] Kernel updates (#48465)** (a85f331)
- **[`Kernels`] Enable functions into kernels registry and allow non inheritance (#48443)** (4a2e450)
- **[`Qwen4 Exp`] Use partial to avoid skipping mask more easily (#48456)** (5c2ede9)
- **Remove deprecated mask functions (#48476)** (7c65cdb)
- **[MTP] Save memory by only capturing the last layer's hidden_states (#48475)** (c5ff582)
- **Allow capturing only necessary hidden_states with capture_outputs (#48081)** (65f70d4)
- **No inherit decorator for NeoMME (#48457)** (cdfdcad)
- **Batch Rebalance Data Sampler (#47340)** (4c0dafa)
- **Grounding dino fp16 dtype [backlog] (#48438)** (eed9b89)
- **Init the process group with a load-scaled timeout for sharded loading (#48228)** (45bd05b)
- **Raise a clear error when a token is both forced and suppressed (#47511)** (be255fb)
- **Clarify device placement in pipelines (#47367)** (764ed51)
- **[MiniCPMV4_6] Update test_small_model_vision_generation_batch expected output (value drift) (#48406)** (a932513)
- **Document image_hidden_states/pixel_values mutual exclusivity for SmolVLM/Idefics2/Idefics3 (#47714)** (0d4dbf5)
- **Re-order a bit for easier navigation (#48434)** (5f7bb43)
- **Deprecated stuff gone (#48367)** (3cf87d8)
- **[Docs]: Update GLM 5.3 (#48401)** (42ca970)
- **Avoid print to stdout that fails the job `check_failed_tests` job (#48391)** (805a9e9)
- **[VibeVoice] Skip generate export tests (flaky) (#48396)** (efed243)
- **skip mtp slow tests for now (#48328) (#48329)** (9aafba6)
- **Update Tailscale action version in workflow (#48394)** (721e492)
- **[Improvement] Make gated delta rule more explicit  (#47625)** (83d46aa)
- **Raise when a paged attention forward is called with no cache (#48297)** (a19f0d7)
- **[docs] Pass ContinuousBatchingConfig and sliding window models  (#48381)** (e972043)
- **[Qwen3VLMoe] Update `test_small_model_integration_test_batch` expected output (value drift) (#48376)** (155b899)
- **[ONNX] Skip affected models on torch 2.13 (two dynamo regressions) (#48191)** (1987631)
- **Retry get_daily_ci_runs on stale GitHub API cache (#48374)** (aa37261)
- **Implement VibeVoice  (#40546)** (640a08a)
- **[Docs] Change 5.3 Flash pos in toc (#48366)** (afa8081)
- **Quiet continuous batching at default verbosity (#48314)** (2e7033e)
- **Wait for the first request in the async continuous batching bootstrap (#48304)** (01846b5)
- **[CB] Fail faster (#48334)** (b6c0bfe)
- **Ignore a stale best checkpoint recorded in a resumed trainer state (#48319)** (36deb0b)
- **Keep MXFP4 weights quantized on XPU when use_kernels is set (#47923)** (e15ea9b)
- **Create the continuous batching CPU group with local synchronization (#48302)** (dabae5f)
- **Resolve continuous batching config against the text config for composite models (#48299)** (598d8ba)
- **[`CI`] Unblock fast CI for now (failing tests) (#48344)** (1e45299)
- **[qwen4_exp] disable torch/onnx export tests due to data-dependent control flow (#48345)** (0f0f7e1)
- **[debug] Trace previous CI run selection in get_previous_daily_ci.py (#48338)** (36bc98e)
- **Normalize HunYuanVL's legacy field aliases via attribute_map (#48261)** (2c889bf)

### Docs
- **docs: fix docstring parameter names that do not match signatures (#48575)** (3822359)
- **docs: remove phantom parameters from docstrings (#48576)** (345e787)

### Chore
- **CI: Point test fixtures at hf-internal-testing copies we already host (#48521)** (892044a)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/huggingface/transformers?utm_source=github-action)._