## huggingface/transformers — v5.17.0…v5.18.0

_242+ commits._

### Features
- **model: Add GTE to Transformers (#48416)** (57296d1)
- **Add Strix Halo (gfx1151) Atlas Inference Hub-kernel path for Qwen3.5/3.6/3.8 Gated DeltaNet (#49127)** (8520b15)
- **[CI] ssh-runner: add optional cache_type input to switch between bucket and EFS runners (#49159)** (22a2f92)
- **[Executorch] Add MLX recipe (#48910)** (5880561)
- **Add image processing tester init (#48829)** (89b6b17)
- **Add `Trainer.end` (#48875)** (98d3982)
- **Add support for Nemotron Omni (#46509)** (4b28d51)
- **Add MPS maintainer (#49052)** (002e1ed)
- **Add Nemotron3Diarization (#49056)** (8f080cb)
- **Add workflow for building XPU CI Docker images (#49039)** (560acb7)
- **Update tokenizer gguf support  (#48656)** (f7ca943)
- **[CB] [Major] Upgrade the cache to support different attention types (#47809)** (d94b0f7)
- **Add pose estimation keypoint preprocessing to Sapiens2ImageProcessor (#47199)** (e86881e)

### Fixes
- **[Nemotron3Diarization] fix streaming last stft frame dropped (#49167)** (88536f2)
- **Fix missing router_logits in Qwen3.5-MoE and other MoE models (#49179)** (5cd2877)
- **Fix MiniMax M3 partial 3D vision rotary embeddings (#49164)** (4fcb1ff)
- **Fix MPS GQA version gating (#49210)** (a877116)
- **Fix additional_special_tokens data loss with extra_special_tokens (#47848)** (e0299e2)
- **Fix stale _added_tokens_encoder entries in cpmant and wav2vec2 (#47440)** (5aab642)
- **Fix deepstack features for mixed-input (#49177)** (0c18a63)
- **Fix SequenceBiasLogitsProcessor edge cases: token id 0 and prefix equal to context (#49117)** (e57ba2b)
- **QA: fix noisy comment checker (#49184)** (b8bea61)
- **Fix command syntax for optimum-cli export (#46451)** (f339035)
- **Fix Trainer checkpoint resume crashing on CPU with multiple processes (#49123)** (90a491d)
- **Fix backslash handling in generate flag values in transformers chat (#48709)** (e159fce)
- **[CacheHardIntegrationTest] use dedicated safetensors repo to fix Xet bucket cache corruption (#49182)** (1e71827)
- **Fix lost call (#49163)** (a1289da)
- **[imagegpt/vilt/trocr] fix cache fixture tests: use hf_hub_download instead of load_dataset (#49178)** (ebd5de0)
- **[CB] Fix failing tests discovered when using the B200 (#49171)** (a73a2ba)
- **[PerceptionLM] Restore test_inputs_embeds overrides to fix flaky test (#49165)** (f5af320)
- **Fix get_json_schema dropping items/enum for unions of list/dict/Literal types (#49136)** (2e703d6)
- **Fix UMT5 decoder self-attention not being causal (#49135)** (d7ba937)
- **[Gemma] Fix test_model_7b_fp16_static_cache expected value for cuda 8 after #49084 (#49133)** (8d43073)
- **[AMD] Fix some integration tests (#49153)** (0bc057f)
- **[gemma3n] fix audio test fixture: use hf_hub_download instead of load_dataset (#49128)** (07338b6)
- **Fix exporters import on torch < 2.9 (is_contiguous_or_false) (#49124)** (96331a9)
- **[serge] Fix 2 integration tests for model `jamba` failing with `other` (other (2)) (#49044)** (6b14ee9)
- **[serge] Fix 2 integration tests for model `pvt_v2` failing with `other` (#49067)** (f8770a1)
- **[serge] Fix 2 integration tests for model `hy_v3` failing with `output_mismatch` (tensor values differ (2)) (#49068)** (4ad04d4)
- **[serge] Fix 2 integration tests for model `cvt` failing with `output_mismatch` (tensor values differ (2)) (#49076)** (11c1661)
- **[NemotronH-Omni] Fix device mismatch in test tensor creation (#49102)** (5e4d630)
- **Fix MaskFormerSwin attention mask dtype to follow hidden states (#49032)** (72c3b93)
- **Fix RGB early-return skipping PNG tRNS compositing (#49005)** (6da3313)
- **[Zamba] Fix associative scan breaking ONNX export and OOM in integration test (#49092)** (047de76)
- **Fix NaN in Parakeet eager attention with padded batches (#49070)** (93b45aa)
- **Fix odd head_dim validation for RoPE configurations (#48524)** (e37e548)
- **[serge] Fix OOM in Moshi integration tests with MemoryCleanupMixin (#48839)** (c320b44)
- **Fix assisted eos token condition (#49075)** (010f0ca)
- **Fix XPU Docker image build (#49073)** (c47c6ab)
- **fix(ci): harden GitHub Actions workflows (#49057) (#49059)** (20eeb6e)
- **Fix assistant masks for processors (#48793)** (e0882ff)
- **Fix startup failures: drop pull-requests: read from check_failed_tests.yml (#49057)** (267a0c4)
- **Fix startup failures: drop permissions reusable-workflow callers cannot grant (#49055)** (abbc443)
- **QA: Fix llama4 leak (#49042)** (4c9a6a5)
- **Fix mask creation not being skipped under `torch.compile` (#48975)** (d85573b)
- **fix(zamba):  add use_associative_scan config flag to avoid torch.compile slowdown (#48331)** (624b32a)
- **Fix assisted decoding for VLM due to dropping attn mask  (#49019)** (e3a275b)
- **fix noisy comments (#49013)** (39eedeb)
- **Fix RecurrentGemma compiled generation with StaticCache (#48961)** (1194956)
- **[docs] Fix legacy hf CLI references (transformers) (#48988)** (577d8fa)
- **Fix static cache per layer head shapes (#48619)** (00b29fc)
- **Fix mps autocast handling in rotary embeddings (#49006)** (7cd73d9)
- **[vLLM] Fix video token counting for Transformers backend video inputs (Part 2) (#48900)** (d2db88d)
- **Fix Qwen3OmniMoeIntegrationTest OOM (#48987)** (5bb12bb)
- **Fix SwitchTransformers Top1 router: raw logits, expert capacity accounting, and router losses (#48421)** (13d21d5)
- **Fix vibevoice TTS batched audio index (#48902)** (a10365f)
- **Reduce peak memory in examples_torch CI job (OOM fix) (#48983)** (21912ef)
- **Fix typos in DeepseekV4 comments (#48953)** (ae05b6b)
- **Fix image processor class-level size mutation and min_pixels handling (#48916)** (d53e587)
- **Fix more stuff (#48979)** (c210bb3)
- **Fix infeasible cost matrix errors in hugarian matcher losses (#47730)** (e138c4d)
- **Fix AXK2 integration test: update CUDA expected text and rename class (#48941)** (c587bc8)
- **Fix tests due to dropping attn mask (#48903)** (b856a67)
- **fix(vibevoice-asr): use integer ceiling division for audio token count (#48864)** (719d809)
- **Modular conversion small fix (#48934)** (95ad2ed)
- **Fix possessive typo in Whisper long-form warning (#48877)** (bdb4cc0)
- **Fix compile cache tests (#48923)** (9b819e7)
- **Fix T5 tied weights order for lm_head (#48238)** (483b880)
- **🚨 [vLLM] Fix video token counting for Transformers backend video inputs (Part 1) (#48894)** (770e4c4)
- **Fix docstring argument names that don't match signatures (#48474)** (94d1ce8)

### Backend
- **v5.18.0** (a906d3c)
- **QA: restore masking_utils export comment with a noqa (#49209)** (62f28fe)
- **gguf user defined tokens (#49008)** (8a5a4b5)
- **[`Kernels`] Sync mamba version (#49205)** (dbac24c)
- **Map bare list, tuple and dict annotations to the right JSON schema type in get_json_schema (#49145)** (541cf67)
- **Fail fast on eval OOM under `auto_find_batch_size` (#49198)** (8ed81bc)
- **Replace datasets that no longer load with maintained uploads (#49018)** (a0ccef2)
- **[MPS] Let mps sdpa handle grouped query attention directly (#49187)** (6fc8c14)
- **[Parakeet] Convert NeMo's stochastic depth to layerdrop (#49191)** (f36c750)
- **Video processors - general maintenance (#48251)** (528c267)
- **update to torch 2.14 (#49199)** (b1463f5)
- **Finish removing the MPS autocast workaround (#49157)** (d125eb2)
- **Document running the example scripts on Hugging Face Jobs (#49050)** (da4bd0d)
- **Document image_like_kwargs (#49180)** (7087be0)
- **[MPS] Remove cu_seqlens_k clone workaround for metal-flash-sdpa (#49091)** (2ceac52)
- **No more -hf repo names for ESMC (#49158)** (ba9e224)
- **[generation] Encode multimodal data only once (#45783)** (0291458)
- **Deprecate the use_mamba_kernels config flag that no longer has any effect (#49155)** (83e427a)
- **Pick the default flash implementation based on the current hardware (#49109)** (0a13bb7)
- **Auto generate model inits (#47829)** (33aaea5)
- **Skip flash tests that fall back to a hub kernel when kernels is missing (#49129)** (89310e2)
- **Restore the Unicode whitespace set in the GPT-SW3 tokenizer (#48912)** (885320e)
- **qwen3 models map to wrong tokenizer class on the hub (#49116)** (a4fe1d5)
- **Remap the legacy Gemma 1 hidden_act in the config post-init (#49084)** (27166ea)
- **Use grouped_mm on TPU devices under torch.compile (#49097)** (6e4bcc5)
- **Keep the eos ids the config declares (#49082)** (e825bfc)
- **Use namespaced dataset ids in docs and PyTorch examples (#49017)** (f0e39b8)
- **🚨 [ROCm] gpt-oss: route FA3 to aiter-flash-attn, generate ROCm fixtures (#46837)** (763a150)
- **Keep `attention_mask` as `None` in OPT's causal mask creation (#49002)** (83445dd)
- **Bump huggingface_hub upper bound to <3.0 (transformers) (#49083)** (ff4612b)
- **Keep already-decoded array arguments unchanged (#49062)** (6d43ab4)
- **Honor `config.output_router_logits` in the MoE VLM wrappers (#48885)** (f324707)
- **[Nemotron3Diarization] nit: hub pr merged to main (#49065)** (c8b81b6)
- **[docs] Loading behavior (#49023)** (c00f362)
- **Register the remaining mamba-ssm kernel layers on XPU (#49035)** (979041c)
- **Only seed numpy in BigBird block-sparse attention during training (#49037)** (2adb62d)
- **Pin GitHub Actions to commit SHAs (#49049)** (a008a65)
- **[Chat] Loading GGUF models served with the Chat CLI (#49031)** (1599d20)
- **Scope GITHUB_TOKEN permissions per job (#49046)** (5692873)
- **PEFT x Dtensor-based TP integration (#48485)** (935f7ab)
- **Shorten noisy OpenVINO SDPA comment (#49033)** (b3c3f5c)
- **Added more usage of MemoryCleanupMixin (#49011)** (2c4914f)
- **Keep special token ids the tokenizer does not define (#48708)** (46375fc)
- **Deprecate min-max pixels (#49021)** (a12a224)
- **OpenVINO HF Exporter (#47003)** (fd290dc)
- **updated GraniteMoeHybrid expectations (CPU) (#49016)** (14793d4)
- **Summarization examples: download NLTK punkt_tab, not punkt (#49014)** (8c277a3)
- **Enable compressed-tensors FP8 kernels on MPS (torch >= 2.15) (#48985)** (0bc2528)
- **Update ggml kernels path  (#48991)** (235efe7)
- **🚨 Speed up detr image processing (#48066)** (8445b13)
- **QA: applied ruff rule PLW1514 (#48990)** (f441076)
- **[generate] Always correctly restrict assisted decoding with max length/eos token  (#48981)** (87d34bf)
- **🚨 Remap indexers layer_type (#48974)** (e2d83fd)
- **[generate] Simplify candidate generators by removing required `update_candidate_strategy` (#48982)** (a2fc752)
- **Let `prefix_allowed_tokens_fn` override model `-inf` and raise an exception on unsatisfiable generation constraints. (#48927)** (66880ec)
- **Contract the mamba2 chunk scan with einsum instead of broadcast-then-sum (#48978)** (a52f659)
- **Better attn default for gguf (#48935)** (abd76d1)
- **Skip an unnecessary image copy in the torchvision image normalization path (#48897)** (207fca7)
- **Keep image processor backends in sync on keys and dtypes (#48739)** (a7d9d94)
- **Resolve the Hub revision once per load instead of passing a private _commit_hash around (#47611)** (d67c729)
- **Align special tokens on the text config (#48847)** (ea31b0c)
- **DOC: Improve documentation for remap_legacy_layer_types function (#48651)** (266cc92)
- **Improve lazy import error messages (#48602)** (7301fbc)
- **Standardise Aria's MoE onto the experts interface (#48907)** (fd204af)
- **Update arXiv citation in NeoMME model doc (#48926)** (1855613)
- **Size the device_map buffer from the largest leaf module (#47211)** (4618eba)
- **Clean up some MoE models' integration tests (#48833)** (6c6bac2)
- **Mark test_generate_with_static_cache as flaky for olmo and bigbird_pegasus (#48856)** (da58b89)

### Docs
- **docs: fix dead doc links in i18n READMEs and the xlnet docstring (#48905)** (3713bd8)
- **docs: fix docstring parameters that do not match signatures (#48904)** (c9ad4a7)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/huggingface/transformers?utm_source=github-action)._