## huggingface/transformers — v5.15.1…v5.16.0

_163+ commits._

### Features
- **Add Qwen4Exp model (#48337)** (276c5c4)
- **Add an opt-in per-frame pixel cap (cap_pixels_per_frame) to the Qwen3-VL video processor (#48071)** (27e7f6c)
- **Add shared ImageProcessingTester (#47745)** (b064ba6)
- **Add MLU support to is_flash_linear_attention_available (#46995)** (6a9e2d3)
- **Revert "Support per-layer cache configuration and attention-mask selection" (#48175)** (6b3a9de)
- **Support per-layer cache configuration and attention-mask selection (#47901)** (c84ac34)
- **🚨[wav2vec2] Support attn_implementation=sdpa dispatch (#46196)** (2b29630)
- **Delete old mlinter review comments before posting new ones (#48107)** (de795d9)
- **feat: add nvfp4 quantization (#47883)** (80c6674)
- **Support `BatchFeature` in length-grouped samplers (#48034)** (24fb898)
- **[new model] step 3.7 (#46658)** (c49429e)
- **[OLMoE] Update expected logits for A10G and add torch.no_grad() (#47989)** (bd5df99)
- **Declare sdpa support in `TimmWrapper` (#47939)** (1cdaf59)

### Fixes
- **Fix video-llama modular conversion (#48336)** (d3cdaec)
- **Add a regression test for force_accelerate_hooks signature preservation  (#48260)** (5fcf605)
- **Fix `scores` type in stopping criteria docstrings (#47676)** (53e3152)
- **Docstring check didn't match some file - fix it (#48121)** (b541699)
- **fix bug for blt model parallel bug (#48327)** (2514a5e)
- **Cpmant fix use cache (#48013)** (0f93813)
- **Fix AutoTokenizer returning TokenizersBackend for DeepSeek-R1-Distill-Qwen models (#48211)** (51edaf8)
- **[docs] Fix failing doctests (#47687)** (2c1376c)
- **Fix build_2d_sinusoidal_position_embedding on MPS (#47897)** (562cfd9)
- **Fix `BayesianDetectorModel.from_pretrained()` by calling `post_init()` (#48254)** (60f2b9f)
- **[Fix] Avoid duplicating tests in CI (#48287)** (a548f26)
- **[`GDN`] Fix recurrent FLA fallback (#48266)** (93fbc55)
- **Fix `tie_word_embeddings` not lifted from `text_config` for some VLM configs (BC regression) (#45857)** (cb1e460)
- **Fix dtype mismatch in grouped_mm_fallback for LoRA training on Mamba+… (#47933)** (4ad9199)
- **Fix nemotron_h save_pretrained emitting singular backbone.embedding.weight (#48075)** (9916b46)
- **fix(data_collator): align TokenClassification numpy_call with torch_call (#48212)** (2745e77)
- **Fix a typo in a use of a local variable field_ in a test (#48184)** (195734e)
- **doc: Fix typo in SigLIP2 Flash Attention code example (#48197)** (fe314d4)
- **[serge] Fix 2 integration tests regressed by commit 16780c86b20a (PR #47622) (#48134)** (17044fa)
- **[xcodec2] Fix flex attention and flash dispatch tests (#48244)** (c7cf04b)
- **[Fix] Export crashes on kernel-decorated function (#47808)** (f440843)
- **[Gemma4] Fix stale expected values in integration tests (#48233)** (3d6c720)
- **[TableTransformer, PI0] Fix stale expected values and OOM in integration tests (#48198)** (d56c55b)
- **Fix stale expected values in integration tests (cuda sm_86 / Aug04 regressions) (#48171)** (3ad85ca)
- **Fix `gpt_oss` runs on GPU (#48118)** (275b524)
- **[EsmFold2] Fix stale expected distogram logit values (#48182)** (d4c297d)
- **Fix Apr 05 integration test regressions (cuda sm_86) (#48170)** (cc33c37)
- **Fix ROCm SDPA-flash skip guard that crashes on RDNA GPUs (#47965)** (3cb4f80)
- **Fix integration test expected values for cuda sm_86 (Mar 15 regressions) (#48168)** (251c2d4)
- **Fix DynamicCache reconstruction during ExecuTorch export (#47900)** (3cd1a23)
- **[LLaVA] Fix pixtral integration tests for cuda sm_86 (#48166)** (945dac9)
- **[Mistral3] Fix batched integration tests: padding_side=left + update expected values (#48161)** (dd6c64b)
- **Fix DeepSeek V2 default vocab size (#48159)** (e4dd56a)
- **[InternVL] Fix stale expected values for Llama integration tests (cuda sm_80) (#48153)** (0c92811)
- **[Whisper] Fix speculative decoding: preserve cleared suppress tokens through super().generate() (#48108)** (b668804)
- **fix bugs for clvp model (#47127)** (9ad4b85)
- **Fix ASR pipeline mono conversion for channels-last audio (fixes #47886) (#47888)** (bb8f235)
- **Fix CpmAnt loading: size lm_head to vocab_size (#48012)** (e872ca2)
- **Fix Qwen2.5-Omni / Qwen3-Omni-MoE generation with a compilable cache (#47872)** (279dbcf)
- **Fix mlinter artifact path (#48088)** (f2ef0b9)
- **[Video] Fix convert_to_rgb channel slicing and alpha blending for RGBA videos (#48053)** (e7e8b7f)
- **[Whisper] Fix decoder position IDs for left-padded batches in longform generation (#48028)** (e12c79c)
- **fix(pipeline): preserve model.generation_config precedence over pipeline defaults (#47752) (#47953)** (bc77726)
- **[serge] Fix 2 integration tests for model `generation` failing with `import_or_config` (other (2)) (#48061)** (2dcc93f)
- **[serge] Fix 2 integration tests regressed by commit b9090ae58cda (PR #47096) (#48060)** (02d9442)
- **[GPT2] Fix encoder_attention_mask being silently discarded in cross-attention (#47946)** (e0aeab4)
- **[Fix] Small FA-related test failures in CB (#47341)** (e4075b9)
- **fix: honor empty processor_kwargs={} in multimodal pipelines (#48044)** (606e501)
- **Fix EOS for candidate generators (#47931)** (5b90fc3)
- **fix failed test cases for muse_glimmer (#48011)** (6663431)
- **Fix sliding window cache index off-by-one on wraparound (#47708)** (8c62d6f)
- **[MoE] Fix Blackwell GPU crash with torch._grouped_mm on torch <= 2.8 (#48014)** (2e97985)
- **[Florence2] Fix two integration test failures caused by torch 2.13 and auto-dtype (#48031)** (07371dc)
- **[CI] Fix startup failure in pr_build_doc_with_comment workflow by adding missing get-pr-number dependency (#47971)** (6f4de58)
- **Fix MTP config when mlp_layer_types is absent (#48015)** (298ae85)
- **[ModernVBERT] Fix integration test checkpoint (404 since April) (#48009)** (e378772)
- **Fix cropping (#48006)** (a93de5e)
- **Fix DFlash candidate token device mismatch with device_map="auto" (#47877)** (18388f3)
- **[Whisper] Fix speculative decoding: UnboundLocalError, cache corruption, and speed regression (#48000)** (a5f41e0)
- **[Whisper] Fix batch decode_with_timestamps in WhisperTokenizer.decode() (#47997)** (64b6595)
- **[Whisper] Fix integration test failures on A10G (dtype, stale values, API changes) (#47995)** (16cbed7)
- **[DeepSeekV2] Fix integration tests OOM: use device_map=auto instead of 8-bit quantization (#47991)** (a61d5f9)
- **[OLMo] Fix OOM in logits tests by adding torch.no_grad() (#47986)** (c7ad728)
- **[GPTNeoX] Fix post_processor not overridden when loading from pretrained (OLMo garbage generation) (#47988)** (ce5c8f5)
- **[AXK1] Fix expected logits for CUDA A10G (#47980)** (96fe6dc)
- **Fix gemma4 video to device (#47896)** (c1ff118)
- **[emu3] 🦮 Black Labrador is back! Fix image generation broken since #37033 (#47948)** (60eefbc)
- **Fix CLIP _init_weights when a child module carries quantized weights (#47921)** (1319013)
- **Potential fix for code scanning alert no. 267: Artifact poisoning (#47949)** (463f684)
- **Fix GatedDeltaNet A_log dtype to prevent -inf under bfloat16 init (#47944)** (95940bf)
- **[serge] Fix 2 integration tests for model `got_ocr2` failing with `other` (other (2)) (#47937)** (ed3edd8)
- **Fix Jinja block endings in CHAT WITH MODELS' Writing a chat template … (#47960)** (6d2a421)
- **[tests] Fix expected output for Qwen2.5-VL batch_wo_image on CUDA (#47968)** (53a4389)
- **[serge] Fix 2 integration tests for model `opt` failing with `other` (other (2)) (#47909)** (00e8e49)
- **[serge] Fix 2 integration tests for model `vivit` failing with `output_mismatch` (tensor values differ (2)) (#47566)** (14080ff)
- **Fix Gemma `sliding_window` being halved on every config save/reload (#47940)** (f4062e2)
- **Fix compressed-tensors loading for KV-cache-only quantized models (#47904)** (9b3b02e)
- **fix: correct checkpoints, config annotations, and create_dummy_models improvements (#47902)** (c962ff8)

### Backend
- **release: v5.16.0** (93d1bcf)
- **Stop pairing two-word pretokenized documents in the python tokenizer (#48279)** (6ec0f83)
- **Refactor audio buffer creation to use NumPy (#48295)** (efdc97d)
- **[docs] Kernel supported models (#48258)** (ea1410d)
- **[Gemma4] Investigate flaky test_generation_beyond_sliding_window_1_eager (#48236)** (50adfa1)
- **Let gradient checkpointing skip layers with every_n_layers (#48200)** (2b4e2b3)
- **Disable daily nightly CI (#48292)** (fe1dfb3)
- **Pipeline parallel naive inference (#47289)** (b0c8d42)
- **gs (#48288)** (da7234a)
- **replace xpu-smi subprocess call in benchmark_v2 (#48083)** (b3d7e8c)
- **ignore mlinter ci file (#48267)** (9bb5f6a)
- **Compute MoE load-balancing loss per layer to avoid giant one-hot materialization −99.7% @ 128k (#48131)** (a353632)
- **deterministic layer_types buffer registration in multiple models (#48162)** (a530903)
- **`force_accelerate_hooks` should not hide the signature it wraps (#48156)** (89a7273)
- **[Optimization] Avoid redundant token length initialization loop in StopStringCriteria (#48195)** (b830f5a)
- **[docs] Cache crop (#47950)** (602f674)
- **Post two CI badges on a PR: CPU PR CI and GPU run-slow (#48190)** (e453228)
- **Assign a reviewer even when a codeowner has left, and route models by modality (#48085)** (e277d1e)
- **[VITS] Un-skip test_model_forward (#46375)** (bb593bf)
- **Apply context parallelism to the evaluation path (#48167)** (695bd68)
- **🚨 TP dtensor API inference + training (#47579)** (861f4c4)
- **Port ESMC and ESMFold2 to Transformers (#46419)** (c8df3fc)
- **[Qwen2.5-Omni] Update stale expected values for cuda sm_86 (#48164)** (cdea840)
- **Retry transient network errors (RemoteDisconnected) in github_utils (#48124)** (9511951)
- **Revert "[Quantization]: Refactor is_quantization_compressed for format-based detection" (#48072)** (238a2db)
- **Always tie embeddings for LongT5 and Pop2Piano (#47620)** (5909a46)
- **Use generator with seed for LengthGroupedSampler in Trainer._get_eval_sampler for deterministic eval order with per_device_eval_batch_size > 1 (#48025)** (94f09cf)
- **Enable mlinter findings artifact for inline PR reviews (#48117)** (decba1d)
- **[Video] Warn and return all frames when num_frames exceeds total_num_frames (#48074)** (b67f702)
- **Accept artifact dir as argument in post_mlinter_review.py (#48106)** (604da44)
- **Improve group images by shape (#47964)** (0e2ce39)
- **[PE] Skip test_sdpa_can_dispatch_on_flash for TimmWrapper-backed models (#48064)** (f9b76f2)
- **[Gemma3] Update integration test expected values for A10G (#48036)** (ff09c38)
- **Fallback from 'lanczos' to 'bicubic' when on cuda (#48026)** (e801869)
- **[CircleCI] Enable CI for private forks, no-op for public repo (#48056)** (bea0343)
- **:rotating_light: Leftover processors (#47924)** (5068901)
- **[Gemma3n] Update integration test expected values for A10G + torch 2.13 (#48035)** (f471539)
- **Cohere compass tests (#47895)** (1863da9)
- **Moving mlinter to 0.1.4 (#47918)** (d032a44)
- **[docs] Muse Glimmer (#47882)** (0650ff3)
- **Proper separation of tests (#47943)** (1162241)
- **unpin `pytest` in the `examples_torch` deps (#48023)** (579f7a9)
- **Let the GPU verify caller turn on the memory probe (#48001)** (ec30f2f)
- **Remove duplicate block_sparse_moe assignment in GraniteMoeDecoderLayer (#47876)** (2831798)
- **Bump default flash-attn2 hub kernel version to v3 (#47863)** (6724a56)
- **Align logit distributions for CandidateGenerators using sampling (#48007)** (240b833)
- **:red_circle: Allow tokenizers 0.23.1 (#46381)** (242f5df)
- **[Gemma] Update expected values for A10G (#47976)** (df0db63)
- **[docs] torchcodec + vision processor features (#47935)** (a597f97)
- **Remove stale (None, None) fallback in qwen2_5_vl batch_different_resolutions test (#47972)** (7acb594)
- **Scan a diff in trufflehog, not the whole repo history (#47945)** (0cdd8a1)
- **Make muse glimmer exportable (#47871)** (d112311)
- **[CohereCompass] Minor docs fixes (#47903)** (fe747d8)

### Docs
- **docs: use relative paths for README language menus and add fa/ro entries (#47777)** (64f3045)
- **docs: fix incorrect PEFT anchor link in fine-tuning section (#47927)** (918dbf1)
- **docs: add installation instructions for NVIDIA Spark (ARM64) devices (#47906)** (dfb1af6)

### Chore
- **CI: gate the hunyuan-moe slow test  (#48330)** (949e301)
- **CI: fix muse OOMs (#48284)** (f2c9d59)
- **CI: add vllm-test-init and vllm-test-transformers jobs on dedicated runners (#47934)** (85577db)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/huggingface/transformers?utm_source=github-action)._