## unslothai/unsloth — v0.1.815-beta…v0.1.900-beta

_441+ commits._

### Features
- **Add a CI gate for code shapes that heuristic antivirus scanners quarantine (#12178)** (b91172e)
- **unsloth start opencode: size the output limit to the context and add --max-tokens (#12111)** (07f004d)
- **Studio: add a Scroll while generating setting (Auto-scroll or Manual) (#12110)** (e7a5b49)
- **Add a Windows probe for code integrity blocks, and audit bundle signatures in CI (#10408)** (5a28557)
- **Studio: keep a chat's start date in the system prompt, note a new date on the latest user turn (#12096)** (7e80a43)
- **Studio: rework Appearance settings and add flavor color themes (#12030)** (73858d5)

### Fixes
- **Fix Studio launcher repair for missing install id (#6933)** (6c25ccc)
- **Studio: fix inductor CantSplit on torch 2.12/2.13 and speed up the Qwen-Image-2.1 VAE (#12059)** (03d3cc4)
- **Studio: fix Decision API validation and TypeSafe SDK metadata (#12186)** (06911b7)
- **Fix training with accelerate 1.15 on torch without a distributed backend (AMD Windows ROCm) (#12162)** (eaab494)
- **Fix Gemma and Gemma2 embeddings scaled twice on transformers 5.4+ (#12116)** (88d35ac)
- **Fix Gemma2 padding masks during batched cached decoding (#12008)** (f7c9c68)
- **Fix rowwise FP8 scale axes in fused LoRA backward (#11799)** (7581d50)
- **fix(studio): mark /api reads no-store so an idle desktop app stops rewriting its disk cache (#12148)** (1dddc14)
- **Point dsh at Unsloth through a --patch overlay instead of settings.yaml (#12145)** (2c03020)
- **Patch the TRL trainers that moved to trl.experimental (ORPO, CPO, Online DPO, GKD, ...) (#12097)** (69623ef)
- **Fix the prebuilt wheel publish step and shorten the release notes to one sentence (#12068)** (bbe44ab)
- **fix: handle strided cross entropy inputs (#10713)** (9159f4e)

### Backend
- **Update pyproject.toml** (8662ab2)
- **Update _version.py** (6c18b49)
- **GKD: right-align left-padded rows before the student forward (#12200)** (05d47bd)
- **Keep xFormers attention causal when the decoder is called without a mask (#12199)** (825c1f2)
- **Unpack the mirrored uv wheel with python3 -m zipfile instead of an inline one-liner (#12194)** (78e634c)
- **Studio: vendor laya 0.3.5 and hold its weights in float16 (#12202)** (4ab5c77)
- **Stop gpt-oss generation at the harmony tool call token (#11449)** (7bacc24)
- **Studio: name the MCP server while the tool call is still streaming (#9212)** (329a7de)
- **Load Mistral-format checkpoints (params.json only) through transformers' Mistral4 (Mistral-Large-3) (#12144)** (78afe44)
- **Do not reinstall llm-compressor when it is already installed (#6806)** (949b94a)
- **Studio: return no-op instead of 500 for an out-of-range scan folder id (#8397)** (94dda36)
- **Unsloth Studio: stop a Mac browser hiding GPU-only models from a remote CUDA / ROCm / Intel server (#8833)** (60bf93c)
- **Studio: compile VAEs by repeated block, one H3 graph per first render, vectorise the HunyuanVideo-1.5 VAE mask (#12061)** (0771bdc)
- **Studio: compile FLUX.1 dynamic, CUDA-graph the SDXL U-Net, decode one-frame Qwen-Image latents in 2D (#12060)** (c6eb2e6)
- **Load block-FP8 checkpoints in 4-bit (NF4) when load_in_4bit=True is passed (#12146)** (91bcb18)
- **Studio: capability tags in the model picker match the vision pill and get their own colours (#12192)** (da85782)
- **Studio: fused int8 MLP kernels and real-arithmetic RoPE for DiT denoisers (FLUX.1 11% less GPU time per step) (#12083)** (8f615ac)
- **Name utf-8 on the SSLKEYLOGFILE writability probe (#12190)** (ea038e7)
- **Store the sandbox credential path names in pieces (#12188)** (9bc35b6)
- **Count only the retry loop's own sleeps in the JSON fallback backoff test (#12189)** (9970dcc)
- **Pass the MXC host-prep script to PowerShell as plain -Command text instead of an encoded command (#12182)** (c551a10)
- **Join the npm scanner's credential path markers from pieces (#12181)** (b47a0cf)
- **Studio: load a downloaded GGUF from disk when the Hub cannot be read or refuses it (#12161)** (866b5bc)
- **Studio: run the isolated Windows Terminal on cmd.exe with stock git when Git Bash cannot start in MXC (#12164)** (a19dae7)
- **Studio: tell a rejected Hugging Face token apart from an unreachable Hub (#12159)** (32f806e)
- **Studio: grant MXC Tier 3 read access to runtime folders once, not per launch (#12121)** (1508a0e)
- **Studio: Triton-fused VAE norms, caches and attention for every image and video VAE (1.7x to 6.3x decode) (#12078)** (e07672b)
- **Read the gradient checkpointing precondition off transformers' own method (#12187)** (505cfda)
- **Refresh WSL shortcut icons through a Windows Python instead of an emitted P/Invoke (#12184)** (407bb3b)
- **Studio: pick the right model size when the file names are lowercase (#11722)** (856d80e)
- **Studio: list a crashed or cancelled run's saved checkpoints on the Export page (#11600)** (7236e91)
- **Studio: read public Hub repos without a saved token the Hub rejects (#12158)** (ecf63dd)
- **Studio: list HF cache models written without symlinks in the local inventory (#12156)** (6b55c9a)
- **Route GKD distillation through the chunked generalized JSD (#12136)** (41b2ff6)
- **Repair chat templates that always append the generation prompt, also on FastModel loads (#12139)** (e417798)
- **Narrow Kimi-K3's zero-padded KDA A_log to num_heads so the fla backward runs (#12132)** (4e6f0a4)
- **Load speech-to-text models (Voxtral, Qwen2-Audio) through FastModel (#12127)** (e3e0a21)
- **Keep Linear layers an FP8 checkpoint stores in bf16 unconverted (#12124)** (b779aca)
- **Keep flash attention off towers the Auto classes do not register (#12129)** (4780f63)
- **Rebuild CohereTokenizer from tokenizer.json when transformers v5 changes its ids (#12130)** (12b6dbf)
- **Dequantize block-FP8 weights with a ragged last block on 16-bit loads (GLM-5.3) (#12133)** (a9ed05a)
- **Studio: never ask for approval to run search_conversation (#12175)** (269b67b)
- **Studio: keep conversation recall working across repeated compactions (#12174)** (6beb057)
- **Turn off the MoE aux loss for dense models that carry a router config (#12076)** (8af84f1)
- **Studio: list at most the 12 most recently active projects and sections in Move to (#12155)** (3bf6002)
- **Refuse K-EXAONE 2.0 on a transformers that ignores its config (#12128)** (0309133)
- **Build the native image processor at defaults when a VLM repo has no preprocessor_config.json (#12125)** (28c117a)
- **Let a loaded Kimi K2.5 / K2.7 processor take processor(text=..., images=...) (#12126)** (c94ae3c)
- **Studio: LTX-2.3 about 4.5x faster per clip (unguided distilled sampling, compile fixes, hosted FP8) (#12067)** (3d45be3)
- **Phi-4-reasoning-vision: give remote multimodal prep an indexable cache view on transformers 5 (#12123)** (3040eca)
- **Keep Nemotron-H mixer.out_proj unquantized under a caller's BitsAndBytesConfig (#12131)** (876faf0)
- **Attention resolver: respect a declared _supports_sdpa = False, and skip flash_attention_2 when a class's compatible flash kernels exclude it (MiMo-V2-Flash) (#12147)** (c86708f)
- **SAC probe: verify Studio identity before sending a password (#12176)** (741ea5c)
- **Studio: round hover for the Settings close button (#12172)** (3a3cca5)
- **Studio: Chats library in the Library (#12122)** (4b98af8)
- **Tests: stop Windows tests tripping Bitdefender and the 16-bit application dialog on real machines (#12167)** (be9e6af)
- **Studio: ignore an SSLKEYLOGFILE the process cannot write instead of failing every HTTPS client (#12166)** (f9150d2)
- **Studio: Select all only picks the models the search shows (#11485)** (87ff26a)
- **Studio: reset the download progress bar when a retry restarts the file (#11593)** (8fb2fa3)
- **Studio: stop runaway tool output from using up memory (#11723)** (a886277)
- **Gemma2: keep softcapping attention under int32 indexing and fall back to eager if compile fails (#12154)** (9c74453)
- **Studio: price MLX loads in the memory panel and fit an unpinned context to available memory (#10287)** (9828490)
- **Studio: stop a hung system node or npm from stalling setup (#12165)** (c176530)
- **Studio: make compile knobs reach the render thread on torch 2.12+ (#12075)** (a735bcc)
- **Studio: show when a prompt was sent while hovering it (#12170)** (869af7c)
- **Keep scan_packages.py from tripping Bitdefender's Python stealer signature (#12169)** (99c3928)
- **Gemma2: use flash_attn_with_kvcache for cached decoding (#12112)** (379f14a)
- **Windows installers find an NVIDIA GPU on the PCI bus, and say which CUDA it can use (#11166)** (268b771)
- **Remove the reflection-emit apparatus and all four emitted types (#11193)** (bd612c8)
- **Recover CUDA compute capabilities without emitting a P/Invoke type (#11173)** (b7d2b33)
- **Refresh a rewritten shortcut's icon where the shell cannot define the type (#11116)** (ecc691b)
- **Ask Python for the process image table when the native helper is unavailable (#11115)** (7e107f8)
- **Unsloth Studio Installer: ask Python for a path identity before giving up on an exact one (#11104)** (72acc86)
- **Studio: TurboQuant KV cache option for MLX inference (#11170)** (eadcab6)
- **Fall back from FBGEMM for rowwise FP8 on GPUs it has no kernel for (RTX PRO 6000 / 5090) (#12098)** (5d6ecc4)
- **Speed up block-FP8 LoRA training: run FP8 linears eagerly, 8 warps for 128-row GEMM tiles (#12027)** (7c8e50a)
- **Settle the permission step's reloads on the pill instead of networkidle (#12153)** (d96e922)
- **Studio: only offer Agent Skills when Code is on (#12118)** (dbe2c6f)
- **Studio: keep checkpoint compaction under --disable-tools (#12119)** (9290d18)
- **Studio: move Managed accounts to the Accounts tab and drop Mark as unread from chat menus (#12120)** (b9716be)
- **Wait for every diffusion run a test started before undoing its runs dir (#12149)** (8625c8d)
- **Batched serving on the MLX path: several replies decoding at once (#10310)** (1232012)
- **Gate the Mllama CUDA forward test on has_real_cuda (#12143)** (839cdbf)
- **Studio: keep the last full-attention layer unquantized in the MLX KV cache (#11083)** (164f9c8)
- **Pass gradients through compressed-tensors activation quantization so W8A8 checkpoints train with LoRA (#11585)** (9d15a59)
- **Count the MLX grammar engine slot in the Apple Silicon step totals (#12135)** (ef63a38)
- **Studio: show the Hub error for an unreadable GGUF repo instead of routing it to Transformers (#12117)** (3ab8cc6)
- **Unsloth Desktop: Create, Edit, and Delete Skills from the Skills Menu (#11800)** (c0ebb4c)
- **Studio: read more chat attachment formats, and hand the rest to the python tool (#11379)** (57a8849)
- **Keep the Triton MoE grouped GEMM in compiled graphs and index weights past 2^31 elements (#12114)** (b5bae41)
- **Keep ORPO / CPO rows within max_length on TRL 0.29+ (#12115)** (88ac167)
- **Trim comments in the DeepSeek-V4 grouped LoRA, remote-code shims and Mistral-format loader code (#12089)** (10f2575)
- **Treat a reaped grandchild as dead in the formatter timeout test (#12107)** (32d57fc)
- **Unsloth Studio installer (AMD/Linux): explain why a newer ROCm gets ROCm 7.2 PyTorch (#11651)** (84e97f1)
- **Studio: cache and coalesce nvidia-smi reads in the backend (#11995)** (0d2905f)
- **Studio: quantize the KV cache of sliding-window MLX models such as Gemma 4 (#11082)** (c4f0b0b)
- **Default GRPO to TRL's dapo loss with beta 0, and cap CISPO weights at 5.0 (#12088)** (a9982dc)
- **Studio: share MLX VLM prompt-cache snapshot buffers and replay exact prompts (#11659)** (32a35ae)
- **Carry the resolved attention implementation to nested configs a remote config baked flash attention into (Nemotron 3 Nano Omni) (#12099)** (f53b103)
- **Read settings.py as UTF-8 in the unknown-palette test (#12101)** (23b0807)
- **Read settings.py as utf-8 in the palette filter test (#12105)** (23e3fe9)
- **Studio: return an error when a model can't use the tools sent to it (#11960)** (dbba692)
- **Upload all GGUF files in one commit so create_pr opens one pull request (#11857)** (65f27c0)
- **Studio: let full-scope keyless callers auto-switch models, and explain the refusal elsewhere (#11180)** (7be1d14)
- **Studio: keep a pinned context as a request limit instead of refusing KV cache quantization (#11084)** (1759e64)
- **Studio: refuse a hand-set Metal context only past the GPU wired limit (#10804)** (e8c5187)
- **Studio: regionally compile Lumina-2 and HiDream-I1, and re-decide their auto precision by measurement (#12039)** (eec898c)
- **Unsloth Studio / Desktop: let the GPUs picker say how much of the model each card gets (#12015)** (3cbe541)
- **Load the repo AutoProcessor for AutoModel-only repo-code VLMs (#12037)** (7c8a51f)
- **Keep Llama 3.2 Vision off flash attention (vision and cross attention have no is_causal) (#12033)** (7b33e85)
- **Studio: skip the cuDNN benchmark search for the MiniMax-H3 audio VAE on A100 / B200 / RTX PRO 6000 (first render up to a minute faster, 25-29 GiB lower peak) (#12040)** (3a092d8)
- **Studio: keep an eagerly decoded image VAE contiguous on NVIDIA (#12035)** (186b486)
- **Studio: chat attachment cards, chips and the Library viewer (#12017)** (0db1cc5)
- **Studio: document viewer for PDF, Word, Excel and PowerPoint, with origin links in Library (#12001)** (23a87c0)
- **Studio: honor response_format on the MLX backend with grammar-constrained decoding (#10180)** (a00fcbb)
- **Studio: decode the SDXL VAE in fp16 on fp16 GPUs (#12036)** (83fbe4a)
- **Studio: stop generating when the client disconnects on a non-streaming request (#11961)** (b3f8ffe)
- **Studio: reject echo, suffix and best_of on /v1/completions that llama-server ignores (#11964)** (4697865)
- **Studio: stop the dense-quant probes pinning a CUDA context on every card of a multi-GPU host (#11954)** (5592fc5)
- **Studio: don't stop long running tool calls that print nothing (#11969)** (1d8c891)
- **Studio: follow redirect pages when reading a web page (#11856)** (405894c)
- **Unsloth Studio installer (AMD/Windows): don't mistake ZLUDA for an NVIDIA GPU (#11736)** (ecbb0d9)
- **Studio: find models in a Hugging Face cache folder added as a location (#11970)** (49d6c01)
- **Studio: tell Safari which account the Studio password belongs to (#11858)** (ced1753)
- **Stop unsloth chat and unsloth inference switching a loaded GGUF to a different quant (#11855)** (1f186c6)
- **Studio: keep MLX elapsed time and duration when the final eval event omits them (#11967)** (305b055)
- **Studio: keep the Images Cancel load pill off the model panel divider (#12079)** (db60370)
- **Studio: put Reapply next to Generate on the image and video pages (#11974)** (5d566d1)
- **Keep the padding-free column test off the datasets numpy formatter (#12082)** (08f3dff)
- **Studio: quieter settings headings and one section gap on every page (#12064)** (d54b42a)
- **Give the killed formatter grandchild as long to die as it had to start (#12086)** (d4d224e)

### Chore
- **CI: CodeQL advanced setup that analyses only the languages a PR touches (#11998)** (fdbc867)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/unslothai/unsloth?utm_source=github-action)._