## pytorch/executorch — v1.4.1…v1.5.0-rc1

_706+ commits._

### Features
- **Arm backend: Add static public API manifest for 1.5 (#22735)** (8a98594)
- **Add ExportRecipe support for Arm targets (#22368)** (77fb78d)
- **Cortex-M: add opt-in explicit-layout lowering (#22544)** (9ae78da)
- **Cortex-M: add explicit layout pooling kernels (#22379)** (fcc3eb5)
- **Cortex-M: add explicit layout convolution kernels (#22378)** (0dd8f87)
- **Add quantized stacked-halves RoPE operator (#22423)** (ecd46df)
- **Arm backend: Add pre-decomposition partitioner pipeline (#22514)** (c8d5189)
- **[ET-VK][ops] Extend arange, clamp, and index.Tensor support** (2c1da32)
- **Arm backend: Support rank-3 max_pool2d inputs (#22441)** (1e3b7fb)
- **Add pass to replace input dim order clones with permutations. (#21241)** (7490bd4)
- **Add QAT support and pass hooks to QuantizationRecipe. (#21935)** (ddaabb9)
- **Add metrics to runner (#22283)** (e39bda8)
- **Add oncall to all executorch build files missing one (#20496)** (c65ad53)
- **Arm backend: add TinyStories-42M Ethos-U KV-cache demo (#22433)** (87e0749)
- **Arm backend: Add xfail artifact to dump location** (0264212)
- **Support extra ops modes for LLM Models (#18670)** (80a07e6)
- **Add more CI testcases for Exynos Backend (#18664)** (4f689f3)
- **Arm backend: Add bounded U55 unfold_copy support** (8c21609)
- **Add batch runner (#22116)** (1074d73)
- **Add scheduler for batched requests (#22035)** (9133153)
- **[executorch][native] Add Method + ValueRole to the in-memory IR** (e3a53ea)
- **[executorch][native] Add Graph index lists to the in-memory IR** (8f1fcaf)
- **[executorch][native] Add Node to the in-memory IR** (2cfc4b2)
- **[executorch][native] Add Argument tagged union to the in-memory IR** (5e63f93)
- **[executorch][native] Add Value + index-list ids to the in-memory IR** (b842800)
- **[executorch][native] Add Scalar value type to the in-memory IR** (76f238e)
- **[executorch][native] Add TensorMeta + ScalarType in-memory IR types** (2e2a515)
- **[executorch][native] Add native runtime Program reader + DOT visualizer** (9f078b6)
- **Arm backend: Add TOSA reduction op node visitors (#22340)** (337d320)
- **Add pass to remove non-contiguous dim order from a graph. (#21057)** (fdae101)
- **Support multi-method models in the export pipeline (#21723)** (0c97728)
- **Add bf16 x fp32 gemv decode (#22255)** (0b0e6ea)
- **Add bool coverage for roll kernel (#21728)** (8abd640)
- **[ExecuTorch][WebGPU] Add thin partitioner frontend (#22353)** (ebbc7bc)
- **Arm backend: Add runner CMake API manifest** (d466ed5)
- **Add JNI/Kotlin bindings for the ImageProcessor (#21830)** (6e8f02a)
- **Add ETDump profiling to the Apple frameworks and SwiftPM package (#22205)** (8b76a3e)
- **[ET-VK][ops] Add batch support to q8ta convolutions** (16d3b9f)
- **[ET-VK][runtime][2/3] Add typed dispatch size abstractions** (c99fcd4)

### Fixes
- **Fix Arm recipe BUCK dependency (#22574)** (898795d)
- **Qualcomm AI Engine Direct - Fix a QNN crash on AMD (#22543)** (9036d84)
- **Fix the Samsung MobileBert test setup so it can run (#22550)** (89e108c)
- **Arm backend: Fix ConvTranspose2d batch norm fusion (#22517)** (8ea353d)
- **Fix SLEEF preprocessor macro name to match ATen vec headers** (902fcf5)
- **Fix nondeterministic mypy CI dependencies and lint error exposed (#22482)** (2abf9a6)
- **Fix MLX C++20 errors (#22424)** (a6b115b)
- **Fix contiguous dim order target dependencies** (b213481)
- **Forward fix for apple event tracer header import** (89e4dd4)
- **Fix for llama test (#22364)** (f5102b8)
- **Fix RISC-V YOLO26 layout regression (#22362)** (6067c7a)
- **Fix the Apple event tracer header imports (#22363)** (99f8a17)
- **Fix native static attention sliding-window masks** (4fd1610)
- **Arm backend: fix unbound ARM_SETUP_CURL_PROGRESS_ARGS under bash 3.2 (#22337)** (fba44cc)
- **Fix mutated-buffer tagging after indirect mutation detection change (#22296)** (ff50edd)
- **Arm backend: Fix corner case in convert int64 pass (#22342)** (60cb889)
- **Fix grouped conv1d conversion after decomposition (#22213)** (4866659)

### Backend
- **[RELEASE ONLY CHANGES] Finalize ExecuTorch 1.5 dependencies (#22724)** (f7140a4)
- **Stop the backend suites logging captured output (#22581)** (5cdcc02)
- **Declare dynamic quantization transform dependency in test harness (#22576)** (c570b6c)
- **Lower the CMake version floor so 3.26 through 3.28 can build again (#22570)** (4fc2acc)
- **Cortex-M: expose explicit-layout AOT (#22545)** (bf88c64)
- **Move the shared reusable workflows to linux_job_v3 (#22246)** (f624c34)
- **Detect ATen when a target names it by its resolved label (#22553)** (b903c2a)
- **Fine-tune MobileBert on smaller batches in its test (#22565)** (32722a3)
- **Read a single tensor element as the type the caller asked for (#22567)** (758a896)
- **Run a recipe's edge-manager passes after partitioning (#22474)** (6975be1)
- **Qualcomm: bounds-check the delegate (#22237)** (372aa3b)
- **Arm backend: profile SmolLM2 Ethos-U KV cache on FVP (#22560)** (738f197)
- **[ET-VK][q8ta] Route im2col convolution through unsigned dot** (de3f49d)
- **[ET-VK][q8ta] Route pointwise convolution through unsigned dot** (0c3f3a9)
- **[ET-VK][q8ta] Avoid dynamic im2col vector stores** (df1402a)
- **NXP backend: Building MCUXpresso example** (090f5de)
- **Arm backend: Pass memory mode to Corstone** (e6d1319)
- **Arm backend: handle Python SymInt mod and div (#22557)** (9363f00)
- **Arm backend: Allow FP64 operators to decompose (#22551)** (6ac72d7)
- **Arm backend: Reject complex dtypes (#22555)** (44111b7)
- **Arm backend: Reject unsupported comparisons in FP profile (#22558)** (696a144)
- **Arm backend: Enable per-delegate profiling in perf_monitor (#22511)** (f5d3e54)
- **Bump the PyTorch pin to 2.14 (#22501)** (af6454e)
- **Stop shipping unusable MKL search paths in the Linux wheel (#22541)** (5410b1a)
- **Lower Conv1d atomically through TOSA Conv2d (#22282)** (c65b23a)
- **Honor range-learned scales in the tied embedding quantization path (#22476)** (02ee1ca)
- **[ET-VK][runtime] Resolve invariant quantization shaders once** (fe5d8d6)
- **[ET-VK][runtime] Inline resize update checks** (8d33e93)
- **[ET-VK][runtime] Check resize inputs before outputs** (cb0db9e)
- **[ET-VK][runtime] Replace resize update set with generation stamps** (d3e5f1b)
- **Arm backend: Delegate reflection padding on U55** (843f77e)
- **Arm backend: Clean up of TOSA dialect operators (#22515)** (7e9ac9b)
- **Arm backend: Cast integer comparisons in FP profile (#22503)** (c7bbe0b)
- **Qualcomm AI Engine Direct - [doc] QNN ExecuTorch on Windows (#22502)** (8f50849)
- **Bump MLX pin to v32.2 (#22495)** (135a109)
- **Optimize Enn runner  (#18735)** (5bcf795)
- **Read the query length at build time where it is known (#22494)** (33ed3c5)
- **Give causal attention the mask PyTorch asks for (#22443)** (d5e4ae9)
- **Run attention on MLX wherever the fused kernel can compute it (#22419)** (4747ab7)
- **Enhance workflow to label external PRs (#22381)** (457a2a8)
- **Skip impure ops in constant_prop_pass (#22418)** (684d4bd)
- **Arm backend: Updated docs and checks for VGF (#22440)** (291a6cc)
- **Arm backend: Output path fixup from test name** (73b3aa5)
- **Arm backend: Match PyTorch nearest upsample semantics (#22437)** (1d432f8)
- **Arm backend: Reject identity permutes from TOSA delegation (#22436)** (b55076b)
- **NXP backend: handle tests of mlperf tiny classification (#21811)** (0963441)
- **Arm backend: Materialize symbolic shapes (#22407)** (8b47c60)
- **Save the etrecord the lowered paths already generate (#22303)** (834a4fb)
- **Stop a build without the QNN SDK downloader from failing a lowering (#22428)** (47f6ac2)
- **Correct the runner comment, and tighten the moshi installer (#22425)** (793a079)
- **Run CUDA delegates on the per-thread stream (#22318)** (0334561)
- **Keep the CUDA memory pool warm between delegates (#22312)** (9161ce1)
- **Run the Qualcomm SDK setup on the paths that need it (#22395)** (51f9b07)
- **Drop redundant recompiles before ARM pass retracing (#22162)** (5d8bef4)
- **Qualcomm: reject a delegate whose QNN graph I/O does not match its si… (#22239)** (bb041e8)
- **Give the QNN test suite enough memory to finish (#22421)** (30ca5e7)
- **Refresh the apt index before installing ffmpeg for the Mimi test (#22420)** (85ae3ab)
- **Qualcomm: let the 8a8w QAT activation spec share an observer with per… (#22238)** (e3ddb5b)
- **Adds batching executor (#22115)** (e4e7a69)
- **Report NotFound when a CUDA delegate's weights blob is missing (#22311)** (f2ff3f9)
- **exir: serialize non-finite floats in a form flatc accepts (#22152)** (1b1d4b5)
- **Build the flatc host tools for macOS on an iOS build (#22304)** (448fbfe)
- **Pass the Apple event tracer's C++ handle opaquely across the framework boundary (#22396)** (e7cf828)
- **Report the MLX backend as unavailable on the iOS simulator (#22336)** (44dfa52)
- **Enable use_sdpa_with_kv_cache in the qwen3_5 example config (#22390)** (5b829f9)
- **Fold non-matching end permutes instead of discarding the region (#22293)** (0c6ecea)
- **Stop reading a backend config field that no longer exists (#22298)** (2cd8f26)
- **Export an MLX model for the iOS demo build (#22297)** (2b3a32d)
- **manually unroll block 64 to keep bf16 gemv accumulators in registers (#22346)** (beb5b16)
- **Stop the Qualcomm backend setup from running during an import (#22392)** (519b740)
- **ensure qparams match for slices (#21949) (#21949)** (39d623e)
- **Avoid per-node GraphModule retracing in view/permute propagation (#22066) (#22066)** (283043b)
- **Arm backend: Adjust loglevel in noop-removal (#22397)** (a822ab2)
- **Arm backend: Resolve inferred view dims (#22341)** (5428092)
- **Skip redundant recompile in identical-input transform fusion (#22217)** (90452c6)
- **Move _unittest.yml to linux_job_v3 (#22245)** (f665a3b)
- **Cortex-M: keep the activation on a max pool that will not lower (#22232)** (f939419)
- **Stabilize random model tests (#22262)** (54f41b8)
- **hoist bf16 gemv dispatch (#22256)** (fc2465b)
- **Make MLX work when model is loaded on one thread and executed on another (#22294)** (b46bd04)
- **Correct comments on the Core ML wheel library (#22242)** (68f8880)
- **Transforms: let a backend declare what counts as a permute (#22125)** (2f6d0de)
- **Transforms: let the shared driver decline a reduction it cannot cross (#22124)** (f6babd5)
- **Arm backend: Test minimal Classic ML runner in CI** (9d0d22f)
- **Qualcomm AI Engine Direct - [GenAI Pipeline] PR6: Compilation & inference strategy implementations (#22284)** (c27baa8)
- **Optimize the custom SDPA kernel for BF16 (#21789)** (9b558d9)
- **Stop the Qualcomm backend from doing setup work at import time (#22270)** (7a6c1c4)
- **Import transformers only where the LLM runner needs it (#22268)** (1d280bb)
- **Strip local symbols from the macOS wheel's binaries (#22272)** (f715ad2)
- **Check the wheel is usable with only the dependencies it declares (#22273)** (f2d148b)
- **Publish only the package name in the wheel's top_level.txt (#22267)** (dbddba4)
- **Stop two shipped files from deleting directories they do not own (#22271)** (5e46a63)
- **[ET-VK][runtime][3/3] Enforce device-safe dispatch geometry** (74d024c)

### Docs
- **docs: add WebGPU backend documentation (#22360)** (1cfcdf8)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/pytorch/executorch?utm_source=github-action)._