## qualcomm/aimet — 2.37.0…2.38.0

_70 commits._

### Features
- **Support spinquant R1 with online embedding rotation** (8960f1b)
- **Support Windows ARM64 in setup-uv-env (#7672)** (391c45d)
- **Add ONNX FP8 (e4m3, e5m2) QuantSim Python integration (#7560)** (29263ff)
- **reindex create-new-action (#7666)** (acd2ff7)
- **Support external embedding rotation for models with no LM head** (9a7726e)
- **Support simulating int32 accumulator truncation for MatMul/Conv** (cc65cdc)

### Fixes
- **Disable credential persistence in release checkout to fix public push (#7702)** (3166759)
- **Fix incorrect LSTM weight shapes in test_lstm_no_optional_outputs (#7681)** (f9a35c9)
- **Skip all nightly and weekly regression jobs on the public mirror (#7667)** (4649f59)
- **Fix CODEOWNERS team reference to qcom-ai-hub/aimet-maintainers** (98516f4)
- **Fix minor type error in case peft library isn't present** (f99d521)
- **Fix deprecated download-artifact@v3 in release workflows (#7633)** (a568796)
- **Move GenAI scorecard and regression tests to new runners (#7625)** (71d9619)
- **Move CI onto static runners and fix the internal-repo gate (#7619)** (9b63884)
- **Fix missing pypirc in PyPI publish job (#7624)** (278c819)
- **Fix sin/cos output range to -1..1** (2c95329)
- **Fix sequential MSE failure in deepspeed 0.19.4** (cee185a)
- **Fix VLM visual prefill to yield dicts of named inputs (#7604)** (c3499a5)
- **Patch extra generation config EOS ids into ONNX exported config (#7594)** (c184f2f)

### Backend
- **Place online R1 rotation at final residual if no lm_head (#7696)** (b63789f)
- **Update release notes for 2.38.0 (#7691)** (d06b6b5)
- **Skip overflow-underflow scale adjustment if weight quantizer is frozen** (b841e4a)
- **Prune unnecessary onnx QDQ export test cases** (98127e4)
- **Print top 10 longest-running test modules in CI pipeline** (4161df2)
- **Re-enable export-time backward propagation of Pad encoding** (a58515f)
- **disable breeze ci code review (#7687)** (69f7bdd)
- **Downscale excessive num_iterations in adascale unit test** (649c762)
- **Revert Pad/Where from grid-equivariant ops (#7680)** (9aca552)
- **Thread the FP8 CPU QDQ kernels through IForLoopRunner (#7678)** (1f61053)
- **Speed up the FP8 CPU fake-cast with frexp/ldexp (#7677)** (90b7acd)
- **create breeze integration for code review and issue comment (#7676)** (2070117)
- **Implement export-time encoding propagation of clipping ops** (2d151b3)
- **Route Windows CI jobs by aimet-windows-{amd64,arm64} labels (#7670)** (cb45b29)
- **Temporarily exclude aten.copy from grid-preserving op** (0b80bbe)
- **Skip loading encodings if quantizer is frozen** (1feeb61)
- **Raise hrnet_pose accuracy-drop threshold (#7668)** (f50e4ea)
- **Update xfail message for _safe_softmax aten-qnn ir alignment test** (57a7789)
- **Move Pad/ScatterElements/ScatterND to grid-equivariant op** (e8ce349)
- **Implement simulation-time encoding tying of grid-equivariant ops in aimet-torch** (35ce986)
- **Install patchelf from PyPI to satisfy auditwheel >= 6.8.1** (49a86d0)
- **Accept QSpec argument to set_param_type** (e29e8d9)
- **Implement simulation-time encoding tying of grid-equivariant ops in aimet-onnx** (926bf6f)
- **Implement export-time encoding propagation of grid-equivariant ops** (89bf3b7)
- **Treat aten::group_norm as QNN IR operator** (9c0ca1b)
- **Load LPBQ encodings directly to QcQuantizeOp** (354a33e)
- **Extend definition of grid preserving op into multi-input grid-equivariant op** (00aed7a)
- **Define set_qspec API for QcQuantizeOp** (495702e)
- **Extend ATen-QNN IR alignment test cases** (4356a48)
- **Treat torch quantize/dequantize ops as quantization encoding in ONNX QDQ export (#7630)** (d76bb1e)
- **Reduce usage of large pretrained weights/datasets in llm tests** (a32a82b)
- **Resolve implicit merge conflict** (f0556eb)
- **Update action versions (#7626)** (0576a4c)
- **Gate workflows on repository owner instead of server host (#7624)** (742e103)
- **Use the public PAT for the repo-sync mirror push (#7623)** (b9c496e)
- **Preserve graph signature throughout aimet_torch.export** (443f638)
- **Change secrets to vars (#7625)** (0d111e1)
- **Exclude getitem from grid-preserving op** (0c294c0)
- **Re-map output encoding of UnsafeBarrier node to its producer output** (cebedfa)
- **Refactor GroupedBlockQuantizeDequantize class into scale_quantizer attribute** (1e1ac74)
- **Resolve false positive deepspeed unit test failure in edge case** (5226a83)
- **Refactor llm_topology tests: dedicated unit file + extracted fixtures (#7573)** (041170a)
- **Pin torch thread counts in unit tests and revert CPU bump** (7429a6a)
- **Export transposed linear weight as param encoding in v1.0.0 format** (2411e7f)
- **Delete enable_recompute API permanently** (006db48)
- **Switch to relative imports (#7612)** (b51c295)
- **Ensure non-negative relu input encoding in static ATen calibration** (165180f)
- **Request 64 CPUs for unit test runners** (727c8b7)
- **Avoid session rebuild for GenAILab prefill** (924f5da)
- **Dedup shared initializers at the ONNX float-export boundary** (30df35a)
- **Insert barrier around RotaryEmbedding to suppress encoding propagation** (0d653dc)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/qualcomm/aimet?utm_source=github-action)._