## qualcomm/aimet — 2.39.0…2.40.0

_37 commits._

### Features
- **Add physical reasoning datasets and metrics to GenAI Lab (#7840)** (17f4d5f)
- **Add audio (ASR) model support to GenAI Lab (#7783)** (342178e)
- **Add all kv-cache inputs/outputs to LlmTopology** (68e24d2)
- **Introduce session_options option while creating QuantizationSimModel (#7800)** (c173926)
- **Add encodings_to_onnx_qdq for QDQ conversion (#7759)** (fcfc439)
- **Add CUDA FP8 QDQ kernels (#7758)** (644c7bd)
- **Add native KV cache generator mixin and hooks (#7645)** (89ab3cb)

### Fixes
- **Fix onnx calibration state restoration** (df58875)
- **Fix device mismatch bug upon .to() for qmodules with pre-quantized parameters** (4111e41)
- **Stop pinning regression to onnxruntime-gpu 1.19.2** (4ff2a6d)
- **Fix fp4-int8 dynamo export bug** (68588d4)
- **Fix false negative bug in depthwise conv detection** (7729c97)
- **Grace grader OOM fix (#7754)** (9eb4b3f)

### Backend
- **Temporarily pin onnx<1.23.0** (70091e7)
- **Update release notes for 2.40.0 release** (ef119ad)
- **Document encodings_to_onnx_qdq (#7843)** (70b7f8a)
- **Report ONNX sensitivity by node name and ship an LLM driver (#7855)** (a9268e6)
- **Take SpinQuant rotation placement from an LlmTopology argument (#7829)** (00c6089)
- **Absorb graph output name restoration into _add_onnx_qdq_nodes (#7826)** (ced6287)
- **Take AdaScale block boundaries from an LlmTopology argument (#7823)** (c548be0)
- **Throw error when exporting float8, float4, or [u]int4 tensors with onnx<1.19** (73bca97)
- **Default to using all float outputs in make_psnr_eval_fn** (5c41452)
- **Publish a single wheel per platform instead of +cpu/+cuXXX variants to GitHub release** (4fc12d8)
- **Pass named precisions and providers to ONNX QuantSim (#7790)** (ac6d21c)
- **Remove ConnectedGraph-flavored LlmTopology from llm_topology (#7817)** (092a9dc)
- **Run SpinQuant on onnx_ir instead of ConnectedGraph (#7812)** (7b9aef6)
- **Drop ignored connected_graph param from get_decoder_block_boundaries (#7815)** (2d71c85)
- **Auto-upcast fp16 AdaScale block to bf16 in torch (#7776)** (1f03b7c)
- **Analyze LLM topology on onnx_ir instead of ConnectedGraph (#7801)** (23633b1)
- **Exclude depthwise convolution from aimet-torch supergroup** (2a6b723)
- **Render AIMET progress bars as widgets in notebooks, summary lines when not on a tty (#7718)** (576546c)
- **remove greedy sampling (#7799)** (6360587)
- **Use GH token to pull AIHM and use install external repos (#7779)** (8209b14)
- **Implement set_weight_quantizer_to_nvfp4_int8 API** (05cc606)
- **Update creator username to use QTI when launching jobs (#7784)** (ad9820d)
- **Adjust test cases to PyTorch 2.14-style clamping and restore CI pytorch to 2.14** (cd1831c)
- **Update schema to remove default precision values (#7739)** (9b6403b)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/qualcomm/aimet?utm_source=github-action)._