## avifenesh/memra — v0.97.0…v0.98.0

_143 commits._

### Features
- **rung 4 hygiene: clippy clean on the new code + unblock crate-wide clippy** (6ba3952)
- **feat: dsv4-mint — 0731 MXFP4->NVFP4 lossless cast + 0731 census gates (mint lane; one new bin, no shared-module edits; input_scale OMITTED per measured preview calibration; vestigial num_nextn_predict_layers derived from compress_ratios)** (46b9c97)
- **feat: lane-7 gate-a swiglu_limit saturation cell (recon intel: vLLM sm12x fork lost the clamp) + recon notes banked (CUTLASS #3096 = perf-rung landmine; no CUTLASS in the correctness arms)** (eea0541)
- **feat: lane-7 native quantized expert GEMM arms — act_quant-fp8 codes kernel (per-128 pow2-ceil, kernel.py law), dsv4_fp4_gemm (NVFP4 per-16xscale_2 / MXFP4 per-32-e8m0, fixed-tree f32, CPU-mirrorable), byte-row gather; MEMRA_DSV4_EXPERT_ARM=native seam on BOTH the GPU forward and the CPU oracle (expert act-quant emulation, M:604-606 order); dsv4-native-gemm-gate (bit-exact mirror + f64 bound + class-shift vs rung); class-aware C(d_b,d_q) thresholds in gpu-gate/decode-gate/oracle-check/greedy-verify; moe_x capture + mtp capture pass-through; shared-expert tail factored (stays bf16 under both arms)** (24aa7ab)
- **feat: lane-6 true decode path — per-layer ring/pending/store caches (reference decode state machine), prefill handoff population, single-token block_decode + decode_step (PP bounce per step); gate binaries dsv4-gpu-decode-gate (a1/c/d/e, corrected pair bounds + m-floor context) and dsv4-decode-oracle-check (a2/b vs CPU oracle); dsv4-decode-probe bisect instrument; indexer_score gains lim0 (decode causality lives in the store)** (d61e027)
- **feat: dsv4 MTP GPU path — MXFP4 expert slabs (detected, refused on surprise), MtpDev on the last stage, shared-head mtp_logits_last; gate un-skips the 15th array + MXFP4 sub-gate samples** (1e4723a)
- **feat: dsv4 GPU trunk bring-up — 2-card layer split, oracle-parity kernels (-fmad=false), bf16 GEMM rung, gate + greedy + CPU-verify binaries** (7026e6e)
- **feat: dsv4-mtp-probe — decisive counting probe for the NextN token-shift convention** (e6e2d2f)
- **feat: dsv4 CPU paired-forward + dsv4-forward fixture gate (forward lane)** (3a5c500)
- **feat: dsv4-census gate — full-artifact header census + sample decode + structure checks** (ff9bf0a)
- **feat: DeepSeek-V4-Flash arch config + tensor census + CPU quant decoders (loader lane)** (1fa8048)

### Fixes
- **Merge branch 'lane/dspark-harvest-fix-20260820' into release/v098-candidate** (becc35b)
- **dspark(q38) slice 3 fix: parity-proof ssm stash via indirect-source copy** (727daa8)
- **ci journal: perf-quick rows for lane/dspark-harvest-fix-20260820 (31b cells flat: plain 42.29/39.84, spec 108.07@0.798 / 102.42@0.817, 0 fail 0 warn; correctness stage green)** (77ffb69)
- **dsv4 it5 F-side arms on the it5 tip: MEMRA_DSV4_ROUND_PROFILE instrument + MEMRA_DSV4_DSPARK_CHAIN=device (transport-only, default host) + MEMRA_DSV4_DSPARK_MARKOV=rowblk (bit-exact row-blocked markov GEMV, default base) — banked patch scripts applied verbatim, all anchors held over the fp8 commits** (b15447c)
- **fix: EAGLE3 draft rope width goes through resolve_rope_dim_count, not a third implementation** (7ac984d)
- **rung 4d: post-hygiene regression gate PASS both arms (accept shas match the pre-hygiene run); bank the branch-ref deletion hazard + the worktree-reflog recovery** (23199f3)
- **data: 0731 re-gate ALL GREEN — task A (output-sample 14/14 REF+clamp on the mint, tf-gate 158/160 in-band-only both arms, decode 52/52 + CPU teacher-forcing 256-258/260 all in-band, preview witness byte-identical after the probe fix, byte-inertness shas exact) + task B f32x rung green pending ratification (257/260 in-band, -5.1 to -6.1 ms/step informational); logs+JSONs banked** (3b3bb9c)
- **lane-9 rung E verdict: revert sink-scores take-3 (211.9 us/inst, 3.1x regression banked) and gemv 2-row (13.44 vs 13.04) to the lane-8 shapes; fp4_sel tables kept; negative results banked in-source** (eb31f6f)
- **lane-9 rung E (bit-exact, ncu-guided): sink scores take-3 = 4-slot ILP + hoisted products (ncu: 0.2% DRAM / 8.3% SM latency bound); gemv_bf16 2-row blocks (wo_a class was 48% DRAM); fp4_gemm_sel back to smem tables (ALU decode regression 37.8->45.9 us banked in-source)** (b7a73df)
- **lane-9 rung A take 2: K1 scores = lane-8 shape + float4 q vectors (same f64 chain), K2 zero-skip reverted; take-1 smem-phase regression banked (177.7 us/inst); K3 8-wide chunks kept (73.5 -> 31.9 us/inst)** (7d3c6b7)
- **lane-8 receipts: rung A/B/B' gate results banked (hostmath byte-identity chain, device-arm v1 gates + perf regression finding) + interleaved A/B driver** (daa296e)
- **lane-8: route/sinkhorn kernel v2 — parallel transcendentals + row/col thread ownership (bit-identical values to v1; v1 single-thread forms measured as the device-arm regression: 55.5 vs 44.2 ms/step hostmath)** (2855d68)
- **data: lane-7 ALL GATES PASS — final receipts: 21/21 bit-exact kernel gate + saturation cells, output-sample 15/15 native (logits 4.996 vs thr 107.6) + bf16 regression identical, greedy 157/160 all in-band, decode a1 52/52 + a2/b green, CPU-only class-shift pin 3.68 max-abs; speed shape 91 vs 101 ms/step (informational); support claim + remains** (1dc08d0)
- **fix: decode gate (c) short-probe criterion per the banked lane-7 protocol (n_new>=40 + block/wrap/coarse boundary checkpoints when s<1024; long-probe criterion unchanged when reachable)** (5232a09)
- **data: lane-7 gate (b) PASS native 15/15 (final logits max-abs 4.996 vs thr 107.6 = 0.05x of the class bound; top1/top5 exact, 0 out-of-band; layer0_attn_out canary 1.832e-2 == lane-4 exactly) + bf16-arm regression PASS numerically identical to the lane-4 banked table** (ffada78)
- **fix: lane-7 gate-a zero-sign canonicalization — __nv_cvt emits -0.0=0x80 for tiny negatives that RNE to zero, house CPU encoder canonicalizes to +0 (decoded values identical); dsv4 act-quant kernel canonicalizes (c&0x7F)==0 -> 0; named-mismatch diagnostics in the gate; run-1 log + analysis banked before rerun** (ed2940f)
- **data: lane-6 gates ALL PASS — (a1) 63/63 vs re-prefill at pair bounds, (a2) 62/62 vs CPU oracle at lane-4 thr, (b) 2164/2176 all-in-band, saturation reached, byte-determinism, cache math exact; fp4/act_quant grid.y ceiling fix (crash at s=1024 prefill) with lane-4 gate rerun identical; 102→112 ms/step decode vs ~1100 re-prefill (informational)** (1fa31c7)
- **fix: lane-4 derived top-20 boundary-band rule (replaces fixed overlap floor) — banked before rerun; MTP run-3 log analysis** (a44682c)
- **fix: lane-4 gate-policy corrections after run 1 (extreme-value factor, e2m1 flip bounds, top-20 rule) — receipts banked before rerun** (043a83a)
- **fix: compressor census — two measured shape classes keyed on compress ratio** (d51cb17)

### Backend
- **release: v0.98.0 — the ratified dspark bundle (version + internal pins)** (bc466cc)
- **Merge release/v098-candidate @ 22cde04bb9 — the owner-ratified dspark+dsv4 bundle train (v0.98.0)** (3adcd14)
- **dspark(q38): tap-sink pool return holds on the Err path too (review carry-over ×3)** (22cde04)
- **dsv4_gpu.cu: the NVTX include is optional (__has_include) — CI's minimal toolkit has no nvtx3** (1d51330)
- **dsv4-decode-probe: refuse on the RESOLVED dense arm, not the literal env** (e5e59d4)
- **ci journal: full rig battery rows for the it5 composite (rc=0: correctness all green incl. serve-smoke 0 failed, perf 9/9 above rolling median, 31b-plain-short 42.24 [OK]; attempts 5-6 were killed by co-tenant rig GPU contention — recorded)** (475e130)
- **boundary: prune the 33 TRANSIENT historical-blob pins (v0.96 §5 seam)** (9ddea26)
- **merge lane/dsv4-it5 @ e813d0f2e5 — coordinator ruling: item-3 cells green on box7, it5 joins the v0.98 train** (d1e0595)
- **ci journal: perf-quick rows post-fmt for release/v098-candidate (correctness green; 31b cells flat: plain 42.2x/39.8x, spec 108.0x@0.798 / 102.31@0.817, 0 fail 0 warn)** (31b8f42)
- **ci journal: full rig battery rows for release/v098-candidate (+fmt on the merge guard)** (749f287)
- **dspark(q38): ratified default flips — harvest strategy-keyed, verify-window confidence-slot tau=.5** (b0a53ec)
- **merge lane/surface-oracle-fixes-20260820 — oracle manifest: config.json-first + harvest field** (2f319c5)
- **Merge branch 'lane/nrot-safetensors-20260820' into release/v098-candidate** (8bb2140)
- **merge lane/dspark-engine-bundle-20260820 — reconcile with v0.97.0 ROUND-STREAM arm** (3f2d498)
- **Merge branch 'lane/dspark-vt-confidence-20260820' into release/v098-candidate** (4a9a21a)
- **data: cross-engine cell ornith15 — vLLM 249 both probes, memra-spec 234.6/208.8 (0.94x/0.84x, only engine using the checkpoint MTP head), llama.cpp 205.6/217.6, memra-plain 198.5 (untuned); N=6 both orders, protocol attached** (78d55ab)
- **ci journal: perf-quick rows for lane/dspark-engine-bundle-20260820 close (correctness green; 31b cells flat: plain 42.29/39.85, spec 108.11@0.798 / 102.58@0.817)** (f62c273)
- **dspark(q38) slice 3: default OFF (opt-in MEMRA_DSPARK_VERIFY_GRAPH=1) — measured disposition** (21a8594)
- **dspark(q38) slice 3: instantiate segment graphs with UPLOAD, not AUTO_FREE_ON_LAUNCH** (fe987d3)
- **dspark(q38) slice 3: graphs persist across generations (model-owned ctx)** (01e86a1)
- **ci journal: perf-quick rows for lane/dspark-engine-bundle-20260820 slice 3 (correctness green; 31b cells flat: plain 42.21/39.85, spec 108.07@0.798 / 102.45@0.817 — the graphs ctx is bin-arm dspark-only, MTP untouched)** (6cb50c6)
- **dspark(q38): engine-bundle slice 3 — per-(segment, vt) verify graphs for the GDN linear runs** (82598d1)
- **ci journal: perf-quick rows for lane/dspark-engine-bundle-20260820 slice 2 (correctness green; 31b cells flat: plain 42.19/39.83, spec 108.03@0.798 / 102.56@0.817 — the core_stream vtok_dev seam is inert for MTP)** (77b2b31)
- **dspark(q38): engine-bundle slice 2 — defer the chain readback, one sync/round** (15f94be)
- **ci journal: perf-quick rows for lane/dspark-engine-bundle-20260820 slice 1 (correctness green; 31b cells flat: plain 42.32/39.91, spec 108.63@0.798 / 103.43@0.817, accept + tok/round unchanged — the batched cols restore is exercised by the MTP spec cells)** (01655af)
- **dspark(q38): engine-bundle slice 1 — batched GDN state snap/commit copies** (2e166ea)
- **fp8 dense arm: drop the bf16 dual residency (it5 ledger item 3) — trunk bf16 slabs hold HOST staged residency under the arm (DenseBf16::Host, the hybrid/moe-cache staged idiom: same f32_to_bf16_exact bytes, staged H2D per prefill pass, transient frees stream-ordered at pass close); device decode/verify reads stay on the fp8 twins (dwsel), legacy stays a boot refusal, drafter/MTP blocks keep resident bf16 (no twins, .dev() fails loudly if a future rung stages them), dsv4-decode-probe refuses the fp8 arm at boot (hermes shape — its instrument reads resident slabs); loaded_bytes counts device bytes only; census arithmetic ~-2.7 GiB/card vs bf16 arm, ~-5.4 vs dual-resident builds — nvidia-smi cells OWED (2-card window)** (e813d0f)
- **data: Ornith-1.5-35B-A3B artifact + MTP-head training receipts (lane/ornith15-st-nvfp4-20260819, squashed)** (05baebc)
- **ci journal: perf-quick rows for lane/dspark-vt-confidence-20260820 (default-off seam: correctness stage green, perf 0 fail 0 warn, touched-trunk cells flat)** (a85cfb4)
- **dspark(q38): H4 confidence-driven verify windows — MEMRA_DSPARK_VT={confidence|confidence-slot} (+_TAU), ladder default-unchanged** (6673147)
- **dspark(q38): harvest-convention seam MEMRA_DSPARK_HARVEST={dflash|dspark} (default dflash until B1 ratifies) + DSPARK-strategy oracle reference + toothed convention gate** (9745062)
- **dspark q38 oracle: config.json-first fail-closed manifest, written BEFORE the dumps, with library versions** (2f66cdc)
- **graph_update: set_exec_params is unsafe, with the driver's contract written down** (5bb810e)
- **surfaces: one body for per-request sampling defaults + worker-truth four-surface parity teeth** (3ace9ad)
- **fp8 dense arm: unroll-by-2 with early weight loads in both GEMV twins (load scheduling only — per-thread accumulation order verbatim, bit-identical; MLP for the halved-width uint2 reads)** (01bbd99)
- **dsv4: illegal env combos are boot refusals, never post-build aborts (hermes fingerprint a4e3d9a8eab4cf17)** (8ba15b6)
- **fp8 dense arm: smem e4m3 LUT in both GEMV twins (bit-inert decode transport — table values are dsv4_e4m3's own; arithmetic decode cost exp2f+divide per element and measured ~7% plain where halved bytes price ~30%)** (0bfa60d)
- **dsv4: scope the ratified f32x dots default to the device decode path (legacy + unset env resolves f64 — the ratification's own wording; explicit f32/f32x on legacy still refused). Box4 find: dsv4-gpu-gate panicked at load under the flipped default.** (82a754f)
- **dsv4 iteration 5 (box4): FP8 dense arm (MEMRA_DSV4_DENSE_ARM=fp8) + owner-ratified f32x drafter exit-head default** (28d1eeb)
- **boundary allowlist: pin 10 historical receipt blobs carrying the dead box1 instance id (redacted at tip; grandfathering deliberate, surfaced to owner)** (546f0d2)
- **boundary: redact EC2 instance id from lane receipts (conditions banked in darklanes box-mirror); cargo fmt on it4 bins** (692efb8)
- **dsv4 iteration 4: verify-depth knob + SPS/confidence telemetry + fmad pricing instrument** (03d4bf2)
- **iteration 3 rung 4 receipts: batched verify GATED (bit-exact vs sequential), drafted A/B measured on owner corpora (1.383x / 57.0 tok/s), graphs decision = NO** (bcfa0c1)
- **rung 4c: drafter exit-head dots hoisted (bit-exact) + measured f32x arm fork** (4dd459d)
- **rung 4: bench WIDE bands + per-round emitted banked** (5299a9e)
- **iteration 3 rung 4: owner-corpora drafted cell + drafted-round profile bracket** (ea5c87e)
- **iteration 3 rung 4: batched T=k+1 device verify + §3.1 device commit/rollback (BIT-EXACT vs sequential)** (0629622)
- **iteration 3: §3.1 ring-hazard gate ALL GREEN x2, f32x ratified as default (A/B x5: 38.6->24.7 ms/step), GPU drafter gated identity-clean** (13230ca)
- **iteration 3 receipts: §3.1 ring-hazard decision record (hybrid transient-kv + snapshot-rollback + hwm + accepted-only drafter writes), drafter VRAM plan arithmetic (dev1-resident, ~600-700K ctx cap with drafter ON, split disproven by the free-VRAM sum), GPU port + perf design** (c3887b8)
- **batched gate: MEMRA_DSPARK_BATCH_SMOKE truncation knob (fast machinery feedback, banner-marked, never a verdict)** (2e3a0a4)
- **iteration 3 rung 1: §3.1 ring-hazard machinery on the CPU oracle — batched T=k+1 verify (transient window kv + recorded compressor advance), snapshot-rollback with bounded replay + store high-water marks, TrunkOracleBatched seam + run_spec_greedy_batched (sequential-parity round/budget accounting), dsv4-dspark-gate 'batched' mode (per-round bit-level state-class compare vs a sequential twin); decision banked in CompCkpt/decode_batch docs; unit-pinned rollback mechanics** (4a9ac81)
- **iteration 3: merge lane/dsv4-dspark (drafter oracle, spec_oracle seam, dsv4_decode stateful port) into the loader lane @ 3b3bb9ca8d** (f8ba684)
- **lane-10 receipts: q38 box2 correction bound to the acceptance profile (correctness observable only, no wall-speed claim)** (69160d4)
- **0731 re-gate: NextN-vs-DSpark probe uses the RAW tensor name mtp.0.e_proj.weight (stem probe missed — preview witness caught the NextN block not loading; measured key sets banked in-comment)** (ad9f0ad)
- **0731 re-gate: derivations banked BEFORE runs (REF GPU-class fork rule, tf-gate band, boundary protocol, f32x fork class + gate order)** (3bea90c)
- **0731 re-gate lane: GPU bring-up on the 0731 lineage + the f32x extension rung (code, gates pending)** (de8de1c)
- **lane-10 final receipts: full gate table ALL GREEN (clamp 45/45 structural, ref 0/159 trunk flips + 9 in-band id near-ties, greedy LITERAL spec==plain identity 160/160 x2 across seam refactor), wall clock, end-state** (b81e600)
- **gate-formula correction (derivation banked before rerun): REF logits threshold = measured contract fork / 3; ref-only fp4 one-lane flip allowance on index_score** (70485b3)
- **dsv4 0731 oracle lane: derive nextn from compress_ratios (vestigial num_nextn_predict_layers trap), trunk-only fixture sets (optional mtp_top20, gate skips mtp_logits_last when unbanked)** (4b61cb2)
- **lane-9 receipts: informational lane-7 column at 8200 (90.4/99.6/109.6; cumulative 3.03x/3.12x/2.95x)** (9b5de6f)
- **lane-10 receipts: Gate D2 run1 PASS — literal spec==plain identity 160/160, zero corrections, 3.865 accepted/round** (217a233)
- **lane-9 receipts: THE HEADLINE A/B x5 (n_new=8200) — lane-8 stack vs lane-9 stack 1.43-1.46x (42.60 -> 29.80 ms/step, 33.6 tok/s s~200; 27.0 at 8k), byte-stable arms x5, conditions line; lane summary (claim + bar math + contradictions + remains)** (065b74f)
- **lane-10 receipts: clamp arm PASS x2 (procedure #2 ratified), ref-arm x2 records** (8ff156d)
- **lane-10 seam refactor (owner directive): spec_oracle module = family-generic propose-verify arbitration loop (TrunkOracle/OracleDrafter traits, run_spec_greedy); dsv4 adapters (DsparkOracleAdapter, TrunkOracleAdapter); greedy gate rides the generic loop** (e1ff0c1)
- **lane-10 decision procedure #2 ratified (clamp arm 45/45 @1e-3*absmax, worst 1.2e-4): ref-arm float comparisons demoted to measurements; draft-id in-band budget 12** (fa5877e)
- **lane-9 receipts: rung record (A/B/C/D takes + negative results), fork gate results (255/260 all in-band), tip profile table, graphs verdict (not adopted: mass 28.93 > 25 flip condition, gap 1.83ms; flip comes with drafter or extended f32 ruling)** (b11a539)
- **lane-10 receipts: full torch bank + ref-run1 record + PRE-REGISTERED decision procedure #2 (clamp arm adjudicates before its evidence lands)** (4c73795)
- **lane-10: greedy cross-check applies the banked flip policy (in-band flip rounds adjudicated with margins, accepts checked vs own prefix); Round struct** (1624a7b)
- **lane-9 rung D (bit-exact trio): route parallel argmax tree (host tie rule, associative), rowsq/rmsnorm 8-wide load batching (per-thread order unchanged), fp4_gemm_sel ALU decode (probe-proven bit-equal e4m3/e2m1 constructions replace conflict-serialized smem LUTs)** (9d88cc3)
- **lane-10 receipts: smoke records + logs** (203e88f)
- **lane-10: gate doc header updated to the corrected two-sided doctrine** (cbe4e6e)
- **lane-10 gate-formula correction #1 + flip policy (banked): two-sided components gate (clamp arm structural @1e-3*absmax, ref arm flip-noise @fork), trunk realization-flip adjudication with banked torch margins, greedy in-band corrections** (fffdd24)
- **lane-9 rung C: f32-accumulation island-dots serving arm (owner-gated fork) behind MEMRA_DSV4_DOTS_ARM — dsv4_dots_f32acc kernel (gemv-class fixed tree, vectorized), device-path-only routing (dots_dev), legacy/prefill pinned to f64; refuses f32 without the device path** (e8bd3dc)
- **lane-10 receipts: plan of record + gate formulas banked before runs** (19f8b94)
- **lane-10 dspark: CPU decode-state port (ring/pending/stores per model.py decode branches) + DSpark drafter oracle module + dsv4-dspark-gate (components + spec==plain greedy e2e); dot/par_rows made pub** (581d531)
- **lane-9 rung B: fp4_gemm_sel bit-exact 4-column blocks (amortized smem tables + L1-hot activations; per-column group ownership, in-group order, and halving tree unchanged)** (7477792)
- **lane-9 rung A: sink-trio bit-exact reshape (K1 smem-staged coalesced q tile, K2 zero-skip denominator, K3 8-wide head-chunks) + rung-0 fresh profile table + owner ruling banked (dots_f32 unblocked as gated f32 serving arm) + rung C fork derivation** (96d5c09)
- **lane-9 plan of record: rung plan (sink trio bit-exact reshape, fp4_gemm_sel floor, launch fusions, graphs measure-dont-assume) + gate protocol, banked before the fresh profile** (8470ff7)
- **gate-formula correction (derivation banked before rerun): REF logits threshold = measured contract fork / 3; ref-only fp4 one-lane flip allowance on index_score** (b13baff)
- **dsv4 0731 oracle lane: derive nextn from compress_ratios (vestigial num_nextn_predict_layers trap), trunk-only fixture sets (optional mtp_top20, gate skips mtp_logits_last when unbanked)** (ef9afa9)
- **lane-8 final receipts: greedy-mode verification, 8k near-flat shape table, lane summary (claim + gate chain + ranked remains + bar math)** (ac477a9)
- **lane-8 receipts: seam attribution A/B x5 at final tip (1.40-1.44x path term; rung-A term 1.52x)** (95ee1bb)
- **lane-8 receipts: THE HEADLINE A/B x5 — lane-7 baseline vs lane-8 device stack: 2.12-2.19x (90.4->42.6 ms/step, 23.5 tok/s at s~200), byte-stable arms, conditions line** (6138b35)
- **lane-8 receipts: rung-B' seam A/B x5 table + rung C gates banked (GEMV contradiction: matches cuBLASLt m=1, -0.7ms)** (b52ccf6)
- **lane-8 receipts: device v2 bit-identity + second profile (gap layer 1.8ms -> graphs condition not met, banked; clock check; remaining ranking)** (eae6fe7)
- **lane-8 rung C: deterministic fixed-tree bf16 GEMV replaces the cuBLASLt m=1 GEMMs on the device decode path (class-II f32-reorder fork, to be decode-gate + oracle adjudicated; nvjet measured ~10ms/step at 2.3x off weight bandwidth)** (e05782d)
- **lane-8 rung B': sink-attention three-kernel split for the device path (bit-exact: same f64 dot chains, same slot-order denominator/output accumulation, pads/underflow skip proven bit-inert; scores/evals/den ride the arena)** (042d542)
- **lane-8 rung B: device-resident decode step behind MEMRA_DSV4_DECODE_PATH (arena StepWs, device build_idx/topk/sinkhorn/route/head-gate + hostmath byte-identity arm, indirect fused expert dispatch = one launch per projection, peer-copy PP boundary + events, device argmax greedy mode; tid2eid validated at load; native-arm requirement asserted)** (eeeb6b9)
- **lane-8 rung 0+A: banked nsys profile (step is kernel-underparallelization dominated — contradicts the host-round-trip premise; 153 D2H drains cost only the ~12ms gap layer) + BIT-EXACT reparallelizations: indexer_score block-per-(t,j)/thread-per-head (22.4ms/step -> sub-ms class), fp4_gemm vectorized loads + smem decode tables (same products, same order); lane-8 device-step kernels + FFI staged (sinkhorn/route/gemm_sel/combine/build_idx/topk/argmax/rope_at)** (0a0a53a)
- **lane-8 bench: cudaProfilerApi bracket (steps 16-48) for nsys rung-0 capture** (77ef389)
- **lane-8 plan of record (structural inventory of the decode step, rung plan, gate protocol, A/B banking) + dsv4-decode-bench instrument** (4312c40)
- **data: lane-7 gate (d) PASS — decode gate native (a1 52/52 at pair bounds, worst 3.10; short-probe boundaries covered; determinism; cache math exact) + oracle-check (a2 all checkpoints, worst 6.37 vs ~214 pair bound; b 252/260 all 8 in-band, margins 0.036-0.75); run-1 c-criterion log banked** (ea6cd60)
- **data: lane-7 gate (c) PASS — native greedy 160 tokens x2 byte-identical; CPU quantized-oracle teacher-forcing 157/160, all 3 disagreements in-band at margins 0.064-0.446 (inside even the bf16-class band)** (23309b8)
- **data: lane-7 gate (a) PASS — 21/21 projections BIT-EXACT (both recipes, both lane-1 pins), f64 bounds all pass, class-shift measured 0.47-0.51x of the (u_q/sqrt3)·||x.w||2 prediction, swiglu_limit saturation cells PASS (clamp engaged, h at expf-ULP, GEMM legs bit-exact); run-1 FAIL + run-2 diag banked (zero-sign canonicalization)** (1e87f91)
- **data: lane-6 run-1 bisect — reference not realization-stable (m-floor 0.18-3.08 measured, control table banked); gate-policy corrections derived and banked before rerun** (6da5121)
- **data: lane-4 tip rerun cert — gate table identical across builds; push-override process note** (3ce248a)
- **data: lane-4 MTP gate PASS (15/15 arrays, 7/7 dequant sub-gate) + lane summary and next-lane handoff** (d18fcf6)
- **data: lane-4 gates (b)+(c) PASS — 160/160 greedy agreement CPU==GPU, byte-identical determinism across runs, VRAM stable; receipts + logs banked** (22ae3fe)
- **data: lane-4 gate (a) PASS — 14-array table, expert-dequant sub-gate 5/5 bit-exact, placement checkpoints; run logs banked** (d2a4c96)
- **data: lane-4 plan of record — placement math, quant rungs, bf16 threshold derivation, gate protocol (banked before loading)** (d05a2dc)
- **data: receipts — determinism rerun at 6f310ea71, both gate outputs numerically identical** (f78a0cd)
- **data: dsv4 forward lane receipts — both fixture gates PASS 15/15, contract = artifact clamp-only, MTP V3 shift confirmed** (4124ca0)
- **data: dsv4 loader lane receipts — census gate PASS, decode oracle bit-exact, geometry findings** (f85d0c2)

### Docs
- **docs(flags): row for MEMRA_DSPARK_BATCH_SMOKE (unmasked by the carve-out narrowing)** (3571829)
- **docs(flags): catalog MEMRA_DSPARK_DEFER_READBACK + MEMRA_DSPARK_VERIFY_GRAPH; narrow the scope-note carve-out to MEMRA_DSV4_*** (2302742)
- **docs(readme): v0.98.0 truth pass — strategy-keyed drafter harvest + confidence-window scheduling** (b23c845)
- **docs: lane-7 plan of record banked before build — kernel-match survey (no engine arm matches; cited arithmetic), native NVFP4/MXFP4 arm design per kernel.py fp4_gemm law, FP8-linear stay-bf16 decision, u_q=2^-4 class coefficient + full gate derivations** (25ac3c9)
- **docs: lane-6 decode design banked before build — cache layout (ring/pending/stores), VRAM f(seq) table, decode-vs-reprefill equivalence doctrine, gate protocol** (957a408)
- **docs: gate header reflects the derived top-20 band rule** (7006199)

### Chore
- **chore: sync Cargo.lock (memra-engine sha2 dep from the lane-4 gate binaries — the lock update had lived uncommitted on the box)** (c47c7bc)
- **chore: dsv4 forward lint pass — zero lane-file clippy warnings, no arithmetic change** (6f310ea)
- **chore: drop unused muts in dsv4 census closures** (37e2665)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/avifenesh/memra?utm_source=github-action)._