Table of contents

Generator resblock r0, emitted from the checkpoint

Kokoro-Hexagon 0a03be39Updated 2026-09-23

decoder.module.generator.resblocks.3 built through the QNN C API from weights read out of kokoro-v1_0.pth, finalized and executed on SM8550 / Hexagon V73. No ONNX, no Python, no compile step. src/runspace/R0Emit.ps1.

Correctness

Shape C=128 T=7681   maskSum=6721  k=1.142836
Ops=189   FinalizeRc=0   ctxBytes=5,926,912
snrDb=62.10   maxAbs=1.388E-001   nonFinite=0

Scored against oracle_r0.f32 from the exported reference, on the real in_z input and the real mask. The block is three passes of AdaIN -> Snake -> Conv(dilation d) -> AdaIN -> Snake -> Conv(dilation 1) -> residual add, with d = 1, 3, 5.

Snake uses ElementWiseSin and a self-multiply, never Pow, which is miscomputed on this part. AdaIN uses ElementWiseRsqrt. k = W / sum(mask) folds to a constant because the mask is fixed for a phrase. The per-channel sc scaling cancels algebraically — (x-mu)/sc / sqrt(var/sc^2 + eps/sc^2) is (x-mu)/sqrt(var+eps) — and exists only to keep the variance in fp16 range.

Speed, and a tiling null result

tile=128 (untiled)   Ops=189    ctx 5,926,912   meanMs=280.4   62.10 dB
tile=16              Ops=1335   ctx 9,261,056   meanMs=278.7   62.10 dB

Channel tiling bought nothing here, against 1.52x in the isolated measurement (channel-tile-speedup-20260923.md). The difference is structural: this graph concatenates the tiles back to [1,128,T] before every convolution, so intermediates still materialise at full width and DDR traffic is unchanged. In the isolated case nothing re-formed the wide tensor.

Tiling pays only if data stays tiled across the whole chain. Getting that here means tiling the convolution as well — output-channel tiles reading the full input — so the concat happens once per block rather than four times. The projection in channel-tile-speedup-20260923.md assumed the tiling would carry through and should not be relied on until that is built.

Identical SNR at both tile sizes does confirm the tiling is bit-exact in practice.

The real gap

One resblock takes 280 ms while the whole compiled c64 generator takes 342 ms. About 130 full-tensor ops each moving roughly 4 MB is ~520 MB of traffic, which accounts for the time at realistic bandwidth. The emission is correct and unfused; the compiler fuses these chains so intermediates never materialise. Fusion, not tiling, is the gap between our emission and theirs.