Generator resblock r0, emitted from the checkpoint
decoder.module.generator.resblocks.3 built through the QNN C API from weights read out of
kokoro-v1_0.pth, finalized and executed on SM8550 / Hexagon V73. No ONNX, no Python, no
compile step. src/runspace/R0Emit.ps1.
Correctness
Shape C=128 T=7681 maskSum=6721 k=1.142836
Ops=189 FinalizeRc=0 ctxBytes=5,926,912
snrDb=62.10 maxAbs=1.388E-001 nonFinite=0
Scored against oracle_r0.f32 from the exported reference, on the real in_z input and
the real mask. The block is three passes of AdaIN -> Snake -> Conv(dilation d) -> AdaIN ->
Snake -> Conv(dilation 1) -> residual add, with d = 1, 3, 5.
Snake uses ElementWiseSin and a self-multiply, never Pow, which is miscomputed on this
part. AdaIN uses ElementWiseRsqrt. k = W / sum(mask) folds to a constant because the
mask is fixed for a phrase. The per-channel sc scaling cancels algebraically —
(x-mu)/sc / sqrt(var/sc^2 + eps/sc^2) is (x-mu)/sqrt(var+eps) — and exists only to keep
the variance in fp16 range.
Speed, and a tiling null result
tile=128 (untiled) Ops=189 ctx 5,926,912 meanMs=280.4 62.10 dB
tile=16 Ops=1335 ctx 9,261,056 meanMs=278.7 62.10 dB
Channel tiling bought nothing here, against 1.52x in the isolated measurement
(channel-tile-speedup-20260923.md). The difference is structural: this graph concatenates
the tiles back to [1,128,T] before every convolution, so intermediates still materialise
at full width and DDR traffic is unchanged. In the isolated case nothing re-formed the wide
tensor.
Tiling pays only if data stays tiled across the whole chain. Getting that here means tiling the convolution as well — output-channel tiles reading the full input — so the concat happens once per block rather than four times. The projection in channel-tile-speedup-20260923.md assumed the tiling would carry through and should not be relied on until that is built.
Identical SNR at both tile sizes does confirm the tiling is bit-exact in practice.
The real gap
One resblock takes 280 ms while the whole compiled c64 generator takes 342 ms. About 130 full-tensor ops each moving roughly 4 MB is ~520 MB of traffic, which accounts for the time at realistic bandwidth. The emission is correct and unfused; the compiler fuses these chains so intermediates never materialise. Fusion, not tiling, is the gap between our emission and theirs.