Table of contents

Model-lowered R0Sub0 versus LLVM

Kokoro-Hexagon 0a03be39Updated 2026-09-24

Date: 2026-09-24
Device: retail Samsung Galaxy S23, SM8550, Hexagon V73
Scope: the pinned R0Sub0 elementwise benchmark kernel only

Compared artifacts

The lowered schedule uses eight independent lanes and 30 HVX vector registers. Each lane preserves the graph's floating-point operation order. Hardware timing uses c31:30 around the compute region; host timing covers the FastRPC invoke.

Counterbalanced physical runs

Each competitor received one cold invoke followed by 12 warm invocations. The second run reversed execution order.

Execution order Lowered median ticks LLVM median ticks Tick speedup Lowered median invoke LLVM median invoke Invoke speedup
LLVM, PowerShell 32,726 72,416 2.213x 5.522 ms 7.851 ms 1.422x
PowerShell, LLVM 34,524 77,890 2.256x 6.548 ms 8.777 ms 1.340x

Across both orders, the slowest lowered hardware sample was 40,036 ticks and the fastest LLVM sample was 63,739 ticks. Output parity and both performance gates passed in both runs. The host startup files were restored and hash-checked after each run.

External raw results:

Claim boundary

This demonstrates that the model-aware eight-lane specialization outperforms the pinned LLVM-generated implementation for this kernel on this device. It is not a claim about unrelated kernels, other devices, or LLVM in general.