Table of contents

R0.0 HVX Physical-Substrate result (R0SUB0-HVX-POINTWISE-R001)

Kokoro-Hexagon 0a03be39Updated 2026-09-23

Date: 2026-09-23

Target: Samsung Galaxy S23 (SM8550 / Hexagon V73)

Host: AndroidSMA Preview Host (dev.mansfieldplumbing.androidsma.preview)


1. Proven Scope vs. Architectural Boundary

What This Result Proves:

  1. PowerShell-Authored Native HVX Emission: Pure PowerShell named encoders in src/emit/Hexagon.ps1 emitted genuine V73 HVX instructions (vload-post, vstore-post, vadd-sf, vmpy-sf, vsplat).
  2. Physical Hardware Execution on SM8550: The in-memory synthesized ELF shared object loaded and executed on the physical Qualcomm Snapdragon 8 Gen 2 cDSP.
  3. 128-byte HVX Vector Geometry Verified: The loop geometry \(983,168 / 32 = 30,724\) vector iterations matches physical HVX width (128 bytes = 32 FP32 lanes).
  4. Whole-Path Streaming Floor: Pure 128-byte vector streaming moved/read/wrote 7.86 MB (3.93 MB in, 3.93 MB out) with a warm median latency of 5.212 ms (minimum 4.704 ms), establishing an effective whole-path throughput floor of ~1.51 GB/s under the current admission path.
  5. In-Register Pointwise Vector Arithmetic Executed: Adding an 11-instruction vector floating-point arithmetic body (affine scale, shift, polynomial evaluation, squaring, inverse scaling) raised wall-clock latency to 8.142 ms (minimum 7.251 ms)—an addition of +2.930 ms of execution time across all 983,168 elements.
  6. Scalar Path Superseded: The prior 25.998 ms scalar affine result (kokoro-affine-emitted-20260923.md) is superseded (\(25.998 / 8.142 \approx 3.19\times\) wall-clock speedup) despite the HVX path performing substantially more arithmetic.

What Is NOT Yet Demonstrated (Not Full r0.0 Parity):


2. Hardware Measurements & Comparative Benchmark

Specimen / Kernel Execution Mode Measured Hardware Time (Warm Median, 12 calls) Code Size Output Verification
Prior Scalar Affine Baseline Scalar R-register (sfmpy + sfadd) 25.998 ms 1,216 bytes Bit-exact reference
HVX 128-byte Vector Streaming 128-byte HVX (vload-post + vstore-post) 5.212 ms (min 4.704 ms) 428 bytes VectorStreamingMatch=True (exact byte echo)
HVX In-Register Pointwise Arithmetic 128-byte HVX Pipeline (11 vector ops/iter) 8.142 ms (min 7.251 ms) 508 bytes OutputComputed=True (non-zero arithmetic output)

Arithmetic vs. Memory Traversal Profile:

Whole-path streaming floor:     5.212 ms
Pointwise arithmetic execution: 8.142 ms
Difference:                     2.930 ms

Analytical Instruction Estimate:


3. Admission & Platform State


4. Provenance & Artifact Hashes