Kokoro-Hexagon
Kokoro-Hexagon runs the Kokoro-82M text-to-speech model on the Hexagon DSP of
Qualcomm Snapdragon devices. PowerShell reads the model checkpoint, authors
and lowers the computation, and emits the Hexagon machine code directly. No
vendor inference runtime, compiler or Python component is part of the build
or the device execution path. QNN, ONNX Runtime, PyTorch and LLVM are used
only as reference implementations to check results.
Product boundary
| Part |
Contents |
| Base appliance |
A model-less Android APK (NativeActivity, CoreCLR, PowerShell), under 40 MiB, built by a separately maintained fork of Pwsh's setup.ps1 |
| Model DLL |
A versioned, weight-bearing managed assembly: phoneme and text path, tensor catalog, graph, control contract and lowered hot paths |
| Model store |
Admits a model DLL only against a signed release manifest (length, SHA-256, assembly identity, compatibility), then switches an atomic active pointer; the new model loads at the next process start |
| Device execution |
Managed code plus directly emitted Hexagon code; PCM plays through a resident AAudio stream |
Model build: pinned Kokoro checkpoint -> PowerShell reader and lowering -> verified model DLL + Hexagon code
Host build: setup-kokoro.ps1 -> model-less APK
Device: admitted model DLL -> Hexagon execution -> PCM -> AAudio
Status
| Area |
State |
| Checkpoint reader |
548 FP32 tensors read from the pinned PyTorch checkpoint without PyTorch, embedded in a managed DLL and read back |
| Phoneme contract |
114-character vocabulary, 510-phoneme limit and voice-row selection validated in an emitted DLL |
| FP32 reference stages |
ALBERT, text encoder, duration, F0/N, decoder core, generator and inverse STFT re-authored as bounded PowerShell references from the pinned model source; full-path numerical comparison with the stock model is open |
| Emitted Hexagon kernels |
The fused AdaIN/Snake kernel emitted by PowerShell runs on SM8550 and SM8635 with output byte-identical to the pinned reference and to Hexagon Clang 19 output, at 2.21-2.30x the speed of the Clang build in on-DSP timer ticks |
| Model store |
Signed, transactional admission implemented and tested on Windows; not yet wired into the appliance |
| End-to-end speech on the owned path |
Planned |
The next milestone is one admitted phoneme string to audible PCM on both
devices through the same model DLL.
Findings
- Kokoro's AdaIN normalizes over the whole utterance. Sliding 64-frame windows
with overlap reached 5 dB SNR against the full-length model; one fixed
window per phrase with statistics masked to valid frames reached 24.2 dB.
Segmentation therefore happens at phrase boundaries, not fixed windows.
- On Hexagon V73 in an unsigned process domain, the supervisor cycle counter
s31:30 traps (exception 0x1B01); the QTimer pair c31:30 (19.2 MHz) is
readable and is used for on-DSP timing.
- For the fused AdaIN and Snake loop, Hexagon Clang 19 did not auto-vectorize
plain C at
-O3 (scalar code, about 26 ms); with HVX intrinsics it used all
32 vector registers. The first PowerShell-emitted kernel used 22 and ran at
0.87x the Clang build's speed with identical output
(2026-09-23). The lowered
version ran at 2.21-2.26x on SM8550 and 2.29-2.30x on SM8635, byte-identical
(2026-09-24).
In this documentation
Source: github.com/MansfieldPlumbing/Kokoro-Hexagon.