Table of contents

Sherpa-onnx Kokoro streaming-method audit

Kokoro-Hexagon 0a03be39Updated 2026-09-26

Source was already mirrored in a local upstream mirror; no clone or checkout mutation was needed. This audit reads upstream commit 040afe360a38e25daaa325ce8889abf93ea02609 by immutable Git object. It is a scheduling comparison only. No sherpa-onnx code, ONNX model, runtime, or frontend is a Kokoro-Hexagon product input.

What the callback actually means

Kokoro-Hexagon consequence

Use completed legal utterance units as one optional scheduling mode after the exact full-utterance phoneme-to-PCM path works. Keep the model and AAudio stream resident; hand PCM to a bounded playback queue with cancellation and explicit drain. Never call a completed-unit callback intra-model streaming. Compare short-first-unit and steady larger-unit schedules against exact one-shot execution on identical admitted text, voice, speed, and weights. Kokoro's context and full-span AdaIN statistics make a segment cut potentially audible, so segmentation is not stock-equivalent to the one-shot utterance. Preserve the one-shot path and gate any segmented mode by measured first-audio latency, gaps/underruns, memory, and listening results.

This was a source audit, not a benchmark. No sherpa-onnx timing, Kokoro-Hexagon speech timing, device playback, or quality result is inferred from it.