Table of contents

Speech from PowerShell

Kokoro-Hexagon 0a03be39Updated 2026-09-27

Future planning research only. This document does not describe a working speech build or authorize any change to stock Kokoro semantics. The direct PowerShell phoneme-to-PCM model and physical speaker gate take precedence.

SMA may own text planning. A language model may supply authored text, communicative objective, discourse annotations, or token surprisal, but it does not write the phoneme stream and it does not mix stage directions into the transcript.

authored text + objective + optional model evidence
    -> coordinate-preserving SMA parse
    -> speech-act and boundary candidates
    -> spoken normalization
    -> admitted phoneme adapter (checked against a pinned oracle)
    -> Kokoro inputs
    -> device acoustic and listening gates

speaker is a stable discourse identity; voice is the selected Kokoro style asset. They are separate on purpose: a character can be recast without changing the dialogue plan, and one voice can serve multiple anonymous roles.

Representations

Multiple speakers

Every speech plan and phrase bundle carries both speaker and voice. A speaker turn selects a verified voice/style resource without reloading the model DLL. The intended pipeline keeps the model session and one bounded AAudio stream alive across phrases. Device results name both fields so a listening result remains attributable; this multi-speaker path is not yet device-proven.

This is a target, not evidence-free: each admitted voice asset must be pinned by hash, each phrase still passes the FP32 oracle gate, and the first multi-speaker claim requires a physical speaker result. Voice resources belong in the verified model DLL, not a host export step at inference time.

The initial cue vocabulary is deliberately small: breath, pause, laugh, cough, clear_throat, sigh, and gasp. Pauses require an explicit duration from 20 to 5000 milliseconds. breath denotes a planning boundary to be rendered as a measured short pause, not an inhalation sound. Other nonverbal cues are not proven audio outputs. Asterisks and ordinary square brackets have no special meaning.

Borrowed mechanisms

Two local research repositories provide useful, bounded mechanisms:

Neither repository proves a speech model. Their search and admission patterns are the reusable parts.

Evidence admitted so far

Primary studies support the architecture, but not a universal fixed syllable limit:

These sources justify candidate features and test design. They do not establish that Kokoro exposes the corresponding controls or that any proposed mutation sounds better on this device.

Admission sequence

  1. Preserve authored text and reject malformed cue cards.
  2. Produce speech-act, boundary, state, and prominence candidates without changing authored text.
  3. Lower one candidate through a reversible mutation.
  4. Verify pronunciation and tensor bounds.
  5. Measure the physical device and play the result through the speaker.
  6. Compare blinded or paired listening judgments and retain the full outcome ledger, including rejected candidates.

The current implementation proves only the first two steps. Contour labels and word-count planning budgets are explicitly provisional until later gates pass.