Table of contents

Corpus benchmark

Kokoro-Hexagon 0a03be39Updated 2026-09-22

Id Capacity Frames Seconds FrontMs GenMs GenP95 ColdFrontMs ColdGenMs SnrDb


p01 64 56 1.40 5.80 342.40 344.00 22.50 342.00 18.43 p02 96 76 1.90 7.30 585.80 586.90 25.30 585.60 6.06 p03 96 79 1.98 7.50 585.10 586.10 24.60 582.80 22.33 p04 96 92 2.30 7.40 586.70 587.90 24.10 583.80 22.03 p05 128 108 2.70 9.20 1160.60 1161.60 27.00 1125.30 23.44 p06 160 131 3.27 15.50 1369.60 1371.60 31.10 1339.80 24.06 p07 128 101 2.52 9.10 1161.40 1162.70 24.90 1146.40 22.61 p08 128 103 2.58 9.10 1160.60 1162.50 24.30 1149.00 20.60 p09 160 130 3.25 14.80 1370.80 1374.50 31.80 1337.00 23.27 p10 128 125 3.12 9.10 1161.70 1163.20 26.90 1126.80 22.11

Phrases : 9 AudioSeconds : 23.12 GeneratorRtfMean : 0.372 GeneratorRtfP50 : 0.4182 GeneratorRtfP95 : 0.46 GeneratorRtfMax : 0.46 PreparedToAudioMeanMs : 996.7 SnrMeanDb : 22.1 AllPlayed : True

Method

10 phrases (bench/corpus.json, af_heart), each exported at the smallest capacity bucket in {64, 96, 128, 160} that holds its frame count. Per phrase: 1 cold run, then 10 warm repeats of the front and generator contexts, then AudioTrack playback gated on PlaybackHeadPosition. All 9 passing phrases played through the speaker. PreparedToAudioMeanMs is cold front + cold generator, i.e. prepared inputs to first PCM, and excludes the PC-side frontend and har_source.

Notes