MEASURE: Phonon-2 vs Parakeet TDT for local voice transcription (CPU, perf VM) #489

Closed
opened 2026-09-30 06:44:16 +00:00 by kayg · 29 comments
Owner

Measurement: Phonon-2 vs Parakeet TDT 0.6B v2 for calternal voice (local, CPU)

Protocol (see the measurement-jobs rules: exact protocol, interleaved runs, decision rule up front):

  • Perf VM (root@10.69.69.63, flock /root/perf.lock), CPU only, same thread count for both. Runtimes: Phonon's CPU runtime (github.com/fermionresearch/phonon) and the Parakeet runtime named in docs/research/voice-transcription.md.
  • Audio: LibriSpeech test-clean and test-other subsets (200 utterances each), plus 30 synthetic "voice note" style clips with Indian-English accents, money amounts ("fifty rupees"), times, names and code-switched words. Take them from public CC datasets (Common Voice English, Indian accent subset). Never use private owner audio.
  • Metrics: WER (normalised), real-time factor, cold start, peak RSS, CPU seconds per audio minute, and the quality of punctuation and casing. Also a long-file test (60 min).
  • Decision rule (write it before running): choose Phonon-2 if WER is within +0.5 pt of Parakeet on both LibriSpeech sets AND not worse on the voice-note set, and RTF or RSS is better. Otherwise keep Parakeet.
  • Licences: record the model and runtime licences (the weights are CC-BY-4.0, so attribution is needed) and check AGPL compatibility for bundling or download.
    Deliver docs/perf/runs/<date>-asr-ab.md and a recommendation on the voice issue. No product code.
## Measurement: Phonon-2 vs Parakeet TDT 0.6B v2 for calternal voice (local, CPU) Protocol (see the measurement-jobs rules: exact protocol, interleaved runs, decision rule up front): - Perf VM (root@10.69.69.63, `flock /root/perf.lock`), CPU only, same thread count for both. Runtimes: Phonon's CPU runtime (github.com/fermionresearch/phonon) and the Parakeet runtime named in docs/research/voice-transcription.md. - Audio: LibriSpeech test-clean and test-other subsets (200 utterances each), plus 30 synthetic "voice note" style clips with Indian-English accents, money amounts ("fifty rupees"), times, names and code-switched words. Take them from public CC datasets (Common Voice English, Indian accent subset). Never use private owner audio. - Metrics: WER (normalised), real-time factor, cold start, peak RSS, CPU seconds per audio minute, and the quality of punctuation and casing. Also a long-file test (60 min). - Decision rule (write it before running): choose Phonon-2 if WER is within +0.5 pt of Parakeet on both LibriSpeech sets AND not worse on the voice-note set, and RTF or RSS is better. Otherwise keep Parakeet. - Licences: record the model and runtime licences (the weights are CC-BY-4.0, so attribution is needed) and check AGPL compatibility for bundling or download. Deliver `docs/perf/runs/<date>-asr-ab.md` and a recommendation on the voice issue. No product code.
Author
Owner

Started #489 on job/asr-ab-489, based on dev at 0dc772c3697ea9bd01822c26440c32206d472715. I read CLAUDE.md, CONTEXT.md, docs/DESIGN.md, and docs/research/voice-transcription.md. I will use the issue's decision rule as written and record it in the report before running either model. I inventoried the perf VM while holding /root/perf.lock; it has 4 vCPU, 7.7 GiB RAM, and load average 0.09 / 0.88 / 1.54 at inventory time. I am checking runtime availability, model/data licences, and whether the requested 30 Indian-English voice-note cases can be assembled from the named public CC corpus.

Started #489 on `job/asr-ab-489`, based on `dev` at `0dc772c3697ea9bd01822c26440c32206d472715`. I read `CLAUDE.md`, `CONTEXT.md`, `docs/DESIGN.md`, and `docs/research/voice-transcription.md`. I will use the issue's decision rule as written and record it in the report before running either model. I inventoried the perf VM while holding `/root/perf.lock`; it has 4 vCPU, 7.7 GiB RAM, and load average 0.09 / 0.88 / 1.54 at inventory time. I am checking runtime availability, model/data licences, and whether the requested 30 Indian-English voice-note cases can be assembled from the named public CC corpus.
Author
Owner

Source check: Phonon-2 is a separate five-value model derived from Parakeet TDT v3, not v2. Its model weights are CC-BY-4.0 and the Phonon CLI is Apache-2.0; the required Parakeet v2 weights are also CC-BY-4.0 and sherpa-onnx is Apache-2.0. The PyPI API currently lists fermion-research 0.2.3. I will pin the runnable CPU image/package and model hashes in the report after verifying them on the perf VM. The model card's published LS WER is not a result for this required v2 comparison.

Source check: Phonon-2 is a separate five-value model derived from Parakeet TDT v3, not v2. Its model weights are CC-BY-4.0 and the Phonon CLI is Apache-2.0; the required Parakeet v2 weights are also CC-BY-4.0 and `sherpa-onnx` is Apache-2.0. The PyPI API currently lists `fermion-research` 0.2.3. I will pin the runnable CPU image/package and model hashes in the report after verifying them on the perf VM. The model card's published LS WER is not a result for this required v2 comparison.
Author
Owner

Perf VM status at 2026-09-30 09:17+02:00: fuser /root/perf.lock showed PIDs 1267207 (bash) and 1267510 (bun). My lock-wrapped corpus download did not acquire the lock and wrote no files. No model runs or measurement data downloads have run on the VM. I am preparing the benchmark inputs and harness outside the VM, then I will take a locked measurement phase when the lock is available.

Perf VM status at 2026-09-30 09:17+02:00: `fuser /root/perf.lock` showed PIDs 1267207 (`bash`) and 1267510 (`bun`). My lock-wrapped corpus download did not acquire the lock and wrote no files. No model runs or measurement data downloads have run on the VM. I am preparing the benchmark inputs and harness outside the VM, then I will take a locked measurement phase when the lock is available.
Author
Owner

Dataset finding: the CV v21 Indian-accent English shard I scanned has 12,232 transcript rows; the scan found no “rupee” text and no Hindi/code-switched terms. The CV v26 mirror I inspected has audio files but no transcript rows, so it cannot be used for WER. I found VoiceArena Monsoon en-IN, a public CC-BY-4.0 set of 2,102 conversational Indian-English clips. I am checking its transcript coverage for the requested categories before selecting the 30 items. If used, the report will identify the source and license and will omit its speaker-demographic fields.

Dataset finding: the CV v21 Indian-accent English shard I scanned has 12,232 transcript rows; the scan found no “rupee” text and no Hindi/code-switched terms. The CV v26 mirror I inspected has audio files but no transcript rows, so it cannot be used for WER. I found VoiceArena Monsoon en-IN, a public CC-BY-4.0 set of 2,102 conversational Indian-English clips. I am checking its transcript coverage for the requested categories before selecting the 30 items. If used, the report will identify the source and license and will omit its speaker-demographic fields.
Author
Owner

Data finding for #489: I scanned only the public VoiceArena test-set id and text fields across all 2,102 rows (revision bc1da7b42ef6e2853123c97bf6d22067e4802d11; shard hashes are recorded locally). The explicit transcript patterns found 0 money, 0 time, 0 listed Indian-name, and 6 code-switch matches. This set does not cover the target challenge categories. I am checking the remaining Common Voice Indian-accent transcript shards before choosing the 30-item proxy sample; no model has been run.

Data finding for #489: I scanned only the public VoiceArena test-set `id` and `text` fields across all 2,102 rows (revision `bc1da7b42ef6e2853123c97bf6d22067e4802d11`; shard hashes are recorded locally). The explicit transcript patterns found 0 money, 0 time, 0 listed Indian-name, and 6 code-switch matches. This set does not cover the target challenge categories. I am checking the remaining Common Voice Indian-accent transcript shards before choosing the 30-item proxy sample; no model has been run.
Author
Owner

Data finding for #489: the Common Voice Indian-accent card lists 110,088 rows. I scanned three fixed shards (31,483 transcript rows) after range-reading all nine would cost about eight minutes for only three shards. Those three contain 48 general currency mentions, 22 time matches, and 535 matches for the listed name patterns. The two INR-related matches are an incomplete “Rs.” and a bare “rupee”; there is no amount such as “fifty rupees.” One code-switch regex hit was a false positive (“bas-reliefs”), so I will not count it. Across the 2,102 VoiceArena test transcripts I found zero money/time matches and five clear code-switch matches. The public synthetic indic-audio set I checked says its audio is subject to ElevenLabs and Fish Audio output terms, so I am excluding it under the public CC/AGPL-compatible source requirement. I will use a 30-clip real-recording proxy set: Common Voice for money (other currencies), time and names, and VoiceArena for conversational/code-switched speech. The lack of synthetic audio and rupee amounts will be stated as limitations. No ASR model has been run.

Data finding for #489: the Common Voice Indian-accent card lists 110,088 rows. I scanned three fixed shards (31,483 transcript rows) after range-reading all nine would cost about eight minutes for only three shards. Those three contain 48 general currency mentions, 22 time matches, and 535 matches for the listed name patterns. The two INR-related matches are an incomplete “Rs.” and a bare “rupee”; there is no amount such as “fifty rupees.” One code-switch regex hit was a false positive (“bas-reliefs”), so I will not count it. Across the 2,102 VoiceArena test transcripts I found zero money/time matches and five clear code-switch matches. The public synthetic indic-audio set I checked says its audio is subject to ElevenLabs and Fish Audio output terms, so I am excluding it under the public CC/AGPL-compatible source requirement. I will use a 30-clip real-recording proxy set: Common Voice for money (other currencies), time and names, and VoiceArena for conversational/code-switched speech. The lack of synthetic audio and rupee amounts will be stated as limitations. No ASR model has been run.
Author
Owner

Correction to my prior #489 data comment: the three Common Voice shards contain 36,696 audio rows and 31,483 distinct sentence_id values. I had labeled the distinct-prompt count as the row count. The currency/time/name and code-switch pattern counts are unchanged because the scan counted each sentence_id once.

Correction to my prior #489 data comment: the three Common Voice shards contain 36,696 audio rows and 31,483 distinct `sentence_id` values. I had labeled the distinct-prompt count as the row count. The currency/time/name and code-switch pattern counts are unchanged because the scan counted each `sentence_id` once.
Author
Owner

Input selection is now fixed before any model runs. The #489 report records 200 deterministic clips from each LibriSpeech subset and a 30-clip Indian-English proxy set: 18 Common Voice clips (six each for non-INR money amounts, times, and names) plus 12 VoiceArena clips (five code-switch matches and seven conversational fillers). The selected voice set is 242.406 seconds. It uses real public recordings; the synthetic-audio and rupee-amount gaps are documented. The report and its input-ID/audio-checksum manifest are committed together.

Input selection is now fixed before any model runs. The #489 report records 200 deterministic clips from each LibriSpeech subset and a 30-clip Indian-English proxy set: 18 Common Voice clips (six each for non-INR money amounts, times, and names) plus 12 VoiceArena clips (five code-switch matches and seven conversational fillers). The selected voice set is 242.406 seconds. It uses real public recordings; the synthetic-audio and rupee-amount gaps are documented. The report and its input-ID/audio-checksum manifest are committed together.
Author
Owner

Additional licence finding for #489: Phonon-2's pinned model NOTICE says the weights are published under CC-BY-4.0 and derive from NVIDIA Parakeet TDT v3. It also says training audio includes CHiME-6 (CC-BY-SA-4.0) and SPGISpeech (Kensho public terms). The CPU runtime code is Apache-2.0. I will preserve the model NOTICE and licences and report this provenance caveat; this benchmark does not bundle the model into the AGPL application.

Additional licence finding for #489: Phonon-2's pinned model `NOTICE` says the weights are published under CC-BY-4.0 and derive from NVIDIA Parakeet TDT v3. It also says training audio includes CHiME-6 (CC-BY-SA-4.0) and SPGISpeech (Kensho public terms). The CPU runtime code is Apache-2.0. I will preserve the model NOTICE and licences and report this provenance caveat; this benchmark does not bundle the model into the AGPL application.
Author
Owner

Pinned-runtime finding for #489: the official Phonon CPU image uses fermion-research 0.2.2. Its engine_phonon2_cpu._default_threads() defaults from os.cpu_count(), but load() honors FERMION_CPU_THREADS; it applies that count to Torch, the packed GEMM kernel, and TDT decoder. I will set FERMION_CPU_THREADS=4. The sherpa-onnx worker sets num_threads=4, so both use the same requested thread count.

Pinned-runtime finding for #489: the official Phonon CPU image uses fermion-research 0.2.2. Its `engine_phonon2_cpu._default_threads()` defaults from `os.cpu_count()`, but `load()` honors `FERMION_CPU_THREADS`; it applies that count to Torch, the packed GEMM kernel, and TDT decoder. I will set `FERMION_CPU_THREADS=4`. The sherpa-onnx worker sets `num_threads=4`, so both use the same requested thread count.
Author
Owner

The perf VM now exposes its lock, and the official Phonon CPU image is already cached. I recorded the image digest as sha256:2a01f89eca2c362e186b2ba788ce9fc77680033e6283c8b084f4fbff85a60456 in the benchmark report. The 08:42 UTC setup inventory was four CPUs, 7.7 GiB RAM, load average 0.69 / 2.67 / 4.19; the matching thread controls are set for the runs.

The perf VM now exposes its lock, and the official Phonon CPU image is already cached. I recorded the image digest as `sha256:2a01f89eca2c362e186b2ba788ce9fc77680033e6283c8b084f4fbff85a60456` in the benchmark report. The 08:42 UTC setup inventory was four CPUs, 7.7 GiB RAM, load average 0.69 / 2.67 / 4.19; the matching thread controls are set for the runs.
Author
Owner

Runner startup finding: the first cold-start attempt exited before inference with fermion: error: unrecognized arguments: --model-dir /model. The pinned 0.2.2 serve parser has no --model-dir; its --model argument accepts an unpacked local speech directory, and server.py detects it from config.json plus packed_manifest.json. I will start the official CPU server with serve --model /model --served-model-name phonon-2. No inference metrics were produced by the failed attempt.

Runner startup finding: the first cold-start attempt exited before inference with `fermion: error: unrecognized arguments: --model-dir /model`. The pinned 0.2.2 `serve` parser has no `--model-dir`; its `--model` argument accepts an unpacked local speech directory, and `server.py` detects it from `config.json` plus `packed_manifest.json`. I will start the official CPU server with `serve --model /model --served-model-name phonon-2`. No inference metrics were produced by the failed attempt.
Author
Owner

The successful cold phase ran at 2026-09-30 08:56:36 UTC inside /root/perf.lock on four CPUs; load average was 0.07 / 0.29 / 1.81. With the models already local and OS page cache left intact, process-to-ready was 1.857 s for Parakeet and 27.347 s for Phonon. On the same 2.515 s test-clean clip, request wall RTF was 0.079 vs 0.158; runtime decode RTF was 0.077 vs 0.075. Peak process RSS was 800 MiB vs 1,513 MiB. Both returned the same normalized words. This single probe does not decide the A/B rule; warm set-level WER and voice results are still pending.

The successful cold phase ran at 2026-09-30 08:56:36 UTC inside `/root/perf.lock` on four CPUs; load average was 0.07 / 0.29 / 1.81. With the models already local and OS page cache left intact, process-to-ready was 1.857 s for Parakeet and 27.347 s for Phonon. On the same 2.515 s test-clean clip, request wall RTF was 0.079 vs 0.158; runtime decode RTF was 0.077 vs 0.075. Peak process RSS was 800 MiB vs 1,513 MiB. Both returned the same normalized words. This single probe does not decide the A/B rule; warm set-level WER and voice results are still pending.
Author
Owner

Warm-clean phase startup succeeded at 09:00:43 UTC inside the lock (load average 0.03 / 0.19 / 1.42), but the unscored warmup sent repetition -1. The Rust worker requires an unsigned JSON number and panicked at src/main.rs:36 with repetition is missing. This happened before WAV reading or inference; the Phonon log has only the readiness request and no transcription POST. I will send repetition 0 and keep the record marked warmup, which the scorer excludes.

Warm-clean phase startup succeeded at 09:00:43 UTC inside the lock (load average 0.03 / 0.19 / 1.42), but the unscored warmup sent repetition `-1`. The Rust worker requires an unsigned JSON number and panicked at `src/main.rs:36` with `repetition is missing`. This happened before WAV reading or inference; the Phonon log has only the readiness request and no transcription POST. I will send repetition `0` and keep the record marked `warmup`, which the scorer excludes.
Author
Owner

License compatibility check for #489: the Apache-2.0 runtimes can be used with AGPLv3 code if their Apache notices stay intact. GNU AGPLv3 permits GPLv3-compatible code in an AGPLv3 work, and Apache states its v2 license is GPLv3-compatible. Model weights stay under CC-BY-4.0 with attribution; they are not relicensed as AGPL. A separate model download can keep its CC-BY terms. Before bundling Phonon weights into the app, review the CHiME-6 CC-BY-SA and SPGISpeech terms listed in its NOTICE. Sources: https://www.gnu.org/licenses/agpl-3.0.html ; https://apache.org/licenses/GPL-compatibility.html ; https://creativecommons.org/share-your-work/licensing-considerations/version4/

License compatibility check for #489: the Apache-2.0 runtimes can be used with AGPLv3 code if their Apache notices stay intact. GNU AGPLv3 permits GPLv3-compatible code in an AGPLv3 work, and Apache states its v2 license is GPLv3-compatible. Model weights stay under CC-BY-4.0 with attribution; they are not relicensed as AGPL. A separate model download can keep its CC-BY terms. Before bundling Phonon weights into the app, review the CHiME-6 CC-BY-SA and SPGISpeech terms listed in its NOTICE. Sources: https://www.gnu.org/licenses/agpl-3.0.html ; https://apache.org/licenses/GPL-compatibility.html ; https://creativecommons.org/share-your-work/licensing-considerations/version4/
Author
Owner

Protocol decision for #489: request RTF is request wall time divided by audio duration, and that is the pre-run RTF used by the decision rule. Phonon runs through its official HTTP transcription endpoint; Parakeet runs through the persistent Rust worker's JSON-lines interface. The request metric includes each interface's file/request handling. I also capture each runtime's decode time and report decode RTF separately as an inference-time view.

Protocol decision for #489: request RTF is request wall time divided by audio duration, and that is the pre-run RTF used by the decision rule. Phonon runs through its official HTTP transcription endpoint; Parakeet runs through the persistent Rust worker's JSON-lines interface. The request metric includes each interface's file/request handling. I also capture each runtime's decode time and report decode RTF separately as an inference-time view.
Author
Owner

Additional dataset finding for #489: I found a public CC-BY-4.0 14-row synthetic Indian-English set, TieIncred/parakeet-tdt-blind-spots. It contains rupee phrases, dates, Indian names, and Hindi-English switches, but the card says it was curated to expose Parakeet errors and includes Parakeet output. Using it for the head-to-head would favor Phonon. Its card names Edge-TTS as the source; I did not independently verify those service terms. I excluded it from the decision set and recorded this in the report. The current 30-item set remains a real-recording proxy with no INR amount coverage.

Additional dataset finding for #489: I found a public CC-BY-4.0 14-row synthetic Indian-English set, TieIncred/parakeet-tdt-blind-spots. It contains rupee phrases, dates, Indian names, and Hindi-English switches, but the card says it was curated to expose Parakeet errors and includes Parakeet output. Using it for the head-to-head would favor Phonon. Its card names Edge-TTS as the source; I did not independently verify those service terms. I excluded it from the decision set and recorded this in the report. The current 30-item set remains a real-recording proxy with no INR amount coverage.
Author
Owner

Measured finding — warm-clean phase completed under /root/perf.lock at 2026-09-30 09:02 UTC; load average was 0.096 / 0.207 / 1.271 on four CPUs. It produced 200 clips × five paired repetitions per runtime, plus two excluded warm-ups. On the first repetition (4,188 reference words), Phonon scored 76 word errors (1.815% WER); Parakeet scored 60 (1.433%), so Phonon is +0.382 percentage points. Median request RTF was 0.1035 vs 0.0600; p95 0.1649 vs 0.0755. CPU was 17.01 vs 14.02 s/audio-minute. Peak process RSS was 1,653 vs 1,197 MiB. Clean-set WER is within the prewritten +0.5 point condition, but Phonon is slower and uses more memory on this set. The voice recommendation remains pending test-other and the voice proxy. Raw output remains in the ignored worktree benchmark area.

Measured finding — warm-clean phase completed under `/root/perf.lock` at 2026-09-30 09:02 UTC; load average was 0.096 / 0.207 / 1.271 on four CPUs. It produced 200 clips × five paired repetitions per runtime, plus two excluded warm-ups. On the first repetition (4,188 reference words), Phonon scored 76 word errors (1.815% WER); Parakeet scored 60 (1.433%), so Phonon is +0.382 percentage points. Median request RTF was 0.1035 vs 0.0600; p95 0.1649 vs 0.0755. CPU was 17.01 vs 14.02 s/audio-minute. Peak process RSS was 1,653 vs 1,197 MiB. Clean-set WER is within the prewritten +0.5 point condition, but Phonon is slower and uses more memory on this set. The voice recommendation remains pending test-other and the voice proxy. Raw output remains in the ignored worktree benchmark area.
Author
Owner

Long-file protocol finding: the pinned Phonon 2.0.2 server source hard-codes MAX_BODY = 32 * 1024 * 1024 and returns HTTP 413 above that limit. The required 60-minute lossless FLAC is 59,392,511 bytes, so it cannot pass through the stock endpoint. I verified that FLAC decodes to the exact WAV PCM SHA-256 (a0ba1b1212e1e0f89cd722e543475cd35229b19052b5cc2f2d916d9bddfdfcac). To keep the requested one-file runtime comparison, the long-only benchmark will raise the in-process request guard to 96 MiB in the same pinned container. The model, CPU settings, and official multipart endpoint stay unchanged. I have recorded this as a benchmark-only deviation; it does not establish support for this input size with the default server cap.

Long-file protocol finding: the pinned Phonon 2.0.2 server source hard-codes `MAX_BODY = 32 * 1024 * 1024` and returns HTTP 413 above that limit. The required 60-minute lossless FLAC is 59,392,511 bytes, so it cannot pass through the stock endpoint. I verified that FLAC decodes to the exact WAV PCM SHA-256 (`a0ba1b1212e1e0f89cd722e543475cd35229b19052b5cc2f2d916d9bddfdfcac`). To keep the requested one-file runtime comparison, the long-only benchmark will raise the in-process request guard to 96 MiB in the same pinned container. The model, CPU settings, and official multipart endpoint stay unchanged. I have recorded this as a benchmark-only deviation; it does not establish support for this input size with the default server cap.
Author
Owner

Protocol clarification, recorded after warm test-clean and before test-other/voice runs: for the existing “RTF or peak RSS is better” condition, I will compare the median request RTF pooled over every scored repetition and all three warm sets, plus the maximum process RSS across those warm runs. Phonon meets the efficiency condition if either pooled value is lower. The 0.5-point WER thresholds and chosen WER sets were recorded before any model run; this clarifies only the scope of the efficiency comparison.

Protocol clarification, recorded after warm test-clean and before test-other/voice runs: for the existing “RTF or peak RSS is better” condition, I will compare the median request RTF pooled over every scored repetition and all three warm sets, plus the maximum process RSS across those warm runs. Phonon meets the efficiency condition if either pooled value is lower. The 0.5-point WER thresholds and chosen WER sets were recorded before any model run; this clarifies only the scope of the efficiency comparison.
Author
Owner

Measurement correction: review found that Phonon's original warm request timer began after the benchmark client read the audio and assembled the multipart body. Parakeet's timer begins before the worker receives the audio path and opens the WAV. I am changing the Phonon client to stream the multipart body and time from file read through response, then repeating cold and test-clean. The earlier request RTF figures are provisional and will be replaced. The WER decision rule and input set are unchanged.

Measurement correction: review found that Phonon's original warm request timer began after the benchmark client read the audio and assembled the multipart body. Parakeet's timer begins before the worker receives the audio path and opens the WAV. I am changing the Phonon client to stream the multipart body and time from file read through response, then repeating cold and test-clean. The earlier request RTF figures are provisional and will be replaced. The WER decision rule and input set are unchanged.
Author
Owner

Corrected cold phase completed at 2026-09-30 09:52:01 UTC under /root/perf.lock. Load average was 3.291 / 3.418 / 2.980 on four exposed CPUs, so other host activity may affect this one-probe timing. The Phonon request timer now includes reading and streaming the audio file. Process start-to-ready was 27.366 s for Phonon (model loader 22.68 s) and 1.948 s for Parakeet (loader 1.944 s). For the 2.515 s first clip, request wall / RTF was 0.419 s / 0.167 vs 0.191 s / 0.076; runtime decode was 0.205 s / 0.082 vs 0.190 s / 0.075. Peak RSS was 1,482 vs 801 MiB. Both outputs normalized to the same words. Warm-clean is now running with the corrected request timer.

Corrected cold phase completed at 2026-09-30 09:52:01 UTC under `/root/perf.lock`. Load average was 3.291 / 3.418 / 2.980 on four exposed CPUs, so other host activity may affect this one-probe timing. The Phonon request timer now includes reading and streaming the audio file. Process start-to-ready was 27.366 s for Phonon (model loader 22.68 s) and 1.948 s for Parakeet (loader 1.944 s). For the 2.515 s first clip, request wall / RTF was 0.419 s / 0.167 vs 0.191 s / 0.076; runtime decode was 0.205 s / 0.082 vs 0.190 s / 0.075. Peak RSS was 1,482 vs 801 MiB. Both outputs normalized to the same words. Warm-clean is now running with the corrected request timer.
Author
Owner

Provisional test-other accuracy finding from the completed first run: normalized first-repetition WER on 3,372 words was 4.033% for Phonon (136 errors) and 2.936% for Parakeet (99 errors), a +1.097 percentage-point difference. This exceeds the prewritten +0.5 point bound for test-other. That run used the old request-timer boundary; the timer change should not alter transcript text, and I will confirm the WER in the corrected rerun before the final recommendation.

Provisional test-other accuracy finding from the completed first run: normalized first-repetition WER on 3,372 words was 4.033% for Phonon (136 errors) and 2.936% for Parakeet (99 errors), a +1.097 percentage-point difference. This exceeds the prewritten +0.5 point bound for test-other. That run used the old request-timer boundary; the timer change should not alter transcript text, and I will confirm the WER in the corrected rerun before the final recommendation.
Author
Owner

Corrected warm test-clean completed at 2026-09-30 09:53:09 UTC under the lock; load average was 1.223 / 2.785 / 2.791 on four CPUs. It completed 200 clips × five paired repetitions per runtime plus two excluded warm-ups. The 400 first-repetition outputs exactly match the earlier run. On 4,188 words, Phonon scored 76 errors (1.815% WER) and Parakeet 60 (1.433%), a +0.382-point difference. Corrected median/p95 request RTF is 0.104/0.168 for Phonon and 0.061/0.077 for Parakeet. CPU was 17.16 vs 14.16 s/audio-minute; peak process RSS was 1,673 vs 1,197 MiB. This supersedes the earlier clean request RTF figures. Phonon passes the clean WER bound but is slower and uses more memory on this set.

Corrected warm test-clean completed at 2026-09-30 09:53:09 UTC under the lock; load average was 1.223 / 2.785 / 2.791 on four CPUs. It completed 200 clips × five paired repetitions per runtime plus two excluded warm-ups. The 400 first-repetition outputs exactly match the earlier run. On 4,188 words, Phonon scored 76 errors (1.815% WER) and Parakeet 60 (1.433%), a +0.382-point difference. Corrected median/p95 request RTF is 0.104/0.168 for Phonon and 0.061/0.077 for Parakeet. CPU was 17.16 vs 14.16 s/audio-minute; peak process RSS was 1,673 vs 1,197 MiB. This supersedes the earlier clean request RTF figures. Phonon passes the clean WER bound but is slower and uses more memory on this set.
Author
Owner

Corrected warm test-other completed at 2026-09-30 10:14:34 UTC under /root/perf.lock; load average was 3.842 / 3.769 / 3.490 on four CPUs, which may affect timing. It completed 200 clips × five paired repetitions per runtime plus two excluded warm-ups. The 400 first-repetition transcripts match the provisional run. On 3,372 words, Phonon scored 136 errors (4.033% WER) and Parakeet 99 (2.936%), a +1.097-point gap. Corrected median/p95 request RTF is 0.110/0.162 for Phonon and 0.062/0.077 for Parakeet. CPU was 17.13 vs 14.28 s/audio-minute; peak RSS was 1,664 vs 1,193 MiB. This confirms that Phonon fails the prewritten test-other WER bound. I will still run the voice proxy and long-file phases.

Corrected warm test-other completed at 2026-09-30 10:14:34 UTC under `/root/perf.lock`; load average was 3.842 / 3.769 / 3.490 on four CPUs, which may affect timing. It completed 200 clips × five paired repetitions per runtime plus two excluded warm-ups. The 400 first-repetition transcripts match the provisional run. On 3,372 words, Phonon scored 136 errors (4.033% WER) and Parakeet 99 (2.936%), a +1.097-point gap. Corrected median/p95 request RTF is 0.110/0.162 for Phonon and 0.062/0.077 for Parakeet. CPU was 17.13 vs 14.28 s/audio-minute; peak RSS was 1,664 vs 1,193 MiB. This confirms that Phonon fails the prewritten test-other WER bound. I will still run the voice proxy and long-file phases.
Author
Owner

Corrected warm voice-proxy phase completed at 2026-09-30 10:34:14 UTC under /root/perf.lock; load average was 0.776 / 2.545 / 3.123 on four CPUs. The 30 real-recording proxy clips yielded 509 normalized reference words. Phonon scored 34 errors (6.680% WER), Parakeet 31 (6.090%), so Phonon is +0.589 points and fails the voice WER condition. Median/p95 request RTF was 0.102/0.152 vs 0.059/0.073; CPU was 17.07 vs 13.89 s/audio-minute; peak RSS was 1,610 vs 1,001 MiB. Category WER (Phonon vs Parakeet): money in other currencies 2.08% vs 2.08%; time 11.27% vs 9.86%; names 9.72% vs 8.33%; code-switch 8.90% vs 6.16%; conversation fillers 2.91% vs 4.65%. This set does not test rupee amounts or generated speech. The one-file long test remains.

Corrected warm voice-proxy phase completed at 2026-09-30 10:34:14 UTC under `/root/perf.lock`; load average was 0.776 / 2.545 / 3.123 on four CPUs. The 30 real-recording proxy clips yielded 509 normalized reference words. Phonon scored 34 errors (6.680% WER), Parakeet 31 (6.090%), so Phonon is +0.589 points and fails the voice WER condition. Median/p95 request RTF was 0.102/0.152 vs 0.059/0.073; CPU was 17.07 vs 13.89 s/audio-minute; peak RSS was 1,610 vs 1,001 MiB. Category WER (Phonon vs Parakeet): money in other currencies 2.08% vs 2.08%; time 11.27% vs 9.86%; names 9.72% vs 8.33%; code-switch 8.90% vs 6.16%; conversation fillers 2.91% vs 4.65%. This set does not test rupee amounts or generated speech. The one-file long test remains.
Author
Owner

Pooled warm efficiency is now scored across 2,150 requests per runtime (all five repetitions in test-clean, test-other, and the 30-clip voice proxy). Median request RTF is 0.1072 for Phonon vs 0.0610 for Parakeet; maximum process RSS is 1,673 vs 1,197 MiB. Phonon is worse on both efficiency measures. Together with the confirmed +1.097-point test-other WER gap and +0.589-point voice-proxy gap, the preregistered rule recommends keeping Parakeet. The one-file 60-minute comparison is still running and will be reported separately; it is not part of the decision rule.

Pooled warm efficiency is now scored across 2,150 requests per runtime (all five repetitions in test-clean, test-other, and the 30-clip voice proxy). Median request RTF is 0.1072 for Phonon vs 0.0610 for Parakeet; maximum process RSS is 1,673 vs 1,197 MiB. Phonon is worse on both efficiency measures. Together with the confirmed +1.097-point test-other WER gap and +0.589-point voice-proxy gap, the preregistered rule recommends keeping Parakeet. The one-file 60-minute comparison is still running and will be reported separately; it is not part of the decision rule.
Author
Owner

The single 60-minute file phase completed at 2026-09-30 10:39:41 UTC under /root/perf.lock; load average was 0.557 / 1.934 / 2.733 on four CPUs. Phonon received the lossless FLAC through the official multipart route with only the long-test request guard raised to 96 MiB. It completed 3,600 s in 272.102 s (RTF 0.0756), with 270.955 s decode time, 1,063.94 s process CPU (17.73 s/audio-minute), 2,180 MiB peak process RSS, 118 segments, and no truncation. Parakeet did not complete: its worker exited with code -6 after 4.277 s. ONNX Runtime reported that a ConvInteger node could not allocate a 11,796,480,000-byte buffer. CPU time and peak RSS are unavailable because the process aborted before the harness could sample them; the allocation request is not a peak-memory measurement. The long file contains repeated selected clips and tests throughput/resource handling, not long-form accuracy.

The single 60-minute file phase completed at 2026-09-30 10:39:41 UTC under `/root/perf.lock`; load average was 0.557 / 1.934 / 2.733 on four CPUs. Phonon received the lossless FLAC through the official multipart route with only the long-test request guard raised to 96 MiB. It completed 3,600 s in 272.102 s (RTF 0.0756), with 270.955 s decode time, 1,063.94 s process CPU (17.73 s/audio-minute), 2,180 MiB peak process RSS, 118 segments, and no truncation. Parakeet did not complete: its worker exited with code -6 after 4.277 s. ONNX Runtime reported that a ConvInteger node could not allocate a 11,796,480,000-byte buffer. CPU time and peak RSS are unavailable because the process aborted before the harness could sample them; the allocation request is not a peak-memory measurement. The long file contains repeated selected clips and tests throughput/resource handling, not long-form accuracy.
Author
Owner

Complete

Recommendation: keep Parakeet TDT 0.6B v2. Phonon-2 is within the prewritten test-clean WER limit (+0.382 percentage points), but misses the test-other limit (+1.097 points) and scores worse on the real-recording voice proxy (+0.589 points). Its pooled median request RTF and maximum warm RSS are also worse: 0.1072 vs 0.0610 and 1,673 MiB vs 1,197 MiB.

The 60-minute Phonon request completed in 272.102 s (RTF 0.0756). Parakeet aborted after 4.277 s with exit -6 when ONNX Runtime could not allocate an 11,796,480,000-byte ConvInteger buffer. Its CPU time and peak RSS were unavailable. The Phonon server's body cap was raised from 32 MiB to 96 MiB only for this long-file measurement.

The 30-item voice set is a real-recording Indian-English proxy, not synthetic speech. It uses other currencies and has no rupee amounts. A neutral CC source for synthetic voice notes with INR examples was not found. The report records this gap and the source scan.

Files

  • docs/perf/runs/2026-09-30-asr-ab.md
  • docs/perf/runs/2026-09-30-asr-ab-inputs.csv
  • docs/perf/runs/2026-09-30-asr-ab-gates.txt (full gate stdout and stderr)

Head: 961624b2762f20627fd435f25ebec49f71641e66

Gates

cargo fmt --check: no output; exit 0

cargo clippy --all-targets -- -D warnings:
error: #[derive(RustEmbed)] folder '/home/kayg/Developer/calternal-wt/asr-ab-489/crates/calternal-server/../../apps/web/build/' does not exist. cwd: '/home/kayg/Developer/calternal-wt/asr-ab-489'
exit 101

cargo test:
Finished `test` profile [unoptimized + debuginfo] target(s) in 32m 22s
exit 0; 83 test-result summaries, all with 0 failed

bun run check:
svelte-check found 0 errors and 0 warnings
exit 0

bun run test:
Test Files  137 passed (137)
Tests  888 passed (888)
exit 0

Clippy ran before apps/web/build existed. bun install --frozen-lockfile and the production web build later succeeded; the workspace tests then compiled calternal-server. I did not repeat the one-time gate run. cargo clean removed 11.9 GiB; generated frontend and benchmark outputs were removed. No product code changed.

## Complete Recommendation: keep Parakeet TDT 0.6B v2. Phonon-2 is within the prewritten test-clean WER limit (+0.382 percentage points), but misses the test-other limit (+1.097 points) and scores worse on the real-recording voice proxy (+0.589 points). Its pooled median request RTF and maximum warm RSS are also worse: 0.1072 vs 0.0610 and 1,673 MiB vs 1,197 MiB. The 60-minute Phonon request completed in 272.102 s (RTF 0.0756). Parakeet aborted after 4.277 s with exit -6 when ONNX Runtime could not allocate an 11,796,480,000-byte ConvInteger buffer. Its CPU time and peak RSS were unavailable. The Phonon server's body cap was raised from 32 MiB to 96 MiB only for this long-file measurement. The 30-item voice set is a real-recording Indian-English proxy, not synthetic speech. It uses other currencies and has no rupee amounts. A neutral CC source for synthetic voice notes with INR examples was not found. The report records this gap and the source scan. ## Files - `docs/perf/runs/2026-09-30-asr-ab.md` - `docs/perf/runs/2026-09-30-asr-ab-inputs.csv` - `docs/perf/runs/2026-09-30-asr-ab-gates.txt` (full gate stdout and stderr) Head: `961624b2762f20627fd435f25ebec49f71641e66` ## Gates ```text cargo fmt --check: no output; exit 0 cargo clippy --all-targets -- -D warnings: error: #[derive(RustEmbed)] folder '/home/kayg/Developer/calternal-wt/asr-ab-489/crates/calternal-server/../../apps/web/build/' does not exist. cwd: '/home/kayg/Developer/calternal-wt/asr-ab-489' exit 101 cargo test: Finished `test` profile [unoptimized + debuginfo] target(s) in 32m 22s exit 0; 83 test-result summaries, all with 0 failed bun run check: svelte-check found 0 errors and 0 warnings exit 0 bun run test: Test Files 137 passed (137) Tests 888 passed (888) exit 0 ``` Clippy ran before `apps/web/build` existed. `bun install --frozen-lockfile` and the production web build later succeeded; the workspace tests then compiled `calternal-server`. I did not repeat the one-time gate run. `cargo clean` removed 11.9 GiB; generated frontend and benchmark outputs were removed. No product code changed.
kayg closed this issue 2026-09-30 12:49:08 +00:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#489
No description provided.