Interspeech 2026 · Sydney

Probing Low Frame Rate Degradation in Neural Audio Codecs

Alex Gichamba1   Moise Busogi1

1 Carnegie Mellon University Africa, Rwanda

Abstract

Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where generation cost scales linearly with sequence length. Recent work shows codecs can operate at 12.5 Hz and below, but the mechanisms behind low frame rate degradation remain insufficiently understood. We investigate them through a controlled frame rate ablation. We reproduce a quality cliff at 6.25 Hz reported in prior work and evaluate two candidate explanations, phonemic collision and codebook saturation, neither of which shows evidence of a fundamental barrier. The cliff is instead caused by a training misconfiguration: a fixed clip duration yields too few tokens at low frame rates, starving the decoder of inter-token context. Once corrected, word error rate degrades smoothly with phonemic load down to 3.1 Hz and 1.6 Hz, suggesting the efficiency gains of low frame rate codecs are more accessible than previously assumed.

Listen

Hearing the frame rate cliff

Standard training keeps a fixed clip duration, so at 6.25 Hz each clip is only two tokens long and the decoder never learns to span token boundaries. Matching the token count instead recovers intelligibility at the same frame rate and bitrate. You can hear the difference.

6.25 Hz · same model, same bitrate (750 bps), one training change

Both clips reconstruct the same utterance at 6.25 Hz. Only the training-time token count differs.

Fixed T clip = 0.38 s

Standard DAC recipe. Two tokens per clip during training → collapse.

WER 107.4% · STOI 0.46 · MCD 20.17

Fixed K = 19 tokens

Matched sequence length. Near parity with the 12.5 Hz baseline at half the bitrate.

WER 15.4% · STOI 0.89 · MCD 3.72

Listen · the full sweep

Reconstructions across frame rates

Matched-token (K = 19) DAC reconstructions of one utterance from LibriSpeech test-clean. Intelligibility falls off smoothly with phonemic load all the way to 1.6 Hz (192 bps).

Referenceground truth original 16 kHz audio
100 Hz12 kbps WER 5.0% · STOI 0.98 · SPK 0.96
50 Hz6 kbps · DAC WER 5.4% · STOI 0.97 · SPK 0.93
25 Hz3 kbps WER 5.8% · STOI 0.95 · SPK 0.90
12.5 Hz1.5 kbps WER 7.2% · STOI 0.93 · SPK 0.82
6.25 Hz750 bps WER 15.4% · STOI 0.89 · SPK 0.62
3.125 Hz375 bps WER 29.4% · STOI 0.84 · SPK 0.48
1.6 Hz192 bps WER 63.2% · STOI 0.76 · SPK 0.32

Cite

BibTeX

@inproceedings{gichamba2026probing,
  title     = {Probing Low Frame Rate Degradation in Neural Audio Codecs},
  author    = {Gichamba, Alex and Busogi, Moise},
  booktitle = {Proc. Interspeech 2026},
  year      = {2026},
  address   = {Sydney, Australia}
}