Super Sonique / listening room

Less noise.
More voice.

50/50 versus synthetic-only. Compare the original Super Sonique model, trained on equal real-pair and synthetic-input branches, with a model trained entirely on synthetic inputs. Same 40 real-noise clips, 10 artificially noised clips, regular Auphonic references and Sidon baseline.

50/50 step 20,000 · EMASynthetic-only step 40,750 · EMABoth models Heun 8 · 15 evaluationsAudio 48 kHz · full clips

Different training corpora and training durations: this is a checkpoint comparison, not a controlled ablation of the data mixture. The previous 50/50 audio and clip order are preserved.

The artificial-noise speech was held out for the original 50/50 run. Overlap with the external synthetic-only training corpus is unknown: its full training manifests were not included in the checkpoint archive.

How to read the scores

MOS score recovery.

Recovery = 100 × (denoised MOS − noisy MOS) / (reference MOS − noisy MOS). Scores are predicted by Microsoft DNSMOS P.835, not human ratings. Negative recovery means degradation; above 100% means a score above the reference. A clean–noisy gap below 0.1 is marked n/a. Summary percentages use cohort mean scores, not the mean of clip percentages.

Audio metrics · overall scores and methodology

NISQA and DNSMOS predict overall speech quality (higher is better). SpkSim measures speaker-embedding cosine similarity to the actual noisy input, using WavLM Base Plus SV—not similarity to the clean reference. The input’s self-similarity is shown as a dash. These are the restoration metrics in Sidon §4.2; the paper also uses MMS word error rate for English and character error rate for multilingual speech.

ASR-reference WER, not verified WER: multilingual Parakeet TDT 0.6B v3 transcribes each regular clean reference and each comparison waveform with automatic language detection. Word error rate measures disagreement against that automatic reference text, which may itself contain errors. This replaces the paper’s MMS recognizer at your request; lower is better. The clean reference’s zero is a self-comparison, not a transcription-accuracy result. Empty reference text is marked n/a.

This panel includes multiple languages. Interpret learned quality scores cautiously where the language or recording conditions differ from the scoring model’s training coverage; NISQA is a prediction, not a language-independent human rating.

Scoring protocol and comparison limits

All tables below score uncompressed, pre-playback float waveforms. “Actual input” is what the denoisers received; on real recordings, it is not the codec reconstruction in the noisy player. The MOS recovery cards and sort order retain their original scoring basis (public MP3 for real clips; float WAV for artificial clips), so their MOS values can differ from the DNSMOS table.

Cohort tables summarize all 40 real or 10 artificial clips and are not changed by source filters. NISQA, DNSMOS and SpkSim are arithmetic means; ASR-reference WER divides total substitutions, deletions and insertions by total reference words, not a mean of clip percentages. These metrics are not a reproduction of Sidon’s published benchmark. The paper does not specify exact NISQA or DNSMOS checkpoint revisions, and our ASR model and automatic reference text differ from its protocol. Reference scores are context, not an upper bound; higher SpkSim does not by itself imply better denoising.

ReferenceRegular Auphonic
NoisyReal codec / artificial input
50/50EMA step 20,000
Synthetic-onlyEMA step 40,750
Sidon v0.1Original Sidon

Loading the listening room…

Real recordings

More real-noise clips

The remaining test pairs, ordered by the original 50/50 model's MOS recovery. The noisy listening baseline is a KVAE reconstruction, as in the previous site; all denoisers receive the original noisy waveform. Recovery MOS is calculated on the public MP3s.

Controlled corruption

Artificial noise · +5 dB

Ten held-out regular clean speech clips, mixed with recorded noise using the Sidon augmentation pipeline. Noise comes from the training noise bank: this tests held-out speech under a matched noise distribution, not unseen noise. No reverb, bandwidth, codec, clipping augmentation or packet loss is enabled. The mixer’s final saturation still applies; +5 dB is measured before saturation.

The noisy player is the exact saved corrupted input from the previous comparison. Following the evaluation guide, MOS uses pre-playback float WAVs for this section. Expand codec diagnostics to hear noisy and clean KVAE reconstructions. Best Super Sonique MOS recovery first.

ASR-reference WER is diagnostic: the artificial-noise panel has unverified language labels and may include languages outside Parakeet's supported set.