I tested 11 Mandarin speech models. The most accurate hid the mistakes
I labelled 168 clips of my Mandarin text-to-speech podcast by ear and ran 11 open speech-recognition models against them. Qwen3-ASR-1.7B was the most accurate, but the most accurate models also silently fixed my AI voice's mispronunciations. What that means for choosing a quality checker, which two-pass setup I am switching to, and what no transcript can catch.
TL;DR
- 📊 Qwen3-ASR-1.7B was the most accurate of 11 Mandarin ASR models on my podcast audio: about 1 character in 40 wrong.
- ⚠️ The most accurate models silently "fixed" words my text-to-speech voice mispronounced, so the mistake vanished.
- 🚨 My current second-pass checker gets about 1 character in 7 wrong and floods review with false alarms.
- ✅ Swapping it for Qwen3-ASR-1.7B catches as many defects and cuts wasted listening by more than half.
- 👂 Pauses and tone never show up in any transcript, so a human ear stays in the loop.
I check every Mandarin episode of my podcast with speech recognition, and the check has a blind spot I could finally measure. The podcast library is voiced by a text-to-speech engine. Two speech-recognition (ASR) passes transcribe each episode, a language model compares the transcripts with the script, and anything doubtful comes to me as a short clip to judge by ear. Over two weeks I labelled 168 of those clips: 136 where I confirmed the audio says exactly the script, and 27 where I heard a real defect. That gave me something public benchmarks don't have: real podcast audio with a verified transcript, plus a list of mistakes that a person actually caught.
📊 The accurate models are very accurate
On audio I confirmed correct, the best model gets about one character in 40 wrong; my current second pass gets about one in seven wrong.

| Model | Characters wrong | Same, numbers normalised | Correct clips falsely flagged (of 136) |
|---|---|---|---|
| Qwen3-ASR-1.7B | 2.4% | 2.1% | 1 |
| Qwen3-ASR-0.6B | 2.8% | 2.7% | 1 |
| FireRedASR2-AED | 2.9% | 2.9% | 4 |
| Fun-ASR-Nano | 3.0% | 2.9% | 5 |
| Paraformer (my current first pass) | 4.4% | 4.3% | 3 |
| SenseVoice-Small | 4.4% | 4.3% | 5 |
| Whisper-large-v3 | 7.1% | 3.8% | 13 |
| Belle-Whisper-turbo-zh | 7.3% | 7.5% | 10 |
| Cohere Transcribe | 8.2% | 5.1% | 15 |
| Zipformer-large | 13.8% | 13.4% | 29 |
| Zipformer-middle (my current second pass) | 14.3% | 14.0% | 33 |
The top four beat my current first pass by a margin that holds up under resampling. Whisper-large-v3 and Cohere Transcribe look worse mainly because they write numbers as digits: count "V3" and "V三" as the same and Whisper's error rate nearly halves. My second pass, a small Zipformer model, missed 80 of 84 English terms on correct audio. It writes "Attention" as Chinese syllables.
🚨 Accuracy is not detection
The models that know Mandarin best hear what the speaker should have said, and that hides the defect I needed to find.

My TTS voice read 省 as "xing" instead of "sheng". Qwen3-ASR-1.7B and FireRedASR2-AED both wrote the correct character, so the transcript matched the script and the mistake disappeared. When the voice said "AutoML" as "Auto-L", FireRed and Qwen3-ASR-0.6B wrote "AutoML". Cohere Transcribe did the opposite: it wrote "Altomel", and "行" for the misread 省, which is exactly what it heard.
A quality checker needs at least one pass that transcribes the sound, not the meaning. A strong language model inside the recogniser is a feature for subtitles and a liability for QA.
🔍 My second pass mostly made noise
Replacing my second pass with Qwen3-ASR-1.7B catches as many defects and sends me fewer than half as many false leads.

My pipeline sends a spot to me when both passes disagree with the script, or when one pass shows a meaning change or a different reading of a character with several pronunciations. I scored all 55 possible pairs of the 11 models against that rule:
| Pair | Defects caught (of 10) | Correct spots sent to me (of 218 checked) | Correct clips falsely flagged (of 136) |
|---|---|---|---|
| Paraformer + Zipformer-middle (current) | 8 | 36 | 33 |
| Paraformer + Qwen3-ASR-1.7B (next step) | 8 | 14 | 3 |
| Qwen3-ASR-0.6B + Qwen3-ASR-1.7B | 8 | 11 | 1 |
| Qwen3-ASR-1.7B + Cohere Transcribe | 9 | 12 | 15 |
| Whisper-large-v3 + Qwen3-ASR-1.7B | 9 | 13 | 13 |
The new pair catches as many defects, though not the same ones: it misses a misread 调 that the old model caught, and catches a garbled "Llama" the old model missed. Two pairs reach 9 of 10, but with only ten defect spots, a one-catch edge is not a difference I trust yet. The case against the current second pass does not depend on that small sample: it rests on the 136 correct clips.
🧭 What a transcript will never tell you
Only 7 of the 19 audio defects I heard reliably change a transcript, so the human listening step stays.

Nine were pauses or a flat question tone, and three were the right character read the wrong way. Text ASR has no way to see those. My plan:
- Swap the second pass for Qwen3-ASR-1.7B.
- Keep my own ears on tone, pauses and the characters that have more than one reading.
- Keep labelling until I have 30 more word-level defects, enough to settle the 8-versus-9 question.
Choose the recogniser in a quality gate by what it notices, not by its leaderboard score.
How I measured
- Dataset. 168 clips (25 minutes) from the podcast's Mandarin episodes, each judged by me. A clip I marked correct uses its script text as the reference transcript. Clips from one early review page were cut longer than their card text, so those 13 are left out of the main numbers.
- Scoring. Character error rate after the same text normalisation the production checker uses, plus a pinyin-level error rate that does not count same-sound characters such as 它/他 as errors. Confidence intervals come from resampling the clips.
- Caveats. The audio is synthetic speech from two voices, not human speech, so the ranking may not transfer to human recordings. The defect results rest on ten spots. One Cohere Transcribe output dropped most of a clip.
- Models tested: Qwen3-ASR-1.7B and 0.6B, FireRedASR2-AED, Paraformer, Cohere Transcribe, Fun-ASR-Nano, SenseVoice-Small, Whisper-large-v3, Belle-Whisper-large-v3-turbo-zh, and two WenetSpeech Zipformer models.