Most speech-to-text is benchmarked on audio that looks nothing like a WhatsApp voice note.
The standard evaluation sets are read speech, broadcast news, or recorded interviews: single speaker, decent microphone, one language, quiet room, speaker aware they are being recorded. A WhatsApp voice note is close to the opposite on every axis. I have spent a while building around this, and the gap turned out to be wider than I expected.
Acoustics
Phone held at arm's length while walking, in a car, in a kitchen, on a street. Distance-to-mic varies wildly within a single recording, which breaks a lot of assumptions about consistent gain.
Then there is the codec. Voice notes are Opus at low bitrate — efficient, but it discards exactly the high-frequency detail that helps disambiguate fricatives. /s/ versus /f/ versus /th/ get genuinely harder, and those distinctions carry real meaning.
Register
Conversational, not read. False starts, self-corrections, filler, trailing
Discussion
Say something first
It all starts with you—share your thoughts now.