A model can get almost every spoken word right and still give a Deaf person an incomplete picture of what happened.
DeafBench’s non-speech v1 test made that visible. OpenAI Whisper turbo scored 2.0% word error rate and preserved 95.0% of the critical spoken information. Those numbers look strong.
It captioned 0 of 19 environmental sound events.
One score answered the wrong question
Word error rate asked how closely the transcript matched the spoken words. It did not ask whether the model reported the alarm, knock, siren, phone ring, or door closing around those words.
That is not a complaint about the math. It is a reminder to match a metric to the experience we actually care about.
Use more than one view
DeafBench keeps word error rate because it is useful. Then it adds separate measures for critical information and non-speech events. A model does not get to hide one kind of failure inside a good average from another.
This test is small, and it should be described that way. It is evidence of a specific failure mode, not proof that one model is always better or worse than another.