A model can get almost every spoken word right and still give a Deaf person an incomplete picture of what happened.

DeafBench’s non-speech v1 test made that visible. OpenAI Whisper turbo scored 2.0% word error rate and preserved 95.0% of the critical spoken information. Those numbers look strong.

It captioned 0 of 19 environmental sound events.

One score answered the wrong question

Word error rate asked how closely the transcript matched the spoken words. It did not ask whether the model reported the alarm, knock, siren, phone ring, or door closing around those words.

That is not a complaint about the math. It is a reminder to match a metric to the experience we actually care about.

Use more than one view

DeafBench keeps word error rate because it is useful. Then it adds separate measures for critical information and non-speech events. A model does not get to hide one kind of failure inside a good average from another.

Measured result: 12 samples, 2.0% WER, 95.0% critical-information recall, and 0.0% non-speech recall. The full report is in the DeafBench repository.

This test is small, and it should be described that way. It is evidence of a specific failure mode, not proof that one model is always better or worse than another.