Word Error Rate

Word Error Rate

The word error rate measures how many mistakes a speech recognition system makes: it compares the recognized text with what was actually said, and expresses the result as a percentage. The lower the word error rate, the more accurate the recognition.

When a program converts spoken language into text, it makes mistakes along the way. It can misrecognize words, leave them out, or invent ones that were never said at all. The word error rate — WER for short — is the metric used to measure how many of these errors occur. It sets the number of incorrect words in relation to the total number of words in the original and expresses the result as a percentage. A word error rate of 0% means: every word was recognized correctly. The higher the value, the more went wrong.

What the word error rate is measured against

Without a uniform measure, any manufacturer could claim their system is the best — without any proof. The word error rate creates comparability. You can use it to test two systems on the same audio files and directly see which one is more accurate.

For many applications, there are rough benchmark values: under 5% is considered very good in research. Humans transcribing spoken language among themselves achieve around 4%. Modern systems reach this range with clear audio quality and standard language — but with background noise, strong accents, or technical jargon, the value can quickly rise to 20% or more.

Importantly, the word error rate treats all errors equally. A misrecognized filler word like “uh” counts just as much as a misrecognized name in a medical report. This is a well-known weakness of the metric — it says nothing about whether an error was harmless or consequential.

Three types of errors behind the number

The word error rate combines three different kinds of errors. First, substitutions: the system recognizes a word, but the wrong one — “bears” becomes “bares,” for instance. Second, deletions: a word that was spoken doesn’t appear in the recognized text at all. Third, insertions: the system adds a word that was never said. All three types count as errors, are added together, and divided by the total number of words in the reference text.

The reference text principle is crucial here. You always need a template — a carefully transcribed, error-free text of what was spoken. This is called a transcript. The system compares its output word for word against this transcript and counts the deviations. Without such a reference text, the word error rate cannot be calculated.

A small calculation example: someone says five words. The system misrecognizes one and omits one. That amounts to two errors out of five words — a word error rate of 40%. If only one error had occurred, the value would be 20%. This allows quality to be read off directly.

Word error rate in products and headlines

The word error rate crops up everywhere speech is automatically recognized. Voice assistants like Siri or Google Assistant, automatic captions on YouTube or in video conferences, dictation features on smartphones — all of these are speech recognition systems, and all of them are evaluated, among other things, using the word error rate.

The metric appears in tech news whenever companies unveil new speech models. A typical headline reads: “New model achieves WER of 3.5% on the LibriSpeech dataset” — where LibriSpeech is a widely used collection of audio files against which systems are compared by default. Such benchmarks, meaning standardized tests, make it possible to compare progress over the years.

In areas where errors are especially costly — such as medical documentation or automatic captioning of news broadcasts — the word error rate alone is often not enough. There, additional criteria are applied, for example how often key terms specifically are misrecognized. Still, the word error rate remains the first and most widely used point of reference when it comes to comparing speech recognition systems.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.