Sales
8 Minutes reading time

Testing Voice-to-CRM: Why the Word Error Rate (WER) Isn't Enough

The word error rate tells you how many words are wrong - not whether the values in your CRM are correct. Here you'll learn how to test Voice-to-CRM with critical names, numbers, and product codes, measure errors separately, and define acceptance criteria.
Key Takeaways
In This Article

AI This article was created with the help of AI.

Key takeaways

  • WER counts substitutions, omissions, and insertions equally: one wrong name produces the same error count as one filler word.
  • Build your test set from confusable pairs like Meier/Mayer, fifteen/fifty, and 1.5/15, plus anonymized product-range terms.
  • Separate four checkpoints: understood text, normalized value, correct field, correct record - and measure correction time.
  • Phrase lists can weight names and domain terms during recognition; clarify with the vendor what they actually offer for this.

Why does the word error rate tell you too little?

When you evaluate Voice-to-CRM, the discussion almost always starts with the word error rate. The metric has a clear definition: WER is the sum of substitutions, omissions, and insertions, divided by the word count of the human-created reference transcript.[1] That's a clean metric for speech recognition models. For your CRM data quality, it's not enough.

Three limitations you need to know

  • WER counts every word equally. A wrong filler word and a wrong amount produce the same error count. In your CRM, the consequences are completely different: one goes unnoticed, the other skews your forecast.
  • WER is blind to capitalization, number formats, and name spellings. Yet these are exactly the values missing from your dashboards when they arrive incorrectly.
  • WER says nothing about the destination. Whether 'one point five' ends up in the amount field as 1.5 or as 15 is decided after recognition, during normalization and field mapping.

Important context: this is not a general ranking of speech models. It's about the limits of the metric. A low WER is a necessary but not sufficient condition for clean CRM data. That's why you test Voice-to-CRM on the values that really carry your pipeline: names, amounts, quantities, and product codes.

Which names and numbers belong in your test set?

Your test set needs confusable pairs that have to be distinguished in real-world sales. Fictional but realistic cases are perfectly sufficient: they only need to cover the error patterns that will later hit your real deals.

Case Spoken Critical Because
Last Name Meier Sounds almost identical, but lands in different contact records
Number Word fifteen Order-of-magnitude error in the quantity
Decimal Amount one point five Factor of ten in the amount or discount field
Product Code PX-450 Spelling and digit sequence must come through cleanly
Article Number A-1207 vs. A-1702 Swapped digits create the wrong line item in the quote

‍

For each case, note two things in advance: the expected value and the target CRM field. Only this expected-vs-actual comparison makes the test measurable. Without the notes, you'll end up debating what the assistant 'meant' instead of being able to measure.

  • Record the expected value exactly: 1.5 stays 1.5, not 15.
  • Name the target CRM field: opportunity, amount field, custom field, contact.
  • Add anonymized terms from your real product range: your own product names, internal project codes, typical customer roles.

The last point is the most important one. Entity errors - a wrong name or a wrong number - hurt your data more than half a point of overall WER. A generic test set with standard sentences hides exactly these errors.

How do you test speech under realistic conditions?

A test in a quiet office tells you little about the moment that counts: after the customer meeting, in the car, on the street. You repeat the same cases over the same channel you use in everyday life, under two conditions.

  1. Quiet office: the reference condition to see your baseline accuracy.
  2. Typical post-visit environment: car, street noise, hands-free system. This is exactly where debriefs happen in practice.
  3. Record per run: speaker, connection (voice call, network quality) and exact inputs. Only then can you attribute differences instead of guessing.

For context on sample size: Microsoft requires 30 minutes to 5 hours of representative audio for meaningful accuracy tests.[2] A single test run falls far short of that. Don't derive a representative benchmark from it - treat it as a snapshot. The word error rate (WER) depends heavily on condition and environment; the same cases can be recognized cleanly in the office and incorrectly in a stationary car.

In practice, that means: plan several short runs across several days, not one long marathon. Note per run which cases fail repeatedly. Repeated errors under the same conditions are a real pattern; a one-off outlier is usually just noise.

How do you distinguish recognition errors from field errors?

When a wrong value ends up in the CRM, there can be four different causes. Only once you separate the layers do you know where to intervene: the speech model, the normalization, the field mapping, or the record matching.

Checkpoint Question Example of an Error
Understood Text Was the spoken content recognized correctly? 'Meier' is transcribed as 'Mayer'
Normalized Value Was the value formatted correctly? Spoken 'one point five' ends up as 15 instead of 1.5
Correct Field Does the value land in the intended field? The amount sits in the free-text field instead of the amount field
Correct Record Was the right entry updated? The update goes to the wrong opportunity

‍

The end goal is the structured CRM update from the voice debrief: the data must arrive in the right field and the right format, not just sit in the transcript. That's exactly how the Phone Assistants work: the spoken debrief becomes structured updates that land in the right fields in the CRM.[3]

For each error, note what correction was needed and how long it took. That's your real currency: not the error count, but the remaining correction effort per case. Record matching itself - assigning a conversation to the right CRM entry - is a topic of its own and deserves its own article.

What do domain vocabulary and follow-up questions achieve?

There is a technical way to strengthen critical terms: phrase lists. These are pre-defined term lists that increase the recognition weight of those words without training a custom model. Microsoft describes this for Azure Speech Services: names, places, homonyms and industry-specific terms can be weighted via a phrase list, with the weight adjustable in a range from 0.0 to 2.0.[4]

Azure is an example of technical feasibility here, not proof of a specific product architecture: which provider uses which engine in the background is usually something you can't verify from the outside. But the question you ask your vendor stays the same, no matter which engine is running.

  • Can the vendor store your own terminology - product names, project codes and frequent customer names?
  • Can the assistant spell things out during capture, for example with product codes like 'PX-450'?
  • Does the assistant confirm critical numbers with a follow-up question before the value goes into the CRM?

The third point is often more effective than any model metric. A follow-up question costs seconds. A wrong figure in the forecast costs credibility - at the next pipeline review and in front of management. An assistant that actively confirms critical values takes more off your plate than a transcript that looks fluent but misses the details.

When does Bliro pass your real-world test?

In the end, what counts is an acceptance checklist, not a demo impression. You fill it with your own cases and set the thresholds before the test, together with sales IT and the team. Don't take accuracy guarantees from vendor demos at face value; a threshold proposal is exactly that: a proposal you align on before the test.

Metric What You Measure Example Threshold (Suggestion)
Critical Values Correct Share of test cases where name, amount and product code land correctly in the right field At least 90 percent of critical values correct
Unresolved Errors Cases that remain faulty after the run and have to be corrected manually No unresolved errors for amounts and product codes
Correction Time Average time to correct a faulty case Under 30 seconds per case

‍

You then deliberately repeat the problematic cases from your test set with the Bliro Phone Assistants (Vicky & Tim). The assistants hold a real voice call before and after the customer meeting, structure the debrief into a visit report and fill CRM fields from what was said, instead of being typed.[3] Exactly this path, from the spoken sentence to the filled field, is what your acceptance checklist tests.

So Bliro passes your practical test when the critical values from your real product range reliably land in the right field, unresolved errors stay below your threshold and the correction time stays within bounds. Everything else is a demo.

The next step: run critical values through with Vicky and Tim

You now have two tools: a test set with confusable pairs, expected values and target CRM fields, and an acceptance checklist with three metrics. Both work without vendor support. Take both and test your own names, amounts and product codes, not those of a demo account.

  • Run the critical cases from your real sales routine with the Bliro Phone Assistants (Vicky & Tim): preparation via voice call, debrief after the appointment, CRM update from what was said.
  • On the product page for field sales you can see how the assistants work in your reps' daily routine: Bliro for field sales.
  • Decide at the end with data instead of impressions: do the critical values pass your test, or not.

Sources

  1. https://en.wikipedia.org/wiki/Word_error_rate
  2. https://learn.microsoft.com/en-us/azure/ai-services/speech-service/how-to-custom-speech-evaluate-data
  3. https://help.bliro.io/en/articles/15068158-meet-vicky-tim-voice-ai-assistants-for-field-sales
  4. https://learn.microsoft.com/en-us/azure/ai-services/speech-service/improve-accuracy-phrase-list

A Day in the Life of a Field Sales Rep, Powered by Bliro.

A field sales rep operates Bliro entirely by voice from the car: right after each customer visit he calls Vicky, Bliro's AI voice assistant, and dictates his visit report while driving. Bliro then updates the CRM, schedules the follow-up in his calendar and drafts the follow-up email - voice-to-CRM and the full desk work, with no admin left for the evening.
A Day in the Life of a Field Sales Rep, Powered by Bliro.

By clicking play you agree to load content from YouTube and to marketing cookies.

Your questions, our answers

What is the word error rate (WER) and how is it calculated?
Why is the WER not enough to evaluate Voice-to-CRM?
Which names and numbers should I include in my test set?
How do I test speech recognition under realistic conditions?
What are phrase lists and do they help with names and technical terms?
How do I define acceptance criteria for Voice-to-CRM?

The personal assistant for your field sales team

Vicky & Tim are Bliro's AI voice agents for B2B field sales teams. They prepare conversations, maintain CRM entries, and create follow-ups - by voice, without typing. A transcription of conversations can optionally be used in addition.
Book a Demo