← Ideas & guides

Guide

Who can actually use your AI voice service?

A fluent answer does not establish who can use a voice service. Examine languages, ways of speaking, abandoned attempts and whether the alternative communication route actually works.

By Dr. Sven JungmannPublished: · Reviewed:
A brass acoustic horn stands beside a blank slate board and a piece of chalk.

At a glance

  • Supported languages do not establish that the intended people can complete their tasks.
  • Read-speech benchmarks and real telephone conversations answer different questions.
  • Include abandonment, assistance and the success of alternative communication routes.

A fictional hospital is considering an AI telephone service for arrival information and callback requests. The demonstration is reassuring: a clear question receives a fluent answer. Before extending that demonstration to patients, I would ask who was able to take part in it. Which languages, ways of speaking and communication needs were represented? Which people could complete the task without someone else speaking for them?

Mati Staniszewski’s Stanford conversation prompted this question. He describes voice systems assembled from recognition, language processing and speech generation, while discussing different communication preferences. This is a supplier’s account of the technology. It offers no independent evidence that a particular healthcare service is accessible. [1]

For a medical or technology leader, accessibility belongs in the definition of the service. A channel that works well for one set of callers may offer little help to another. The following approach is my proposal for examining that boundary. It has not been evaluated as a protocol, and the hospital example is constructed.

Describe whose task should become easier

Begin with the people expected to use the service and the work they want to complete. In the example, a caller might need to find the correct entrance or request a callback from a named department. Write down what a successful outcome means in ordinary language, including how the caller will know that the request has been received.

A list of supported languages leaves important questions open. Can someone understand the reply, formulate their request naturally and correct an error? What happens when they switch languages for a department name, spell a surname or pause before continuing? These are questions for the intended interaction, rather than properties that follow automatically from a language label.

I would discuss the initial coverage with reception staff, relevant accessibility representatives and prospective users. Identify language and communication needs that matter locally, then record where evidence is missing. A group absent from the test remains unexamined. Its absence should not quietly become a reason to treat its needs as unusual or unnecessary.

Examine what a recognition score actually covers

Blaschke and colleagues’ 2025 Betthupferl study compared recognition models on professionally read German dialect stories. Errors exceeded those for standard speech; some reference mismatches were valid alternatives. The small benchmark covered read speech. Telephone conversations and healthcare tasks were outside its scope. [2]

This leads to two separate checks in the proposed service. First, did the system preserve what the person meant? Second, did the interaction produce the intended outcome? A grammatically polished transcript could contain the wrong destination. A transcript that preserves dialectal wording could still express the request correctly. The assessment needs someone who can judge the utterance in its language and context.

For the fictional hospital, I would prepare ordinary requests and foreseeable variations with the people involved. Let participants express them in their own way. Ask them afterwards what they intended, and compare that with what the service recorded and did. A disagreement belongs in the record even when the final answer sounds plausible.

Include people whose speech the demonstration did not represent

The 2026 HeyJay! paper tested Whisper large-v3-turbo on read English commands from 36 speakers with neuromotor speech impairments, finding widely varying word error rates. It provides no German telephone-service validation; spontaneous dialogue and several forms of atypical speech were absent. [3]

For local planning, I would keep language knowledge, pronunciation and speech impairment distinct. They may overlap in a person’s experience, but they are different reasons to examine whether an interaction works. The cited studies do not supply a ready-made distribution of needs for a German clinic. That requires local information and participation.

Invite relevant participants under an agreed, accessible process, with their consent and appropriate support. Do not ask staff to imitate a disability or manufacture an accent and then treat the result as evidence about the people concerned. Artificial examples can expose obvious design faults; they cannot establish the lived usability of the service. Participation should also allow someone to stop without losing access to the ordinary service.

Record only the characteristics needed for the agreed evaluation, with the relevant safeguards. Keep the evaluation focused on agreed communication barriers, without inferring health characteristics from a caller’s voice. Where the evidence is thin, describe its limits explicitly instead of producing a confident percentage for a very small group.

Count the people who never reach a completed conversation

A review limited to successful recordings can miss the decisive failure. Someone might abandon the call after repeated requests to rephrase, never reach the department they need, or ask a relative to take over. In the constructed example, those outcomes belong alongside completed requests when judging access.

I would record the intended task, its outcome, repetitions, corrections, assistance and abandonment. Keep the number of people visible as well as the number of attempts: many successful attempts by one comfortable speaker do not establish broad coverage. Distinguish an untested situation from a tested failure and from a supported interaction. Avoid collapsing all three into a single overall success rate.

Compare patterns across the locally relevant needs without pretending that a small, voluntary sample represents everyone. If one pattern repeatedly blocks completion, investigate the interaction and involve the affected participants in the next design decision. The question is what prevented access and what change might remove that obstacle. A lower overall error average would not settle it by itself.

Test the alternative route with the same care

The service should make another route available before a caller has to prove that speaking to the system is failing. Depending on the setting, this might involve a person, a usable text channel or another established communication option. Announcing that an alternative exists is only the beginning: the person must be able to find it, use it and understand what happens next.

A text form may suit someone who cannot use the voice channel. It may be difficult for someone who chose voice because typing is hard. A transfer to a person could help with an unfamiliar pronunciation, while still leaving a language barrier unresolved. I would test the actual alternative with the participants for whom it is intended, including waiting, repeated explanation and whether the original task is eventually completed.

Before launch, the hospital’s decision should therefore name the interactions supported by evidence, the remaining gaps and the workable routes available around them. Keep those boundaries visible in the service description. An accessibility claim should become more specific as evidence improves. For me, that is the meaningful promise of a voice service: a particular person can reach the help they sought, using a way of communicating that actually works for them.

Sources and further reading

  1. Staniszewski: voice systemsStanford Online

    Public starting point.

  2. Betthupferl: dialect recognitionInterspeech 2025

    Read dialect stories.

  3. HeyJay!: atypical speechScientific Data

    Read English speech; version of record September 3, 2026.

The starting point for this reflection

Staniszewski: voice systems

Perspective and interests

This article was developed with AI assistance. The hospital example is fictional. The proposed approach is an original organisational inference and has not been evaluated as a protocol.

I am the founder and CEO of aiomics and have a commercial interest in responsible AI adoption in medicine. This article provides no treatment recommendations and does not replace specialist accessibility assessment.

Keep reading

What should your event make possible?

Tell me about your audience, occasion and timing. We can shape a talk around the questions that matter to them.

Enquire about a talk