# Speech recognition, TTS & voice AI FAQs

> Answers about ASR, transcription, speech data, consent, WER, code-switching, TTS listening tests, pronunciation, voice agents, and monitoring.

Published: 2026-08-20. Updated: 2026-08-20. Last verified: 2026-08-20. Review interval: 183 days. Disclosure: none.

## How to use these answers

Start with the question that matches your decision about speech and voice systems. Each answer defines the term, then names the operational consequence that a brief, workflow, test, or contract should make explicit.

## The rule behind the page

A speech metric is meaningful only for a defined user, language variety, acoustic condition, task, and consequence. Listen to failures and segment results before turning an average into a release decision.

- Define the audience and purpose
- Name the languages, locales, scripts, and modalities
- Separate requirements from preferences
- Assign decision and escalation authority
- Keep source, output, and evidence versioned

## Frequently asked questions

### Scope

#### What is the difference between ASR and transcription?

ASR is technology that generates text from speech. Transcription is the delivered written record and may include human creation or correction, speaker labels, timestamps, formatting, and project-specific conventions.

#### What is the difference between TTS and voice-over?

TTS synthesizes speech from text. Voice-over is a production service that can include script adaptation, casting, performance, direction, recording, editing, and mixing, whether the voice is human or synthetic.

#### What belongs in a speech localization brief?

Specify use case, languages and varieties, speakers, channels, acoustic conditions, latency, accessibility, privacy, allowed transformations, output formats, and acceptance evidence. Separate transcription, translation, subtitling, dubbing, ASR, TTS, and voice-agent work rather than calling all of it voice.

#### How should dialect and accent requirements be written?

Name the target audience and operational need, provide representative samples, and distinguish language variety from individual accent. Avoid vague native or neutral labels. Define whether variation should be preserved, normalized, recognized, synthesized, or used only for evaluation.

#### When is verbatim transcription appropriate?

Use verbatim output when disfluencies, repetitions, discourse markers, or speech events are evidence, such as research, legal, or model-data tasks. For captions or readable records, an edited specification may serve users better. State the rule and exceptions.

#### How should multilingual and code-switched audio be scoped?

Identify expected language combinations, switching granularity, script and transliteration rules, speaker behavior, and how unknown language should be marked. Sample real code-switching during pilots because monolingual averages conceal boundary failures.

### Data

#### How much speech data do we need?

There is no universal hour count. It depends on whether you are training, adapting, or evaluating; the model and task; language and accent coverage; recording conditions; label depth; rights; and the performance gap you need to close.

#### What makes speech data representative?

It covers the users, language varieties, devices, channels, environments, speaking styles, demographics relevant to the task, code-switching, noise, and difficult entities—not only studio recordings from convenient speakers.

#### What consent is needed for voice data?

Consent and other lawful bases depend on jurisdiction and use. The program must clearly define collection, model or evaluation use, reuse, sharing, retention, withdrawal limits, synthetic voice implications, and any biometric or sensitive-data treatment.

#### How much speech data is enough?

There is no universal duration. Need depends on task, model, language variation, speakers, channels, noise, label complexity, and target error. Use learning curves and error coverage to decide whether another hour adds representative evidence rather than chasing a round number.

#### How should speech consent be documented?

Record what was captured, purposes, recipients, retention, model or product uses, commercial use, withdrawal limits, and whether voice identity may be replicated. Consent language should match the actual lifecycle and applicable employment, biometric, publicity, and privacy obligations.

#### What makes a speech dataset representative?

Representation covers relevant languages, varieties, speaker demographics where lawful and useful, devices, rooms, noise, speaking styles, health or accessibility conditions, and task situations. Report coverage and gaps rather than claiming a dataset represents everyone.

### ASR evaluation

#### Is word error rate enough to evaluate ASR?

No. Add semantic and entity accuracy, commands or task success, critical-term errors, speaker attribution, punctuation if relevant, latency, confidence behavior, and segmented results by language, accent, condition, and user group.

#### How do code-switching and names affect ASR?

They expose language-identification, vocabulary, tokenization, pronunciation, and context limitations. Build targeted sets with real switching patterns and important names rather than translating monolingual scripts.

#### Why is word error rate not enough?

WER weights substitutions, deletions, and insertions but can hide named entities, negation, numerals, punctuation, speaker attribution, language switching, and unequal subgroup failures. Pair it with task-specific error slices and human review of consequential examples.

#### How should ASR be evaluated across language varieties?

Create representative slices with enough evidence, use appropriate normalization and tokenization, report uncertainty, and review error types with speakers or experts from those varieties. Do not pool a dominant variety into one reassuring language score.

#### How should names, numbers, and terminology be tested?

Build focused sets from the actual domain, including plausible confusions and formatting. Evaluate semantic correctness as well as surface spelling, and test any contextual biasing without allowing it to hallucinate terms that were not spoken.

#### What should real-time ASR testing include?

Measure partial and final accuracy, stability, endpointing, latency distribution, interruptions, packet loss, background speech, correction behavior, and downstream impact. A strong offline transcript score does not prove a usable live conversation.

### TTS evaluation

#### How should TTS quality be evaluated?

Combine listening studies with task checks for intelligibility, naturalness, pronunciation, prosody, speaking rate, emotion if required, audio defects, voice consistency, locale fit, and safety. Use qualified listeners from the target population.

#### What is wrong with using only mean opinion score?

MOS compresses subjective responses into an average and depends on the exact prompt, scale, audio setup, language, and raters. It can hide pronunciation or subgroup failures that break a product.

#### How should synthesized speech be evaluated?

Evaluate intelligibility, pronunciation, naturalness, prosody, meaning, voice suitability, consistency, artifacts, and task success with target-language listeners. Include long-form and difficult inputs; a few polished samples are not a production evaluation.

#### How should pronunciation dictionaries be governed?

Store the term, language and variety, phonetic or textual guidance, context, source, approval, exceptions, and version. Test in sentences and across voices. A pronunciation that works for one model or grammatical form may fail elsewhere.

#### What should be checked when a TTS model speaks translated content?

Review the translation before synthesis, then check pronunciation, segmentation, emphasis, numbers, abbreviations, names, code-switching, and timing. Separate text defects from voice-rendering defects so the right asset or system is corrected.

#### How should synthetic-voice consistency be measured?

Sample scripts, emotions, lengths, recording conditions, updates, and languages, then assess voice identity, loudness, pacing, prosody, pronunciation, and artifacts. Define which variation is acceptable and retain versioned reference outputs.

### Voice agents

#### What must be localized in a conversational voice agent?

Recognition, language detection, prompts, TTS voice and pronunciation, dialogue strategy, interruption, confirmations, error recovery, tool entities, cultural norms, accessibility, escalation, and safety behavior.

#### How should a voice agent handle low recognition confidence?

Use task-appropriate confirmation, constrained choices, reprompting, spelling or alternate channels, and human escalation. Do not expose a raw score or repeatedly blame the speaker.

#### How should a multilingual voice agent choose language?

Use explicit user choice when possible, preserve it across turns, and treat automatic detection as revisable evidence. Handle code-switching and low-confidence detection without trapping the user, and avoid equating location, name, or accent with preferred language.

#### What should happen when a voice agent does not understand?

Use a clear repair prompt, offer slower repetition, spelling, keypad, text, interpreter, or human transfer where appropriate, and preserve context during escalation. Repeatedly guessing can turn an ASR error into a consequential action.

#### How should interruption and turn-taking be localized?

Test pause length, overlap, backchannels, politeness, confirmation, and barge-in with target-language users. Conversation timing and signs of completion differ; simply translating the words can make an agent feel rude or unusable.

#### When must a voice agent confirm an action?

Confirm before consequential, costly, privacy-sensitive, or difficult-to-reverse actions and whenever recognition ambiguity changes intent. Read back the critical fields in a locally understandable form, then capture explicit acceptance or route to a person.

### Quality

#### Who should review speech output?

Use target-language listeners with relevant variety and domain knowledge; add phoneticians, audio engineers, accessibility specialists, or safety reviewers when the task requires them. One bilingual manager is not a representative listener panel.

#### How should speech annotation disagreements be resolved?

Use written rules, shared calibration audio, explicit uncertainty labels, and adjudication by qualified people. Preserve difficult examples and update the guideline. Forced agreement without learning hides ambiguity in the data and inflates quality claims.

#### What is an effective speech QA sample?

Combine random coverage with risk, subgroup, language, device, low-confidence, novel, and known-error slices. Sample complete interactions as well as clips so reviewers can see whether errors recover or compound.

#### How should audio quality defects be categorized?

Separate source capture, segmentation, labeling, language, translation, synthesis, mixing, and delivery defects. Add severity based on intelligibility, meaning, accessibility, user action, and reach, not merely technical annoyance.

#### How should speech output be tested for accessibility?

Test with people who use the relevant captions, transcripts, audio description, controls, or assistive technology. Evaluate comprehension, navigation, timing, identification, and recovery in the real channel; technical conformance alone does not prove an equivalent experience.

#### How should privacy affect speech quality review?

Minimize reviewer access, redact or mask where feasible, use secure tools, limit downloads, log access, and define retention. Privacy controls must not silently remove the context needed for a valid judgment; adjust the evaluation design deliberately.

### Operations

#### What belongs in a pronunciation lexicon workflow?

Term source, language and locale, spelling variants, phonemic or platform format, audio evidence, status, approver, model or engine, test sentence, and version history.

#### How do we monitor speech quality after launch?

Track task failures, corrections, fallback, abandonment, latency, critical entities, language and device segments, user reports, and sampled audio under approved privacy controls. Feed confirmed failures into data, lexicons, prompts, and regression sets.

#### How should speech assets be versioned?

Link audio, transcript, translation, timing, speaker metadata, pronunciation assets, model or voice version, edits, and approvals to stable IDs. A changed transcript can invalidate downstream subtitles, synthesis, or evaluation even when the audio file is unchanged.

#### What should be monitored after a speech system launches?

Monitor task failure, recognition and repair patterns, latency, language switching, transfers, complaints, severe examples, subgroup coverage where lawful, cost, and drift. Sample content under privacy controls because aggregate telemetry cannot reveal every meaning failure.

#### How should a voice-model update be released?

Re-evaluate representative languages, speakers, acoustic conditions, pronunciations, safety flows, latency, and accessibility; compare failures to the current version; stage traffic; and keep rollback. Document provider changes that cannot be pinned.

#### What belongs in a speech delivery package?

Include final audio or transcripts, specifications, locale and speaker metadata, timecodes, pronunciation assets, known limitations, QA results, releases or consent records as appropriate, technical formats, and a manifest linking every item to its version.

## Methodology

Admas wrote these answers from delivery practice and checked definitions, standards, protocols, accessibility requirements, and professional guidance against the primary references listed on the page. Terms vary across companies and regions, so contracts and project specifications should define any term whose interpretation changes scope, quality, price, or acceptance.

## Source register

- [Web Content Accessibility Guidelines 2.2](https://www.w3.org/TR/WCAG22/) — World Wide Web Consortium; verified 2026-08-20.
- [Multidimensional Quality Metrics terminology](https://www.themqm.org/files/Terminology_Full_MQM_Formatted-for-the-website_02_26_Updated-Version.pdf) — MQM Council; verified 2026-08-20.
- [Artificial Intelligence Risk Management Framework: Generative AI Profile](https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence) — U.S. National Institute of Standards and Technology; verified 2026-08-20.
