# Speech & voice AI services

> Speech data, TTS, and STT systems built for natural output and dependable recognition across languages.

Localize voice AI for the language, speaker, culture, and real acoustic conditions of use.

Speech quality depends on more than audio volume: phonology, speaker design, prompts, recording, metadata, evaluation, and product context all matter.

## Services

- [Speech data programs](https://admas.net/capabilities/speech/speech-data/index.md): Speaker, prompt, recording, transcription, and validation systems for multilingual speech development.
- [Text-to-speech development](https://admas.net/capabilities/speech/text-to-speech/index.md): Language design, pronunciation resources, voice evaluation, and product QA for natural synthesized speech.
- [Speech recognition evaluation](https://admas.net/capabilities/speech/speech-recognition/index.md): Language-aware ASR testing, error analysis, and improvement plans across speakers, accents, domains, and environments.
- [Conversational voice-agent localization](https://admas.net/capabilities/speech/conversational-voice-agent-localization/index.md): End-to-end localization of spoken agent experiences, from recognition and dialogue policy to synthesized voice and repair.
- [Speech quality evaluation](https://admas.net/capabilities/speech/speech-quality-evaluation/index.md): Perceptual and task-based evaluation for recognition, synthesis, translation, enhancement, and spoken dialogue systems.

## Signals that this work matters

- A voice is intelligible but does not sound natural or local
- Recognition performance collapses across accents and environments
- Low-resource languages lack reliable speech data
- Teams cannot explain why speech metrics and human judgment disagree

## Signals field brief

- [Production Voice AI Across Languages](https://admas.net/signals/production-voice-ai-across-languages/index.md): Voice quality is not a transcript score plus a natural-sounding demo. Production systems must listen, interpret, speak, take turns, recover, and hand off across real languages, speakers, devices, and environments.

## Outcomes

- **Representative speech data:** Speakers, prompts, recordings, and metadata match the intended use.
- **Natural voice output:** Pronunciation, prosody, pacing, and persona are evaluated in language context.
- **Robust recognition:** Errors are measured by user impact, language pattern, and acoustic condition.

## Frequently asked questions

### What does speech & voice ai cover?

Speech data, TTS, and STT systems built for natural output and dependable recognition across languages. Admas treats it as a connected practice spanning Speech data programs, Text-to-speech development, Speech recognition evaluation, Conversational voice-agent localization, and Speech quality evaluation. A project can start with one focused service and expand only where the evidence shows a dependency.

### Who is speech & voice ai for?

This work is usually shared by speech, product, data, accessibility, conversation-design, and market teams building or localizing voice experiences. The exact team depends on who owns the affected user journey, data, system, content, market decision, and release risk.

### When should a team start speech & voice ai work?

Start before a launch is locked when possible. Common signals include a voice is intelligible but does not sound natural or local, recognition performance collapses across accents and environments, low-resource languages lack reliable speech data, and teams cannot explain why speech metrics and human judgment disagree. A focused diagnostic can still help when the work has already become a recovery project.

### What inputs does a speech & voice ai engagement need?

Useful starting inputs are target languages, dialects, and speaking contexts, audio or model access, speaker and consent requirements, acoustic conditions and devices, product tasks, scripts, prompts, and and quality thresholds. They do not need to be complete: unknowns should be recorded as assumptions, risks, or discovery questions rather than silently filled in.

### How does speech & voice ai connect to other localization and internationalization work?

The practice rarely stands alone. Product architecture affects localization; data affects model behavior; language quality affects release decisions; and program design affects whether improvements persist. Admas maps those handoffs explicitly so each specialist can work from the same acceptance criteria.

### What should AI automate in speech & voice ai, and what should people own?

Speech models can draft transcripts, synthesize candidates, segment audio, and surface likely errors. Native listeners, phoneticians, voice specialists, conversation designers, and engineers are still needed to judge pronunciation, prosody, intelligibility, demographic coverage, and real interaction failures.

### How is quality measured in speech & voice ai?

Use evidence tied to the intended decision, not one universal score. Typical measures include intelligibility and naturalness, task and recognition accuracy by cohort, pronunciation and prosody, speaker and acoustic coverage, latency, and accessibility and failure recovery. Results should be segmented by language, market, content or task type, and risk so an average cannot hide a serious local failure.

### How is speech & voice ai priced?

Pricing depends on recording or evaluation hours, languages, dialects, and speaker profiles, studio and equipment needs, transcription and annotation depth, model or integration work, and quality and consent controls. A defensible estimate separates repeatable production units from discovery, engineering, review, management, pass-through costs, and contingency. Admas scopes the acceptance criteria and review path before treating a volume number as a quote.

## Start a project

- [Build a speech & voice ai project brief](https://admas.net/start-a-project/index.md?focus=speech): Share the product or model, modalities, languages, timing, and current constraint.
