# Localization quality, review & testing FAQs

> Answers about QA, QC, LQA, functional testing, MQM, severities, review calibration, sampling, automation, release gates, and defect learning.

Published: 2026-08-20. Updated: 2026-08-20. Last verified: 2026-08-20. Review interval: 183 days. Disclosure: none.

## How to use these answers

Start with the question that matches your decision about language and locale quality. Each answer defines the term, then names the operational consequence that a brief, workflow, test, or contract should make explicit.

## The rule behind the page

Quality is fitness for a defined purpose, not the absence of reviewer comments. Connect each check to a requirement, each severity to user impact, and each metric to a decision.

- Define the audience and purpose
- Name the languages, locales, scripts, and modalities
- Separate requirements from preferences
- Assign decision and escalation authority
- Keep source, output, and evidence versioned

## Frequently asked questions

### Definitions

#### What is the difference between QA, QC, and evaluation?

QA focuses on whether the process can meet requirements; QC monitors or checks work against them; evaluation produces evidence about output quality for a decision. Localization teams often use the labels interchangeably, so define the activity and output.

#### What is linguistic QA?

A structured assessment of target-language content against requirements such as accuracy, completeness, terminology, language quality, locale conventions, style, and formatting. It may score issues or simply drive correction and release decisions.

#### What is functional localization testing?

Testing the localized product's behavior: layout, input, links, variables, fonts, directionality, formats, search, navigation, and locale-specific workflows. A linguist and tester may work together because language and function interact.

#### What is the difference between quality assurance and quality control?

Quality assurance designs processes that make requirements achievable; quality control checks work and process outputs against those requirements. Both matter. A final spellcheck cannot compensate for missing context, unsuitable suppliers, or broken source files.

#### What is localization testing?

Localization testing evaluates a localized product or asset in context, including language, layout, fonts, directionality, formats, input, functions, links, media, accessibility, and locale behavior. The exact test scope should be named rather than hidden under QA.

#### What does fit for purpose mean in localization?

It means the output meets defined audience, task, content, channel, risk, and lifecycle requirements. It does not mean good enough without criteria. Two outputs can require different levels of polish while both being fit for their declared uses.

### Planning

#### How should quality requirements be defined?

Name the audience, purpose, content and locale scope, critical terms, risk classes, quality dimensions, severities, sampling, acceptance thresholds, environments, evidence, and decision owner before production starts.

#### Should every string receive the same review?

No. Route by risk, visibility, novelty, content lifespan, automation confidence, supplier evidence, and reversibility. Critical legal, safety, payment, or high-traffic content may justify deeper review than transient internal text.

#### What is a locale test matrix?

A risk-based selection of languages, scripts, platforms, content, and user journeys used to cover meaningful internationalization and localization failure modes without testing every combination equally.

#### How should localization quality requirements be written?

Translate goals into observable criteria by content class and locale: meaning, terminology, tone, functionality, accessibility, formats, severe-error tolerance, checks, sample, evidence, and approver. Avoid asking for perfect, native, or high quality without operational meaning.

#### How should risk determine QA depth?

Increase independent review, context, sampling, specialist involvement, functional testing, and release evidence as potential harm, reach, irreversibility, novelty, and uncertainty rise. Low-risk content can use lighter gates with monitoring and easy correction.

#### What is a localization test plan?

It maps scope, platforms, builds, locales, environments, devices, journeys, content, test techniques, data, roles, entry and exit criteria, issue handling, schedule, and evidence. It also declares exclusions so untested areas are not mistaken for approved ones.

### Issue handling

#### What makes a useful localization bug report?

Build and locale, exact location, source and target, screenshot or media timecode, steps, expected and actual behavior, issue category, user impact, severity rationale, reproducibility, and related identifiers.

#### How should severity be assigned?

Based on consequence: meaning or task failure, safety or legal impact, reach, recoverability, visibility, and release risk. Grammar-category labels do not determine severity by themselves.

#### How should preferences be separated from defects?

A defect violates an agreed requirement or clearly harms meaning or use. A preference is a valid alternative without that evidence. Record preferences in style or terminology only after an authorized decision, rather than retroactively scoring them as errors.

#### What information should a localization bug contain?

Include build, locale, platform, exact location, steps, observed and expected behavior, source and target, screenshot or recording, impact, reproducibility, and relevant IDs. Protect personal or confidential content in evidence.

#### How should localization defect severity be assigned?

Use impact on meaning, task completion, safety, legality, accessibility, reach, and recoverability. Severity should not depend on reviewer preference or the effort needed to fix the issue, though priority can consider schedule and exposure.

#### How should linguistic disagreements be resolved?

Check the brief, source meaning, terminology, style, target-market convention, and evidence. Give a named language owner final authority for defined matters, record rationale for reusable decisions, and distinguish an acceptable alternative from a defect.

### Review

#### Why do bilingual and in-context reviews find different problems?

Bilingual review is strong for source-target meaning and consistency. In-context review exposes screen, flow, media, layout, variable, and user-intent problems. Mature QA uses each where its evidence is strongest.

#### How should reviewers be calibrated?

Give them the same brief and examples, run a representative sample, compare categories and severities, resolve disagreements, adjust guidelines, and monitor drift. Calibration should focus on decisions, not forcing identical wording.

#### What is the difference between bilingual review and monolingual review?

Bilingual review checks the target against the source for accuracy and completeness. Monolingual review assesses the target as user-facing content for clarity, naturalness, consistency, and effect. They answer different questions and may both be needed.

#### When should review be independent?

Use an independent qualified reviewer when risk, contract, regulation, evaluation integrity, or learning requires a second judgment. Independence should include enough separation and authority to identify a defect without pressure to confirm the first pass.

#### How should reviewer changes be audited?

Retain the submitted target, proposed change, category, severity, reason, reviewer, disposition, and final version. Analyze accepted and rejected changes to improve guidance and consistency without turning raw edit counts into a performance score.

#### How can client or in-country review avoid endless preference changes?

Define the reviewer's purpose and authority, require contextual reasons for changes, use terminology and style ownership, separate defects from preferences and new source, set a review window, and adjudicate recurring disputes.

### Metrics

#### What is MQM used for?

MQM helps teams define issue categories and severities for translation-quality evaluation. Tailor it to the content and decision; an enormous taxonomy can create inconsistent tagging and distract from user impact.

#### Is an error score a complete measure of quality?

No. Scores depend on taxonomy, weights, sample, annotators, normalization, and thresholds. Pair them with severe examples, task success, user impact, confidence, and segmented analysis.

#### How large should a quality sample be?

Base it on the decision, expected defect rate, language and content variability, risk, and confidence needed. A small random sample may miss rare severe errors; targeted challenge sampling and full checks of critical content can complement it.

#### Which localization quality metrics are useful?

Use a small decision-linked set such as severe defects, acceptance or rejection, correction effort, escaped defects, task success, turnaround, recurrence, and coverage. Segment by content, language, route, and risk so averages do not hide failures.

#### Why can an average quality score be misleading?

It can combine incomparable languages and content, dilute rare critical errors, reward over-reporting, and hide sample uncertainty. Show distributions, severity, slice coverage, and examples alongside any aggregate.

#### How should post-editing effort be measured?

Measure active correction time under a clear protocol, rejection and retranslation, error type, accepted output, and total workflow overhead. Edit distance alone cannot show research, cognitive effort, missed errors, or whether the final result met requirements.

### Automation

#### What can automated localization QA reliably check?

File validity, missing or extra tags and variables, numbers, forbidden terms, empty targets, duplicates, length rules, encoding patterns, and some consistency issues. Every check needs false-positive handling and an owner.

#### Can LLMs evaluate translation quality?

They can support triage and pattern discovery, but results can vary by prompt, model, language, reference, and error type. Validate them against qualified human judgments and never hide severe-language failures behind an average model score.

#### Which localization QA checks are safe to automate?

Deterministic checks for missing content, tags, variables, numbers, prohibited terms, encoding, length rules, file structure, locale metadata, and known patterns are strong candidates. Every check needs scope, false-positive handling, and ownership.

#### What can automated QA not reliably decide alone?

It cannot generally determine full meaning, audience effect, legal sufficiency, cultural suitability, natural voice, accessibility experience, or whether an exception is justified. Use it to surface evidence, not to manufacture confidence.

#### How should false positives in automated QA be managed?

Record justified exceptions with scope and expiry, refine rules using recurring examples, rank findings by impact, and measure reviewer burden. If people must ignore most alerts, real failures will also be ignored.

#### How should an AI quality estimator be validated?

Compare it with qualified human decisions on representative languages, content, systems, severities, and difficult cases. Test calibration, subgroup error, drift, manipulation, and whether using the score improves the intended routing or release decision.

### Release

#### What belongs in a localization release gate?

Scope and version, completed checks, severe open issues, locale coverage, regression results, accessibility and functional status, accepted deviations, rollback readiness, approver, and links to evidence.

#### What evidence should a localization release gate require?

Require source and target versions, locale scope, completed checks, severe-issue status, functional and accessibility evidence as applicable, approvals, known exceptions, rollback, and a named release decision. Evidence should be retrievable after launch.

#### Who can accept localization risk?

A named person with authority over the affected product or market should accept residual risk using visible evidence. Translators and QA staff can explain risk, but they should not be made to approve business, legal, safety, or release exposure outside their role.

#### When should a localization release be blocked?

Block when agreed critical criteria fail: serious meaning, safety, legal, payment, privacy, accessibility, data, or core-journey defects; missing required locales or evidence; or no safe fallback. Define this before deadline pressure arrives.

#### How should known localization defects be documented at launch?

List issue, locale, reach, impact, workaround, owner, correction date, monitoring, and accepting authority. Do not close a known defect merely because it was accepted for one release.

#### What should happen after a localized hotfix?

Verify the correction in context, check adjacent and reusable content, update translation memory or terminology where appropriate, rebuild the correct version, rerun relevant regression checks, and close the issue with release evidence.

### Improvement

#### How should escaped localization defects be used?

Confirm the defect, classify user impact and root cause, fix and verify it, then update the source, terminology, guidance, tool check, test, supplier feedback, or release control that could prevent recurrence.

#### How should localization teams learn from escaped defects?

Contain and correct the issue, then trace source, instructions, assets, workflow, automation, review, and release decisions. Improve the earliest effective control and add a regression case without reducing the analysis to individual blame.

#### How often should quality criteria be reviewed?

Review them when content, markets, users, regulation, workflow, suppliers, or models change and at a defined periodic cadence. Use defect and reviewer evidence to remove unused rules and clarify ambiguous ones.

#### How can review feedback improve translation memory and terminology?

Validate the feedback, apply it to the approved target, update the authoritative asset with context and ownership, find affected content, and communicate the change. Do not bulk-promote every reviewer edit into reusable memory.

#### What should a localization quality retrospective cover?

Review outcomes, severe and recurring defects, estimates, rework, reviewer consistency, automation burden, supplier fit, source problems, incidents, and user signals. Assign a small number of owned changes with verification dates.

#### How can quality improve without adding review to everything?

Improve source content, context, terminology, components, routing, supplier fit, deterministic checks, reviewer calibration, and feedback reuse. Concentrate expert review on novel, uncertain, or consequential work and monitor lighter paths.

## Methodology

Admas wrote these answers from delivery practice and checked definitions, standards, protocols, accessibility requirements, and professional guidance against the primary references listed on the page. Terms vary across companies and regions, so contracts and project specifications should define any term whose interpretation changes scope, quality, price, or acceptance.

## Source register

- [Internationalization Glossary](https://www.w3.org/TR/i18n-glossary/) — World Wide Web Consortium; verified 2026-08-20.
- [Glossary of Unicode Terms](https://www.unicode.org/glossary/) — Unicode Consortium; verified 2026-08-20.
- [Web Content Accessibility Guidelines 2.2](https://www.w3.org/TR/WCAG22/) — World Wide Web Consortium; verified 2026-08-20.
- [Multidimensional Quality Metrics terminology](https://www.themqm.org/files/Terminology_Full_MQM_Formatted-for-the-website_02_26_Updated-Version.pdf) — MQM Council; verified 2026-08-20.
