# Speech, annotation & evaluation tools compared

> Compare 12 labeling and research tools by data boundary, modality, annotation schema, model assistance, quality workflow, integration, and export.

Published: 2026-08-19. Updated: 2026-08-20. Last verified: 2026-08-20. Review interval: 90 days. Disclosure: none.

## Use this to eliminate, not crown

Start with a real data, research, and model-evaluation teams building multilingual datasets workflow. Eliminate candidates that fail a deployment, format, locale, accessibility, or exit requirement; then test the survivors with your own content.

- Write three non-negotiable gates before opening the table.
- Choose the view that matches the decision in front of you.
- Open the linked documentation for any field that changes the shortlist.
- Run one representative import, production, review, and export cycle.

## Proof-of-work protocol



```
1. Import a representative source with placeholders, markup, plurals, and right-to-left text.
2. Route it through translation, review, QA, and correction.
3. Exercise automation or integration under realistic permissions.
4. Export every asset and target file needed to leave.
5. Record defects, manual steps, plan dependencies, and owner responses.
```

## Comparison data

[Download the versioned JSON dataset](https://admas.net/resources/data/tools.json).

## Frequently asked questions

### Which annotation tool is best for multilingual model evaluation?

Choose by annotation unit, modality, ontology, reviewer qualifications, adjudication, data boundary, automation, and export. A flexible UI is insufficient if it cannot preserve locale metadata or an auditable decision record.

### What locale metadata belongs with every annotation?

Keep the content language tag, script when relevant, asserted region or variety, modality, annotator qualification, guideline version, suggestion provenance, timestamps, and adjudication state.

### How do we avoid automation bias in model-assisted labeling?

Blind a controlled sample to predictions, measure acceptance and correction by language and label, rotate review order, preserve suggestion provenance, and compare against an independently annotated reference set.

## Methodology

Admas reviewed current first-party product documentation against criteria specific to this tool category. A missing claim is recorded as not established, not scored as a failure. Entries describe documented product capabilities, not plan entitlement, implementation quality, security posture, or suitability for a particular organization. Re-run the proposed proof-of-work and export test before procurement.

## Source register

- [Label Studio documentation](https://labelstud.io/guide/) — HumanSignal; verified 2026-08-20.
- [Argilla documentation](https://docs.argilla.io/) — Argilla; verified 2026-08-20.
- [Prodigy documentation](https://prodi.gy/docs) — Explosion; verified 2026-08-20.
- [ELAN documentation](https://archive.mpi.nl/tla/elan/documentation) — Max Planck Institute for Psycholinguistics; verified 2026-08-20.
- [CVAT documentation](https://docs.cvat.ai/docs/) — CVAT.ai; verified 2026-08-20.
- [doccano documentation](https://doccano.github.io/doccano/) — doccano; verified 2026-08-20.
- [INCEpTION user guide](https://inception-project.github.io/releases/latest/docs/user-guide.html) — INCEpTION; verified 2026-08-20.
- [Praat manual](https://www.fon.hum.uva.nl/praat/manual/Intro.html) — Praat; verified 2026-08-20.
- [Labelbox documentation](https://docs.labelbox.com/docs/overview) — Labelbox; verified 2026-08-20.
- [SuperAnnotate documentation](https://doc.superannotate.com/) — SuperAnnotate; verified 2026-08-20.
- [Encord Annotate documentation](https://docs.encord.com/platform-documentation/Annotate/annotate-projects/annotate-manage-annotation-projects) — Encord; verified 2026-08-20.
- [Dataloop annotation documentation](https://docs.dataloop.ai/docs/annotations-help) — Dataloop; verified 2026-08-20.
