Audio & Speech Annotation

Transcribed and labelled audio for speech recognition and audio understanding — built to a transcription convention chosen for your model, not a generic one.

Verbatim or clean: the decision that shapes the dataset

The first question on any transcription project is what to do with the things people actually say. Real speech contains false starts, filler, repetition, and self-correction. Verbatim transcription keeps them; clean transcription removes them. Both are legitimate, and they produce meaningfully different models.

A recogniser trained on cleaned transcripts learns that disfluency does not exist, then encounters it constantly in production. One trained on verbatim transcripts handles real speech better but produces output that needs post-processing before a human reads it. We decide this deliberately with your team, apply one convention throughout, and record which was used — a dataset that mixes both is worse than either.

What we annotate in audio

Our audio and speech work covers:

  • Speech transcription — verbatim or clean-read, to a documented convention.
  • Timestamping and forced alignment review — at utterance or word level.
  • Speaker diarisation — who spoke when, including overlapping speech.
  • Speaker attribution — consistent identity labels across a recording or a series.
  • Acoustic event labelling — non-speech sounds, background conditions, and environment tagging.
  • Audio classification — clip or segment-level tagging.
  • Emotion and paralinguistic annotation — where the label scheme can be defined precisely enough to be reproducible.

Accent, dialect, and conversational speech

Speech models fail unevenly. A recogniser can post strong average accuracy while performing badly for particular accents, age groups, or speaking styles, because those groups were thinly represented in training — and an aggregate figure conceals it completely.

Where it matters to your application, we report transcription quality broken down by the speaker characteristics present in your data rather than as a single number. Overlapping speech and rapid turn-taking are handled by explicit convention, since they are both common in real recordings and a frequent source of inconsistency between transcribers.

Quality assurance

Every delivery passes through a layered review rather than a single pass. Annotators work to a written specification agreed before production starts; a second reviewer checks a defined proportion of each batch; and a final quality gate samples the batch independently against the same specification. Disagreements are not silently overwritten — they are resolved against the guideline, and where the guideline turns out to be ambiguous, the guideline is amended and the affected items are reworked.

That last step matters more than it sounds. Most systematic annotation error is not carelessness; it is a specification that failed to anticipate a real case. Treating every disagreement as a possible guideline defect is what keeps error rates from plateauing partway through a project.

For transcription we measure word error rate against a reviewed reference, separating substitution, insertion, and deletion errors — they have different causes and different fixes. A high deletion rate usually indicates audio quality or overlap problems rather than transcriber performance, and treating it as the latter wastes effort.

Data handling

Client data is encrypted in transit and at rest, access is limited to the specialists assigned to your project, and every member of the annotation workforce works under a signed confidentiality agreement. Where a dataset carries additional handling constraints, those are agreed in writing before any data is transferred.

Common questions

Should we choose verbatim or clean transcription?

It depends on the model. Verbatim keeps disfluencies and suits speech recognition that must handle real conversational audio; clean removes them and suits downstream text applications. The important thing is choosing one and applying it consistently — a dataset mixing both conventions is worse than either.

Can you handle overlapping speakers?

Yes. Overlapping speech is handled by an explicit convention agreed before production, because it is common in real recordings and one of the largest sources of inconsistency between transcribers when left undefined.

How do you report transcription accuracy?

As word error rate against a reviewed reference, split into substitution, insertion, and deletion errors, and broken down by speaker characteristics where that matters to your application. Aggregate accuracy hides uneven performance across accents and speaking styles.

Discuss your audio & speech project

Tell us what you are building and what your data has to support. We will come back with a specification, a pilot scope, and a realistic timeline.