Text & NLP Annotation

Annotated language data built on guidelines precise enough that two independent annotators reach the same answer — and measured to prove they do.

Text annotation is a guideline problem first

Image annotation disagreements are usually about a few pixels. Text annotation disagreements are usually about meaning, and they are far harder to detect after the fact. If two annotators disagree about whether a sentence is negative or merely critical, nothing in the output looks wrong — the dataset simply contains a class boundary the model cannot learn because it is not consistently drawn.

That makes the annotation guideline the primary deliverable of the setup phase, not paperwork attached to it. We build guidelines against real samples from your corpus, test them by having multiple annotators independently label the same batch, and revise wherever agreement falls short before scaling up.

Tasks we support

Our text and NLP annotation covers:

  • Named entity recognition — including domain-specific and nested entity schemes.
  • Text classification — single-label, multi-label, and hierarchical taxonomies.
  • Sentiment and stance annotation — with explicitly defined boundaries between adjacent categories.
  • Intent classification — for conversational and search applications.
  • Relation and event extraction — linking entities into structured relationships.
  • Coreference resolution — resolving references across a document.
  • Span and rationale marking — identifying the passage that supports a label, not only the label.
  • Instruction and preference data — prompt and response construction, and ranked comparisons between candidate outputs.

Data for language model training

Training and aligning language models needs a different kind of annotation from classical NLP. Preference data — where an annotator ranks candidate responses against a rubric — depends almost entirely on how well that rubric distinguishes genuinely better answers from merely longer or more confident ones.

We build these rubrics explicitly, covering factual accuracy, instruction adherence, tone, and refusal behaviour as separate axes rather than a single quality score. Where annotators are asked to write rather than to rank, we track provenance so you know which portion of a dataset is human-written and which is human-edited.

Measuring agreement

For subjective tasks, a raw accuracy figure against a reference set is misleading, because the reference set embodies one annotator's judgement. We measure inter-annotator agreement instead, using chance-corrected statistics — Cohen's kappa for two annotators, Fleiss' kappa for more — so that agreement which would have occurred at random is discounted.

Agreement is measured per class, not only overall. A dataset with strong headline agreement routinely contains one or two categories where annotators are effectively guessing, and those are precisely the categories a model will fail on in production.

Every delivery passes through a layered review rather than a single pass. Annotators work to a written specification agreed before production starts; a second reviewer checks a defined proportion of each batch; and a final quality gate samples the batch independently against the same specification. Disagreements are not silently overwritten — they are resolved against the guideline, and where the guideline turns out to be ambiguous, the guideline is amended and the affected items are reworked.

That last step matters more than it sounds. Most systematic annotation error is not carelessness; it is a specification that failed to anticipate a real case. Treating every disagreement as a possible guideline defect is what keeps error rates from plateauing partway through a project.

Multilingual work

Multilingual annotation is not translation of a guideline. Category boundaries that are natural in one language frequently do not survive into another, and a schema that works in English can force annotators in another language to choose between two equally wrong labels.

Where a project spans languages, we validate the schema separately in each one and raise cases where a category does not transfer cleanly, rather than absorbing the mismatch silently into the labels.

Data handling

Client data is encrypted in transit and at rest, access is limited to the specialists assigned to your project, and every member of the annotation workforce works under a signed confidentiality agreement. Where a dataset carries additional handling constraints, those are agreed in writing before any data is transferred.

Common questions

How do you measure quality on subjective tasks like sentiment?

With chance-corrected inter-annotator agreement — Cohen's kappa for two annotators, Fleiss' kappa for more — reported per class rather than only in aggregate. Accuracy against a single reference set is misleading for subjective work, because the reference itself reflects one annotator's judgement.

Can you produce preference data for language model alignment?

Yes. We build ranked comparisons against an explicit rubric that separates factual accuracy, instruction adherence, tone, and refusal behaviour, rather than collapsing them into a single quality score. Provenance is tracked so you know which portion is human-written and which is human-edited.

Do you work in languages other than English?

Yes. For multilingual projects we validate the label schema separately in each language rather than translating one guideline, because category boundaries often do not transfer cleanly and forcing them produces silently inconsistent labels.

What if our annotation guideline turns out to be ambiguous?

We raise it rather than resolve it silently. Ambiguity discovered in production is treated as a guideline defect: the guideline is amended with your team and affected items are reworked, which is what stops error rates from plateauing mid-project.

Discuss your text & nlp annotation project

Tell us what you are building and what your data has to support. We will come back with a specification, a pilot scope, and a realistic timeline.