Medical Image Annotation

Annotated medical imaging for clinical AI development — where the label is a clinical judgement, and the review process has to reflect that.

Medical annotation is a clinical judgement, not a drawing task

On most annotation projects, a well-written guideline lets a trained non-specialist produce correct labels. Medical imaging is different: the boundary of a lesion, the distinction between two tissue types, or whether a finding is present at all are clinical determinations, and inter-observer variation exists among qualified clinicians themselves.

That has a direct consequence for how quality is defined. A single expert's annotation is not ground truth — it is one expert's reading. Where a project needs a defensible reference standard, that requires multiple independent readings and a documented adjudication process, and we scope it that way rather than presenting a single-reader dataset as definitive.

Modalities and tasks

Our medical annotation work covers:

  • CT and MRI — volumetric segmentation, slice-level annotation, and multi-planar review.
  • X-ray — finding detection, localisation, and classification.
  • PET — uptake region delineation, including fused PET/CT review.
  • Ultrasound — structure identification and measurement point placement.
  • Digital pathology — whole-slide region annotation, cell-level labelling, and tumour segmentation.
  • Organ and anatomical structure segmentation — for planning, measurement, and normative modelling.
  • Lesion detection and tumour delineation — with explicit rules for boundary and inclusion.
  • Anatomical landmark placement — for registration, alignment, and measurement.

Privacy and de-identification

Medical imaging carries patient identity in more places than most teams expect. DICOM headers contain extensive identifying metadata, and burned-in text on the image itself survives header stripping entirely — as do identifiable facial features in head CT and MRI volumes, which can be reconstructed into a recognisable face.

De-identification is therefore treated as a distinct, verified step rather than an assumption. We confirm with your team what has already been removed before transfer, and what remains our responsibility, rather than proceeding on the assumption that incoming data is clean.

Client data is encrypted in transit and at rest, access is limited to the specialists assigned to your project, and every member of the annotation workforce works under a signed confidentiality agreement. Where a dataset carries additional handling constraints, those are agreed in writing before any data is transferred.

Quality assurance

Every delivery passes through a layered review rather than a single pass. Annotators work to a written specification agreed before production starts; a second reviewer checks a defined proportion of each batch; and a final quality gate samples the batch independently against the same specification. Disagreements are not silently overwritten — they are resolved against the guideline, and where the guideline turns out to be ambiguous, the guideline is amended and the affected items are reworked.

That last step matters more than it sounds. Most systematic annotation error is not carelessness; it is a specification that failed to anticipate a real case. Treating every disagreement as a possible guideline defect is what keeps error rates from plateauing partway through a project.

For segmentation we report Dice similarity against the agreed reference, and where a project uses multiple independent readers we report the spread between them as well. That spread is informative in itself: a task where qualified readers diverge substantially is one where a model trained to a single reader will not generalise, and knowing that before training is considerably cheaper than discovering it after.

What this data is not

Annotated datasets support the development of clinical AI; they are not themselves a medical device, a diagnosis, or clinical advice. Regulatory validation of any system trained on this data remains the responsibility of the developer, and we scope annotation work as an input to that process rather than a substitute for it.

Common questions

Are your medical annotations reviewed by clinicians?

Medical annotation tasks are scoped to require the appropriate clinical expertise for the determination being made, and the review process is agreed during project setup. Where a defensible reference standard is required, that means multiple independent readings with a documented adjudication process rather than a single reader.

How is patient data protected?

De-identification is treated as a verified step rather than an assumption. We confirm what has been removed before transfer — including DICOM header metadata, burned-in image text, and identifiable facial structure in head volumes — and what remains in scope for us. Data is encrypted in transit and at rest with access limited to the assigned project team.

What accuracy measure do you use for segmentation?

Dice similarity against the agreed reference. Where multiple independent readers are used, we also report the spread between them, because a task with high inter-reader variation tells you something important about how well a model trained on it will generalise.

Can annotated data be used to certify a medical device?

No. Annotated datasets are an input to clinical AI development, not a substitute for regulatory validation. Certification of any system trained on this data remains the responsibility of the developer.

Discuss your medical annotation project

Tell us what you are building and what your data has to support. We will come back with a specification, a pilot scope, and a realistic timeline.