The problem

The data that could train the model was the data they couldn't touch.

Their research teams needed raw clinical material, medical images, scans and free text, to build the next generation of diagnostic algorithms. The company owned all of it. It just could not legally use it, because every file held names, identifiers and confidential patient details protected by HIPAA.

Doing it by hand was not an option. Reviewing millions of scans and documents to black out identifying details would have taken thousands of hours and cost a fortune, and it still would not have kept up with research.

The naive automated tools made a different mistake: they redacted too aggressively, stripping out the clinical measurements and history that made the data worth training on in the first place. So the choice on the table was slow and safe, or fast and useless. Development fell months behind competitors.

What we did

Hide the patient, keep the medicine.

An engine that removes the identity and protects the clinical signal, with a human guaranteeing zero leaks.

Purpose-built

Language and vision together

We built a de-identification engine that combines language and computer-vision models, trained to recognise more than 50 kinds of identifying detail inside complex medical text and inside the scans themselves.

Context-aware

Mask the identity, not the finding

The engine removes or masks only the direct identifiers, names, dates, addresses and faces, while leaving the clinical context, the measurements and the findings fully intact and usable for research.

Keeps the value

No more over-redaction

Because it understands what it is reading, it stops destroying the very data that makes a file worth training on, the failure mode of the blunt tools they had tried before.

Human-in-the-loop

Zero leaks, verified

A review layer checks the output. Files cleared at full confidence are released to research immediately; anything uncertain goes to a specialist for a fast check, so the guarantee is zero exposure, not almost zero.

At scale

Terabytes, not samples

The pipeline runs across terabytes of images and documents, turning a locked archive into a working dataset rather than a hand-picked sample.

Compliant by design

HIPAA and GDPR, provably

Privacy protection is built into the flow, with the audit trail to prove compliance with the strictest healthcare and privacy standards.

The result

A locked archive became the training set.

The data got safe, and research got fast.

Live

Development cycle down about 60%

Freeing terabytes of previously locked clinical data gave the models the material they had been starved of, and cut the development cycle for new AI diagnostic products by roughly 60%, so the company reached the market faster and spent far less preparing data.

And compliant

Zero patient-data leaks

The work met the strictest regulatory standards in full, HIPAA and GDPR, with no incident of patient data being exposed or leaked. The data that had been a legal liability became the company's primary training asset.

Why it holds

Anonymisation is only useful if the data still is.

Anyone can redact a file into uselessness. The hard part is removing the identity while keeping the clinical signal, and being able to prove the identity is truly gone. That takes models that understand what they are reading and a human who stands behind the guarantee. It is the same principle as our data-readiness work: the value is in making data usable, safely, not just moving it around.

More case studies

Related work.

Sitting on data you can't legally use?

Book a strategy call Bring the archive legal won't let you open. Thirty minutes, no slides, or see more case studies.