To train new diagnostic models, the engineers needed terabytes of real clinical data. It was sitting right there, and legal had locked it: the files were full of patient identities, and redacting them by hand was never going to happen at that scale. New products stalled for months.
Their research teams needed raw clinical material, medical images, scans and free text, to build the next generation of diagnostic algorithms. The company owned all of it. It just could not legally use it, because every file held names, identifiers and confidential patient details protected by HIPAA.
Doing it by hand was not an option. Reviewing millions of scans and documents to black out identifying details would have taken thousands of hours and cost a fortune, and it still would not have kept up with research.
The naive automated tools made a different mistake: they redacted too aggressively, stripping out the clinical measurements and history that made the data worth training on in the first place. So the choice on the table was slow and safe, or fast and useless. Development fell months behind competitors.
An engine that removes the identity and protects the clinical signal, with a human guaranteeing zero leaks.
We built a de-identification engine that combines language and computer-vision models, trained to recognise more than 50 kinds of identifying detail inside complex medical text and inside the scans themselves.
The engine removes or masks only the direct identifiers, names, dates, addresses and faces, while leaving the clinical context, the measurements and the findings fully intact and usable for research.
Because it understands what it is reading, it stops destroying the very data that makes a file worth training on, the failure mode of the blunt tools they had tried before.
A review layer checks the output. Files cleared at full confidence are released to research immediately; anything uncertain goes to a specialist for a fast check, so the guarantee is zero exposure, not almost zero.
The pipeline runs across terabytes of images and documents, turning a locked archive into a working dataset rather than a hand-picked sample.
Privacy protection is built into the flow, with the audit trail to prove compliance with the strictest healthcare and privacy standards.
The data got safe, and research got fast.
Freeing terabytes of previously locked clinical data gave the models the material they had been starved of, and cut the development cycle for new AI diagnostic products by roughly 60%, so the company reached the market faster and spent far less preparing data.
The work met the strictest regulatory standards in full, HIPAA and GDPR, with no incident of patient data being exposed or leaked. The data that had been a legal liability became the company's primary training asset.
Anyone can redact a file into uselessness. The hard part is removing the identity while keeping the clinical signal, and being able to prove the identity is truly gone. That takes models that understand what they are reading and a human who stands behind the guarantee. It is the same principle as our data-readiness work: the value is in making data usable, safely, not just moving it around.
A booking platform's search buried the best matches and its AI invented facilities that were never there. Human-verified data, relevance tuning and fact-check agents lifted top-result relevance 35% and cut invented-info reports 80%.
Read the case → Data readiness · InsuranceA leading Israeli insurance group's AI kept hallucinating. We traced it to a ~30% mismatch in the data, rebuilt the pipeline, and put a hard 85% confidence floor under every answer.
Read the case →