The problem

The science was ready. The law said no.

A model this powerful needs the widest possible range of patients, but the obvious approach, pool everyone's scans and genomes into one global database, was flatly illegal. Data-sovereignty rules meant patient data could not cross the border it was collected in, and trying to build a central data lake was blocked before it began.

Even setting the law aside, the logistics were impossible. Moving enormous medical imaging files, MRI and CT scans, between continents ran straight into bandwidth limits that made a central collection unworkable.

And the partners were not set up for it. There was no uniform, managed infrastructure inside the hospitals and labs to run complex algorithms locally, so the compute could not simply be pushed out to where the data lived either.

What we did

Bring the model to the data, not the data to the model.

Train inside each country, and move only what the model learned, never a patient record.

Edge nodes

Compute inside each jurisdiction

We stood up a network of distributed server nodes placed locally in each country, sitting inside the boundary of its own regulator, so the processing happens where the data legally has to stay.

Federated learning

The model trains in place

A federated-learning architecture trains the shared model separately inside each country, on that country's local patient data, so no one ever has to gather the data into one place to learn from it.

Weights-only fabric

Only the maths travels

A private data fabric moves just the model's mathematical weights between countries, never real clinical data, so the global model updates and improves while every patient record stays home.

Privacy by design

Compliance built into the architecture

Because no sensitive record ever crosses a border, GDPR, HIPAA and every local data-sovereignty rule hold by construction, not by a policy bolted on afterwards.

Bandwidth solved

Megabytes, not terabytes

The heavy imaging never moves; only the algorithm's learnings, a few megabytes, travel between sites, which turns an impossible cross-continental transfer into a routine one.

Diversity without pooling

One model, twenty countries

The unified global model still learns from the full genetic diversity of more than 20 countries, capturing the breadth that makes rare-disease detection work, without ever centralising the data.

The result

The breakthrough, without the breach.

Compliant by design, and more accurate for it.

Live

Compliant, and unprecedented accuracy

The consortium met every data-sovereignty and privacy law, cleared the international project with the regulators, and started within months. The unified model reached unprecedented accuracy on rare diseases precisely because it learned from 20 countries' genetic diversity.

And leaner

95% less cross-border traffic, zero leaks

Because the heavy scans stayed put and only the model's insights travelled, international data traffic fell over 95%, and not one piece of sensitive patient information ever left its home country.

Why it holds

Move the learning, not the data.

Pooling sensitive data to train a model is the assumption that breaks on privacy law and on bandwidth. The value here is inverting it: keep every record where the law requires, train locally, and share only the mathematics the model learned. Compliance becomes a property of the architecture rather than a promise. It is the same privacy-first discipline behind our medical-anonymisation work, aimed at training across borders rather than de-identifying a single dataset.

More case studies

Related work.

Data you can't move, but need to learn from?

Book a strategy call Bring the data locked behind privacy law or borders. Thirty minutes, no slides, or see more case studies.