On a platform running millions of searches a day, the result at the top of the page is the whole game. Theirs was often the wrong one, and the AI layer meant to help sometimes described a pool, a view or a price that simply did not exist. So travellers left.
Their search knew how to match exact words, but not what a person actually meant. Ask for "a quiet hotel for families with a heated pool" and the ranking fell apart, because it read keywords and missed the intent behind them. The right property was in the catalogue; it just never reached the top of the page.
To patch it, the company had bolted on a generative answer layer. It made things worse in a new way: it confidently stated facilities, distances and prices that were not real. A traveller who booked on the strength of an invented detail did not stay a happy customer for long.
Underneath both problems sat the same gap. The models had been trained on automatic signals with no human check on quality, and there was no verified, labelled reference set to retrain against. Without that anchor, accuracy quietly slipped over time and no one could measure by how much.
A verified reference set, models tuned for intent, and fact-checks that never sleep.
We built a review process where classification specialists worked through hundreds of thousands of real search queries, scored the quality of the results, and fed precise, structured feedback back into the system.
That labelled work became a verified reference set: the highest-quality data the organisation had ever held, and the anchor for training the models and measuring accuracy over time.
Using the verified data, we retrained the search and recommendation models to connect a traveller's intent to the best match in the catalogue, so a complex, natural sentence returns the right place, not the closest keyword.
We wired in agents that watch the generative answer layer in real time and cross-check what it produces against the company's factual database, catching an invented detail the moment it appears.
Anything the checks cannot trace to a trusted source is blocked before a user ever sees it. The system would rather show less than show something that is not true.
The same checks run continuously, so relevance holds and the models do not quietly drift back to the old behaviour once the launch is over.
Better matches and no invented details turned browsers into bookings.
Accuracy and relevance of the results on the first page rose roughly 35%, which lifted engagement straight away. More than 200,000 queries were labelled and rated along the way, becoming the organisation's strongest data asset for everything that comes next.
Reports of the site showing invented or wrong details fell 80%, and the conversion rate from a search to an actual booking climbed with it. People found what they meant, trusted what they saw, and completed the booking.
You cannot fine-tune your way to relevance on data no one has checked. The lift came from the verified reference set first, then the tuning, then the fact-checks that keep the generative layer honest in production. It is the same lesson behind our data-readiness work: the model was rarely the thing that was broken. The data was.
A medtech firm's clinical data was locked away by privacy law. An AI de-identification engine masked the patient details but kept the clinical signal, freeing terabytes safely and cutting the development cycle about 60%, at full HIPAA compliance.
Read the case → Data readiness · InsuranceA leading Israeli insurance group's AI kept hallucinating. We traced it to a ~30% mismatch in the data, rebuilt the pipeline, and put a hard 85% confidence floor under every answer.
Read the case →