The problem

The best match was there. No one saw it.

Their search knew how to match exact words, but not what a person actually meant. Ask for "a quiet hotel for families with a heated pool" and the ranking fell apart, because it read keywords and missed the intent behind them. The right property was in the catalogue; it just never reached the top of the page.

To patch it, the company had bolted on a generative answer layer. It made things worse in a new way: it confidently stated facilities, distances and prices that were not real. A traveller who booked on the strength of an invented detail did not stay a happy customer for long.

Underneath both problems sat the same gap. The models had been trained on automatic signals with no human check on quality, and there was no verified, labelled reference set to retrain against. Without that anchor, accuracy quietly slipped over time and no one could measure by how much.

What we did

Fix the data first, then the model.

A verified reference set, models tuned for intent, and fact-checks that never sleep.

Human-verified

People rated the results

We built a review process where classification specialists worked through hundreds of thousands of real search queries, scored the quality of the results, and fed precise, structured feedback back into the system.

Reference set

A trusted source of truth

That labelled work became a verified reference set: the highest-quality data the organisation had ever held, and the anchor for training the models and measuring accuracy over time.

Intent, not keywords

Tuned for what people mean

Using the verified data, we retrained the search and recommendation models to connect a traveller's intent to the best match in the catalogue, so a complex, natural sentence returns the right place, not the closest keyword.

Fact-check agents

Autonomous checks on every claim

We wired in agents that watch the generative answer layer in real time and cross-check what it produces against the company's factual database, catching an invented detail the moment it appears.

Grounded or blocked

Unbacked answers never ship

Anything the checks cannot trace to a trusted source is blocked before a user ever sees it. The system would rather show less than show something that is not true.

Holds over time

Continuous, not one-off

The same checks run continuously, so relevance holds and the models do not quietly drift back to the old behaviour once the launch is over.

The result

The right result, at the top, and true.

Better matches and no invented details turned browsers into bookings.

Live

Top-result relevance up about 35%

Accuracy and relevance of the results on the first page rose roughly 35%, which lifted engagement straight away. More than 200,000 queries were labelled and rated along the way, becoming the organisation's strongest data asset for everything that comes next.

And trusted

Invented information down 80%

Reports of the site showing invented or wrong details fell 80%, and the conversion rate from a search to an actual booking climbed with it. People found what they meant, trusted what they saw, and completed the booking.

Why it holds

Relevance is a data problem before it is a model problem.

You cannot fine-tune your way to relevance on data no one has checked. The lift came from the verified reference set first, then the tuning, then the fact-checks that keep the generative layer honest in production. It is the same lesson behind our data-readiness work: the model was rarely the thing that was broken. The data was.

More case studies

Related work.

Search that shows the wrong thing?

Book a strategy call Bring your worst query. Thirty minutes, no slides, or see more case studies.