Buying AI tools that can show their accuracy

Most AI-powered UX tools lack reliability and accountability in their outputs, so the purchasing question is what accuracy the vendor will put in writing.

A vendor specification with the accuracy figure left blank

A demonstration is not an accuracy claim.

Nielsen Norman Group, drawing on work from the Baymard Institute, makes a point that belongs in a procurement conversation rather than a design one: most AI-powered UX tools lack reliability and accountability in their outputs.

The practical consequence is that a team evaluating one of these tools is usually evaluating a demonstration, and a demonstration is not an accuracy claim.

Asking the question that separates vendors

One question does most of the work: against what ground truth was this measured, and what was the result.

A vendor with a defensible answer will describe a benchmark set, who labelled it, and how the tool scored against it. A vendor without one will describe the model, the training approach or the size of the dataset, none of which is an accuracy figure.

The follow-up matters as much. What does the tool do when it is unsure. A system that returns a confident answer at all times is not more accurate than one that declines, it is less honest about the same error rate.

Testing on work the team already has

The cheapest evaluation available to any team is to run the tool over a study it has already analysed by hand and compare. The answers are already known, the disagreements are visible immediately, and the exercise takes an afternoon.

What to look at is not the overall hit rate. It is the character of the misses. A tool that misses evenly is a tool with a known error rate. A tool that misses systematically on one kind of finding has a bias the team would be importing into every future study.

A system that returns a confident answer at all times is not more accurate. It is less honest about the same error rate.

BrilliantUX editorial principle

Writing the requirement into the contract

Where the tool will influence decisions rather than just save clerical time, the accuracy expectation belongs in the agreement alongside uptime.

That is a normal commercial request and it has a useful side effect: a vendor unwilling to state a figure in writing has answered the evaluation question anyway.

Source

Nielsen Norman Group, Demand Accuracy in Your AI Tools: Lessons from Baymard Institute.

nngroup.com/articles