Measuring an assistant

Task completion and time on page stop meaning anything once the product acts on the user behalf. What replaces them.

Acceptance, edit distance and reversal shown side by side

The replacement measures have to be ones a finance team recognises.

Task completion assumes the user performs the task. Time on page assumes that less time is better. Neither assumption survives a product that acts on the user's behalf.

When a system drafts the reply, reconciles the ledger or assembles the report, a fast session can mean the assistant worked or that the user gave up. The same number covers both, which makes it useless as a signal.

Measuring the outcome instead of the interaction

The measures that still mean something are the ones attached to the job rather than to the screen.

Acceptance rate is the first: what share of what the assistant produced was used without modification. It is blunt and it is honest, and it splits usefully by task type.

Edit distance is the second: of the outputs that were modified, how much was changed. A high acceptance rate with heavy editing is a different product problem from a low acceptance rate, and a single satisfaction score hides the difference.

Reversal rate is the third and the one teams instrument last. How often did a person undo something the system did, and how long after the fact did they notice. A reversal discovered twenty minutes later is a usability finding. A reversal discovered a week later by someone else is a trust problem.

Keeping a measure the business recognises

Nielsen Norman Group's read on 2026 is that demonstrated business impact is what keeps UX funded, and that applies with particular force here, because assistant features are expensive to run and the cost is visible on a bill.

The bridge measure is usually time to outcome rather than time on task. How long from a request being made to the work being finished and accepted, including the review. That number is comparable to the pre-assistant baseline, it is legible to a finance team, and it does not reward a fast interaction that produced something nobody could use.

A fast session can mean the assistant worked, or that the user gave up. The same number covers both.

BrilliantUX editorial principle

Instrumenting before launch rather than after

Acceptance, edit distance and reversal all require events that have to exist in the product before they can be counted, and retrofitting them means a quarter with no baseline.

The short version of the instrumentation list is: log what the system proposed, log what shipped, log the difference, and log every undo with its latency. Four events, defined before the feature is built, and they answer most of the questions a team will have in the first six months.

Teams that skip this end up defending an assistant feature with a satisfaction survey, which measures how people feel about the idea of the feature rather than whether it did the work.

Source

Nielsen Norman Group, State of UX 2026: Design Deeper to Differentiate.

nngroup.com/articles/state-of-ux-2026