Measuring an assistant
Task completion and time on page stop meaning anything once the product acts on the user behalf. What replaces them.

The replacement measures have to be ones a finance team recognises.
Task completion assumes the user performs the task. Time on page assumes that less time is better. Neither assumption survives a product that acts on the user's behalf.
When a system drafts the reply, reconciles the ledger or assembles the report, a fast session can mean the assistant worked or that the user gave up. The same number covers both, which makes it useless as a signal.
Measuring the outcome instead of the interaction
The measures that still mean something are the ones attached to the job rather than to the screen.
Acceptance rate is the first: what share of what the assistant produced was used without modification. It is blunt and it is honest, and it splits usefully by task type.
Edit distance is the second: of the outputs that were modified, how much was changed. A high acceptance rate with heavy editing is a different product problem from a low acceptance rate, and a single satisfaction score hides the difference.
Reversal rate is the third and the one teams instrument last. How often did a person undo something the system did, and how long after the fact did they notice. A reversal discovered twenty minutes later is a usability finding. A reversal discovered a week later by someone else is a trust problem.
Keeping a measure the business recognises
Nielsen Norman Group's read on 2026 is that demonstrated business impact is what keeps UX funded, and that applies with particular force here, because assistant features are expensive to run and the cost is visible on a bill.
The bridge measure is usually time to outcome rather than time on task. How long from a request being made to the work being finished and accepted, including the review. That number is comparable to the pre-assistant baseline, it is legible to a finance team, and it does not reward a fast interaction that produced something nobody could use.
A fast session can mean the assistant worked, or that the user gave up. The same number covers both.
BrilliantUX editorial principle
Instrumenting before launch rather than after
Acceptance, edit distance and reversal all require events that have to exist in the product before they can be counted, and retrofitting them means a quarter with no baseline.
The short version of the instrumentation list is: log what the system proposed, log what shipped, log the difference, and log every undo with its latency. Four events, defined before the feature is built, and they answer most of the questions a team will have in the first six months.
Teams that skip this end up defending an assistant feature with a satisfaction survey, which measures how people feel about the idea of the feature rather than whether it did the work.
Nielsen Norman Group, State of UX 2026: Design Deeper to Differentiate.



