Skip to content
Field notes
AI UXMetricsProduct Design

Assistant usage is not the metric. Resolution is.

Teams measure AI features by how often people use them. That rewards the wrong design. Five behaviours that tell you whether an intelligence layer is doing its job, and how to instrument them from the first release.

Matija Vojvodic · 3 min read

Ask a product team how their AI assistant is doing and you will hear a usage number: sessions, messages, weekly actives. It is the easiest thing to count and the least useful thing to know. An assistant that people open ten times for the same question is not succeeding. It is failing slowly, with engagement.

The design goal of an intelligence layer is that the user's need gets resolved, ideally without a conversation at all. So the metrics have to be about resolution, and they have to be designed in before the first release, because they are much harder to add afterwards.

If the assistant is working, people should need it less, not more.

Five behaviours worth measuring

Take KNIME's AI-assisted data onboarding as the example. The need is simple to state: get my data into a configured node. Here is what tells you whether the AI is helping.

  • In-context resolution. Time to first configured node, and the share of users who get there from the Input File card without opening a conversation. A suggestion applied in three steps counts; a chat session that ends with the same question counts against.
  • Evidence use. How often people open Info next to Apply, and whether they do so before or after applying. Low use is not automatically bad, but a drop after a model change is the earliest warning you will get.
  • Modify and dismiss rates. Because every suggestion has state, you can see how often a draft is applied as-is, edited first, or rejected. Rising edits on one setting means the model or the default is wrong for that source.
  • Proposal to action. The time between a suggestion appearing and the user applying or dismissing it. If it grows, the suggestion is not legible or not trusted.
  • Repeat attempts. Users re-importing the same file or re-opening the same node within a session. This is the number that catches confident wrong suggestions.

Why usage misleads

Usage rises when the product is confusing and when the assistant is good at explaining confusion. It also rises when the assistant is placed everywhere, because a prompt on every card will be clicked. Neither tells you the user got what they came for. Resolution metrics are indifferent to how the need was met, which is the point: a pre-filled draft applied in one click and a well-handled conversation score the same if the user leaves with a working node.

Instrument from the first release

Each of the five needs a definition of “resolved” for the object in question, and that definition is a design decision, not an analytics one. For an Input File it is “a node was applied from this card and not replaced within the session”. For a Suggestion it is “applied or modified, not dismissed”. Write those definitions next to the design of the component, and ship the events with the feature. The object model makes this cheap: suggested, applied, modified and rejected are already states. Retrofitting them six months later means a quarter of data you cannot trust.

What I would show a leadership team

One chart per behaviour, by source type, and a single line at the top: how many datasets reached a configured node this month without a conversation. If that number grows while assistant sessions stay flat or fall, the architecture is working. If sessions grow and resolution does not, you have built a very engaging way to not help people.

Building something complex?

Let's turn it into a system people trust.

If this resonates, I help teams bring the same clarity to their SaaS and AI products – from research and object models to the shipped interface.