Blog

How to tell whether your AI agent is actually working

A satisfaction survey hears from the few customers who bother to answer it. Once an agent answers first, every conversation can be scored — but more numbers is not the same as knowing what's going on. Here's how to read them without fooling yourself.

Playbook · · 7 min read

Anchor AI

Support used to be measured by asking. A survey at the end of the chat, a thumbs-up, a star rating — and a report built from whichever customers felt strongly enough to click. That number was never wrong, exactly. It was thin: the satisfied and the furious answered, and the large, quiet middle didn’t.

When an AI agent takes the first reply, the measurement changes with it. Every conversation is a transcript the system saw from start to finish, so every one can be scored, counted, and traced to the document that answered it. This post is about reading that page — Anchor’s Analytics — in an order that tells you something, and about the two traps that make more data feel like less.

The funnel, read top to bottom

The Overview tab shows five numbers. They aren’t five opinions about the same thing. They’re a funnel, and each one only means something in the light of the one before it.

MetricWhat it countsWhat a low number usually means
Bot involvementConversations the agent took part inThe agent isn’t being reached: the widget isn’t live, or customers go straight to a person.
Answer rateInvolved conversations it actually answeredIt’s being asked things it has no content for. Look at Knowledge gaps.
Resolution rateConversations resolved without a humanAnswers aren’t landing, or the handoff list is set too broad.
Automation rateHandled end to end by the agentThe strict version of resolution: no human touched it at any point.
CX positive shareConversations scored 4–5Customers got answers but didn’t leave happy. Read the Experience tab.

The order matters because each number caps the next. A strong resolution rate on a weak involvement rate is a good agent nobody’s talking to. Read the top of the funnel first; a low number there makes everything below it small-sample noise.

The one to watch first

If you only look at one, look at Resolution rate— “resolved without a human” — because it’s the number that pays. Every point of it is a conversation your team didn’t type. But read it beside the Escalation tab, not alone: a resolution rate that falls after you add a handoff topic hasn’t got worse, it’s routing more on purpose. Hard escalations are policy. Soft escalations, the bot-failure signals, are the ones that should be falling month on month. (How to draw that line is its own post.)

A score for every conversation, and what it isn’t

The CX number deserves a plain description, because it’s easy to mistake for a survey. Anchor scores every conversation with a model — a proxy for how the customer likely felt, built from the transcript — and reports the share scored 4 or 5. It is not customers clicking stars.There’s no rating widget. What it buys you is coverage: the quiet middle is in the number, not just the people who felt like clicking.

The Experience tab breaks it into a Score distribution from 1 to 5, a Sentiment mix, and Reason drivers. The interesting band is the 3s. A 1 is a problem you’ll hear about anyway. A 3 is a customer who got an answer and left unimpressed — usually because the answer was correct but long, or correct but two messages later than it should have been. That band is where a support experience goes from adequate to good, and surveys never showed it to you because those customers don’t answer surveys.

Set targets from your own data, not someone else’s

Every vendor publishes a resolution percentage, and none of them is yours. Their number came from their customers’ documentation and their customers’ questions. Yours will come from how much of what people ask you have actually written down.

So start with a baseline, not a target. The first month — the trial’s thirty-odd conversations, or the first billing period — is measurement. Then set the target as a direction on the same questions: resolution up on the intents that matter most, soft escalations down. The Segments tab gives you Resolution by intent — order status, returns and refunds, shipping and delivery, billing, and so on — so the target can be specific: “returns resolution up, because that’s where the volume is”.

From a number to a fix

A metric that doesn’t tell you what to do next is decoration. Three lists on the Worklists tab do the translating.

  • Content gaps— your knowledge sources ranked by volume, with the ones that get asked about a lot and resolve little flagged “Needs attention”. That’s a document to rewrite, and it’s linked.
  • Failure modes— why the agent fell short: Misunderstood query, Wrong or outdated answer, Missing knowledge, Out of scope, and the rest, each with up to five phrases customers actually typed. The phrases are the brief for what to write.
  • Review queue— conversations flagged as low-confidence or ambiguous, for a person to read and settle.

The loop is short: read the list, fix the document, and watch that intent’s resolution the following week. Because the score covers every conversation, you can see whether the fix worked without waiting for a survey to fill up.

Two ways to fool yourself

First, reacting to this morning. Anchor’s metrics rebuild nightly and the trailing two days are provisional, so a dip you see at 9 AM is often a count that hasn’t settled. Look at weeks, not days.

Second, celebrating a high automation rate that came from a handoff list set too narrow. If customers with real problems are being answered politely and never reaching a person, automation is up and CX is down — and the second number is the one telling the truth.

What good looks like

  • You read the five numbers as a funnel, top first.
  • Resolution rate is the headline, read beside the hard/soft escalation split.
  • The CX score is understood as a full-coverage proxy, not a survey, and the 3s get attention.
  • Targets are directions on your own intents, set after a month of baseline.
  • Every week, one worklist item becomes one document change.

Where Anchor fits

Everything named here is on Anchor’s Analytics page today: the five-number Overview with its funnel, Resolution by intent and by channel, the Experience tab’s score distribution, the Escalation tab’s soft/hard split, and the three Worklists. Export to CSV if you’d rather read it in a spreadsheet. The free trial is 30 days and about 30 conversations — a baseline, which is where the measuring should start.

Common questions

What is a good resolution rate for an AI support agent?

The honest answer is "higher than yours last month on the same questions". Published benchmarks come from other companies' documentation and other companies' customers. Measure your own baseline for a month, then set a direction: resolution up on the intents with the most volume, soft escalations down.

What's the difference between resolution rate and automation rate?

Resolution rate counts conversations resolved without a human. Automation rate is the stricter version: conversations the agent handled end to end with no human involved at any point. A conversation a person glanced at and handed back can count as resolved but not automated.

Is an AI-generated CX score the same as CSAT?

No. CSAT is customers answering a survey, which most don't. Anchor's CX score is a model's estimate of how the customer likely felt, built from the transcript, so it covers every conversation rather than the few who clicked. Treat it as a coverage tool for finding patterns, not as a replacement for hearing from customers directly.

How often should I check AI support analytics?

Weekly for decisions, daily only for emergencies. Anchor's metrics are rebuilt nightly and the last two days are provisional, so a morning dip is often a count that hasn't settled. A weekly look at resolution by intent and the Worklists tab is enough to find the next document to fix.

← All posts

Published

See it answer from your own content

Point Anchor at your help center and watch it answer your real questions — 30 days free, no card.