Operations

How to measure an AI voice agent: the KPIs that matter

How to measure an AI voice agent: which KPIs to track, how to listen to a sample of calls and how to set your own success thresholds from real data.

The AgentVocal AI team · Published 3 October 2026, updated 3 October 2026 · about 8 min read

In this guide8 sections
  1. 01Which indicators should you track?
  2. 02How do you define success for each scenario?
  3. 03How do you set your own thresholds, without borrowed benchmarks?
  4. 04How do you listen to a sample of calls?
  5. 05What does the portal give you for measurement?
  6. 06What does a 30-minute weekly review look like?
  7. 07Which mistakes should you avoid?
  8. 08The next step

You measure an AI voice agent by looking at the outcome of each scenario, not at the number of calls. Start with a few simple indicators (calls answered, outcomes achieved, calls missed, escalations), add listening to a sample of calls, and set your own success thresholds from your own data. There is no “good” number that holds for everyone, and anyone who promises you one without knowing your scenario is selling you a figure, not a measurement.

This guide covers which KPIs to track, how to define success per scenario, how to set thresholds without borrowed benchmarks, how to listen to calls efficiently, and a 30-minute weekly review you can actually keep up.

Which indicators should you track?

Choose a few indicators and follow them consistently. Below are the most useful, with the question each one answers.

Indicator How you define it What it asks
Answer rate Calls picked up by the person out of all attempts (for initiated calls) Are we reaching people? Are the hours right?
Missed calls Received calls that never became a conversation, or attempts that went unanswered Are we losing people? When?
Resolution or confirmation rate Calls with the desired outcome out of conversations held (confirmed, booked, document promised) Is the agent doing its job?
Average length Average time of a conversation, per scenario Are calls too short (people hang up?) or too long (the agent gets lost?)
Escalations Calls handed over to the team, with the reason What does the agent not know? What are people asking for?
Transcript quality How many calls in a sample contain misunderstandings or wrong answers What needs fixing in the instructions?
Complaints and do-not-call requests Their number and reasons Are we being a nuisance? What are we saying wrong?

A few clarifications that prevent confusion:

  • The answer rate says nothing about the agent’s quality. It reflects your cadence: hours, days and the quality of the list. You fix it in the calling hours and the list, not in the script.
  • Resolution is defined per scenario. For order confirmation it means an order clearly confirmed or clearly declined, not just “call ended”. For reception it means a booking made or information given correctly.
  • Average length is not a goal in itself. A shorter call is not better if the person hung up unhappy.
  • Calls nobody answers are not charged, so you can track them without worrying about cost. Voicemail, on the other hand, is charged as a normal minute, so it is worth tracking separately.

How do you define success for each scenario?

A different scenario means a different kind of success. Before launch, write half a page for each scenario:

  1. The desired outcome. What does a “successful call” mean?
  2. Acceptable outcomes. A clear refusal or a rescheduled date are still useful outcomes.
  3. Outcomes that need a person. Complaints, unusual requests, confused callers.
  4. What is not success. A call that ends with no conclusion, wrong information, an annoyed person.

Examples, using generic scenarios:

  • Cash on delivery order confirmation: order confirmed, cancelled or changed (address, product). The outcome lands in your ERP, so you also measure whether the note actually got there.
  • Appointments: bookings made, moved, cancelled; how many received the confirmation message.
  • Missing documents: people who understood what is missing and promised a date; documents actually received afterwards.
  • 24/7 reception: after-hours calls that were answered, with requests correctly noted for colleagues.
  • Surveys: surveys completed to the end, not just calls started.

See the ready-made scenarios to start from their typical outcomes.

How do you set your own thresholds, without borrowed benchmarks?

You will not find target figures here, and you should not accept them from anywhere without checking them. The answer rate on a list of existing customers is a different thing from the rate on a cold list. A scenario with simple questions resolves differently from one full of exceptions. Your thresholds are built from your own data, in four steps:

  1. Measure without expectations. For the first two or three weeks, record the indicators without judging them. That is your baseline.
  2. Compare with the real alternative. What was your team doing before? What can you realistically sustain? Compare the same scenario over the same period. For a wider perspective, see AI voice agent or call center.
  3. Set a minimum and a target. The minimum is the level below which you investigate immediately. The target is the direction you are working towards. Write them down and date them.
  4. Revisit the thresholds. After big changes (a new list, a new scenario, a rewritten script) the baseline resets.

A threshold only makes sense if it triggers an action: “below level X, I listen to every call from that day”, “above level Y of escalations with the same reason, I rewrite that passage”. A number without an action is decoration.

How do you listen to a sample of calls?

Numbers tell you what is happening. Calls tell you why. No metric replaces listening.

What goes into the sample:

  • random calls, to see how things go normally;
  • very short calls (the person may have hung up after the first sentence);
  • very long calls (the agent may have got lost, or the person had a real problem);
  • calls with an unusual outcome or no outcome;
  • calls with an escalation or a do-not-call request.

How to listen efficiently:

  1. Read the summary and outcome in the portal first; then open the transcript and listen only to the problem passages.
  2. Check the opening first: the recording announcement, the introduction as a virtual assistant and a clearly stated purpose.
  3. Look for moments where the person repeated a question, interrupted or raised their voice.
  4. Tag each problem: missing information, weak wording, pronunciation, missing escalation rule, calling hours or list problem.
  5. Put the tags in the same sheet every week. Patterns emerge on their own.

Listen without looking for someone to blame. A bad call almost always points to a script or a rule that needs fixing, not to a flaw in whoever handled it.

Sample size depends on volume and risk. At launch, listen to almost everything. As the indicators settle, move to a fixed sample that you choose, plus every call flagged as a problem.

What does the portal give you for measurement?

The AgentVocal portal gives you the raw material for every indicator above. For each call you get:

  • the audio recording, if you enabled it, with the person informed;
  • the full transcript, word for word;
  • a short summary and the outcome of the call;
  • the call status (for example answered, missed, voicemail).

At the overall level you have campaigns and lists (who has been called, who is next, who asked not to be called again) and reports with the number of calls, outcomes and usage. To check against reality in your business, also verify that the outcome reached the place where your team works, your CRM or ERP, through integrations.

What the portal does not do: tell you on its own what “good” means for your business. Thresholds and definitions of success remain yours.

What does a 30-minute weekly review look like?

A small routine beats a big audit done rarely. A suggested agenda, to adapt:

Step What you do Indicative time
1 Look at the week’s indicators and compare them with last week and with your thresholds 5 minutes
2 Read the reasons for escalations and complaints 5 minutes
3 Listen to the agreed sample and the problem calls 15 minutes
4 Decide on 2 or 3 changes (script, rule, calling hours, list) 5 minutes

Ground rules:

  • same person, same day, same format;
  • only a few changes per week, so you know what made the difference;
  • every change logged with a date (see also the guide to the knowledge base);
  • after big changes, go back to intensive listening for a few days;
  • complaints and do-not-call requests are handled immediately, not at the weekly meeting.

Which mistakes should you avoid?

  • A single number. A resolution rate without length and complaints can look good while people leave unhappy.
  • Averages that hide things. One very good scenario can mask a weak one. Measure per scenario.
  • Unfair comparisons. Do not compare a list of new customers with an old one, or a peak-season month with a quiet one.
  • Too many changes at once. You will not know what helped.
  • No action after measuring. If nothing changes, measurement is just bureaucracy.

The next step

Pick four or five indicators, write the definition of success for each scenario, fix the time of your weekly review and start with a baseline. If you are preparing to launch, the launch checklist includes a monitoring plan for the first week. You can see what a call with a summary and outcome looks like in the demo.

Diagram · a call record

What every call leaves behind, taken apart

Every call has a record in the portal. Open it and you see everything: what was said, what the agent understood and what it did afterwards.

1Headerdirection, line, time and duration

Appointment confirmationOUTBOUND · THU 10:42 · 02:41 02:41

2Recordinglisten to the whole call; the person was informed

3Transcriptwhat was said, word for word

AGENTHello, I’m calling about tomorrow’s appointment. This call is recorded.

CLIENTYes, 10 o’clock still works.

4Summarytwo lines, so you don’t have to listen to it all

The customer confirms the appointment tomorrow at 10:00. No questions.

5Outcomea clear status, the same in reports

Confirmedno follow-up needed

6Actionswhat it wrote and sent after the call

Calendar updatedConfirmation SMS sentNote in CRM
the agentthe customerfictional example

Questions

Frequently asked questions

What is a good resolution rate for a voice agent?

There is no number that holds for everyone. It depends on the scenario, the contact list and what “resolved” means for you. Set your own threshold from the first weeks of real calls, then track it over time.

How many calls should I listen to?

Enough to see patterns. At first you listen to almost every call, then you move to a regular sample plus every call flagged as a problem. You decide the sample size based on your volume.

What do I see in the portal for each call?

The audio recording, the full transcript, a short summary and the call outcome. On top of that, reports show the number of calls, outcomes and usage.

How often should I review the agent's performance?

We recommend a short weekly review in the same format, at least for the first months. Afterwards you can move to a slower rhythm if the indicators stay stable.

The line is free

Hear it, then decide.

Sign up, see the estimated cost in the form and test the agent on your own phone before it calls anyone.