First Contact Resolution and AHT: Measuring Voice AI

Ryan Stevens
Close-up of a woman in a teal uniform wearing a cream headset, with blurred computer monitors behind her.

Evaluate Voice AI by the request outcome and the work required to resolve it.

First contact resolution (FCR) measures whether customers get their issues resolved in the first interaction. Average handle time (AHT) measures how long it takes to handle each contact. Together, they help operations teams evaluate Voice AI, as long as the measurements use consistent definitions.

A shorter call is useful only if it leaves the customer with the right outcome. This guide shows how to compare resolution and handling time, account for repeat work, and build a scorecard for a Voice AI pilot. Read the two measures together, then check the workload, customer feedback, and cost behind them.

How to measure first contact resolution

FCR tracks the percentage of customer issues resolved during the first interaction. An unresolved call can lead to another interaction: a customer calls about a billing issue, leaves without a clear answer, and calls the next day again. The business has now spent time handling the same problem twice.

Repeat contacts add work. Reducing avoidable repeats frees staff to handle other requests, though staffing needs also depend on demand, coverage, and the complexity of the remaining calls.

For phone-only reporting, first-call resolution focuses on the first call; keep the same eligibility and follow-up rules explicit. Compare FCR with your own baseline for the same call types and measurement rules. An external benchmark is useful only when its definition of resolution, channel coverage, and follow-up window match yours.

For this pilot, count each new customer issue once. The first-contact resolution formula is the number of eligible issues resolved during the first interaction with no same-issue repeat in the agreed follow-up window, divided by all eligible new issues, multiplied by 100. Document which issues are eligible before the pilot and report excluded volumes separately. An unresolved issue remains in the denominator.

Choose a follow-up window appropriate to the request and wait until every included issue has had that full window. Match repeat contacts by customer and issue across the channels you can observe. Report gaps in that coverage. Confirm completion in the relevant business record and check a sample with quality review or customer feedback; no callback alone does not prove resolution.

ICMI recommends combining repeat-contact tracking with customer feedback when assessing resolution.

For this scorecard, an AI-to-human transfer that resolves the issue within the same interaction may count toward end-to-end FCR, but it does not count as AI-only resolution. Report transferred outcomes separately and keep this convention consistent with your baseline. A promised callback or booking is not a completed resolution until the agreed completion condition has been met.

Where average handle time fits in

AHT measures the average time required to handle a voice contact, including talk time, hold time, and after-call work. Calculate it by dividing total talk, hold, and after-call work time by the number of handled contacts.

Compare AHT for the same type of request, using the same rules for each component. Keep AI-only, human-only, and transferred contacts visible separately. Document how transfers and multiple agent segments are counted, and measure total handling effort across the full issue as well. Report customer waiting time separately rather than silently changing the AHT definition.

Longer human-handled calls use more staff time per contact. In an AI workflow, also track human follow-up, capacity limits, and usage costs before assuming that a shorter call creates usable capacity.

An AI agent may reduce lookup or documentation time when it has access to the right information and the required systems. Measure whether the configured workflow actually reduces handling time while preserving a correct outcome. Read AHT alongside the contact outcome.

Why a shorter call can create more work

Shorter calls, shorter queues, and faster responses can look appealing. Consider an illustrative example. In one path, the customer's issue takes six handling minutes and is resolved. In another, the first contact takes two minutes but leaves the issue open; a second contact takes five more minutes to resolve it.

The second path uses seven handling minutes across two contacts. Its average of 3.5 minutes per contact looks better than six minutes, but the total handling time per resolved issue is higher.

These figures are hypothetical and assume the same handling-time definition for both paths. They do not establish a cost saving: actual cost also depends on who handles each contact, the platform charges, and any further work. The goal is to resolve the customer's issue during the conversation.

Read FCR and AHT together

A lower AHT can hide repeat work if callers leave without a complete answer. A higher FCR with a lower AHT is encouraging, but it does not establish that AI caused the change. Compare similar requests over comparable periods and check whether routing, staffing, or the mix of calls changed.

A longer call can also be worthwhile if it resolves the issue and prevents further contact. Evaluate the total work required to resolve the customer's request, not just the duration of one interaction. Confirm the effect on repeat contacts, customer satisfaction, and total cost before drawing a business conclusion.

Build a Voice AI pilot scorecard

Use the same definitions for the baseline and pilot. Enter observed results in the blank columns and retain the underlying counts so that changes in volume, coverage, and unresolved work stay visible.

Measure

Fixed definition and evidence

Baseline

Pilot

Decision it informs

Eligible new issues

Unique customer-issue records after agreed exclusions; report exclusions and observation coverage

___

___

Are the cohorts comparable?

FCR

First-interaction resolutions that pass the full follow-up window ÷ eligible new issues × 100; verify task records

___

___

Are more requests resolved at first contact?

AHT

Talk + hold + after-call work minutes ÷ handled voice contacts; same counting rule

___

___

Is handling time changing within each call type?

Same-issue repeat rate

Issues with at least one repeat in the agreed window ÷ eligible new issues × 100

___

___

Is unresolved work returning? This is not automatically 100% minus FCR.

Handoff quality

Audit whether a required transfer happened, reached the right team, and supplied the agreed context; show audit sample size

___

___

Is human support available when needed? A lower transfer rate alone is not success.

Customer satisfaction

Same survey question, scale and timing; show responses and response rate

___

___

Are customers reporting a better experience?

Cost per resolved issue

Attributable operating cost for the issue cohort, including repeat work, ÷ issues resolved by the reporting cutoff; show unresolved count

___

___

Is the whole request becoming less expensive to resolve?

Run a comparable pilot

  1. The operations lead selects a call type and writes down completion criteria, exclusions, transfer rules, and the follow-up window.
  2. The analyst fixes baseline and pilot dates and records volumes, call mix, language, hours, routing changes, and staff coverage. Compare the same segments; a concurrent comparison group is useful where practical. Do not present a simple before/after difference as proof of causation.
  3. The analyst links call records to task outcomes and repeat contacts, and separates AI-only, transferred, and human-only work. Missing outcome data should remain visible as unknown.
  4. A quality reviewer checks a sample of outcomes and handoffs. Include failures and unresolved requests rather than auditing only calls labeled successful.
  5. Operations and finance review the scorecard together. For cost, include AI/platform usage, telephony, staff handling and follow-up, and relevant supervision; show setup/integration cost separately or explain its allocation. Report both periods on the same basis.
  6. Expand only after the agreed outcome, service-quality, and economic criteria are met. If AHT falls while FCR worsens, inspect failed completions; if FCR improves while AHT rises, check whether fewer repeats reduce total work.

For a starting use case involving interactive voice response (IVR), see Newo's guide to replacing IVR, then define the completion and handoff criteria for your chosen workflow.

What improves resolution in practice

Improving FCR requires more than answering calls. Customers need useful, accurate answers and completed actions. Factors to review include:

  • Accurate information
  • Consistent responses
  • Understanding of customer context
  • Fast access to knowledge
  • Proper escalation
  • Reliable workflows
  • Ongoing quality monitoring

When reviewing repeat contacts, check for incorrect or incomplete information, incomplete actions, and cases that require a later step. Assign an owner to the knowledge, policy, or workflow issue you find.

How Newo supports the workflow

Newo separates conversation guidance from action execution. Newo uses an Observer agent that analyzes each conversational turn, tracks information the conversational agent needs, and generates directions for the next response. It also evaluates conversation quality and can trigger a configured recovery path.

The Supervisor agent and Tool Caller check the agent's response for commitments and trigger the corresponding enabled action, such as sending a message or creating a booking. Available actions depend on the tools and workflows configured for that deployment.

These controls address missed steps and unfulfilled actions. You still need to measure their effect on first-contact resolution and handle time in the workflow you deploy.

Choose one repeatable call type, agree on what a completed request looks like, and compare the pilot with a comparable baseline. Use the results to identify the next knowledge, workflow, or handoff problem to fix before expanding.

To evaluate first contact resolution alongside handle time in your own workflow, explore Newo for contact centers and define the outcomes your pilot needs to demonstrate.

More from Newo

Book a demo