XMACNA
AI support metrics: silence is not a solution

AI support metrics: silence is not a solution

AI support metrics must prove outcome, return contact, and handoff — not just silence or conversation ended. See the scorecard.
XMACNA Team

8 min read

Analysis

AI support metrics need to prove that the customer reached the expected result, not just that the conversation ended. A reliable resolution combines final state, action evidence, absence of recontact within a set window, signs of frustration, and a contextual handoff when the Digital Employee reaches its limit.

A company can show high deflection yet still create an invisible queue. The customer gets a response, doesn’t insist at that moment, and returns hours later for the same reason. The first contact seems successful on the dashboard. For the customer, the problem remains open.

At XMACNA, institutional experience with more than **600 Digital Employees operating in Brazil** teaches a simple rule: activity only gains value when it ends in an observable state. In support and sales, conversation, action, and record need to tell the same story.

This is the point of a news release by Smarsh in 3 September 2026. More important than the highlighted percentage is the method used to ask if the interaction truly solved the need.

What does the Smarsh case show about resolution with AI?

Smarsh reported 405 interactions in the second quarter of 2026, with 72% self-service deflection and a manual confidence score of resolution of 2,6 in 3 for its support agent. These are results reported by the company itself in a statement, not an independent experiment or universal reference for other businesses.

The useful detail is in the definition. According to Smarsh, the maximum score requires high confidence that the question was fully resolved, no escalation in the following 48 hours, and no signs of frustration or contestation during the conversation. The review therefore looks beyond the 'session ended' event.

The case also describes an internal agent who provides professionals with an account snapshot including responsible party, status, and support history. The lesson is not to eliminate people. It is to reserve human judgment for what requires experience and deliver them a case ready to advance.

Why can deflection and conversation ended be misleading?

Deflection answers an operational question: how many interactions did not reach a person? Resolution answers another: how many customers achieved the intended outcome?

Both measures can move together but are not synonymous. Salesforce itself updated its deflection and abandonment evaluation because events like click to end and timeout don’t always reflect resolution or frustration. Zendesk also recommends excluding abandonment, partial response, closure without solution, and repeated contact from automated resolution calculation.

For an operation on WhatsApp with continuous support, the risk is even more concrete. The customer may interrupt the conversation to search for a document, talk to someone else, or just give up. If the system turns any silence into success, it optimizes the wrong indicator.

A low handoff rate can also seem efficient while hiding cases retained too long. The goal is not to prevent handoff. It is to forward at the right time, for the right reason, and with sufficient context.

What should count as verified outcome?

Before choosing a tool or setting the dashboard, the company needs to write the resolution contract for each recurring intention. For 'second copy sent,' for example, the response may not suffice: the correct document must have been located, delivered to the authorized contact, and recorded. For 'qualified lead,' the conversation must generate the fields and next step the sales team actually uses.

A minimum scorecard has six fields:

FieldOperational questionPossible evidence
IntentionWhat did the customer want to accomplish?category and objective confirmed
Final stateWhat change proves completion?order updated, appointment created, document delivered
EvidenceWhere can the action be checked?protocol, record, event, or confirmation
RecontactDid the same reason return within the defined window?correlated intention in 24, 48, or 72 hours
ExperienceWas there frustration, contestation, or request for a person?textual signal, evaluation, or sampled review
HandoffDid the exception reach the right person with context?responsible, reason, summary, history, and next step

The window is not universal. Forty-eight hours made sense in the method published by Smarsh. An appointment confirmation may require observation until the scheduled time. A billing question may need to track payment clearance. The horizon comes from the process, not the provider.

How to measure support and sales without creating vanity metrics?

Intercom separates in its reports deflection, confirmed resolution, presumed resolution, experience, and escalation. It also differentiates what is outcome in support and in sales: resolving a request is not the same as qualifying or disqualifying an opportunity.

This distinction is essential for the Digital Seller and the SDR process with AI. In sales, 'responded to lead' measures speed. 'Created opportunity with interest, context, and next step' measures progress. 'Scheduled and recorded' measures execution. The indicator should follow the state useful to the team, not the message volume.

In support, quality must combine resolution, accuracy, effort, satisfaction, cost, and impact on the team. A single number invites shortcuts. A coherent set allows discovering if automation resolved, transferred well, or just postponed work.

How should human handoff work on WhatsApp?

Handoff is a safety and experience capability, not an AI failure. It should occur when data, confidence, permission, or access is lacking; when there is a policy exception; when the person shows frustration; or when judgment and empathy are part of the outcome.

The standard described by Salesforce for human handoff preserves intention, urgency, history, and relevant data. The person takes over without asking the customer to repeat everything again.

In XMACNA's public architecture, the Conversation Portal allows the team to monitor and intervene. The Digital Employee pauses, and the responsible person gets the context. The Intelligent Dashboard preserves contact, stage, objection, opportunity, and history. When closing, the Intelligent Analysis records what happened and feeds the Intelligence Cycle.

The handoff can then be evaluated with objective questions:

  • Is the reason for transfer clear?
  • Did the right person receive the case?
  • Does the summary separate fact, request, and action already tried?
  • Is the next step explicit?
  • Did the customer need to repeat information?
  • Did the final result return to the same record?

How to review quality without reading all conversations manually?

Start with stratified sampling. Review cases classified as resolved, transferred, and abandoned; include high volume, high risk, and low confidence intents. Compare automatic classification with a human conclusion and record the reason for divergence.

Next, track cuts that reveal the problem: intent, channel, flow version, handoff type, recontact, and responsible party. A global average may improve while an important reason worsens.

Models can help classify the entire session, as shown by the Salesforce update, but the automatic judge also needs explicit criteria and auditing. It does not turn a bad definition into a good one. The team remains responsible for deciding what “resolved” means.

The goal is an improvement cycle: identify false positives, find the cause, correct knowledge or action, test again, and observe recontact. The review stops hunting for bad conversations and becomes process maintenance.

How to implement the metric in a real flow?

Choose a frequent and well-defined intent. Avoid starting with the entire service.

  1. Define the expected outcome in business language.
  2. Name the final state and its evidence.
  3. Choose the appropriate recontact window.
  4. List signals that require a person, approval, or stop.
  5. Design the minimum context package for the handoff.
  6. Record conclusion, exception, and responsible party in the same flow.
  7. Review a sample and treat divergences as operational improvement.

A Digital Employee does not exist to produce nice conversations. It exists to PERFORM a clear part of the work, respect limits, and return the result or exception ready for the carbon and silicon team.

In summary

  • closed conversation does not prove resolution;
  • deflection must be read along with abandonment, recontact, and frustration;
  • each intent demands its own final state and evidence;
  • correct handoff preserves context and can be a quality result;
  • service and sales need different definitions of success;
  • the same record should link conversation, action, evidence, and next responsible party.

Want to find out where your operation measures activity instead of results? Start with XMACNA Assessment using a real process. Don’t believe it? Try it.

Frequently asked questions

What are AI customer service metrics?

They are indicators that show if AI led the customer to the expected outcome, with accuracy, evidence, proper experience, and handoff when necessary. They go beyond volume, speed, and closed conversations.

What is the difference between deflection and resolution?

Deflection measures how many interactions did not reach a person. Resolution measures how many problems were actually concluded. A customer who abandons and returns for the same reason may count as deflection without having received a solution.

How long should I observe recontact?

The window depends on the process. It can be 24, 48, or 72 hours, until the scheduled appointment, or until confirmation of an action. The rule should reflect when the result can be verified.

Does human handoff reduce AI customer service quality?

No. A handoff at the right time, with reason, history, and next step, protects the customer and the business. The problem is transferring late, to the wrong person, or without context.

How to start measuring verified resolution?

Choose a high volume intent, define final state, evidence, recontact window, frustration signals, and handoff rules. Then compare automatic classification with sample human review.