Sellium · Metrics

KPIs to measure an AI sales agent:what to track and how to read it

An AI agent selling on WhatsApp leaves more data behind than any sales rep: every conversation is logged, along with its outcome. The risk is not running out of numbers, it is watching the ones that do not matter. Here are the four KPIs that tell you whether the agent is selling, how each one is calculated and how to read them without fooling yourself.

The short answer

What KPIs should you track for an AI sales agent?

An AI sales agent is measured with four KPIs: resolution without a person (the share of conversations that end well with no human involved), conversion (how many end in a quote, an order or a payment), response time (the first reply and the time to resolution) and cost per conversation (the agent, Meta's messaging fees and team hours, over the conversations handled).

None of them is read on its own or against a universal benchmark. You read them against your baseline, the weeks before the agent went live, and alongside two quality indicators that keep you honest: pricing or inventory errors, and escalations the customer asked for but did not get in time.

The KPIs

The four KPIs of an AI sales agent

Each KPI answers a different business question. Resolution tells you how much work the agent absorbs; conversion, whether that work sells; response time, whether it arrives in time; and cost per conversation, whether it pays off. Looking at them together keeps you from improving one at the expense of another.

KPIWhat it answersHow it is calculatedHow to read it
Resolution without a personHow much work the agent absorbsConversations resolved by the agent over all conversations handledHigher is not always better: an escalation that needed to happen is also a win
ConversionWhether conversations sellConversations with a quote, order or payment over conversations with purchase intentCompare by use case, and for the same customers before and after
Response timeWhether the customer gets the answer in timeTime to first reply and time to resolution, as a median and at the 90th percentileSplit it by hours: after-hours is where it changes the most
Cost per conversationWhether the channel pays offThe agent, Meta's messaging fees and team hours on escalations, over the conversations handledCompare it with the cost of handling the same conversation with people

The definitions (what counts as resolved, what counts as purchase intent) are agreed before launch and do not change midstream.

Resolution

Resolution without a person: the most misread KPI

The temptation is to chase the highest possible number. But an agent that never escalates is not a good agent: it is one that holds on to conversations that needed someone, like a price negotiation, a complaint or an order outside the rules. Useful resolution is the kind that ends well for the customer.

Customers notice. In Gartner's consumer survey, the top concern about AI in customer service is that it will become harder to reach a person (Gartner, 2024). That is why resolution is always read together with the quality of the handoff to a sales rep.

  • Resolved: an inquiry answered with data from your systems, a quote sent, an order entered or a payment collected, without the customer asking the same thing again through another channel.
  • Not resolved: a conversation the customer drops halfway, one where they asked for a person and did not get one, or one that comes back the next day about the same issue.
  • Correct handoff: the agent passes what belongs to a person over to them, with the full history. It is a win, and it is measured separately.
Conversion

Conversion: from conversation to sale

Resolving is not selling. An agent can answer every inquiry well and close none of them if it does not propose the quote, ask for the order or offer the payment link. Conversion measures exactly that: the share of conversations with purchase intent that end in a quote, an order or a payment.

Measure it by stage, because each one fails for different reasons; the quote-to-order stage, for example, depends on quote follow-up. To estimate how much you lose to unanswered or late-answered inquiries, use the WhatsApp lost sales calculator.

StageWhat is measuredWhat holds it back
Inquiry to quoteConversations with purchase intent that receive a quotePrices the agent cannot look up, or discount rules nobody defined
Quote to orderQuotes that turn into ordersNo follow-up: the quote sits in the chat and nobody touches it again
Order to paymentPaid orders, when payment is part of the flowPayments that leave the chat, or payment links that arrive late
Average order valueRevenue per order for the same customers, before and afterAn agent that does not suggest add-ons or volume tiers

The cleanest attribution compares the same customers over the same time of year, before and after the agent.

Speed

Response time: the first reply and the time to resolution

With an agent, the first reply arrives in seconds, so that number stops being the interesting one. What matters is the time to resolution: how long it takes from the first message to the quote sent, the order entered or the escalation picked up by a person.

Speed matters because interest cools fast. In Harvard Business Review's study of web leads, companies that tried to contact the customer within the first hour were nearly seven times as likely to qualify the lead as those that waited even an hour longer (Harvard Business Review, 2011).

  • Median and 90th percentile: the average hides the conversations that waited for hours. The 90th percentile tells you how long the worst-served customer in every ten waited.
  • By hours: split business hours from after-hours. That is where the number changes the most, and where the lost sales used to be.
  • Escalations separately: measure how long it takes a person to pick up what the agent escalated. If the agent replies in seconds and the escalation waits until the next day, the bottleneck has moved, not gone away.
Cost

Cost per conversation: how it adds up and what to compare it with

Cost per conversation adds up everything it takes to handle WhatsApp with the agent and divides it by the conversations handled. It has three components, and the one most often forgotten is the time of the people who take the escalations.

The number alone says little: compare it with what it costs to handle the same conversation with people. To run that math with your own data, the sales reps vs AI agent calculator sizes the team you would need to reply on time and works out the cost at which the agent pays off.

Quality

The quality indicators that keep you honest

The four KPIs can look good with an agent that sells badly: high resolution because it never escalates, high conversion because it promises stock you do not have. That is why they come with quality indicators, reviewed just as often.

  • Pricing or inventory errors: orders with a price from another list, or with items out of stock at picking time. They should be close to zero if the agent queries your systems; how to prevent them is in how to keep an AI agent from making up prices or inventory.
  • Missed escalations: customers who asked to talk to a person and did not get one in time.
  • Satisfaction: a short question at the end of the conversation, such as CSAT, compared with the one for the channel handled by people.
  • Sample review: a set of conversations read every week by someone from sales. It finds what no number shows: a tone that does not match the brand, or an answer that is correct but incomplete.
How to read them

How to read an AI agent's KPIs and what to adjust

KPIs are only useful if they trigger a decision. Measure before you start: a few weeks of baseline are enough to compare later with numbers instead of impressions. With the agent live, these are the most common signals and what they usually call for.

SignalWhat it usually meansWhat to adjust
Low resolution and many escalations about priceThe agent has no access to a price list or a discount ruleConnect the missing data or write the rule
High resolution and low conversionThe agent answers but does not move the sale forwardHave it propose the quote, ask for the order or offer the payment link
Many abandoned conversationsToo many questions before giving anything usefulShorten the qualification and answer what the customer asked first
Escalations that wait for hoursNobody owns them, or there is no on-call scheduleDefine who takes each type of escalation and during which hours
Rising cost per conversationMore marketing templates or more hours on escalationsReview the templates being sent and the escalation rules
Pricing or inventory errorsData the agent reads from a copy instead of the systemHave it query the ERP in real time, not a spreadsheet

Review the KPIs weekly during the launch and monthly after that, always with the same calculation rules.

Step by step

How to build your KPI dashboard in six steps

Before you switch the agent on, not after: that way the numbers from the first weeks can be compared.

1

Define what resolved means

Write down what counts as a resolved, escalated or abandoned conversation, and what counts as purchase intent. Without definitions, every team measures something different.

2

Take the baseline

A few weeks of data from the current channel: response time, unanswered inquiries, WhatsApp orders and sales, and team hours.

3

Log every outcome

Each conversation goes into the CRM with its outcome: answered, quoted, order entered, paid or escalated. Every KPI comes from there.

4

Split by use case and hours

Inquiries, orders and quotes do not behave the same, by day or by night. An overall average hides where to adjust.

5

Review a sample every week

Someone from sales reads a set of conversations and flags errors, late escalations and answers that could be better.

6

Decide with the numbers

Adjust rules, connect the missing data or expand the launch when the KPIs and quality hold up.

Benchmarks

Public data to measure with judgment

Third-party figures with published sources, to understand why each KPI matters. None of them is a Sellium result: every implementation is measured against the company's own baseline.

7x

More likely to qualify a lead when contacted within the first hour than an hour later

Source: Harvard Business Review (2011)

64%

Of customers would prefer companies did not use AI in customer service; their top concern is that reaching a person gets harder

Source: Gartner (2024)

95%

Of generative AI pilots at companies stall, with little to no measurable impact on the bottom line

Source: MIT NANDA, The GenAI Divide (2025)

70%

Of sales reps' time goes to tasks that are not selling, according to the reps themselves

Source: Salesforce, State of Sales (2024)

Caveats: the Harvard Business Review figure measures web leads contacted by phone or email, not WhatsApp conversations, and is used as a benchmark for speed. The Gartner survey covered 5,728 customers and measures customer service in general, not sales. The MIT NANDA report is based on 150 interviews, a survey of 350 employees and an analysis of 300 public deployments. The Salesforce figure is based on what sales reps themselves report.

Frequently asked questions

Frequently asked questions about AI sales agent KPIs

What sales leadership and finance ask when they need to evaluate the agent with numbers.

What is a good resolution rate for an AI sales agent?

There is no universal number: it depends on the industry, the use cases and how many conversations need a person by nature, like a negotiation. The useful reference is your own baseline and the month-over-month trend, read together with the quality of the escalations. An agent that never escalates is not better: it holds on to what it should not.

How do you calculate an AI agent's cost per conversation?

Add up the agent's monthly cost, what Meta charged for messages and the team hours spent on escalations, then divide by the conversations handled that month. Then compare it with the cost of handling the same conversation with people; the sales reps vs AI agent calculator does that math.

How do I know the agent is selling and not just answering?

By looking at conversion by stage, not just resolution: the share of conversations with purchase intent that end in a quote, order or payment, and whether average order value for the same customers holds or grows. If resolution is high and conversion is low, the agent answers but does not move the sale forward.

How often should you review the KPIs?

Weekly during the launch, when every adjustment shows up quickly, and monthly once things settle. Keep the review of a conversation sample weekly no matter what, because it catches problems the numbers do not show yet.

Where does the data to measure the agent come from?

From your own systems. Sellium logs every conversation with its history in your CRM, and orders land in your ERP, so the KPIs are calculated with your data and can be matched against actual sales. Agree on the indicators and the baseline during implementation, before you switch the agent on.

How much does Sellium cost?

Sellium is priced in two parts: an initial implementation quoted for each company and a monthly subscription with AI credits that scale up, with no lock-in. You pay for WhatsApp messages directly to Meta, on your own account and with no markup from us, so you see that part of the cost per conversation as is. For your number, request a quote.

Let's define how to measure your agent

Tell us how you sell on WhatsApp and which numbers you track. We'll show you how Sellium would work with your data and which KPIs to agree on before launch.