KPIs to measure an AI sales agent:what to track and how to read it
An AI agent selling on WhatsApp leaves more data behind than any sales rep: every conversation is logged, along with its outcome. The risk is not running out of numbers, it is watching the ones that do not matter. Here are the four KPIs that tell you whether the agent is selling, how each one is calculated and how to read them without fooling yourself.
What KPIs should you track for an AI sales agent?
An AI sales agent is measured with four KPIs: resolution without a person (the share of conversations that end well with no human involved), conversion (how many end in a quote, an order or a payment), response time (the first reply and the time to resolution) and cost per conversation (the agent, Meta's messaging fees and team hours, over the conversations handled).
None of them is read on its own or against a universal benchmark. You read them against your baseline, the weeks before the agent went live, and alongside two quality indicators that keep you honest: pricing or inventory errors, and escalations the customer asked for but did not get in time.
The four KPIs of an AI sales agent
Each KPI answers a different business question. Resolution tells you how much work the agent absorbs; conversion, whether that work sells; response time, whether it arrives in time; and cost per conversation, whether it pays off. Looking at them together keeps you from improving one at the expense of another.
| KPI | What it answers | How it is calculated | How to read it |
|---|---|---|---|
| Resolution without a person | How much work the agent absorbs | Conversations resolved by the agent over all conversations handled | Higher is not always better: an escalation that needed to happen is also a win |
| Conversion | Whether conversations sell | Conversations with a quote, order or payment over conversations with purchase intent | Compare by use case, and for the same customers before and after |
| Response time | Whether the customer gets the answer in time | Time to first reply and time to resolution, as a median and at the 90th percentile | Split it by hours: after-hours is where it changes the most |
| Cost per conversation | Whether the channel pays off | The agent, Meta's messaging fees and team hours on escalations, over the conversations handled | Compare it with the cost of handling the same conversation with people |
The definitions (what counts as resolved, what counts as purchase intent) are agreed before launch and do not change midstream.
Resolution without a person: the most misread KPI
The temptation is to chase the highest possible number. But an agent that never escalates is not a good agent: it is one that holds on to conversations that needed someone, like a price negotiation, a complaint or an order outside the rules. Useful resolution is the kind that ends well for the customer.
Customers notice. In Gartner's consumer survey, the top concern about AI in customer service is that it will become harder to reach a person (Gartner, 2024). That is why resolution is always read together with the quality of the handoff to a sales rep.
- Resolved: an inquiry answered with data from your systems, a quote sent, an order entered or a payment collected, without the customer asking the same thing again through another channel.
- Not resolved: a conversation the customer drops halfway, one where they asked for a person and did not get one, or one that comes back the next day about the same issue.
- Correct handoff: the agent passes what belongs to a person over to them, with the full history. It is a win, and it is measured separately.
Conversion: from conversation to sale
Resolving is not selling. An agent can answer every inquiry well and close none of them if it does not propose the quote, ask for the order or offer the payment link. Conversion measures exactly that: the share of conversations with purchase intent that end in a quote, an order or a payment.
Measure it by stage, because each one fails for different reasons; the quote-to-order stage, for example, depends on quote follow-up. To estimate how much you lose to unanswered or late-answered inquiries, use the WhatsApp lost sales calculator.
| Stage | What is measured | What holds it back |
|---|---|---|
| Inquiry to quote | Conversations with purchase intent that receive a quote | Prices the agent cannot look up, or discount rules nobody defined |
| Quote to order | Quotes that turn into orders | No follow-up: the quote sits in the chat and nobody touches it again |
| Order to payment | Paid orders, when payment is part of the flow | Payments that leave the chat, or payment links that arrive late |
| Average order value | Revenue per order for the same customers, before and after | An agent that does not suggest add-ons or volume tiers |
The cleanest attribution compares the same customers over the same time of year, before and after the agent.
Response time: the first reply and the time to resolution
With an agent, the first reply arrives in seconds, so that number stops being the interesting one. What matters is the time to resolution: how long it takes from the first message to the quote sent, the order entered or the escalation picked up by a person.
Speed matters because interest cools fast. In Harvard Business Review's study of web leads, companies that tried to contact the customer within the first hour were nearly seven times as likely to qualify the lead as those that waited even an hour longer (Harvard Business Review, 2011).
- Median and 90th percentile: the average hides the conversations that waited for hours. The 90th percentile tells you how long the worst-served customer in every ten waited.
- By hours: split business hours from after-hours. That is where the number changes the most, and where the lost sales used to be.
- Escalations separately: measure how long it takes a person to pick up what the agent escalated. If the agent replies in seconds and the escalation waits until the next day, the bottleneck has moved, not gone away.
Cost per conversation: how it adds up and what to compare it with
Cost per conversation adds up everything it takes to handle WhatsApp with the agent and divides it by the conversations handled. It has three components, and the one most often forgotten is the time of the people who take the escalations.
The number alone says little: compare it with what it costs to handle the same conversation with people. To run that math with your own data, the sales reps vs AI agent calculator sizes the team you would need to reply on time and works out the cost at which the agent pays off.
- The agent: the subscription and AI credits for the month. How that price is built is in how much a WhatsApp AI agent costs.
- Meta's messaging fees: what Meta charges based on each message's category and the customer's country. The model is in how much the WhatsApp Business API costs, and you can estimate it with the WhatsApp API cost calculator.
- Team hours: the time people spend on escalated conversations and on reviewing samples. If it climbs month after month, the agent is escalating too much.
The quality indicators that keep you honest
The four KPIs can look good with an agent that sells badly: high resolution because it never escalates, high conversion because it promises stock you do not have. That is why they come with quality indicators, reviewed just as often.
- Pricing or inventory errors: orders with a price from another list, or with items out of stock at picking time. They should be close to zero if the agent queries your systems; how to prevent them is in how to keep an AI agent from making up prices or inventory.
- Missed escalations: customers who asked to talk to a person and did not get one in time.
- Satisfaction: a short question at the end of the conversation, such as CSAT, compared with the one for the channel handled by people.
- Sample review: a set of conversations read every week by someone from sales. It finds what no number shows: a tone that does not match the brand, or an answer that is correct but incomplete.
How to read an AI agent's KPIs and what to adjust
KPIs are only useful if they trigger a decision. Measure before you start: a few weeks of baseline are enough to compare later with numbers instead of impressions. With the agent live, these are the most common signals and what they usually call for.
| Signal | What it usually means | What to adjust |
|---|---|---|
| Low resolution and many escalations about price | The agent has no access to a price list or a discount rule | Connect the missing data or write the rule |
| High resolution and low conversion | The agent answers but does not move the sale forward | Have it propose the quote, ask for the order or offer the payment link |
| Many abandoned conversations | Too many questions before giving anything useful | Shorten the qualification and answer what the customer asked first |
| Escalations that wait for hours | Nobody owns them, or there is no on-call schedule | Define who takes each type of escalation and during which hours |
| Rising cost per conversation | More marketing templates or more hours on escalations | Review the templates being sent and the escalation rules |
| Pricing or inventory errors | Data the agent reads from a copy instead of the system | Have it query the ERP in real time, not a spreadsheet |
Review the KPIs weekly during the launch and monthly after that, always with the same calculation rules.
How to build your KPI dashboard in six steps
Before you switch the agent on, not after: that way the numbers from the first weeks can be compared.
Define what resolved means
Write down what counts as a resolved, escalated or abandoned conversation, and what counts as purchase intent. Without definitions, every team measures something different.
Take the baseline
A few weeks of data from the current channel: response time, unanswered inquiries, WhatsApp orders and sales, and team hours.
Log every outcome
Each conversation goes into the CRM with its outcome: answered, quoted, order entered, paid or escalated. Every KPI comes from there.
Split by use case and hours
Inquiries, orders and quotes do not behave the same, by day or by night. An overall average hides where to adjust.
Review a sample every week
Someone from sales reads a set of conversations and flags errors, late escalations and answers that could be better.
Decide with the numbers
Adjust rules, connect the missing data or expand the launch when the KPIs and quality hold up.
Public data to measure with judgment
Third-party figures with published sources, to understand why each KPI matters. None of them is a Sellium result: every implementation is measured against the company's own baseline.
7x
More likely to qualify a lead when contacted within the first hour than an hour later
64%
Of customers would prefer companies did not use AI in customer service; their top concern is that reaching a person gets harder
95%
Of generative AI pilots at companies stall, with little to no measurable impact on the bottom line
70%
Of sales reps' time goes to tasks that are not selling, according to the reps themselves
Caveats: the Harvard Business Review figure measures web leads contacted by phone or email, not WhatsApp conversations, and is used as a benchmark for speed. The Gartner survey covered 5,728 customers and measures customer service in general, not sales. The MIT NANDA report is based on 150 interviews, a survey of 350 employees and an analysis of 300 public deployments. The Salesforce figure is based on what sales reps themselves report.
Sellium: every conversation, logged with its outcome
Sellium logs every conversation in your CRM, with its history and its outcome: the inquiry answered, the quote sent, the order entered or the handoff to a sales rep. That way your KPIs come from your own systems, not from a vendor report.
Sellium
AI sales agent for WhatsApp
Sellium is Vantegrate's AI sales agent for WhatsApp: it runs on the official API, answers with your real data, quotes, takes orders, collects payments and hands off to your team with the full history, logging everything in your CRM.
Explore SelliumSecure AI agents
Sellium runs 100% on Salesforce and Oracle Cloud infrastructure, certified SOC 2 Type II and ISO 27001. Vantegrate enables, the customer operates and certifies.
See the security modelISV and Consulting PartnerRuns inside your Salesforce
Sellium reads and writes your objects, respects your profiles and permissions, and works alongside Agentforce and your existing flows.
Vantegrate and SalesforceFrequently asked questions about AI sales agent KPIs
What sales leadership and finance ask when they need to evaluate the agent with numbers.
What is a good resolution rate for an AI sales agent?
What is a good resolution rate for an AI sales agent?
There is no universal number: it depends on the industry, the use cases and how many conversations need a person by nature, like a negotiation. The useful reference is your own baseline and the month-over-month trend, read together with the quality of the escalations. An agent that never escalates is not better: it holds on to what it should not.
How do you calculate an AI agent's cost per conversation?
How do you calculate an AI agent's cost per conversation?
Add up the agent's monthly cost, what Meta charged for messages and the team hours spent on escalations, then divide by the conversations handled that month. Then compare it with the cost of handling the same conversation with people; the sales reps vs AI agent calculator does that math.
How do I know the agent is selling and not just answering?
How do I know the agent is selling and not just answering?
By looking at conversion by stage, not just resolution: the share of conversations with purchase intent that end in a quote, order or payment, and whether average order value for the same customers holds or grows. If resolution is high and conversion is low, the agent answers but does not move the sale forward.
How often should you review the KPIs?
How often should you review the KPIs?
Weekly during the launch, when every adjustment shows up quickly, and monthly once things settle. Keep the review of a conversation sample weekly no matter what, because it catches problems the numbers do not show yet.
Where does the data to measure the agent come from?
Where does the data to measure the agent come from?
From your own systems. Sellium logs every conversation with its history in your CRM, and orders land in your ERP, so the KPIs are calculated with your data and can be matched against actual sales. Agree on the indicators and the baseline during implementation, before you switch the agent on.
How much does Sellium cost?
How much does Sellium cost?
Sellium is priced in two parts: an initial implementation quoted for each company and a monthly subscription with AI credits that scale up, with no lock-in. You pay for WhatsApp messages directly to Meta, on your own account and with no markup from us, so you see that part of the cost per conversation as is. For your number, request a quote.
Let's define how to measure your agent
Tell us how you sell on WhatsApp and which numbers you track. We'll show you how Sellium would work with your data and which KPIs to agree on before launch.