The short answer

One question decides this. What did it do to the number, measured how, against what baseline. If a provider cannot answer that with a specific figure and a named measurement period, nothing else on their site is worth reading. Gartner reviewed the market in June 2025 and found roughly 130 genuine agentic vendors among the thousands describing themselves that way.

Why does that one question do all the work?

Most of the other questions are proxies for it.

A provider can answer confidently about their model and their deliverability infrastructure without ever having improved anyone's pipeline. Those answers describe a product. The baseline question describes an outcome.

The number question cannot be answered from a feature list. It requires that somebody measured something before the system was switched on. Most did not, which is why Gartner's survey of 227 chief sales officers found that 31% name difficulty proving the return on AI tools as a top challenge for 2026.

Ask it in this exact form. Which metric moved, by how much, over what period, compared with what it was doing before. Then ask for the before number and watch what happens.

The checklist: what to ask, and why each one matters

Question to askWhy it mattersWhat a weak answer sounds like
What was the baseline before you started, and who measured it?Without a before number, every after number is an assertion."Customers typically see a 3x increase."
Define "qualified" in writing, in the contract.On per-meeting pricing, that definition decides the invoice."We use standard qualification criteria."
What is your no-show policy, and where is it written?I have not found one published. If it is not in the contract, it does not exist."We would work with you on that."
What is the total price, including data, inboxes and domains?Hidden pricing shifts the cost of the negotiation onto you."Pricing depends on your requirements."
Who is the person in the loop, and what do they approve?It tells you whether you are buying software or an outsourced team."It is fully autonomous."
Does the system disclose it is AI, and at what point?Luo et al., Marketing Science 2019, found disclosing AI identity before a sales call cut purchase rate 79.7%. Timing is a commercial decision and a compliance one."That is handled automatically."
Show me a redacted audit log for one real account.Whether every action is recorded decides whether you can debug a bad month or dispute an invoice."We can provide reporting."
What do you refuse to do?A provider with no refusals has no method."We can handle pretty much anything."
What happens to the domains and the data if I leave?Sending reputation is an asset you may only be renting."We can discuss offboarding closer to the time."
What is the shortest contract you will sign?Contract length is a proxy for their confidence in month-two results."Twelve months, that is standard."

Why does per-meeting pricing need a written definition?

Per-meeting pricing sounds like it protects the buyer, and in one respect it does. It moves the risk of a quiet month onto the provider.

It also moves the entire definitional risk onto you. Without a written definition, every disagreement about whether a meeting counted is settled by the party issuing the invoice.

Put it in the contract before signature: seniority, company size, the specific need the prospect expressed, whether they knew what the meeting was about, and what happens when they do not attend. The documented mid-market band for appointment setting runs $300 to $600 per qualified meeting, and ProspectOut publishes $100 per qualified appointment with no retainer. A spread that wide is only interpretable once you know what each provider is counting.

What does nobody in this market publish?

Two absences are consistent, and both favour the seller.

The first is the no-show policy. Meetings that do not happen are the most common source of dispute in this category, and the standard practice is to resolve it case by case, which means resolving it in the provider's favour on their timetable. Ask for a written make-good term. The reaction to the request tells you as much as the answer.

The second is price. Artisan, Qualified and Salesloft publish nothing. Other providers publish openly. Nielsen Norman Group tested 79 participants across 179 B2B sites and found pricing ranked highest in buyer priorities, 29% above product availability. Withholding the thing buyers rank first is a decision about who controls the conversation.

What should I ignore?

Model names and version numbers. Which foundation model sits underneath explains almost nothing about whether the system produces meetings, and the answer changes every few months anyway.

Integration counts. Fifty integrations you will never connect are worth less than one that writes cleanly to your CRM.

Personalisation as the headline mechanism. The measured premium for personalisation in outreach is roughly 0.43 percentage points across 76,977 sends. That effect is real and it is small. A provider selling deep personalisation as the reason their system works is selling you a rounding difference.

Reply rates quoted without a denominator, and case studies without a named measurement window.

Should you be buying an AI SDR at all?

Two checks before you shortlist anyone.

There is no measured floor for outbound viability. Acquisition economics do not fall off a cliff at any single deal size. Benchmarkit's 2025 benchmarks, from 583 companies reporting full-year financials, put the new-customer cost-of-acquisition ratio worst in the $25,000 to $50,000 band at about $2.40 spent per dollar of new ARR, and roughly $2.20 in the $10,000 to $25,000 band below it. Efficiency improves above $50,000 and keeps improving above $250,000. The hard band is the middle: deals large enough to need a human selling motion, too small to pay for one. Work it through at $15,000 ACV: ten booked meetings a month, four in five actually held, at a 10% close rate produces about $144,000 of revenue against a fully loaded SDR at $110,000 to $127,000, before any AE time. You have run a business to break even and called it growth.

Then the direction of the evidence. Fang and colleagues, across seven randomised trials covering one retail platform, found a pre-sale chatbot raised sales by 16.3% among 44,614 consumers while push messaging to 13.7 million produced nothing measurable. Outbound is firm-initiated by definition. Cover inbound response first, where the same technology has an evidence base behind it. What that evidence measured is whether a buyer who raised their hand got answered — not how many minutes it took. Outbound earns its place once the cheaper win is banked.

What is not known

There is no independent benchmark of AI SDR providers. No shared definition of a meeting exists, no audited no-show rate has been published, and no head-to-head trial has been run. Every comparison table you find, this checklist included, is assembled from vendor disclosure and observed market pricing rather than from measurement.

So treat any ranking as an opinion with footnotes, and put your trust in the one number you can verify yourself: what your pipeline did in the ninety days before, and what it did in the ninety days after.

See how Atrium builds this →