The short answer
Automate the expensive omission, not the visible irritation. On the current evidence the first place to start is inbound response, where a pre-sale chatbot raised sales by 16.3% among 44,614 consumers in one of seven randomised experiments run at a single retail platform (Fang et al., working paper, October 2025). The last place to start is firm-initiated outreach: push messaging to 13.7 million consumers in the same programme returned +1.6%, which was not statistically significant.
Most teams pick the opposite end. There is a structural reason for that.
Why does the annoying task so rarely turn out to be the expensive one?
Annoyance is a measure of friction and repetition. Cost is a measure of consequence. They come apart because the failures that cost real money are failures of omission, and omissions have no friction at all.
The report you rebuild every Monday irritates you for forty minutes a week. The enquiry that sat unanswered for nine hours produced no sensation whatsoever. Nothing happened, and nothing kept happening, quietly.
So the rule is simple to state and uncomfortable to apply. If it makes you sigh, it is annoying. If you did not know it happened, it is expensive.
Where should you actually start? A ranking by evidence
| Rank | Function | Evidence grade | Best available study | Measured effect |
|---|---|---|---|---|
| 1 | Inbound response to customer-initiated contact | Proven | Fang et al., a randomised trial of 44,614 consumers | +16.3% on sales |
| 2 | Onboarding and early activation | Proven in one field experiment, though the intervention was human-delivered | Retana, Forman and Wu, M&SOM 2016, 366 treated of 2,673 | First-week churn halved, week-one support questions down 19.6%, eight-month usage up 46.6% |
| 3 | Quote and pricing support | Promising | Karlinsky-Shichor and Netzer, Marketing Science 2024, 17 reps, 67,851 quotes | +7.8% profit for the hybrid, +4.9% for full automation |
| 4 | CRM hygiene and record capture | Promising, unmeasured | No independent effect study exists | Unknown |
| 5 | Firm-initiated follow-up and outbound | Weak | Fang et al., same 7 trials | No measurable effect |
| 6 | Outreach personalisation | Weak | Field test, 76,977 recipients | +0.43 percentage points |
| 7 | Proactive at-risk retention contact | Documented harm | Ascarza, Iyengar and Schleicher, JMR 2016 | Churn rose from 6% to 10% |
The ranking follows the strength of the evidence. The order most teams build in follows what is interesting to build. Those two orderings are close to unrelated.
Why does inbound response rank first?
Because the effect was isolated in a randomised trial and the AI was the thing being varied. The quote-pricing result in row three has the same property; inbound ranks above it on the size of the measured effect. Fang and colleagues ran seven separate randomised experiments at a single cross-border retail platform, with arms ranging from about 30,000 to 13.7 million consumers. The pre-sale chatbot, tested on 44,614 consumers, raised sales by 16.3%. Push messaging, tested on 13.7 million, returned +1.6% and was not statistically significant. The authors report the experiments separately and do not use the terms customer-initiated or firm-initiated; that reading of the pattern is mine.
That asymmetry is the most useful pattern in the set. Someone who has just raised their hand has already spent their attention. You are not buying it, you are answering it.
Inbound response is also where omission is cheapest to detect and most expensive to leave alone. You already paid for the enquiry, and an enquiry nobody replies to is spend you do not collect.
One distinction worth keeping straight, because most vendors blur it. What Fang measured is whether an enquiry gets answered at all. How fast it gets answered is a separate claim, and the familiar five-minute rule behind it rests on two studies from 2007 and 2011 funded by a company selling lead-response software, never independently replicated. Build for coverage. Treat any specific minute threshold as a working assumption of your own.
Why does onboarding rank second?
Retana, Forman and Wu ran a field experiment at a cloud infrastructure provider, published in M&SOM in 2016: 2,673 new customers, of whom 366 were treated. Proactive education delivered in the first days after signup halved first-week churn, cut week-one support questions by 19.6%, and raised accumulated usage across eight months by 46.6%. The intervention was delivered by people, before current AI tools existed; what it establishes is that the moment matters, not that software can occupy it.
Three separate outcomes moved in the right direction from one intervention, and the accumulated usage gain was still visible at eight months, though the authors state the direct effect decays within a week. Very little in revenue operations has that shape of evidence behind it.
It ranks second rather than first only because it is one experiment in one setting, while the inbound result pools seven.
Note also what the intervention was. Education delivered at the moment of highest confusion, before the customer had asked for anything. That is proactive help, which is a different thing from proactive selling. The distinction matters, and the seventh row of the table is where it shows.
Why is quote and pricing support only promising?
The result is good but narrow. Karlinsky-Shichor and Netzer, in Marketing Science 2024, studied 17 representatives and 67,851 quotes. The hybrid, where the model recommends and the representative decides, returned 7.8% higher profit. Full automation returned 4.9%. The assistance was worth $14.85 per quote for low-expertise representatives against $5.21 for high-expertise ones.
Seventeen representatives is a small number of people, however large the quote count. That is why this is promising rather than proven.
Buyers agree that the area matters. Nielsen Norman Group's B2B research, with 79 participants across 179 sites, found pricing ranked highest in buyer priorities, 29% above product availability. The demand side is not in doubt. The measured lift is simply based on one firm.
Why does CRM hygiene rank fourth when nothing has been measured?
Because the mechanism is obvious and the evidence is absent, and both of those facts should be said out loud.
Capture and record maintenance sit underneath everything above them. A routing rule cannot fire on a contact nobody entered. A forecast cannot weigh a stakeholder who appears nowhere. But no independent study measures how much automated capture improves data completeness, or what that improvement is worth in closed revenue.
Rank it fourth, build it, and measure it yourself. Do not buy it on the strength of a claim nobody has tested.
What should you not automate first?
Firm-initiated outreach, on the same evidence that puts inbound first. In Fang and colleagues' seven trials, firm-initiated AI produced nothing. The personalisation premium in a separate field test of 76,977 recipients came to 0.43 percentage points, which is real and far too small to build a first project around.
Proactive retention contact is worse than neutral. Ascarza, Iyengar and Schleicher, in the Journal of Marketing Research 2016, ran a retention intervention on at-risk customers and watched churn rise from 6% to 10%. Reminding people they were considering leaving helped them decide.
Lemmens and Gupta, in Marketing Science 2020, showed the repair. Ranking customers by profit-weighted response rather than by churn probability produced at least a 4% increase in firm profit at identical spend. The outreach was not the fault. The targeting rule was.
How do you choose between the top two?
Count what you already pay for. If you spend on demand generation and some share of enquiries routinely goes without a substantive reply, start with inbound response. You are buying attention and then failing to collect it. Pick your own reply-time target; no independent study sets one.
If enquiries are already answered quickly, and the leak is customers who sign and then go quiet in the first fortnight, start with onboarding. Pull last year's churn and look at when it happened. If a meaningful share left in the first month, the experiment above is the closest thing to a tested answer available.
What this ranking does not tell you
It does not tell you the size of your effect. Every figure above was measured in someone else's firm, with their traffic and their buyers.
It also does not tell you whether the problem is automatable at all. Working through this often surfaces a pricing problem or a positioning problem wearing an operations costume. That is a good month.
Pick the one number that should move. Build one thing. Read the number in two weeks.