The short answer

More than it does for a large team, in percentage terms. Four separate studies now find the same pattern: the gain concentrates where existing capability is lowest, and it shrinks or reverses at the top of the distribution. A two-person team sits at the bottom of the capacity distribution by construction, which is the strongest argument available. The caveat is that all four measured people already doing the job full-time. Nobody has measured what happens when the work simply is not being done.

Why does team size change the answer?

Because the finding that has replicated most often is about skill distribution.

Brynjolfsson, Li and Raymond, published in the Quarterly Journal of Economics in 2025, tracked 5,179 customer support agents on a live deployment. Average productivity rose 14%. Novices rose 34%. Experienced agents gained close to nothing.

Dell'Acqua and colleagues ran a field experiment with 758 consultants and found a 43% quality improvement for below-average performers, against a much smaller gain for those already performing well.

Karlinsky-Shichor and Netzer, in Marketing Science 2024, studied 17 reps and 67,851 quotes and put a currency figure on the same effect. The pricing system was worth $14.85 per quote for low-expertise reps and $5.21 for high-expertise ones. Nearly three times the value from the same software. The only variable was who held it.

The fourth study runs the pattern to its conclusion. METR, in 2025, gave AI tooling to 16 experienced open-source developers across 246 tasks. They were 19% slower with it. They believed they were 20% faster.

A large sales team buys AI to raise a floor it already has. A two-person team buys it to have a floor at all.

What does the evidence actually say by capability level?

Where the person sitsMeasured effectSource
Novice+34% outputBrynjolfsson, Li & Raymond, QJE 2025, 5,179 support agents
Below-average performer+43% qualityDell'Acqua et al., field experiment, 758 consultants
Low-expertise rep$14.85 value per quoteKarlinsky-Shichor & Netzer, Marketing Science 2024, 17 reps, 67,851 quotes
High-expertise rep$5.21 value per quoteSame study
Average across a mixed population+14% outputBrynjolfsson, Li & Raymond, QJE 2025
Experienced specialist on familiar work19% slower, self-reported 20% fasterMETR 2025, 16 developers, 246 tasks

Four studies, four populations, one direction. The value is inversely proportional to the capability already present.

What is the gap in that evidence?

Every one of those studies measured people who were already doing the job. Full-time support agents. Full-time consultants. The comparison was always AI-assisted work against unassisted work by the same kind of worker.

A two-person team is not that case. The two-person case is that outbound does not happen on Thursday, and the follow-up to the March conversation never went out at all. The studies measured slow work made faster. A founder is asking about absent work made present.

Nobody has published a study on that. The measured comparison is unassisted work against assisted work by the same people: +14% on average, +34% for novices. The comparison a founder cares about is zero against something, and the studies are silent on it.

If value rises as capability falls, the case with no capacity at all should sit at the favourable end of that pattern. Nobody has measured it. The four studies above all ran inside organisations where the work was already being done by someone.

What should a two-person team build first?

The work that does not require you to be right, only to be present.

Coverage of open threads. A pass every morning over every live conversation. Anything that has gone quiet gets a next message drafted into your drafts folder rather than sent to the prospect. You open six drafts, send four, delete two. Four minutes covers what memory covered badly.

Research and list construction. Building a properly researched target list is several hours of work that a two-person team defers indefinitely. It is also work with a checkable output, which makes it a good first delegation.

CRM capture from live conversations, so the record exists without anyone typing it.

Notice what those have in common. Each one is currently scoring zero, and each produces an output you can inspect in under a minute.

What should it not do?

Send anything customer-facing without a human on it.

Dell'Acqua's experiment found the counterpart to the productivity result. On tasks requiring judgment, AI users were 19 percentage points less likely to reach the correct answer, while producing work that looked better. Better-presented and more often wrong is a dangerous combination when your total addressable market is small enough to burn.

And a two-person team burns it faster than anyone. Gartner surveyed 210 chief sales officers in early 2026 and found 25% reporting a return of 50% or more, alongside 20% reporting a negative return of 50% or more. That distribution is not a rounding error. It is two different implementations.

Does the maths work at your size?

Compare against what the alternative costs rather than against zero.

SaaS Capital's 2026 survey of more than 1,000 private B2B SaaS companies puts sales spend at 12% to 15% of revenue, and 12% at the $3-5M ARR mark. The documented mid-market band for outsourced appointment setting is $300 to $600 per qualified meeting, with ProspectOut publishing $100 per qualified appointment on no retainer.

There is no measured cliff below which outbound stops working. Acquisition economics do not fall off a cliff at any single deal size. Benchmarkit's 2025 benchmarks, from 583 companies reporting full-year financials, put the new-customer cost-of-acquisition ratio worst in the $25,000 to $50,000 band at about $2.40 spent per dollar of new ARR, and roughly $2.20 in the $10,000 to $25,000 band below it. Efficiency improves above $50,000 and keeps improving above $250,000. The hard band is the middle: deals large enough to need a human selling motion, too small to pay for one. A $15,000 ACV product at ten meetings a month and a 10% close rate produces about $144,000 in revenue against a fully loaded SDR at $110,000 to $127,000, before any AE time. That is the arithmetic that keeps small teams doing outbound themselves, badly, between other work.

The US Census Bureau's BTOS data (CES-WP-26-25) shows 18% of US firms using AI overall, with adoption at 32% among firms of 100 to 249 employees and under 20%, flat, among firms below 20 employees. Among adopters, sales and marketing is the most common function at 52%. Only 2% report labour reductions.

That last figure matters for a two-person team. The measured effect of adoption is more work done. Headcount did not move. You are not replacing a hire you were never going to make.

What actually goes wrong

Not the software. CBIZ's 2026 Mid-Market Pulse, with more than 500 respondents, found that 48% name lack of internal expertise as the biggest barrier. Wharton and GBK surveyed around 800 enterprise decision-makers and found roughly 44% of generative AI budget goes to people and change management rather than to software.

For a two-person team that translates directly. Whatever you build, one of the two of you has to run it every morning. If neither of you will open the queue, the queue is a subscription.

An agent does not make you a better seller. It makes the selling happen on the days you were not going to sell.

See how Atrium builds this →