What AI actually changes in B2B sales (and what it does not)
Three places where AI produces measurable lift in revenue work, two where it reliably makes things worse, and a test for telling them apart before you buy.
The claim that AI writes better cold emails than a person is mostly false, and it is also the wrong question. The interesting question is which parts of revenue work are bottlenecked by something a language model is genuinely good at.
Three are. Two are not.
Where it works
Research compression
A rep preparing ten accounts properly spends most of a morning reading. Funding history, recent hires, product launches, the technology on the site, what the CEO said on a podcast. It is slow, it is necessary, and it is the first thing dropped when the pipeline is thin — which is exactly when it matters most.
This is the strongest case for AI in the stack, because it is a reading problem rather than a writing problem. Summarising twelve sources into six sentences of relevant context is close to the centre of what these models do well, and a wrong summary is caught immediately by the rep who reads it.
First-draft generation
Not final copy. A draft that a rep edits in ninety seconds instead of writing in nine minutes.
The distinction is not pedantic. Sent unedited, generated email converges on a recognisable register — the same three-sentence structure, the same "I noticed that…" opener, the same false specificity. Prospects have learned to recognise it, and recognition is fatal. Used as a draft with a human edit, the same output saves real time and keeps the voice.
Reply triage
A shared inbox with four hundred unread messages contains perhaps twelve that need a response today. Classifying replies into interested, not now, wrong person, unsubscribe and out-of-office is a genuinely well-suited task: high volume, low stakes per item, and errors that are cheap to correct.
Where it does not
Targeting
"Find me accounts like my best customers" sounds like a machine-learning problem and is usually a data problem. The model will happily produce a list. Whether that list is good depends entirely on whether your definition of "best customer" is encoded anywhere, and in most organisations it is not — it lives in the head of one AE who is about to leave.
Fix the definition first. A clearly specified ICP with three firmographic filters beats a sophisticated model over an unspecified one, every time.
Autonomous sending
An agent that researches, writes and sends without a human in the loop is the demo that sells the product and the feature that damages the domain. The failure mode is not that it writes something offensive. It is that it writes something slightly wrong — the wrong company, a stale job title, a funding round from three years ago — at a volume that a human would never reach, and each one is a burned account.
The value of a human in the loop is not that they catch bad writing. It is that they catch confident errors, which is the failure mode that scales.
A test before you buy
Ask a vendor these four questions. The answers separate real tooling from a wrapper.
- What does it do when it does not know? A system that never says "I could not find anything about this account" is fabricating, and you will find out from a prospect.
- Can I see what it read? Grounded output cites its sources. Ungrounded output is a plausible-sounding guess.
- Is every generation logged against the record it touched? Without an audit trail you cannot answer "why did we send that?" when someone asks, and someone will ask.
- What are the guardrails, and who sets them? Sending limits, forbidden claims, required review thresholds. A system with no configurable guardrails has decided your risk tolerance for you.
The realistic framing
AI in revenue work is not a step change in effectiveness. It is a meaningful reduction in the cost of doing the preparation that good reps already knew they should be doing, and a reduction in the cost of the administrative work around it.
That is a smaller claim than the market makes and a much more defensible one. Teams that adopt it as a compression of preparation time do well. Teams that adopt it as a replacement for judgment produce more output, worse results, and a domain reputation problem within a quarter.