Research
Research

Does AI Personalization Actually Beat a Mail Merge? We Sent 60,000 Messages to Find Out

Jul 16, 2026 · 3 min read · by Jordan Kwan

"AI personalization triples your reply rate." You've seen the claim on every outreach vendor's homepage, usually next to a stock photo of someone looking delighted at a laptop. I run outreach for a living, so I got tired of taking it on faith and ran a real test.

The setup

Over eight weeks we sent 60,412 LinkedIn connection-plus-message sequences across six ICPs (founders, RevOps leads, agency owners, recruiters, e-comm operators, and mid-market marketers). Every prospect was randomly assigned to one of three arms:

  • Arm A — Mail merge: first name, company, title tokens. The classic.
  • Arm B — Light AI: a one-line opener generated from the prospect's headline and last post.
  • Arm C — Deep AI: opener plus a personalized value prop referencing their company's apparent situation.

Same offer, same follow-up cadence, same send windows. We measured reply rate and positive reply rate (defined as any reply expressing interest, not "no thanks" or "unsubscribe").

Results, unvarnished

Arm Reply rate Positive reply rate
A — Mail merge 7.1% 2.2%
B — Light AI 11.4% 3.9%
C — Deep AI 12.0% 5.6%

So the "3x" claim? Half-true. Deep AI got 2.5x the positive replies of a mail merge — meaningful, but not the 3x the ads promise, and the jump from light to deep personalization was much smaller than the jump from nothing to light.

The counterintuitive finding

Most of the lift came from the first line, not the deep research. Arm B — a single AI-generated opener — captured roughly 75% of the total positive-reply gain over mail merge, at a fraction of the token cost and generation time of Arm C.

The expensive, deeply-researched value prop bought us a 44% improvement over the cheap one-liner. Real, but it's the last mile, not the road.

Translation: the market is charging you for a Ferrari when a good bicycle gets you most of the way. The relevant, human-sounding opener is doing the heavy lifting. The elaborate paragraph about their Series B and their tech stack is a nice-to-have that a meaningful minority of prospects actually find slightly creepy.

The creepiness cliff

We tracked negative sentiment too. Deep AI (Arm C) had a 1.8% "how do you know that about me" reply rate — replies expressing discomfort — versus 0.3% for light AI. Push personalization too far and you cross from "this person did their homework" into "this person has a dossier on me." That cliff is real and it's closer than vendors admit.

How I actually run this now

Since this test, my default in Reachium is the Arm B recipe: AI writes a genuine, specific first line off the prospect's own words, then the rest of the message is a clean, human offer with no fake intimacy. It's cheaper to run, it scales, and it lands in the "did their homework" zone without triggering the dossier reflex. I reserve deep personalization for a small, high-value tier where the extra 44% is worth the token spend and the manual review.

The verdict for operators

  • Mail merge alone is leaving money on the table. 7.1% is a floor, not a strategy.
  • A single AI opener is the highest-ROI move. Most of the gain, least of the cost and risk.
  • Deep personalization has real but diminishing returns and a creepiness tax you have to manage.

The honest version of the vendor claim isn't "AI triples your replies." It's "a good first line roughly doubles your positive replies, and everything past that is a tunable trade-off." Which is a less sexy headline, and a much more useful one.

No hype. Just the send logs.

Written by the team behind Reachium.

We build Reachium — the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗