
We Used AI to Improve Cold Outreach. The Emails Weren’t the Important Part.
We analyzed 20,000 company websites and sent 3,570 cold emails to find out what matters more. Better targeting or better personalization. The biggest improvement came before the email was even written.
0.36%
That's an inconvenient number to start with. It's the reply rate the web design agency It'sJustContent got from a cold email campaign to 1,382 Dutch companies. Five positive replies. Three sales calls. Zero new clients. It'sJustContent spent real time on that campaign. The internet answered with a shrug.
The easy explanation is: "Cold email doesn't work." We didn't believe that.
So we had to formulate a better question. Which companies should we even be sending an email to?
Suspects, Prospects, and 184 Hours Nobody Has
There's a useful bit of B2B terminology here. Everyone you could theoretically email is a suspect. A prospect is a suspect who actually matches criteria you've defined in advance. The 1,382 companies were suspects, nothing more. Nobody had looked at a single one of their websites before hitting send. It'sJustContent was shooting in the dark and praying somebody would respond. That's no strategy. That's buying lottery tickets and praying one of them pays out.
Looking at every individual website would have helped. It would also have taken roughly 184 hours. That's about eight minutes per website. Nobody at an agency has an extra 184 hours lying around. It doesn't scale.
This is the part where you'd expect me to say we used AI to automate it. We did. What matters more is what we built it to measure.
Building a Judge Out of 192 Imperfect Clues
A website doesn't come with a "please redesign me" label. It comes with scattered, weak signals. A slow website. A confusing service page. A contact form buried three clicks deep. Text that's hard to read on mobile. None of these alone proves anything. Combine them consistently across thousands of sites, and they start to say something real.
We grouped those cues into three domains: technical-functional system quality, visual presentation, and information quality. Underneath those sit ten families. Things like response performance, navigation and mobile usability, visual hierarchy, typographic readability, and clarity of the service proposition.
In total, 192 individual cues. Three instruments score them together. Google Lighthouse and axe-core handle the technical and accessibility checks. Fast, deterministic, boring in the best way. An AI model handles the visual and content analysis.
Together they form what the research literature calls an automated judge. It combines cues into a score. Think of a doctor combining symptoms into a diagnosis without ever seeing the disease directly.
We ran that judge against 20,000 Dutch company websites. Count every underlying cue and it's 4.66 million data points. The top 20% of websites with the most problems became our prospects.
Building a scoring model is easy to believe in and easy to be wrong about. So we tested it the only way that counts. We sent real emails and counted real replies.
We split companies into three groups. Group A was a random sample from the bottom 80%. Group B came from the top 20%. Actual prospects, receiving the same generic email as Group A. Group C also came from the top 20%, except we personalized the email.
Over five working days, across 24 mailboxes, we sent all 3,570 emails and waited fourteen days.
Here's what came back:
- Group A - unselected: 0.50% positive replies
- Group B - selected, generic: 1.60% positive replies
- Group C - selected, personalized: 2.35% positive replies
Selection worked, and it wasn't close. Companies in our top 20% were roughly 3.2 times more likely to reply positively. That's the headline. An automated judge built from website signals successfully separated a warmer audience from a colder one.
Personalization is the more nuanced part of this story. The number moved in the right direction: 2.35% versus 1.60%. But the confidence interval still crossed zero and the difference was not statistically significant (p = 0.238). The evidence isn't strong enough yet to call personalization proven.
Personalization is nearly free, though. Personalizing 1,190 emails cost about €5. The entire experiment cost €242.25, covering all 3,570 emails end to end. At €5.16 per positive reply, that's a fraction of what It'sJustContent's earlier campaign cost per result. And that earlier campaign didn't even know which companies deserved an email in the first place.
From an Agency Secret to a Global Engine
The data showed what we suspected. Targeting made the real difference, more than the copy itself. Knowing that is one thing. Running 192 technical, visual, and content checks across thousands of websites is another.
We built this automated judge to solve a pipeline problem. The numbers told a bigger story. Agencies around the world spend time and money emailing suspects when they could be identifying prospects first.
Stop Emailing Suspects. Start Closing Prospects.
What started as an attempt to fix a 0.36% reply rate is now a full application. Leadodo is used today by more than 50 agencies across 12+ countries. It's changed how they approach cold outreach.
Remember the title of this post? We Used AI to Improve Cold Outreach. The Emails Weren't the Important Part. We meant it. You can write the best email copy in the world. Send it to a company whose website is already flawless, and your pitch is just noise.
Building this kind of system from scratch takes time and money. We already built it, so you don't have to. Stop buying lottery tickets. Stop praying for replies. It's time to reach the companies that actually need your help.