Case Study Swiss Outbound Agency AI Copy Testing

9 AI copy variants tested across 7,963 cold emails at a 0.53% bounce rate

A Switzerland-based outbound agency entering US B2B tech needed to know which messaging works before scaling. A 9-variant testing harness answered that with clean data and near-zero bounces.

7,963 Emails Sent
9 Copy Variants Tested
0.53% Bounce Rate
2,222 Leads Engaged

The Client

Swiss B2B outbound agency expanding into the US tech market.

The Problem

No proof of which messaging angle lands with US buyers.

The Build

9-variant AI copy test harness on pristine deliverability infrastructure.

The Outcome

Statistically clean read on messaging at 0.53% bounce; the reply data told us the offer, not the copy, was the constraint.

Context

The client

A Switzerland-based B2B outbound agency that builds cold email campaigns for mid-market tech companies. They had a proven track record running campaigns for their clients, but had never turned the same methodology inward, and they needed their own outbound engine to sell lead generation services into the US B2B tech market.

The target was specific: companies with 50 to 1,000 employees, headquartered on the US East Coast plus Florida, Indiana, and the DC metro area. The ICP was decision-makers at B2B tech firms actively scaling their go-to-market function but lacking the infrastructure to do cold outbound well.

Constraints

Why this was hard

The agency needed to demonstrate outbound sophistication to prospects who evaluate outbound sophistication for a living. Every operational choice had to be defensible, because the campaign itself doubled as a live portfolio piece.

Architecture

The system

One engine, four layers. A qualified US B2B tech lead pool flows through a deliverability layer, into a 9-variant test matrix, out across 7,963 sends, and finally into reading what the data actually says.

Hand-drawn system map: a qualified US B2B tech lead pool feeding a deliverability layer at 0.53 percent bounce, then a 9-variant copy test matrix, out across 7,963 sends, ending in reading the data to find the offer, not the copy, was the constraint
The full system, as I'd sketch it on a whiteboard. Click to open full size.
LAYER 01

Company intelligence pipeline

A Clay companies table with 47 enrichment and qualification columns: domain normalization and deduplication, AI ICP-fit scoring against 6 criteria as a pass or fail gate, funding and growth signal detection, and exclusion layers for DNC lists, competitors, and existing clients.

LAYER 02

12-provider email waterfall

The people table ran 59 columns and a 12-provider sequential waterfall (BetterContact, Findymail, Hunter, Prospeo, Kitt, Datagma, Wiza, Icypeas, Enrow, Dropcontact, LeadMagic, SMARTe) with a Findymail verification gate on every address. Result: 0.53% bounce across 7,963 sends, in a market where 2 to 3% is acceptable.

LAYER 03

9-variant AI copy test matrix

A two-agent setup: a signal-finding agent found what makes outbound relevant now, then a writing agent drafted the opener. 9 variants on Step 1, 3 on Step 2, a plain-text bump on Step 3, and 3 breakup variants on Step 4, across a 4-step plain-text sequence with tracking disabled.

LAYER 04

Controlled send and read

A capped cadence of 250 leads per day, Monday to Friday, 9 AM to 6 PM Eastern, so every variant got a clean, comparable read. Then I read the reply data honestly rather than forcing a win out of it.

A clean test beats a pretty number. 0.53% bounce across 7,963 sends meant the messaging read was trustworthy, and the read pointed at the offer, not the copy.

Results

What the test showed

The infrastructure worked exactly as designed: near-zero bounce, nine messaging angles read cleanly against each other. The finding was not a high reply rate. It was that at a 0.77% reply rate with 3 interested replies, the copy was not the bottleneck. The offer and positioning needed repositioning before scale. That is the honest read, and it is the point: I would rather hand a client a clean answer than an inflated one.

0.53% Bounce rate, well under the 2 to 3% acceptable range
9 Copy variants read cleanly against each other
7,963 Emails sent across a 4-step sequence
Metric Result
Total qualified leads loaded2,222
Emails sent (4-step sequence)7,963
Bounce rate0.53%
Reply rate (per unique lead)0.77%
Interested replies3
Reading of that dataOffer, not copy
Sequence completion rate72% (1,600 of 2,222)
A/B variants tested9 (Step 1) + 3 (Step 2) + 3 (Step 4)
Company enrichment columns47
Contact enrichment columns59
Email providers in waterfall12

Want a clean read on your outbound before you scale it?

Book a 15-min call

Fit

Who this is for

Stack

Tools used

Clay Logo
Clay
Data Enrichment
SmartLead Logo
SmartLead
Email Sequencing
Apollo Logo
Apollo
Lead Database
Claude AI Logo
Claude AI
AI Personalization
BetterContact Logo
BetterContact
Email Enrichment
Findymail Logo
Findymail
Email Verification

Ready to test your messaging before you scale it?

See how a clean, multi-variant test on pristine deliverability infrastructure tells you what actually works, before you spend on volume.

Book Your Free Strategy Call