A Switzerland-based outbound agency entering US B2B tech needed to know which messaging works before scaling. A 9-variant testing harness answered that with clean data and near-zero bounces.
Swiss B2B outbound agency expanding into the US tech market.
No proof of which messaging angle lands with US buyers.
9-variant AI copy test harness on pristine deliverability infrastructure.
Statistically clean read on messaging at 0.53% bounce; the reply data told us the offer, not the copy, was the constraint.
Context
A Switzerland-based B2B outbound agency that builds cold email campaigns for mid-market tech companies. They had a proven track record running campaigns for their clients, but had never turned the same methodology inward, and they needed their own outbound engine to sell lead generation services into the US B2B tech market.
The target was specific: companies with 50 to 1,000 employees, headquartered on the US East Coast plus Florida, Indiana, and the DC metro area. The ICP was decision-makers at B2B tech firms actively scaling their go-to-market function but lacking the infrastructure to do cold outbound well.
Constraints
The agency needed to demonstrate outbound sophistication to prospects who evaluate outbound sophistication for a living. Every operational choice had to be defensible, because the campaign itself doubled as a live portfolio piece.
Architecture
One engine, four layers. A qualified US B2B tech lead pool flows through a deliverability layer, into a 9-variant test matrix, out across 7,963 sends, and finally into reading what the data actually says.
A Clay companies table with 47 enrichment and qualification columns: domain normalization and deduplication, AI ICP-fit scoring against 6 criteria as a pass or fail gate, funding and growth signal detection, and exclusion layers for DNC lists, competitors, and existing clients.
The people table ran 59 columns and a 12-provider sequential waterfall (BetterContact, Findymail, Hunter, Prospeo, Kitt, Datagma, Wiza, Icypeas, Enrow, Dropcontact, LeadMagic, SMARTe) with a Findymail verification gate on every address. Result: 0.53% bounce across 7,963 sends, in a market where 2 to 3% is acceptable.
A two-agent setup: a signal-finding agent found what makes outbound relevant now, then a writing agent drafted the opener. 9 variants on Step 1, 3 on Step 2, a plain-text bump on Step 3, and 3 breakup variants on Step 4, across a 4-step plain-text sequence with tracking disabled.
A capped cadence of 250 leads per day, Monday to Friday, 9 AM to 6 PM Eastern, so every variant got a clean, comparable read. Then I read the reply data honestly rather than forcing a win out of it.
Results
The infrastructure worked exactly as designed: near-zero bounce, nine messaging angles read cleanly against each other. The finding was not a high reply rate. It was that at a 0.77% reply rate with 3 interested replies, the copy was not the bottleneck. The offer and positioning needed repositioning before scale. That is the honest read, and it is the point: I would rather hand a client a clean answer than an inflated one.
| Metric | Result |
|---|---|
| Total qualified leads loaded | 2,222 |
| Emails sent (4-step sequence) | 7,963 |
| Bounce rate | 0.53% |
| Reply rate (per unique lead) | 0.77% |
| Interested replies | 3 |
| Reading of that data | Offer, not copy |
| Sequence completion rate | 72% (1,600 of 2,222) |
| A/B variants tested | 9 (Step 1) + 3 (Step 2) + 3 (Step 4) |
| Company enrichment columns | 47 |
| Contact enrichment columns | 59 |
| Email providers in waterfall | 12 |
Want a clean read on your outbound before you scale it?
Book a 15-min callFit
Stack
See how a clean, multi-variant test on pristine deliverability infrastructure tells you what actually works, before you spend on volume.
Book Your Free Strategy Call