← All Growth Drops
·9 min read

The Claude Code Loop That Runs Cold Email Without You Writing Prompts

claude-codelead-generationai

Watch the full video: The Claude Code Loop That Runs Cold Email Without You Writing Prompts

Three out of ten automatic campaigns got leads last month. That is not a failure rate. That is a testing machine.

We send 8 million emails a month and generate 200-300 positive replies per day. The bottleneck is not sending. The bottleneck is building campaigns fast enough to find the three that work. Claude Code solved that with a loop we run on every new ICP.

This is not an install tutorial. This is the operator loop: sample, grade, lock, deploy, schedule.

The Loop in Five Steps

  1. 10-row sample — generate copy for 10 real contacts from the ICP
  2. Grade to zero edits 3x — if you touch the copy, it failed. Run again until three consecutive passes with zero edits
  3. Lock the CTA — once the body is clean, freeze the call-to-action. Do not let the model rewrite it on the next run
  4. Cheap worker — frontier model writes the prompt, cheap model runs the rows
  5. Scheduled goal mode — let Claude Code run overnight on the full list

The ICP prompt loop is documented here: ICP prompt loop post.

Frontier Model Writes, Cheap Model Works

"Frontier model = prompt engineer, cheap model = worker."

Model split post

We do not run 80,000 rows through Opus. We run 10 rows through Opus to get the prompt right, then hand the prompt to a cheap model for the bulk work.

Cost hack: we ran 80,000 rows on the Codex $200 plan. Details: 80k rows on Codex.

The math only works if you stop re-prompting on every row. Write once. Grade once. Deploy everywhere.

Autoresearch Doubled the Positive Lead Rate

Once the loop is running, we layer autoresearch on top. Claude Code generates variant hypotheses, tests them on small batches, and promotes winners.

"Autoresearch doubled positive lead rate. Lock the CTA."

Autoresearch results

The numbers: 20.71 vs 10.71 replies per 1,000 after autoresearch locked in the winning variant. The CTA lock is non-negotiable. If the model rewrites "worth a quick call?" into "would you be open to a 15-minute discovery session?" on row 47,000, you just broke the test.

What 3/10 Automatic Campaigns Looks Like

Fable ran ten campaigns through the loop. Three got leads. Trove, Revolution Parts, and Luxury Presence are real clients — I am not naming which three won because the point is the ratio, not the names.

"3/10 automatic campaigns got leads. Trove, Revolution Parts, Luxury Presence — never name the company."

Fable results post

Seven campaigns failed. That is fine. We spent an hour on each, not a week. The loop killed the losers fast.

Dale (lakefront hotels) was a manual win from the same system: Dale lakefront hotels post.

CES 3565: The Benchmark

We track replies per 1,000 sends as our north star. Campaign CES 3565 is the internal benchmark we compare everything against: CES 3565 post.

If a new campaign cannot beat CES 3565 after autoresearch, it does not graduate to the full list. Simple gate. No politics.

Skills, Not Prompts

Everything above lives in skills, not one-off prompts. Our public repo: coldoutboundskills.

Skills we run on every campaign:

  • ICP definition — who gets the email and why
  • Copy generation — the 10-row sample + grade loop
  • CTA lock — frozen call-to-action after three clean passes
  • List routing — which segment gets which variant
  • Scoring — replies per 1k, positive reply rate, bounce rate

Claude Code reads the skill, calls Clay functions for enrichment, and pushes to Smartlead. You are not copy-pasting between tools.

For the Clay side of this (functions, Audiences, CLI taste), see Stop Building Generic Clay Tables.

Videos Worth Watching

What I Would Do This Week

  1. Pick one ICP you already know converts
  2. Run the 10-row sample in Claude Code with your existing Clay table as context
  3. Grade until three consecutive zero-edit passes
  4. Lock the CTA
  5. Deploy to 500 rows with the cheap model
  6. Measure replies per 1k against your internal benchmark

If you want us to run the first campaign through this loop for free, apply at coldoutbound.com.

Key Takeaways

  • 3/10 is a win — automatic campaign generation means fast failure, not slow perfection
  • 10-row sample + grade to zero edits 3x — the quality gate before any bulk send
  • Lock the CTA — do not let the model rewrite your call-to-action on row 47,000
  • Frontier writes, cheap works — 80k rows on a $200 Codex plan is real
  • Autoresearch doubles reply rates — 20.71 vs 10.71 replies per 1k
  • Skills repo: github.com/growthenginenowoslawski/coldoutboundskills

Want results like these for your business?

Apply for a free test campaign and see positive responses before you commit.

Apply for Free Test →