Asset 20 8 2
Does AI recommend your business? Run the free check →

Join 15,000 business owners, marketers and entrepreneurs. The Sunday newsletter you'll be annoyed only arrives once a week.

Article

Week 17 Flop: My Hands-Free AI Cold Email Automation Was an Epic Failure

Well.

If you want to go deeper on this: AI News This Week for Small Business, 2 August 2026.

This is embarrassing.

My AI cold email automation failure did not produce the case study I had planned. It produced Tom, a bearded stock photograph and a list of mistakes long enough to need its own contents page.

Before I tell you how I managed that, here are the first 16 experiments in this rebuild:

  1. Week 1: I de-indexed 1,300 pages to save my website
  2. Week 2: My open rate crashed to 11%
  3. Week 3: The unglamorous SEO work nobody talks about
  4. Week 4: I triaged 1,219 crawled, not indexed pages
  5. Week 5: I took my open rate from 11% to 70%
  6. Week 6: I conquered my inbox instead of my blog posts
  7. Week 7: I rebuilt the income stream that saved my business three times
  8. Week 8: I let Claude loose on my website and went to do my hair
  9. Week 9: I redesigned my newsletter and the click rate imploded
  10. Week 10: I built a tool to automate WordPress publishing
  11. Week 11: I dug 786 de-indexed pages back out
  12. Week 12: AI inbox management gave me my life back
  13. Week 13: The month my website started paying me again
  14. Week 14: My rebuild in public numbers, the month traffic crossed half a million
  15. Week 15: A stranger proved I was getting found by AI
  16. Week 16: How I set up cold email with AI in 2026

You can browse all of them in the Business Experiments, Breakdowns and Results archive.

And now for Week 17.

Last week I proudly showed you how I set up cold email with AI in 2026.

More than 20 mailboxes. Seven sending domains. Warm-up. List cleaning. AI research. AI-written sequences. Automated sending. Reply monitoring. The lot.

It was the most complicated AI experiment I have built so far. It was also a very good example of what an AI implementation project involves once it leaves the prompt box.

It has also been the biggest flop.

Not because nothing worked.

That would have been easier.

Enough of it worked to start emailing real people under my name while the mistakes hid underneath.

One email went out from a man called Tom.

Several mailboxes showed my name next to a stock photograph of a bearded man in a suit.

A live change sent the same opening email to more than 500 people for a second time.

More than a thousand emails opened with “Hi there” because AI had saved the word “there” as the person’s first name.

My reporting told me I had ignored people I had already answered.

The campaign is still running. The financial result is disappointing. I have not turned it around, and I am not going to put a shiny ending on it because I would quite like to keep some dignity.

This is Week 17 of my rebuild in public.

Week 16 was the build.

Week 17 is the flop.

I nearly did not share it.

I kept opening the draft, reading the first few paragraphs and closing it again.

I am trying to grow my AI implementation business. Am I seriously going to publish an article about an AI implementation that called me Tom?

It did not feel like a brilliant positioning move.

The sensible option was to wait. Fix everything. Get a result. Come back with a beautiful before-and-after story where I look as if I knew what I was doing from the beginning.

That would also be a lie.

I did not run this experiment on a client. I ran it inside my own business because that is where I test the limits before I tell anybody else what is safe.

If I only share the weeks where AI makes me look clever, this stops being a rebuild in public and turns into a marketing brochure.

So you are getting the embarrassing week.

AI cold email automation failure compared side by side, showing the 21-mailbox system that was built against the wrong sender name, duplicate sends and broken greetings that arrived
Everything on the left worked. The column on the right is what a recipient received.

The short answer: can AI run cold email hands-free?

My AI cold email automation experiment says no.

AI handled an enormous amount of the work. It built lists, checked addresses, helped write the emails, connected tools, watched campaign data and found faults much faster than I could have done alone.

I also let it go too far.

I did not check every mailbox. I did not send myself a real email from each one. I did not look at the photo, the name, the address and the full thread the way a recipient would see them.

I was hands-free on purpose.

That is how I run these experiments. I push AI as far as I can, keep myself out of the way and find the point where it needs me.

Most weeks, that approach saves me days.

This time it let a bearded man called Tom sell sponsored posts on my behalf.

I think we can call that finding the point.

Why I ran the experiment this way

I have been testing AI inside my business for months.

I do not want another list of clever prompts. I want to know whether AI can take a real job, inside a real business, and get it done.

That means I try to be as hands-free as I can.

If I step in every five minutes, rewrite every sentence and check every setting, I learn what AI can do with me babysitting it. That is useful, but it is not the experiment I am interested in.

I want to know:

  • What can AI handle without me?
  • Where does it start making mistakes?
  • Which mistakes are harmless?
  • Which ones reach customers, prospects or my bank account?
  • Where does a human need to stay in the process?

That approach has worked well with content, SEO, research, reporting and parts of my inbox. My AI content workflow experiment was far easier to keep hands-free because the work waited in a draft before anybody else saw it.

Outreach is different.

An SEO mistake can sit on a draft page until I catch it.

An email mistake arrives in somebody else’s inbox wearing my name.

That is the bit I underestimated.

I thought checking the work would interfere with the test.

It would have been part of a well-built test.

NIST’s AI Risk Management Framework puts human oversight, testing, monitoring and clear responsibility inside the AI lifecycle. Its practical AI RMF Playbook also treats monitoring and human oversight as part of the system, not an optional tidy-up afterwards. I had the AI. I had the tools. I had a large amount of automation.

I had left myself outside it.

Why this AI cold email experiment became such a mess

Cold email looks simple when you reduce it to “find people and send emails”.

Mine involved:

  • seven domains
  • more than 20 mailboxes
  • DNS records and authentication
  • mailbox warm-up
  • sender names and profile photos
  • thousands of contacts from different sources
  • list verification
  • several campaigns and sequences
  • timed follow-ups
  • reply forwarding
  • Gmail replies
  • campaign reporting
  • order and payment tracking
  • monitoring jobs running in the background
Five-stage AI cold email failure cascade covering identity, data, live changes, handoffs and monitoring
The mistakes appeared between tools: identity, data, live changes, handoffs and monitoring.

Each part had its own settings, limits and idea of the truth.

AI could work on every individual part.

The mistakes appeared between them.

The sending platform knew somebody had replied but could not see that I had answered in Gmail.

The mailbox settings showed my display name but did not show me the stock photograph a recipient saw.

The campaign editor accepted a new sequence but did not stop existing contacts from restarting at email one.

The dashboard said ACTIVE even when the campaign was not sending.

I kept checking whether the pieces worked.

I did not stand at the other end and receive the finished email.

That was the stupid mistake underneath most of the other stupid mistakes.

The full horror show: 16 AI cold email automation mistakes

I wish this list were shorter.

Sixteen AI cold email automation mistakes grouped into identity, data and copy, campaign controls, handoffs and reporting
The 16 mistakes clustered around five weak implementation layers.

1. AI sent an email as Tom Parker

The sending address contained the name Tom Parker. The display name should have been Lilach Bullock.

One mailbox reverted.

A recipient replied:

“Hey Tom.”

I had checked the email copy. I had not checked the identity wrapped around it.

2. My name appeared beside a bearded man’s photograph

Several mailboxes had inherited a stock profile picture.

The display name said Lilach Bullock.

The photograph said middle-aged man who would like to discuss your pension.

I had more than 20 mailboxes and I had not opened a real message from every one.

If I had, I would have found this in minutes.

3. I sent the opening email twice to more than 500 people

I changed a live sequence.

The sending platform treated the replacement as a new sequence and restarted existing contacts at the beginning.

The audit found more than 500 people received the opening email again.

You know that awful feeling when you send one email to the wrong person?

Try multiplying it by more than 500.

4. AI saved “there” as 2,766 people’s first name

Campaign B stored the literal word “there” in the first-name field for 2,766 leads.

So the personal opening was:

“Hi there,”

The next sentence reminded them that we had spoken before.

Apparently not memorably enough for me to know their name.

5. I used a July deadline in emails scheduled for August

The offer said prices would disappear at the end of July.

Campaign B still had 2,712 people waiting to receive email one. With the sending cap and follow-up delays, the sequence would keep talking about a July deadline in the middle of August.

That is not urgency. It is a calendar having a nervous breakdown.

Five AI cold email fields a recipient saw first: sender name, profile photograph, greeting, deadline and unsubscribe route, each showing the wrong value
The copy was checked. The five fields wrapped around it were not.

6. I sent 3,960 emails with one subject line

One subject line.

No split test.

Three thousand, nine hundred and sixty emails.

Calling that an experiment is generous. I committed very hard to the first idea.

7. I protected deliverability and left myself unable to diagnose the result

Open tracking and click tracking were off.

No reply could mean:

  • the subject line failed
  • the email was never seen
  • the offer was weak
  • the person was not interested

All four looked the same in the report.

I had switched both off on purpose because tracking can hurt cold-email deliverability.

That decision protected one part of the experiment and blinded another.

I had automated the sending without choosing a clean way to learn what was failing.

8. I presumed the unsubscribe link would be automatic

I presumed the sending platform would add an unsubscribe link automatically.

Because this was a hands-free test, I did not inspect a finished email to confirm it.

It did not.

The emails did tell people to reply if they wanted me to stop, so there was an opt-out route. There was not the one-click link I thought the platform would add.

That distinction matters. The ICO’s current B2B marketing guidance says businesses must not hide their identity and must give a valid address for people to opt out. The ICO also has detailed electronic mail marketing guidance, while RFC 8058 explains how one-click unsubscribe works inside email systems.

The footer met the first standard.

It did not give people the easier second option.

Google’s current email sender guidelines recommend an easy unsubscribe route for all senders and require one-click unsubscribe for marketing senders above its bulk threshold. The M3AAWG sender best practices also say the process should be clear and easy.

9. I kept chasing people after they replied

Six people in Campaign A had replied but were still moving through the follow-up sequence.

One had paid.

One wanted to talk.

One had said no.

The machine was ready to chase all three because I had not made “any reply stops the cold sequence” a hard rule.

10. My report accused me of ignoring 53 people I had answered

Replies came through the sending platform.

I replied through Gmail.

The sending platform could not see Gmail Sent, so it kept showing the other person as the last sender.

When I joined the two records, I found 53 contacts I had already answered.

AI had not invented the wrong answer. I had given it half the conversation.

11. The AI-built setup left 85 replies uncategorised

I did not know the sending platform had a reply-category field.

The AI-built setup did not create that step or tell me it was missing.

The 85 replies still existed in Gmail and in the campaign records. What I did not have was a clean list separating:

  • not interested
  • come back later
  • question
  • interested
  • ordered
  • paid

The sending platform therefore could not tell me which conversations deserved attention.

This does not mean every reply was ignored. It means the reporting layer did not know the difference between a no, a maybe, a question and money in the bank.

12. I put nine people into two campaigns

Nine contacts appeared in both Campaign A and Campaign B.

One campaign told them they had bought from me before.

The other told them we had spoken but never worked together.

Both cannot be true.

The suppression check lived inside each list. It needed to sit across the whole operation.

13. The reply forwarder crashed 23 times

The script that moved replies into my main Gmail logged 23 crashes.

One timeout or connection reset stopped the run.

No retry. No backoff. No separate warning.

The most important bridge in the setup could fall over and wait for me to notice.

AI cold email reporting split across three systems: the sending platform, Gmail and the revenue record, each blind to what the others knew
No single system could see the whole conversation, so the reporting was confidently wrong.

14. The monitoring went to sleep with my Mac

I built a job to check the campaigns every 15 minutes.

Then I found hours-long gaps overnight because the job stopped when my Mac slept.

My automated night watchman was tucked up in bed beside the laptop.

15. The LinkedIn check died after 3 of 46 profiles

The LinkedIn targeting audit had 46 profiles to verify.

It checked three, rejected all three and crashed on the other 43.

The file ended with the words “THIS IS THE DAMAGE REPORT”.

Cheery.

16. I tracked activity better than money

The sending platform showed sends and replies.

It did not have the full revenue, costs, orders or payment state joined to the right campaign.

That makes it easy to stare at busy-looking numbers while the commercial result remains poor.

I built a machine that could tell me how much it had done.

I had not given it one clean place to tell me whether the work was worth doing.

The mistake that caused all the others

I wanted a pure test.

Could AI build and run a cold outreach operation with me as hands-free as possible?

I thought human checking would weaken the result.

I had that backwards.

A business-ready AI system includes human checks in the dangerous places. Removing them does not prove the AI is capable. It proves the workflow has no brakes.

The ICO’s guidance on human review recommends defined review procedures, test plans, logs and a manual fallback when automated processing falls below an acceptable level.

That sounds boring.

So does checking the face attached to your email.

I would now choose boring.

That human layer does not make the implementation slower. It stops the expensive rework that makes an apparently quick build drag on. My guide to how long AI implementation takes explains why testing and adoption belong in the timeline from the beginning.

AI cold email workflow dividing automated work from human approval at the point where work reaches another person
AI can own the heavy lifting. A human still approves the edge where the work reaches another person.

What I would do if I built it again

I would still use AI.

That may sound unhinged after the previous section, but AI was not useless here. It did weeks of research, data cleaning, drafting, checking and technical work.

I gave it the wrong job boundaries.

My new version would work like this.

AI does the heavy lifting

I would let AI:

  • gather and clean contact data
  • flag duplicates
  • verify and enrich records
  • research companies
  • draft email options
  • check claims against a dated source
  • summarise replies
  • prepare reports
  • watch for abnormal sending, bounces and stopped jobs

These are high-volume tasks where AI saves an absurd amount of time.

My guide to AI lead generation in 2026 already makes this distinction: AI is excellent at the boring layers. It needs a human around the message, the conversation and the judgement calls.

I proved my own point by ignoring it.

For the writing layer, I would also use a controlled cold outreach prompt pack instead of letting one generated version become the campaign by default.

A human owns anything another person sees

I would personally approve:

  • the sender name, address, photograph and signature
  • one real test email from every mailbox
  • the first live batch
  • every claim involving money, dates or performance
  • any change to a live sequence
  • the rules that stop or restart sending
  • replies that involve interest, objections, complaints or payment
  • the point where volume increases

This does not mean reading every one of thousands of emails forever.

It means checking the edges where a mistake leaves the software and reaches a person.

That is also how I now think about AI workflow automation for small businesses. The amount of human involvement should depend on the damage a bad output can cause.

It is also why an AI implementation coach should be looking at the handoffs and failure points, not simply showing a business which buttons to press.

I would send a canary through the whole thing

Before a real prospect received anything, I would add my own email address to each campaign.

Then I would open the message on my phone and laptop.

I would check:

  • who it says the email is from
  • the photograph
  • the address
  • the first name
  • the links
  • the unsubscribe route
  • the signature
  • what happens when I reply
  • what the report says afterwards

One email travelling through the complete workflow would have caught Tom, the bearded photograph, broken greetings and missing reporting.

One.

Seven-point AI cold email canary test checking sender name, photo, greeting, links, unsubscribe, reply stop and reporting
One complete seed email can catch the mistakes that separate tool checks miss.

I would lock live campaigns

Once a sequence sends its first email, nobody edits it.

Any change creates a clone.

Only untouched contacts move into the clone.

That rule would have saved more than 500 people from receiving my introduction twice.

I would make the machine prove it is safe before it grows

The sending volume would rise after several stable days.

The system would stop when bounce or complaint rates crossed a fixed ceiling.

It could not restart itself until the evidence was safe.

I would also sample the real emails every day, not stare at the campaign dashboard and assume a green label meant all was well.

If you want the wider process, my guide to building an AI workflow for a small business explains why you should test one narrow workflow against real examples before adding the next layer.

The same idea appears in the NIST AI RMF Playbook’s monitoring guidance: riskier systems deserve more oversight, and monitoring has to continue after deployment.

I did the opposite here.

I built the octopus first and checked its tentacles afterwards.

Six-step AI outreach failure recovery loop from stopping the campaign to resuming with a small batch
Recovery means testing the complete workflow before volume returns.

What I have fixed and what is still a mess

I am continuing the experiment because stopping it now would only tell me how to build a bad system.

I want to find out whether the repaired version can justify itself.

I have spent most of this week trying to turn it around.

Partly because I hate losing.

Partly because I am mortified.

And partly because more than 20 warmed mailboxes represent time and money I do not want to throw in the bin.

Wanting it to work does not make it work. The campaign still has to earn the right to continue.

Fixed or improved

  • All 21 mailbox display names now show Lilach Bullock.
  • AI now checks the sender names and repairs any that drift.
  • Twelve stock photographs were removed. Ten mailboxes showed the correct initials afterwards.
  • Seed addresses now receive live emails from Campaigns B and C.
  • The remaining addresses are going through verification.
  • The reply report checks Gmail Sent before saying I owe somebody a reply.
  • The reply forwarder now retries temporary failures.
  • Six people who had replied were removed from the follow-up queue.
  • Campaign C’s deadline was corrected before its first send.
  • Editing a live sequence is now banned.

Still open

  • Two mailbox photographs still need a recipient-view check.
  • Some email addresses still contain fake persona names even though the display name is mine.
  • Campaign B still needs a safe clone to remove its July deadline.
  • Click tracking still needs switching on.
  • A one-click unsubscribe link still needs adding.
  • The monitoring still stops when the Mac sleeps.
  • The LinkedIn targeting and reply jobs are unfinished.
  • Reply categories, revenue and full costs still need joining into one report.

The system is less dangerous than it was.

It has not earned a success story.

AI cold email automation repair scoreboard showing ten fixed items such as corrected sender names against eight open items such as the missing one-click unsubscribe
Ten repairs are done. Eight remain, which is why this is not a recovery story yet.

The rule I am taking into every AI experiment now

I will keep running hands-free AI tests.

That is how I find out what these tools can do when nobody is leaning over them.

I am changing where I put the human.

AI can work alone when a mistake is private, reversible and easy to spot.

I stay in the loop when a mistake:

  • reaches another person
  • spends money
  • changes live data
  • makes a claim under my name
  • can damage a relationship
  • multiplies before I can stop it

That is the dividing line.

Dividing line for AI outreach showing when AI can work alone on private reversible work and when a human stays in the loop because the mistake reaches another person
AI works alone when a mistake stays private. A human stays in when it leaves the building.

Content research can wait in a draft.

Cold outreach lands in a real inbox.

I can be hands-free while AI does the work.

I cannot be absent from the parts that carry my name.

If you are deciding where AI belongs in your business, start with what AI implementation means in practice and these AI workflow examples for small businesses. Both include the human review step I managed to prove by leaving it out.

My AI outreach safety check

Copy this before you let AI contact anybody on your behalf.

Three AI outreach safety gates for the first send, increasing volume and changing a live campaign
A campaign earns more freedom only after it passes each safety gate.

Before the first send

  • Send a real test from every mailbox.
  • Check the name, address, photograph and signature.
  • Open it on desktop and mobile.
  • Reply and inspect the full thread.
  • Test every link and the unsubscribe route.
  • Check every number against a dated source.
  • Make sure any reply stops the cold sequence.

Before increasing volume

  • Review a sample of real sent emails.
  • Check bounce, complaint and reply quality.
  • Confirm the campaign sent during the intended window.
  • Confirm the report includes replies sent through every channel.
  • Check that existing customers, previous replies and opt-outs are suppressed everywhere.

Before changing anything live

  • Clone the campaign.
  • Move untouched contacts only.
  • Send the clone to your seed inbox.
  • Keep a written record of the change.
  • Make sure there is a kill switch.

You can use AI to perform many of those checks.

A human still has to look at the finished result and decide whether it should leave the building.

Frequently asked questions about AI cold email automation

Can AI automate cold email?

AI can automate research, list cleaning, enrichment, draft writing, sequence setup, monitoring and reporting. A human should approve the sender identity, the first live emails, performance claims, changes to active sequences and important replies. My hands-free test failed because I removed those checkpoints.

Why did this AI cold email campaign fail?

The campaign combined more than 20 mailboxes, seven domains, several data sources, Gmail, a sending platform and background monitoring. Each tool saw one part of the job. I did not test the complete recipient experience before sending at scale, so identity, duplicate-send and reporting mistakes reached real inboxes.

What was the worst AI outreach mistake?

The most embarrassing mistake was sending as Tom Parker with the wrong photograph. The largest mistake was restarting the opening sequence and sending it again to more than 500 people. Both would have been caught by sending one test email through the complete workflow.

Does this mean businesses should avoid AI outreach?

No. Businesses should avoid unsupervised AI outreach. AI can remove hours of research and admin, but the human needs to own identity, judgement, approval and recovery. My guide to cold emails that get replies explains where AI helps and where it hurts.

Is the campaign working now?

The campaign is active and the setup is safer. The financial result remains disappointing and several repairs are still open. I will publish the final result when the campaigns finish. I will not call replies buyers or turn an unfinished test into a win.

The final word

I wanted this article to be the results post.

I pictured a lovely graph, some neat conversion numbers and a smug paragraph about waking up to orders.

Instead, I got Tom.

I got the bearded man.

I got more than 500 duplicate emails, more than a thousand “Hi there” greetings and a report accusing me of ignoring people I had answered.

I also got the clearest answer I have had from any of my AI experiments.

AI can do far more of the work than most businesses realise.

Once it starts talking to another person under my name, I need to be there.

I am still trying to repair this campaign. It may recover. It may end up being an expensive way to learn that cold outreach is the wrong lane for this offer.

Either way, you will get the real ending.

For now, Week 17 is the flop.

And Tom is unemployed.

If you want help deciding what AI can handle inside your business, where a human needs to stay involved and how to build the checks before a mistake reaches a customer, this is the AI implementation work I do.

If you want the final result of this experiment, join the newsletter. I will share it when the evidence exists.

For the bigger picture, see my full guide to AI marketing.

Related reading: How to Improve SEO on a WordPress Website (Without Wasting Six Months on a Plugin) and Feeling Like Your Content’s Invisible in 2023? Here’s the What, Why, and How to Measure Your Content.

I take guest contributions on this topic, so you can write for us about email marketing.

Published and maintained by the Lilach Bullock team, covering marketing, AI and business growth.
Your buyers are asking AI who to use. Does it say you?

See for free whether ChatGPT, Claude, Perplexity, Gemini and Google name you, and get the plan to become the answer.

Check my AI visibility →
Sundays only

Get the Sunday newsletter.

One email a week. AI experiments, marketing tactics, and the workflows Lilach is building right now in her own business.

Subscribe free

Let’s get your marketing running on AI.

Book a free 30-minute call

We figure out what you need, where AI fits in, and what working together would look like.

Book the call →

Or take the 30-second calculator

You’ll see the hours and the money quietly leaking out of your week, and the three workflows worth building first.

Take the calculator →

Or grab the free AI resource library

Prompt packs, templates, checklists, and swipe files. The exact tools I build for paying clients. Yours, free.

Get the library →
Keep reading

More from the blog.