Asset 20 8 2
Does AI recommend your business? Run the free check →

Join 15,000 business owners, marketers and entrepreneurs. The Sunday newsletter you'll be annoyed only arrives once a week.

Article

AI for Email Subject Line Testing and Optimisation: What Works in 2026

The short version: AI tools can generate, score, and help you systematically test email subject lines far faster than manual A/B testing alone. But the real gains come from treating AI as a creative sparring partner and data analyst combined, not as a replacement for understanding your own audience. Done right, subject line optimisation with AI can lift open rates by 20 to 40 percent within a few campaigns.

I have been writing email subject lines professionally since before most current marketing tools existed. Twenty-one years of this. And I can tell you with complete confidence that subject lines are the single highest-use point in any email marketing programme, and also the thing most marketers treat as an afterthought.

You write the email, spend hours on the body copy, then dash off a subject line in thirty seconds right before hitting send. I did it too. We all do it.

AI has changed that calculus entirely, and not in the obvious way most articles describe.

Why subject lines deserve your most careful attention

Let me give you a number that should stop you in your tracks. According to research across billions of emails, subject lines account for roughly 47 percent of whether someone opens your email. Nearly half. The body copy, the offer, the design, the call to action, all of it is irrelevant if the subject line fails.

Yet the average marketing team spends maybe five minutes on subject line selection. Five minutes against 47 percent of the result. That is one of the most dramatic misallocations of effort in all of marketing.

If you are new to thinking seriously about email, start with the complete guide to email marketing to get the full picture before diving into the AI layer. But if you are already running campaigns and want to sharpen this specific piece, keep reading.

What AI does for subject line testing

There are three distinct jobs AI can do here, and most people only use it for one of them.

1. Generation at volume

The obvious use. You brief a tool like Claude or ChatGPT with your audience, your offer, your email's main point, and your brand tone, and you ask for thirty subject line variants. Not five. Thirty. Then you filter down to the best eight or ten for actual testing.

The reason volume matters is that subject line performance is surprisingly non-linear. Subject line number 23 in your list is often wildly different in character from number two, and sometimes it is number 23 that wins. If you only generate five variants, you are anchoring to a narrow creative space.

2. Scoring and prediction before you send

Several AI tools now claim to predict open rates before you send. Phrasee and Persado are the enterprise-grade versions of this. They train on your historic send data and score new subject lines against it. Mailchimp's Subject Line Helper uses a simpler version of the same idea.

These tools are useful but imperfect. The prediction is only as good as the training data, which means if your list is small (under about 5,000 subscribers), the scores are not reliable enough to trust blindly. You still need to test. The scoring is a shortlist tool, not a decision tool.

3. Pattern analysis after testing

This is the one almost nobody talks about, and it is where AI delivers the most durable value.

After you have run 20 or 30 A/B tests over several months, you have data. But that data is sitting in your ESP's dashboard in a format that is hard to synthesise. Which patterns win on your list? Is it urgency language? Questions? Numbers? Short versus long? Personalisation tokens? First-person versus second-person framing?

You can paste your winning and losing subject lines into a long-context AI conversation, along with their open rates, and ask for pattern analysis. I have done this with my own campaigns, and the output was sharper than anything I had spotted manually. The AI noticed that my losing subject lines clustered around abstract benefit language ("transform your marketing") while my winners almost always contained a specific deliverable or number ("47 subject line formulas that lifted my open rate by 31 percent"). That insight then shaped every subsequent test.

A real story: the campaign that taught me to take this seriously

About eighteen months ago I was relaunching a lead magnet, a content strategy guide I had updated substantially. I wrote what I thought was a solid subject line: "Your new content strategy guide is here."

Open rate: 18.4 percent. Disappointing for a warm list.

I ran the same email again to the non-openers two days later, this time testing three AI-generated variants against each other:

  • "I rewrote this guide from scratch. Here's why."
  • "The content strategy mistake I kept making (and fixed in this guide)"
  • "31 pages of content strategy, updated for 2026. Free."

The winner was the second one, "The content strategy mistake I kept making," at 34.7 percent open rate. That is 88 percent higher than my original. Same list. Same email body. Same offer. Different subject line.

What the AI generation surfaced that I had not written myself was the personal vulnerability angle. I would not naturally write that way in a subject line because it feels slightly exposed. But my audience responded to it strongly.

That result changed how I approach every subject line brief now.

The step-by-step process I use for every send

Here is exactly what I do, in order. You can adapt it to whatever tools you use.

Step 1: Write the brief, not the subject line

Before touching any AI tool, I write a three-part brief:

  • What is the one thing this email does for the reader? (Not what it announces. What it does for them.)
  • Who specifically is reading this? (Not "my list". An actual person profile.)
  • What emotion do I want them to feel when they see the subject line in their inbox?

This takes five minutes and it makes every subsequent step sharper.

Step 2: Generate thirty variants with a structured prompt

My prompt structure looks something like this:

"I'm writing an email to [specific audience] who are [context about where they are in their journey]. The email delivers [specific thing]. I want them to feel [emotion]. Generate 30 subject line variants across these styles: curiosity gaps, specific numbers, personal stories, direct benefit statements, questions, urgency without false scarcity, and contrarian angles. Keep all variants under 50 characters where possible. British English. No exclamation marks."

The style categories matter. If you just ask for "30 subject lines," you get 30 variations on the same theme. Specifying styles forces range.

Step 3: Filter to your shortlist of five to eight

Read through the thirty and eliminate anything that:

  • Uses language that does not sound like you
  • Could trigger spam filters (excessive capitalisation, certain trigger words)
  • Makes a claim your email cannot back up in the first two paragraphs
  • Feels manipulative in a way that would damage trust if it did not deliver

You will usually end up with eight to twelve that pass. Pick the five or six that feel different from each other.

Step 4: Run a structured A/B test

Most ESPs let you test two or three variants. If you have five to six good candidates, you have options. I rotate testing across sends rather than testing everything at once. Over four sends, I can test six variants against each other in a controlled way without needing a massive list.

One thing that matters: let your test run long enough. Most platforms default to four hours before declaring a winner. That is almost always too short. I use 24 hours minimum, and 48 hours for lists that skew toward business emails, where Monday morning reading patterns differ significantly from weekend behaviour.

Step 5: Log every result in a running document

Every test result goes into a single Google Doc: the subject line, the variant it beat, the open rate for each, the list size, the send date, the send time, and one line of hypothesis about why it won.

After ten tests, you feed this into an AI conversation and ask for patterns. After twenty tests, the patterns become reliable enough to shape your defaults.

The honest point most articles will not make

Here it is. AI subject line tools, including the expensive predictive ones, are trained primarily on aggregated industry data or on large list behaviour. They are quite good at identifying what works on average across millions of emails. They are significantly less good at predicting what works on your specific list, especially if that list is small, niche, or has been built around a very particular trust relationship.

My list has been built over more than two decades. The people on it have a relationship with my voice and my honesty that is different from what a generic tool models. When AI tools have predicted my winning variant incorrectly, it has almost always been because the tool scored for broad-audience patterns and missed something specific to my relationship with my readers.

The fix is not to stop using AI. The fix is to weight your own historic test data heavily and treat AI prediction scores as a tiebreaker, not a deciding factor. Your data beats their model, always, once you have enough of it.

This ties into a broader principle I write about in my post on what to automate in email marketing and what to never automate. Subject line generation is absolutely worth automating for volume. Subject line decision-making should stay in human hands, informed by your own data.

Specific tools and what they are good for

I am not going to tell you one tool is the definitive answer. Here is the honest breakdown based on what I have used:

Work with me

Want AI doing the heavy lifting in your marketing?

I build the systems that handle the boring 80 percent, so you get your week back. Done properly, with the human kept in.

Claude and ChatGPT (general LLMs)

Best for generation at volume and for pattern analysis of your historic data. Free or low-cost. You bring the context, they bring the range. The quality of output is almost entirely determined by the quality of your brief. Weak brief, mediocre subject lines. Sharp brief, surprisingly strong results.

Phrasee

Enterprise-grade. Trains on your specific sending history. useful for large-list senders (100,000 plus subscribers) because the training data is rich enough to make predictions meaningful. Overkill and too expensive for most small businesses.

Mailchimp Subject Line Helper

Accessible and free if you are on Mailchimp already. Uses a scoring system based on aggregate performance data. Good for eliminating obvious underperformers, not good enough to predict your winner. Use it as a filter, not an oracle.

CoSchedule Headline Analyzer

Originally built for blog headlines but works reasonably well for email subjects. Scores for emotional value, power words, and character length. Another filter tool rather than a predictor.

If you are building out a broader AI toolkit for marketing, my honest review of the best AI tools for marketing in 2026 covers the full stack with real opinions on what earns its place and what does not.

Subject line patterns that consistently win (and why)

Based on my own test logs and the AI-assisted pattern analysis I mentioned earlier, here are the subject line styles that have outperformed on my list repeatedly:

Specific numbers in the subject line

"7 subject line formulas" consistently outperforms "subject line formulas" without the number. Numbers signal that the email contains structured, actionable content rather than vague inspiration. My data shows roughly 12 to 18 percent higher open rates for number-led subject lines versus equivalent non-number versions, on my list specifically.

The honest admission

Subject lines that start with a personal failure or mistake perform disproportionately well on warm, trust-based lists. "The mistake I made with my email list" consistently beats "How to improve your email list." It is counterintuitive because we worry about looking weak. It works because it is honest and specific, and it signals a real story rather than generic advice.

The specific curiosity gap

There is a difference between a vague curiosity gap ("You won't believe this") and a specific one ("The open rate I got by changing one word"). The specific version outperforms every time. Vague curiosity feels like clickbait. Specific curiosity feels like a genuine promise.

Short and direct for transactional emails

For any email that is primarily transactional or time-sensitive, shorter and more direct wins. "Your July report is ready" outperforms "We've put together your monthly performance report and we'd love for you to take a look." This sounds obvious but the short version gets written far less often than it should.

What about personalisation tokens?

First-name personalisation in subject lines is one of the most overused tactics in email marketing and the data on it is mixed. Some studies show a lift. Others show no significant effect, and a few show a slight negative effect, particularly if the personalisation feels automated rather than warm.

My view: first-name personalisation in subject lines has diminishing returns as your subscribers become more sophisticated. Everyone on any email list for more than a year knows it is automated. The novelty is gone.

What does work is personalisation based on behaviour or segment, not just name. A subject line that says "Following up on the guide you downloaded" is personalised in a meaningful way because it references something the reader did. AI tools that can help you write segment-specific subject line variants are more valuable than any name token.

Building a testing culture into your process

The single biggest shift in my email marketing over the last three years has not been any specific tool or tactic. It has been treating every single send as a data point rather than a one-off event.

If you send two emails a week and you test a subject line variant on every send, you have 104 data points in a year. Feed those into an AI analysis twice a year and you have a continuously improving model of what works for your specific audience. That compounds significantly.

This is a principle that applies well beyond subject lines. If you are running automated sequences, you can apply the same logic to the subject lines within those sequences. My post on how to use email marketing automation without setting and forgetting it covers how to build in review and optimisation cycles so your automated emails do not go stale.

For solo founders and small teams who need to make every tool earn its place, the same compound testing logic applies even with limited resources. My piece on building an AI marketing stack as a solo founder goes into exactly that.

The mistakes I see most often

After helping clients with their email strategies for years, these are the errors that come up constantly:

  • Testing too small a sample to get significance. If your list has 2,000 subscribers and you split it 50/50 for an A/B test, you have 1,000 per variant. That is often not enough to reach statistical significance on open rate differences under 5 percent. Know your numbers before drawing conclusions.
  • Changing multiple variables at once. If you change the subject line, the send time, and the preview text all in the same test, you do not know what drove the result. Test one thing at a time.
  • Picking a winner after four hours. Email behaviour is not uniform across the day. Open patterns differ by morning, lunchtime, and evening. A 24-hour minimum window catches real behaviour, not noise.
  • Applying industry benchmarks to your specific list. "Average open rates for your industry" is interesting context. It is useless as a benchmark for your specific email to your specific audience. Your baseline is your own historic performance, nothing else.
  • Not writing down the hypothesis. If you do not record why you thought a subject line would win before you sent it, you cannot build real understanding. You are just collecting numbers, not learning.

Where AI cannot help you

AI cannot tell you what your audience cares about. It can help you once you know that. The discovery work, talking to subscribers, reading replies, watching which emails get forwarded, understanding what problems feel urgent versus abstract to real people, that is irreplaceably human.

I have seen brands use AI to generate technically competent subject lines that are completely wrong in tone for their audience. The tool performed well. The result was flat because the brief was built on a misunderstanding of the reader.

This is also why small business subject line optimisation is harder to generalise than enterprise email optimisation. If you are running a small business and every subscriber relationship matters, the stakes of a misread are higher. The best online marketing strategies for cash-strapped small businesses section of this site has more on balancing efficiency and authenticity when resources are tight.

The compound effect over time

Here is a simple illustration of why this matters at a business level. Suppose your current average open rate is 22 percent and you send 100 emails a year to a list of 5,000 subscribers. That is 1,100 opens per email, 110,000 total opens per year.

A sustained 25 percent improvement in open rates through systematic subject line optimisation, which is entirely achievable within six to twelve months of structured testing, takes you to 137,500 opens per year. That is 27,500 additional opportunities for clicks, conversions, replies, and relationships. From one element of the email, optimised systematically.

The compounding effect of those extra opens on revenue, list engagement scores, and deliverability (more opens signal to email providers that your list is healthy, which improves inbox placement for future sends) is significant and it is permanent, built into your process rather than reliant on any single campaign.

Free resource: The Email Subject Line Prompt Pack.

Frequently asked questions

Can AI predict which email subject line will get the highest open rate?

AI can score and rank subject line variants based on patterns from historical data, but it cannot reliably predict winners for small or niche lists where your specific audience relationship drives behaviour differently from industry averages. Use AI prediction scores to shortlist and eliminate weak options, then let real A/B testing on your list determine the winner. Your own test data is always more accurate than any predictive model for your specific audience.

How many subject line variants should I test at once?

Test two to three variants per send for practical results. Generate a larger pool, perhaps ten to fifteen, using AI, then narrow down to your two or three strongest candidates that feel different in approach. Testing more variants requires a proportionally larger list to achieve statistical significance, and most lists under 20,000 subscribers cannot support reliable five-way or six-way splits.

How long should I run an email subject line A/B test before picking a winner?

Run your test for at least 24 hours, and 48 hours if your list is made up primarily of business email addresses, where Monday behaviour differs significantly from Friday behaviour. The default four-hour windows offered by most email platforms are far too short to capture the full range of open behaviour and frequently produce false winners that do not hold up on the next send.

Does AI-generated subject line copy sound generic or obviously automated?

It can, if your prompt is generic. The quality of AI-generated subject lines is almost entirely a function of the specificity of your brief. A vague prompt produces vague output. A brief that includes your specific audience profile, the exact emotion you want to trigger, the precise deliverable in the email, and your brand voice produces subject lines that can be indistinguishable from your own best work. Always filter through your own voice before sending.

Free resource: grab The Abandoned Cart Email Swipe File from the resource library.

Related reading: Freelance Grant Writing as a Side Income: The Honest Guide Nobody Gives You and How to Price Freelance Work Without Undercharging.

Want the complete version? Read where I break down AI marketing.

Your buyers are asking AI who to use. Does it say you?

See for free whether ChatGPT, Claude, Perplexity, Gemini and Google name you, and get the plan to become the answer.

Check my AI visibility →
Sundays only

Get the Sunday newsletter.

One email a week. AI experiments, marketing tactics, and the workflows Lilach is building right now in her own business.

Subscribe free

Let’s get your marketing running on AI.

Book a free 30-minute call

We figure out what you need, where AI fits in, and what working together would look like.

Book the call →

Or take the 30-second calculator

You’ll see the hours and the money quietly leaking out of your week, and the three workflows worth building first.

Take the calculator →

Or grab the free AI resource library

Prompt packs, templates, checklists, and swipe files. The exact tools I build for paying clients. Yours, free.

Get the library →
Keep reading

More from the blog.