Muhammad Basim
Pin for Using AI for Email Subject Lines
Ai & Automation

Using AI for Email Subject Lines

Muhammad Basim
Muhammad Basim
·6 min read
Using AI for Email Subject Lines

Generating subject lines was never the hard part. Choosing between them was, and it got harder.

A model will produce twenty options in seconds, and twenty options is genuinely better than three. But the mechanism most people use to pick a winner — testing on open rate — stopped being a reliable measurement when privacy protection began pre-fetching images regardless of whether anyone opened the message.

So the workflow has to change at the selection step, not the generation step.


Why the test no longer works

A reported open now consists of three things: a person opening the message, privacy protection pre-fetching content, and a security gateway scanning before delivery.

The proportions differ by list and cannot be separated from outside. Which means an A/B test showing subject line A at 34% and B at 31% is comparing two numbers that each contain an unknown quantity of machine activity.

Two further problems compound it:

Machine opens are not randomly distributed across variants in any way you can verify — they follow the composition of each split, and splits are rarely composed identically.

Small differences were never significant anyway. A three-point difference on a list of a few thousand is inside the noise even with clean data. Most subject line tests never had the sample size to detect what they claimed to detect, and privacy pre-fetching made a marginal test into a meaningless one.


What to test instead

Test at the destination, and test big differences.

Measure what happened after the click, not what happened in the inbox:

  • Replies
  • Purchases, bookings, sign-ups
  • Unsubscribes, which are unambiguous and worth watching per campaign

Test approaches, not phrasings. A curiosity subject line against a direct one is a difference large enough to detect. "Your order is ready" against "Your order's ready" is not, and testing it burns a send.

Treat a marginal result as no result. If you would not bet on it, do not conclude from it. The honest output of most subject line tests is "no detectable difference", and acting on noise is worse than not testing.


Where AI actually helps

Three places, all before selection.

1. Volume of options. Twenty angles on the same email, quickly. The value is coverage — you will consider approaches you would not have reached alone.

2. Explicit variation. Ask for the same message as a question, a statement, a number, a name, and an incomplete sentence. Structured variation beats twenty similar attempts, which is what an unspecific request produces.

3. Length variants. The same idea at four words, eight and fifteen, so you can see which survives compression. Compression usually improves subject lines, and seeing the short version next to the long one makes the choice obvious.


Filtering the twenty down

A person does this, in about two minutes, using four tests.

1. Does it say something, or does it gesture at something? "An update on pricing" gestures. "Pricing goes up on the 14th" says. Specific beats intriguing far more often than the category admits.

2. Would it survive being true? If the email does not deliver what the subject line implies, the open costs you more than it earns. A subject line writing a cheque the email does not cash produces unsubscribes and complaints, which are the metrics providers actually act on.

3. Does it work with the preview text? They are read together and generated separately, which is how you get a subject line and a preview text saying the same thing twice. How to use the preview text.

4. Does it look like the last six? A list learns your patterns. Twenty AI-generated options tend to cluster around the same register, so the variation you think you have may be narrower than it looks.


What generated subject lines get wrong

Four recurring failures worth recognising on sight.

Manufactured urgency. "Don't miss out", "Last chance", "Act now" — applied to things that are not urgent. Recipients learn the pattern quickly, and it is one of the few content signals that has genuinely been weighed by filters for years.

Curiosity with nothing behind it. "You won't believe what we found" is a promise the email cannot keep, and the cost is charged at the unsubscribe.

Stacked benefits. "Save time, cut costs and grow faster" — three claims, none specific, all unverifiable.

Register drift. Slightly more enthusiastic than you are. This is the one that damages a personal-voice list most, because the mismatch between the subject line and the email is the tell.

The common fix is the same as elsewhere in generated writing: replace the general with the specific. What a subject line is doing.


A workflow

Ten minutes, and it produces better subject lines than an hour of staring.

  1. Write the email first. A subject line for an email that does not exist yet describes an intention, not a message
  2. Extract the single most useful sentence in it. Often the subject line is already written, in paragraph three
  3. Generate twenty variants, asking explicitly for structural variation — question, statement, number, name, fragment
  4. Cut to five using the four filters
  5. Write one yourself without looking at the list. It is frequently the best one, and it is the only one that could not have been generated for a competitor
  6. Pick, and check it against the preview text
  7. If testing, test approaches rather than phrasings, and measure at the destination

Frequently asked questions

Can AI write good email subject lines?
It generates useful variety, which is genuinely valuable — twenty structured options beat three. What it cannot do is choose, because choosing depends on knowing what the email actually delivers and what your list has already seen.

Why can't I A/B test subject lines on open rate any more?
Reported opens now include privacy protection pre-fetching content whether or not anyone opened, plus security gateways scanning before delivery. The proportions differ by list and cannot be separated out, so a three-point difference is comparing two numbers that each contain unknown machine activity.

What should I test subject lines against instead?
Outcomes at the destination — replies, purchases, sign-ups — plus unsubscribes, which are unambiguous. And test large differences rather than small ones: a curiosity approach against a direct one is detectable, while two phrasings of the same idea is not.

How many subject line variants should I generate?
Around twenty, asked for as explicit structural variation — question, statement, number, name, fragment — rather than twenty attempts at the same thing. Then cut to five by hand and write one yourself without looking at the list.

What do AI-generated subject lines get wrong?
Manufactured urgency applied to non-urgent things, curiosity the email cannot pay off, stacked unverifiable benefits, and a register slightly more enthusiastic than your own. The last is the most damaging on a personal-voice list.

Should the subject line be written before or after the email?
After. A subject line written first describes what you intended to send; written after, it can extract the most useful sentence the email actually contains — which is often already sitting in the third paragraph.

The short version

  1. Write the email before the subject lineWrite the email before the subject line
  2. Find the most useful single sentenceFind the most useful single sentence already in it.
  3. Generate around twenty variantsGenerate around twenty variants , requesting structural variation explicitly.
  4. Also request length variantsAlso request length variants u2014 four words, eight, fifteen.
  5. Cut to fiveCut to five : does it say something, would it survive being true, does it work with the preview text, does it look like the last six?
  6. Write one yourselfWrite one yourself without looking at the generated list.
  7. Check the chosen line against the preview textCheck the chosen line against the preview text so they do not repeat.
  8. If testing, test approaches rather than phrasingsIf testing, test approaches rather than phrasings
  9. Measure replies, purchases and unsubscribesMeasure replies, purchases and unsubscribes rather than opens.
  10. Treat a marginal result as no resultTreat a marginal result as no result

Free: The 60-Minute Email Authentication Fix

A no-fluff checklist to set up SPF, DKIM & DMARC correctly and pass Gmail & Yahoo's sender requirements.

Muhammad Basim

About the Author

Muhammad Basim

Digital Marketer & WordPress Developer

Muhammad Basim has worked in digital marketing since 2013, focused on email deliverability and AI-assisted content production. He is the author of The Email Deliverability Playbook and The Email Copywriting Playbook.

Related Articles

Newsletter

Free: The 60-Minute
Email Authentication Fix

A no-fluff checklist from the Deliverability Playbook. In one hour: set up SPF, DKIM & DMARC correctly, check your domain against blocklists, and pass Gmail & Yahoo's 2026 sender requirements.

No spam — that would be ironic. Unsubscribe anytime.