Ask any model for a subject line and you'll get something serviceable. Ask for twenty and you'll get two that are genuinely good, twelve that are fine, and six that are clichés you'd never send.
That ratio is the whole method. AI is good at volume and bad at judgement — so use it for the volume and supply the judgement yourself.
The mistake is asking for one and accepting it.
The short version
- Give it real context — the email, the audience, what you've sent before
- Ask for variants across named frameworks, not "good subject lines"
- Filter ruthlessly — most get cut
- Test two, from different frameworks
- Log the result so you're building knowledge rather than repeating the exercise
Step 1 — Give it something to work with
Generic input produces generic output. The context that changes the result:
The actual email. Paste it. A subject line written for a summarised topic will promise something the email doesn't deliver — and a subject line that overpromises gets the open and loses the trust, which is worse than no open.
Who it's going to, specifically. Not "subscribers" — the situation they're in. Someone who bought last month is a different reader from someone who signed up yesterday.
What you've sent recently. Pattern fatigue is real: if your emails all look the same in the inbox list, the eye stops registering them. Give the model your last five subject lines and tell it to avoid those shapes.
Your voice. Two or three subject lines you're happy with as reference. Style adjectives don't transfer; examples do.
Step 2 — Ask by framework, not by vibe
Weak: "Give me some subject lines for this email."
Strong: "Give me four subject lines each using these frameworks: curiosity gap, direct benefit, question, and contrarian. For each, keep the essential payload in the first 40 characters."
Why naming frameworks matters: unprompted, a model reaches for its default register, which is roughly the same default everyone else's prompt produces. Naming the framework forces genuinely different structural approaches rather than four rewordings of one idea.
Add the constraints explicitly:
- Under about 50 characters, with the point in the first few words — mobile shows roughly 30 to 40
- Sentence case, no ALL CAPS
- No excessive punctuation or emoji clusters
- No "Re:" or "Fwd:" prefixes on anything that isn't a reply
Those last three aren't superstition. They don't trip a word filter — they trigger people, who then ignore or report you, which reaches the same destination by a longer road.
Step 3 — The filter
Twenty variants, and most should die. Cut anything that:
Overpromises. If the email doesn't deliver what the subject line implies, you've bought one open and spent future ones. Clickbait works exactly once.
Could front any email in your category. "5 tips to improve your email marketing" is a subject line anyone could send. If it isn't specific to this email, it's noise.
Sounds like a model. You'll know it. Colon constructions, "unlock," "elevate," "the ultimate guide to."
Doesn't survive truncation. Read the first 40 characters alone. If the point arrives at character 55, mobile readers never see it.
You wouldn't open. The most reliable filter there is. Read each one cold and ask whether you'd open it from someone else.
Typical outcome: twenty in, three or four survive.
Step 4 — Test two, from different frameworks
Two, not five. More variants means smaller samples per variant, and small samples produce noise you'll misread as signal.
From different frameworks, because framework is the biggest swing. Testing two curiosity-gap variants tells you which wording won; testing curiosity against direct benefit tells you something about your audience that transfers to every future send.
Sample size matters. At least 1,000 recipients per variant, ideally 2,500-plus. Below 1,000 you're reading noise. If your whole list is under 5,000, send 50/50 across everyone — you won't get an auto-winner, but you'll get an honest data point.
Record the margin, not just the winner. A 3% difference on a small list is probably random. A 30% difference is a pattern worth acting on.
Step 5 — Keep a log
The step that turns testing into knowledge.
Record: date, email type, both subject lines, which framework each used, the winner, the margin, and one line on why you think it won.
After ten to fifteen tests, patterns emerge — and they're patterns about your audience, which no general guide and no model can tell you. That's the actual asset here. The subject lines are disposable; the knowledge isn't.
Then feed the log back in. "Here's what's won for my audience over the last six months" is far better context than any framework list, and it's context nobody else has.
Does AI know what your audience responds to?
No, and it's worth being clear about why.
A model knows what subject lines generally look like across everything it's read. It has no access to your open rates, your audience, or what you sent last Tuesday.
Which means every "high-performing subject line" claim is about the general case, and the general case is exactly what your audience has been trained to ignore by everyone else sending it.
What closes the gap is context you supply: your past winners, your audience's situation, your voice examples. The model can't discover those. It can only work with them once you provide them.
The practical implication: your logged test results are the highest-value prompt input you have, and they compound. Six months of logged tests turns a generic generator into something that produces variants shaped by evidence about your specific list.
What about generated preview text?
Worth doing in the same pass, since it's the third most visible element in the inbox and almost nobody sets it deliberately.
Ask for preview text alongside each subject line, with the instruction that it should complement rather than repeat — adding a specific detail, answering the objection, or opening a second curiosity layer.
Frequently asked questions
Are AI subject lines any good?
Individually, they're average — which makes sense, since a model produces the statistical centre of subject lines it has read. Where they become useful is as a field to choose from. Twenty variants generated in thirty seconds usually contains two or three genuinely good options you wouldn't have thought of at nine on a Tuesday. The value is in the volume plus your filtering, not in the model's judgement about which is best.
How many variants should I test?
Two, from different frameworks. More variants means smaller samples per variant, and below roughly 1,000 recipients per variant you're reading noise rather than signal. Testing across frameworks rather than within one also teaches you more — comparing curiosity against direct benefit reveals something about your audience that transfers to future sends, while comparing two curiosity variants only tells you which wording won this time.
Does AI know what my audience responds to?
No. A model knows what subject lines generally look like across everything it has read, and has no access to your open rates, your audience, or your sending history. Every "high-performing" claim it makes is about the general case, which is precisely what your subscribers have been trained to ignore. The fix is supplying context it can't discover: your past test winners, your audience's specific situation, and your own voice examples.
What to do next
Before generating anything, gather three things: your last five subject lines, your two best-performing ones, and the actual email you're writing about.
That context is what turns a generic generator into something useful — and it takes two minutes.
Then ask for twenty across four named frameworks, cut it to three, and test two.
Free: The subject line swipe file.
Related guides
- AI for email marketing — the wider division of labour
- Email subject line formulas — the frameworks to name
- Email preview text — the line that finishes the job
- An AI-assisted email workflow — where this step sits
Join the Newsletter
Get practical marketing tactics delivered straight to your inbox.

Written by
Muhammad Basim
Related Articles
What Marketing Automation Actually Costs at Scale
Nobody's automation bill jumps because they added more automations. It jumps because they added steps to workflows that were already running — a filter here, an enrichment lookup there, a second notification — and the billing model charges for every one of them, on every run. That's the mechanic behind almost every "why did this […]
One Workflow, Three Tools: A Build Comparison
Feature tables don't tell you what a tool is like to use. The only thing that does is building the same thing twice and noticing where you got annoyed. So here's one realistic workflow — lead routing with enrichment and conditional assignment — specified once and mapped across Zapier, Make, and n8n. What changes between […]
GoHighLevel: 7 Automation Workflows to Build First
GoHighLevel gives you a hundred things you could automate, which is exactly why most accounts end up with forty half-built workflows and no measurable result. Seven are worth building first. They're the ones tied directly to revenue — catching leads before they cool, recovering appointments that would otherwise vanish, and asking for reviews at the […]