Most AI email features optimise against open rate, and open rate stopped being a measurement in 2021.
When Apple Mail Privacy Protection began pre-fetching images regardless of whether anybody opened the message, a large and unknowable share of reported opens became machine activity. Corporate security gateways add more, fetching images and following links before delivery as a scanning step.
So a model trained on open data is learning something about proxy fetch schedules alongside something about human behaviour, and it cannot tell you which is which.
That single fact disqualifies more of this category than anything else, and almost no vendor page mentions it. What follows separates the features that work from the ones built on a broken signal.
The three things AI does well in email
All three are before the send, and all three are checkable by a person.
1. Producing variants to test. Twenty subject lines instead of three. The generation was always the tedious part, and having more options is a genuine improvement — provided the test that follows can actually distinguish them.
2. Drafting the scaffolding. Structure, transitions, the standard sections. The parts of an email that are true of everyone. Where the line sits.
3. Summarising and classifying at volume. Replies into themes, survey responses into segments, support tickets into the questions worth writing an email about. This is the most under-used AI application in email and the one with the clearest value — it turns a pile of unread replies into a content plan.
Notice what is absent: anything that decides who receives a message, when it is sent, or what it says about a specific person. Those are the features being sold hardest, and they are the ones resting on measurement that no longer works.
The measurement problem, stated properly
Three things now contribute to a reported open, and only one of them is a person.
- A human opening the message
- Privacy protection pre-fetching images whether or not the message was opened
- A security gateway fetching content during pre-delivery scanning
The proportions differ by list and are not separable from the outside, which is the part that matters. You cannot subtract the machine opens because you cannot identify them reliably.
Clicks are cleaner and not clean. Corporate security gateways follow links before delivery to check where they lead, and that visit is logged as a click. On a B2B list this produces clicks from recipients who never saw the email, sometimes on every link in it.
What this rules out:
- Open-rate-based optimisation of any kind — subject lines, send times, engagement scores
- Re-engagement segments built on opens, which will include engaged people and exclude disengaged ones
- Any AI feature whose training signal is opens, which is most of them
What still measures something:
- Replies
- Purchases, bookings, sign-ups — anything that happens after the click
- Unsubscribes and complaints, which are unambiguous
- Delivery and bounce data
The practical rule: optimise against what happened on your site, not what happened in the inbox. What to measure instead.
What damages delivery
Four ways AI use degrades email performance, and none of them is about the writing quality.
1. Volume
AI makes producing campaigns cheap, and cheap production leads to more sends. More sends to a list that is not more engaged means falling engagement rates, and engagement is a signal receiving providers weigh when deciding placement.
The compounding version: engagement falls, so placement worsens, so fewer people see the messages, so engagement falls further. A gradual decline that looks like list fatigue and is partly self-inflicted. How the cycle works.
2. Re-engagement campaigns generated at scale
Automating outreach to your least engaged contacts is the single most reliable way to damage a sending reputation.
That segment contains dead addresses, people who forgot subscribing, and recycled spam traps — former real addresses that providers repurposed precisely because nothing legitimate should still be mailing them. What traps do.
AI does not cause this. It makes it easier to do at volume, which is the same thing in practice.
3. Personalisation that misfires visibly
A wrong personalised detail is worse than no personalisation. A model inferring a customer's industry, role or interest from thin data produces confident, specific, wrong statements — and the recipient reads it as either careless or creepy.
Complaints follow from creepy faster than from generic, and complaints are the metric providers act on directly. Google's bulk sender requirements state complaint rates should stay below 0.3%. What is worth personalising.
4. Sending on someone else's domain
Some AI email tools send from their own domain rather than yours. That builds sending reputation for the tool, not for you, and leaving means starting from nothing with none of the history transferring.
The thing that is not true
Spam filters do not detect AI-written content, and there is no credible basis for the claim that they do.
Filtering decisions weigh sender reputation, authentication, recipient engagement, complaint history, and content signals that have been the same for years — link patterns, image-to-text ratio, known-bad phrases, domain reputation of anything linked.
"AI-written" is not among them, and no provider has published anything suggesting otherwise.
What is true is adjacent and gets conflated: generated marketing prose tends toward a register that overlaps with promotional language filters already weigh — superlatives, urgency, stacked benefit claims. That is a content-signal problem that predates AI entirely, and a human writing the same way triggers the same signals. What actually determines placement.
Where the real gains are
Three, and none is a feature you buy.
1. Turning replies into content. Every list generates replies nobody reads systematically. Classify a year of them and you have a content plan built from what your audience actually asked — which is also material no competitor has.
2. Writing fewer, better emails. The saving from faster drafting is best spent on quality per send rather than more sends. Given that volume is the thing that damages delivery, this is the only use of the saving that does not carry a cost.
3. Testing properly. More variants are only useful with a test that can distinguish them, and most small-list tests cannot — a difference of a few percent on a list of a few thousand is inside the noise. Test big differences, not small ones, and treat a marginal result as no result. What is worth testing.
A defensible position
Six rules that survive whatever launches next.
- Use AI before the send, not at the send. Draft, generate variants, classify. Never decide recipients or timing
- Never optimise against opens. The signal contains machine activity you cannot separate out
- Keep sending on your own domain, always
- Do not increase volume because production got cheaper
- Personalise only from data the customer gave you. Not inferred, not enriched
- A person reads every campaign before it sends. Out loud, once
And measure at the destination. Replies, purchases, sign-ups, unsubscribes. Everything that happens inside the inbox is now partly machine-generated and partly hidden.
Frequently asked questions
Can spam filters detect AI-written emails?
No, and no provider has published anything suggesting they try. Filtering weighs sender reputation, authentication, recipient engagement, complaint history and long-standing content signals. What gets conflated with this is that generated marketing prose tends toward promotional language filters already weigh — which affects human writing identically.
Does AI improve email open rates?
The question cannot be answered reliably, because open rate stopped being a clean measurement when privacy protection began pre-fetching images regardless of whether anyone opened the message. Security gateways add more machine opens. Any tool claiming to optimise opens is optimising a signal it cannot separate from proxy activity.
What should AI not be used for in email marketing?
Deciding who receives a message, when it sends, or what it claims about a specific person. Those features generally train on open data, which now contains machine activity, and a wrong personalised detail produces complaints faster than a generic message does.
Does using AI hurt email deliverability?
Not directly. What hurts is what AI makes easy: more sends to a list that is not more engaged, automated re-engagement to the least engaged segment, and personalisation that misfires. All three damage engagement and complaint rates, which providers weigh.
What is the best use of AI in email marketing?
Classifying replies and support questions into a content plan. Every list generates replies nobody reads systematically, and a year of them is material about your own audience that no competitor has. It is the most under-used application and the clearest value.
Can AI write my whole email campaign?
It can write the scaffolding — structure, transitions, standard sections. What it cannot write is the part that makes the email worth sending, which comes from what you know about your customers. A person should read every campaign before it sends.
Is AI-generated personalisation worth it?
Only from data the customer gave you. Inferred details — industry, role, interest guessed from thin signals — produce confident wrong statements, and recipients read those as careless or intrusive. Complaints follow from intrusive faster than from generic.
Free: The 60-Minute Email Authentication Fix
A no-fluff checklist to set up SPF, DKIM & DMARC correctly and pass Gmail & Yahoo's sender requirements.

Muhammad Basim has worked in digital marketing since 2013, focused on email deliverability and AI-assisted content production. He is the author of the Email Deliverability Playbook and the Email Copywriting Playbook, and has run 100+ email campaigns for ecommerce brands, coaches, and B2B senders. He writes about email, SEO, WordPress, and AI — with a bias toward what can be tested over what sounds good.
Related Articles
What to Give AI Before You Ask
Generic output is almost always a context problem, not a prompt problem. A model with nothing specific to work from produces the average of what has been written on the subject. That is not a flaw to be worked around with better phrasing — it is the correct behaviour given no information. The fix is […]
Connecting Your Tools: Automation That Does Not Break
Every integration you add is a thing that can fail without telling you. That is the property that makes connected systems different from the tools they connect. A broken tool announces itself — you open it and something is wrong. A broken integration produces silence, which is indistinguishable from a quiet week. And the failure […]
AI Email Personalisation: What Is Worth Doing
Personalisation built from what a customer told you works. Personalisation built from what a model inferred about them produces complaints. That line is the entire subject, and it does not move as the technology does. The risk is not that the inference is bad. It is that it is confident, specific and occasionally wrong — […]