The useful question about an AI agent is not whether it can do a task. It is what happens the times it does the task wrong.
Capability questions produce arguments. Failure questions produce decisions, and they stay answerable as the capabilities move.
An agent, for the purposes of this article, is a system that takes actions in a sequence rather than producing a single output — one that uses tools, reads results, and decides what to do next.
The two properties that decide everything
Reversibility and verifiability. Almost every sensible boundary follows from these two.
Reversible means a wrong result can be undone before it costs anything. A draft is reversible. A sent email is not.
Verifiable means you can check the work faster than doing it yourself. A list of URLs is verifiable in a minute. A claim about what a competitor's pricing page says is verifiable only by opening it, which is most of the work.
| Easy to verify | Hard to verify | |
|---|---|---|
| Reversible | Delegate freely. Drafts, research summaries, formatting, tagging | Delegate with spot checks. Long research, competitive summaries |
| Irreversible | Delegate with review before the action. Scheduled posts, queued sends | Do not delegate. Sending, publishing, replying to customers, spending |
The bottom-right quadrant is where the damage happens, and it is exactly where agent products are marketed — because those tasks are the ones that feel most valuable to hand over.
What agents do reliably
Work that is bounded, checkable, and stops before an irreversible step.
- Gathering and structuring. Pulling information into a defined shape you specified
- Drafting into a queue. Producing the messages, posts or pages, and stopping
- Classifying at volume. Tickets into themes, responses into segments, content into categories
- Reformatting. One piece into the shapes each channel needs
- Routine checking. Broken links, missing metadata, inconsistencies against a rule you defined
- First-pass research, with sources attached so you can check the ones that matter
The common property: you can look at the output and tell whether it is right, and nothing has happened yet.
What they fail at, and how
Not randomly. In four specific shapes worth recognising.
1. Confident wrongness. A fabricated statistic, a misattributed quote, a plausible figure with no source. The failure is not the error — it is that the error arrives with the same confidence as everything correct, so a reviewer skimming does not catch it. This is why "with sources attached" matters more than it sounds: it converts an unverifiable claim into a checkable one.
2. Compounding drift. A small misreading in step two becomes the premise of steps three through nine. In a single-output task you see the mistake. In a sequence you see only the conclusion, which is coherent and wrong.
3. Missing what is not there. An agent working from what it retrieved does not know what it failed to retrieve. A competitive summary that omits your largest competitor reads exactly like one that does not.
4. Not knowing what it does not know. The judgement about whether it has enough information to proceed is the judgement it is least reliable at. An agent rarely stops and says the task was underspecified; it proceeds on an assumption it does not flag.
The three things to keep away from them
1. Anything that sends to a real person.
Not because the writing is bad. Because volume and error compound in the one place you cannot afford them.
An agent sending outreach at scale generates complaints, and complaints damage the domain reputation your invoices, receipts and password resets depend on. The reputation damage outlasts the campaign by weeks, and cleaning it up is a separate process from stopping the sending. How that damage works · what removal involves.
Draft-and-queue is the pattern that keeps the value and removes the risk. A person approves the batch.
2. Anything that publishes.
Same argument, different surface. A published error is public before it is caught, and the correction never reaches everyone who saw it.
3. Anything that spends.
Ad budgets, bids, subscriptions. Irreversible, and the failure mode is a bill.
The bit that gets skipped
Reviewing agent output takes longer than people budget for, and the saving is smaller than it looks.
A draft produced in two minutes that takes twenty to verify has saved you the two minutes. That is still worth having when the alternative was an hour of drafting — but it is not the saving the pitch implies, and the review is the step that quietly gets dropped when the outputs have been fine for a fortnight.
The honest accounting:
- Generation time saved: large
- Review time added: moderate, and it does not go away
- Net saving: real, and roughly half what it appears to be
- Risk added: proportional to how much review gets skipped over time
The failure pattern is predictable. Output is good for weeks, review relaxes, and the one bad output goes through unexamined. Build the review as a step somebody owns, not as a habit that depends on diligence.
Disclosure
Expectations about disclosing AI-generated and AI-assisted content are tightening, and the specifics differ by market and by platform.
Three things are stable enough to act on now:
- Do not present generated content as a person's first-hand experience. A review, testimonial or case study written by a model and attributed to a customer is a fabrication regardless of what any regulation says
- Do not let an agent represent itself as a person in a conversation with a customer
- Check the rules for your own market and for each platform you publish on, because they are moving
Where a specific obligation applies to you — sector rules, advertising standards, platform policy, or regional legislation — verify it against the current text rather than an article. This one included. Why this site does not publish figures that expire.
Where this leaves a small business
A defensible position that does not depend on predicting what happens next.
- Use agents for gathering, drafting, classifying and checking
- Keep every irreversible action behind a person, without exception, because the exceptions are how it goes wrong
- Require sources on anything factual, which turns confident wrongness into a checkable claim
- Give one person the review step rather than assuming it
- Re-examine every six months, not every launch
And apply the same test to any agent product you are sold: what happens the times it is wrong, and who finds out first — you, or your customer?
Frequently asked questions
What can AI agents do reliably in marketing?
Bounded, checkable work that stops before an irreversible step: gathering and structuring information, drafting into a queue, classifying at volume, reformatting content, and routine checks against a rule you defined. The common property is that you can see whether the output is right and nothing has happened yet.
Should an AI agent send emails on my behalf?
No. Volume and error compound in the one place you cannot afford them — complaints damage the domain reputation your receipts and password resets depend on, and the damage outlasts the campaign. Draft-and-queue keeps the value and removes the risk.
How do AI agents fail?
In four recognisable shapes: confident wrongness that arrives indistinguishable from correct output; drift, where a small early error becomes the premise of everything after it; missing what was never retrieved; and proceeding on an unflagged assumption rather than stopping to ask.
Do AI agents actually save time?
Real saving, roughly half what it appears to be. Generation time drops a lot, review time is added and does not go away, and the predictable failure is that review relaxes after a few good weeks and the one bad output goes through unexamined.
Do I need to disclose AI-generated content?
Requirements differ by market and platform and are moving, so check the current rules for yours. Three things hold regardless: do not present generated content as a person's first-hand experience, do not let an agent represent itself as a person, and verify any specific obligation against the current text rather than an article.
How do I decide whether to delegate a task to an agent?
Ask two questions: can a wrong result be undone before it costs anything, and can you check the work faster than doing it? Reversible and easy to verify — delegate freely. Irreversible and hard to verify — do not delegate at all.
Free: The 60-Minute Email Authentication Fix
A no-fluff checklist to set up SPF, DKIM & DMARC correctly and pass Gmail & Yahoo's sender requirements.

Muhammad Basim has worked in digital marketing since 2013, focused on email deliverability and AI-assisted content production. He is the author of The Email Deliverability Playbook and The Email Copywriting Playbook.
Related Articles

Does Send Time Optimisation Work?
Send time optimisation models are trained on when recipients opened messages, and open timestamps now include machine activity that has nothing to do with when anyone was reading. Privacy protection pre-fetches message content on its own schedule, regardless of whether the recipient opened it. Security gateways fetch content during pre-delivery scanning, which happens before the […]

Prompting for Marketers: The Part That Actually Matters
Output quality is set mostly by the material you supply, not the phrasing you use. That is the finding people take longest to accept, because it is less interesting than the alternative. A carefully worded request with no context produces a competent generic answer. A plainly worded request with your actual customer language, your positioning […]

What You Grant When You Connect a Tool
Your security perimeter includes every vendor you have ever connected, including the ones you stopped using and never disconnected. A connection is a standing grant. It does not expire because you stopped logging in, it does not lapse because the trial ended, and it does not disappear when you delete the app from your phone. […]

