Searching for a statistic and finding it repeated in six articles is evidence of circulation, not of truth.
That is the trap this whole job turns on. A fabricated figure propagates, because the second article cites the first, the third cites the second, and by the tenth the number has the texture of common knowledge with no origin anywhere in the chain.
Verification means arriving at a primary source. Anything short of that is checking that other people also believe it.
The three fabrication shapes
Errors in generated text are not random. Recognising the shapes speeds the check considerably.
1. A real figure attached to the wrong source. The number exists; the attribution does not. Hardest to catch, because searching the number returns results and they look like confirmation.
2. A real source credited with something it does not contain. A genuine report, a genuine organisation, a claim that is not in it. Only opening the document catches this, which is why it survives most checking.
3. A number with no origin, phrased as widely known. "Studies show", "research indicates", "the average is". The most dangerous of the three, because there is no specific claim to disprove and the search results are other people repeating it.
All three share a property: they arrive with exactly the same confidence as everything correct in the piece. There is no textual signal. Which is why fact-checking has to be a separate pass with its own method, rather than something the eye does while reading.
The method
Step 1 — Build a claim inventory
Go through the piece and list every factual assertion, one per line.
What counts as a factual claim:
- Any number, percentage, date or measurement
- Any attribution — "Google says", "the specification requires", "X found"
- Any statement about how something behaves — a rule, a threshold, a default
- Any statement about what an organisation does, requires or recommends
- Any superlative — most, first, only, largest
What does not count: your own reasoning, your own recommendations, and clearly framed opinion.
A 2,000-word article typically produces fifteen to thirty claims. Seeing them listed is itself informative — most people are surprised how many assertions a piece makes and how few they could source.
Step 2 — Assign a source tier to each
| Tier | What it is | Worth |
|---|---|---|
| 1 | The primary document. The RFC, the specification, the official documentation, the study's own paper | Sufficient on its own |
| 2 | The organisation's own statement. A vendor's documentation about their product, a regulator's published guidance | Sufficient for claims about themselves |
| 3 | A named study, opened and read — not a description of it | Sufficient, with method and sample stated |
| 4 | A reputable article citing a tier 1–3 source, where you follow the link | A route to a source, not a source |
| 5 | An article citing another article | Worthless |
| 6 | A vendor statistic about their whole customer base | Marketing. Not evidence |
The rule: every claim needs a tier 1, 2 or 3 source, or it comes out.
Tier 6 deserves particular suspicion because it dominates marketing writing. "Emails with X get 34% more opens" measured across one platform's customers tells you about that platform's customers, with no control and a selection effect. It is not wrong so much as uninformative, and repeating it borrows a confidence nobody earned.
Step 3 — Check each claim at its source
Open the document. Find the sentence. Do not trust a snippet.
Four things to record while you are there:
- The date. A true statement about 2022 may be false now
- The scope. "Gmail requires" may mean bulk senders only
- The exact wording. Paraphrase drift is how a recommendation becomes a requirement
- The URL, so a future check takes seconds rather than repeating this
The paraphrase drift point is worth dwelling on, because it is the most common way careful people publish something false. A specification saying receivers should not do something becomes "receivers do not". A guidance page saying an approach may help becomes "Google recommends". Each step is small and the destination is a claim the source does not make.
Step 4 — Delete what you cannot verify
Deleting is a legitimate outcome and it should happen on most pieces.
Three options for an unverifiable claim, in order of preference:
- Cut it. Usually the piece is better without it
- Replace it with the mechanism. "Complaint rates above the threshold trigger filtering" is stronger than a percentage you cannot source, and it is more useful to the reader
- State the uncertainty explicitly. "Figures circulate for this and I have not found a primary source for any of them" is a legitimate sentence and it builds more trust than a number would
What not to do: keep it with a hedge. "Some studies suggest" attached to a claim you could not source is the same claim with a disclaimer, and the reader takes away the number.
Step 5 — Record what you refused
Keep a note, published or private, of the claims that did not survive.
Three things it does:
- Stops the same figure being re-added in six months by you or a model drafting the update
- Speeds the next check on a related piece
- Demonstrates the discipline if you publish it
Every article on this site ends with a "deliberately not claimed" note listing figures that were cut and why. It costs a paragraph and it is more persuasive than a citation — a citation shows you found something, and the refusal shows you looked and stopped.
Where a model can and cannot help
It can help with the inventory. Asking for every factual claim in a draft, listed, is a reliable task — extraction rather than knowledge.
It cannot verify. A model checking a model's claims produces agreement, not verification, and asking for sources on an unsourced claim frequently produces plausible citations that do not exist.
Treat any source a model supplies as a lead to check, never as a check performed. The check is opening the document.
Claims that need extra care
Four categories where the failure rate is highest.
Anything with a date. Requirements, deadlines, version numbers, policy changes. A model's training has a cutoff and a statement true then may be false now.
Anything about a specification. RFCs get revised, and the revision may remove exactly the clause being cited. Check the current RFC number, not a remembered one.
Anything attributed to a large company's policy. These change quietly, and the article you are half-remembering may predate the change.
Anything that sounds like a rule of thumb. "Under 100KB", "within 3 seconds", "no more than 5". These are the most-repeated and least-sourced claims in marketing writing, and a large share turn out to trace to a single blog post from over a decade ago.
Frequently asked questions
How do you fact-check AI-generated content?
Build an inventory of every factual claim, assign each a source tier, check each at a primary source by opening the document, and delete what you cannot verify. It is a separate pass with its own method — reading for prose does not stop at a plausible statistic.
Why is searching for a statistic not enough?
Finding it repeated in several articles evidences circulation, not truth. Fabricated figures propagate because each article cites the one before, and by the tenth the number reads as common knowledge with no origin in the chain. Verification means arriving at a primary source.
Can AI fact-check its own output?
No. A model checking a model's claims produces agreement rather than verification, and asking for sources on an unsourced claim often produces plausible citations that do not exist. It can reliably list the claims for you to check, which is genuinely useful.
What should I do with a claim I cannot verify?
Cut it, or replace it with the mechanism behind it, or state the uncertainty explicitly. What not to do is keep it behind a hedge — "some studies suggest" attached to an unsourced claim is the same claim with a disclaimer, and readers take away the number.
What are the most common AI fabrications?
Three shapes: a real figure attached to the wrong source, a real source credited with something it does not contain, and a number with no origin phrased as widely known. All three arrive with the same confidence as correct material, which is why there is no textual signal to catch them.
Are vendor statistics reliable sources?
Rarely. A figure measured across one platform's customer base has no control group and a strong selection effect, so it describes that platform's customers rather than a general truth. It is not so much wrong as uninformative, and repeating it borrows confidence nobody earned.
Free: The 60-Minute Email Authentication Fix
A no-fluff checklist to set up SPF, DKIM & DMARC correctly and pass Gmail & Yahoo's sender requirements.

Muhammad Basim has worked in digital marketing since 2013, focused on email deliverability and AI-assisted content production. He is the author of the Email Deliverability Playbook and the Email Copywriting Playbook, and has run 100+ email campaigns for ecommerce brands, coaches, and B2B senders. He writes about email, SEO, WordPress, and AI — with a bias toward what can be tested over what sounds good.
Related Articles
What to Give AI Before You Ask
Generic output is almost always a context problem, not a prompt problem. A model with nothing specific to work from produces the average of what has been written on the subject. That is not a flaw to be worked around with better phrasing — it is the correct behaviour given no information. The fix is […]
Connecting Your Tools: Automation That Does Not Break
Every integration you add is a thing that can fail without telling you. That is the property that makes connected systems different from the tools they connect. A broken tool announces itself — you open it and something is wrong. A broken integration produces silence, which is indistinguishable from a quiet week. And the failure […]
AI Email Personalisation: What Is Worth Doing
Personalisation built from what a customer told you works. Personalisation built from what a model inferred about them produces complaints. That line is the entire subject, and it does not move as the technology does. The risk is not that the inference is bad. It is that it is confident, specific and occasionally wrong — […]