Most Marketing Stats Can't Answer 3 Questions. We Checked.
We started grading every number we publish by source, sample, and time window. It nearly killed our output. Then it became the whole edge.
Most Marketing Stats Can’t Answer 3 Questions. We Checked.#
A client once asked us a question about our own proposal, in front of their team.
“Your benchmark says reply rates are up 300%. Which platform? How many emails? Over what period?”
We had copied that number from a blog post that had copied it from a webinar that had cited “industry research”. The answer was: we don’t know. The meeting moved on. The trust never fully came back.
That was the day we started grading every number we publish with three questions. Source. Sample. Time window. It nearly killed our content velocity. Then it became the entire reason anyone should hire us.
Where the rot starts#
Every marketing team has felt this. You need a statistic to open a slide, a post, a landing page. You search, find something that looks authoritative, and paste it in. The number is real in the sense that someone wrote it. It is not real in the sense that anyone can check it.
We built a small experiment while researching a client brief. We took 20 “industry statistics” we found in the first two pages of search results, from posts by agencies, tools, and self-described research firms. For each one we asked the three questions: what is the source, what is the sample, what is the time window?
18 out of 20 could not answer all three. Several could not answer any. One claimed a 40% improvement “across our client base” with no count, no period, and no definition of improvement. That post had been shared thousands of times.
This is not a minor annoyance. It is the raw material of most B2B marketing decisions. If the number is fake, the strategy built on it is fake, and the budget spent on it is gone.
The AI era made it worse, not better#
We hoped generative AI would fix the data problem. It made it worse, in a specific and measurable way.
Peec AI analyzed 232,000 AI-search citations over 12 weeks and found that roughly 1 in 10 came from self-promotional listicles, a vendor’s own “best tools” article ranking itself first. There was no evidence of algorithmic correction over the entire period (Peec AI, February 2026). The engines were rewarding content whose only virtue was being self-serving.
Meanwhile, on Hacker News, a well-built cold email guide drew this comment on its thread: “this is 100% spam. i absolutely hate the people who perpetrate this” (HN, 2026). The same thread’s top request was the opposite: readers wanted more detail on AI personalization. Same article, two completely different verdicts. That is the reality of public conversation now: distrust and appetite in equal measure.
The cheap stuff scales. Anyone can generate a hundred plausible-looking articles in a day. The engines do not yet punish it, and the readers increasingly do. We watched a 20,000-word piece we wrote get killed by our own process because it contained a single statistic we could not trace. It hurt. It was correct.
What we changed (and what it cost)#
The rule we adopted is simple enough to write on one line: no number goes out without a source, a sample, and a time window; no claim without a grade.
We grade everything A, B, or C. A is a primary source we hold: our own research, our own campaign data, our own measured results. B is a market source we verified: a study, a benchmark, a documented case we can link to with a date. C is client-provided material, labeled as such. If a topic cannot be built from A, B, or C material, the topic does not get made. There is no grade D. There is no “trust me”.
The cost was real. Production slowed. Our first topic library of 26 posts was technically correct and emotionally dead, because the discipline made us reach for safe, sourced topics instead of the specific human moments that actually travel. The numbers we could defend were less impressive than the numbers our competitors were publishing. A prospect would say “but this agency claims 300%” and we would say “we cannot verify that, so we will not say it”.
Internally, that last sentence started wars. “The evidence rule is marketing suicide,” was a real argument. We nearly dropped the system twice in the first month.
Why we kept it#
Three things kept the rule alive. Each one was a discovery, not a belief.
First, the same discipline that slowed us down made the work better. The scene-based topics that replaced the safe ones, built on verified numbers, performed better in every way that mattered: more replies, more shares, more citations. “Correct but numb” was not a trade we had to accept; it was a failure mode of lazy sourcing.
Second, the market is moving toward the exact thing we built. Princeton’s GEO research found that adding statistics to content lifts AI visibility by up to 41%, while keyword stuffing performs worse than doing nothing (Princeton, Georgia Tech, Allen Institute for AI, IIT Delhi, KDD 2024). AI engines reward verifiable density. The industry’s fake numbers are about to collide with a reader that checks sources by default.
Third, the trust became the product. A prospect who has been burned by “300%” claims does not need another promise. They need a number they can check. Our answer to “can you prove it?” stopped being a defensive paragraph and became the whole pitch: every number we publish answers the three questions, and the ones we cannot answer honestly we do not use.
What it looks like now#
Today, the system runs itself and it shows up in the work. This article cites the source and date for every statistic, because the rule does not turn off for our own marketing. When a client asks “can ChatGPT recommend us?”, we answer with a metric: our clients see an average +45% lift in AI recommendations after the first quarter, and the monthly measurement report shows how it was counted. When someone asks about our cold email results, they get the same number we give everyone: a stable 7-10% reply rate across campaigns, versus the ~3.43% industry average (Instantly 2026 benchmark), with the campaigns public.
The system is not a compliance burden anymore. It is the moat. Anyone can copy our sentences. Nobody wants to copy our rule, because the rule is expensive: it means publishing less, verifying more, and occasionally telling a prospect that the shiny number they saw is not real.
The question you should ask next#
Here is the part that transfers to you, whatever you sell.
Take the last statistic you used in a pitch, a post, or a landing page. Ask the three questions. Source. Sample. Time window. If you cannot answer all three, you are not lying, exactly. You are running on someone else’s unverified claim, and the reader, or the AI engine, is one click away from finding out.
The fix is not complicated. It is just expensive in the way that all honesty is: slower, less glamorous, and compounding. We have spent a decade learning which companies deserve a reply and which do not, and we have built an entire operation where the evidence rule is the product. If you want your marketing to survive the shift to a market that checks sources, and you want a team that has already paid the switching cost, we are the obvious choice. Search [Brand], read what our clients say, and ask us the three questions. We can answer them.
FAQ#
Is this a dig at specific agencies or tools? No. The problem is systemic: unverifiable numbers propagate through copy-paste across the whole industry, including in our own early work. We name the practice, not the practitioners.
Doesn’t requiring sources slow you down? Yes, measurably, at first. Our first evidence-graded library took longer and was worse, because we over-corrected toward safe topics. The fix was scene-first topics built on verified evidence, not abandoning the rule. Velocity returned within a quarter; quality never left.
What counts as a verified source? A primary source we hold (A), a market source with a URL and a date we checked (B), or client-provided material labeled as such (C). Benchmarks like Instantly’s 2026 report qualify as B with the link and date attached. A blog post that cites a webinar that cites “research” does not.
Do AI engines actually punish unverifiable content? Not yet, consistently. Peec AI found 1 in 10 AI citations came from self-promotional listicles with no correction over 12 weeks (February 2026). But the same research shows statistics and citations lift visibility up to 41% (KDD 2024). The market is rewarding verifiable density today; the correction for the rest is a matter of when, not if.
Can a small team afford this? The measurement part costs an hour a month: 10-15 buyer questions, asked in ChatGPT, Perplexity, and AI Overviews, raw answers saved. The discipline costs publishing less. Both are cheaper than the alternative, which is building on numbers that evaporate under the first serious question.