All Four of the Big Four Have Now Published AI-Hallucinated Reports. Here's the One Thing They're All Missing.
Deloitte, EY, KPMG, and now PwC have each published research reports with AI-fabricated citations and claims. The pattern points to one missing architectural layer, not four unrelated mistakes.

Deloitte, EY, KPMG. Now PwC. In just under a year, every single one of the Big Four consulting and accounting firms has published a research report that turned out to contain AI-hallucinated content: invented citations, invented statistics, and in at least one case, invented customers.
(We covered the first three — Deloitte, EY, and KPMG — when the pattern was only three firms deep. PwC completes the set.)
The latest is PwC Middle East. GPTZero, the AI-detection firm behind the earlier EY and KPMG findings, examined four of PwC's "thought leadership" reports published between 2024 and 2026. Its most striking finding is in a 2025 report called "Transforming Governance," which GPTZero rated 84% likely to be almost entirely AI-written (rising to 100% once you exclude the reference section). That report claims Denmark, Saudi Arabia, the US, and Australia have all deployed a PwC civic-engagement product called "Citizen Pulse" and seen real improvements in public service reliability. GPTZero could find no public evidence that any of those governments use the product at all.
The other three reports GPTZero looked at had their own problems: a cited academic paper on Riyadh air quality that appears not to exist anywhere in the journal it was attributed to, a claim about JPMorgan's AI strategy sourced to a teenage blogger with 280 followers on Medium, and the same traffic-safety statistic cited three separate times with three different, unrelated sources. PwC Middle East's response, reported by the Irish Times, was that it "takes the accuracy of our published research seriously" and is "updating a limited number of supporting citations."
That response will sound familiar, because it's close to what EY and KPMG said earlier this year. EY retracted a cybersecurity report in May 2026 after GPTZero found more than 70% of its 27 citations didn't hold up. KPMG pulled an "agentic AI" report in June 2026 after only 5 of 45 citations turned out to be genuine, verifiable sources, a pattern GPTZero named "vibe citing": references that look real because they name real publications, but fail as soon as anyone checks the link. And before either of those, in October 2025, Deloitte Australia had to partially refund the Australian government roughly AUD 98,000 of its AUD 440,000 contract after a commissioned welfare-compliance audit turned out to contain fabricated legal citations, invented experts, and a quote wrongly attributed to a federal court judge, all traced back to GPT-4o.
Four different firms. Four different reports. Four different topics: cybersecurity, agentic AI adoption, welfare-system compliance, civic technology. The one thing they share isn't the AI tool. It's the missing step between "AI drafted this" and "we published this": nobody verified that every citation traced to a real, checkable source before the report went out the door.
That's not a hard problem to solve. It's a design decision, and it's the one that's been skipped four times in a row.
At Foragentis, it's the decision we built our intelligence products around from the start. Every claim in a Foragentis report resolves to a verifiable source row in a dedicated data layer before it's allowed to publish. If a claim can't be traced to something real and checkable, it gets excluded, not hallucinated. Acceptance criteria are pre-registered before a report is drafted, not applied as an afterthought once something looks wrong. Every report includes a published "known limits" section that says plainly what could not be verified, so there's no "Citizen Pulse" problem: a claim that four governments use a product that, as far as anyone can find, none of them use.
None of this is about mistrusting AI as a drafting tool. It's about not publishing AI output as finished research without a verification layer standing between the model and the reader. Four Big Four firms have now shown what happens when that layer is missing. The fix isn't complicated. It just has to actually be there.
FAQ
What did GPTZero find in PwC's research reports?
GPTZero examined four PwC Middle East "thought leadership" reports published between 2024 and 2026. A 2025 report called "Transforming Governance" was rated 84% likely to be almost entirely AI-written, rising to 100% once the reference section is excluded. The other three reports had a citation to an academic paper on Riyadh air quality that does not appear to exist in the journal it was attributed to, a claim about JPMorgan's AI strategy sourced to a blogger with 280 followers on Medium, and the same traffic-safety statistic cited three separate times to three different, unrelated sources.
What is PwC's "Citizen Pulse" hallucination?
PwC's "Transforming Governance" report claimed that Denmark, Saudi Arabia, the US, and Australia had all deployed a PwC civic-engagement product called "Citizen Pulse" and seen real improvements in public service reliability. GPTZero could find no public evidence that any of those governments use the product at all — it appears to be an invented case study built around a real-sounding but nonexistent deployment.
What is "vibe citing"?
Vibe citing is GPTZero's term for the citation equivalent of vibe coding: a generative model stitches together fragments of real sources, invents plausible titles, and paraphrases references until they no longer match the original — producing something that looks exactly like scholarship but fails as soon as anyone checks the underlying link.
Why do Deloitte, EY, KPMG, and PwC all have the same underlying problem?
Four different firms published four reports on four unrelated topics — cybersecurity, agentic AI adoption, welfare-system compliance, and civic technology — but each failure traces to the same missing step: nobody verified that every citation resolved to a real, checkable source before the report was published. It is one architectural gap repeated four times, not four unrelated mistakes.
How does ForIntel prevent this kind of AI hallucination?
Every claim in a Foragentis report resolves to a verifiable source row in a dedicated data layer before it is allowed to publish. A claim that cannot be traced to something real and checkable is excluded, not fabricated. Acceptance criteria are pre-registered before a report is drafted, and every report carries a published "known limits" section stating plainly what could not be verified — so there is no "Citizen Pulse" problem waiting to be discovered by an outside auditor.
ForIntel is a research-first business-intelligence product from Foragentis. Every quantitative claim in a ForIntel brief traces to source data, is cross-corroborated across independent measures, and passes a counter-signal verification step before it ships. To see a sample brief or commission a read, reach the ForIntel desk at forintel@foragentis.com.


