As the world of accounting and finance embraces the adoption of AI, the role of the human preparer of financial information is diminishing to reduce costs, drive efficiencies and increase profits.
This shift has resulted, in some cases, in strategic decisions being based upon machine-generated data. This in turn can lead to a high likelihood of ‘hallucinations’ (incorrect or false information generated by AI models because they are probabilistic rather than deterministic).
The reviewer and the users of the information are blindly trusting it to be accurate
Multiple studies have shown that in 2024 alone, global business losses attributed to AI hallucinations were estimated at US$67bn. In a profession where the stakes are high and accuracy is non-negotiable, regardless of whether in financial reporting, financial statements or forecasting, unverified AI outputs can threaten sound decision-making and public trust.
Major scandals
The past year has seen four major scandals relating to corporate reports issued by the Big Four. In October 2025, Deloitte Australia had to pay a partial refund to the Department of Employment and Work Relations after submitting a report to the Australian government containing citations caused by hallucinations. A month later, a similar incident occurred at Deloitte Canada. In both cases, Deloitte acknowledged the errors and subsequently submitted amended reports.
In May 2026, EY Canada recalled a 2025 report due to hallucinations of up to 60% of the report’s references, as reported by one investigation. In addition to hallucinated citations, the main report also featured inconsistencies. For example, the executive summary mentioned the global loyalty points market as exceeding US$200bn, out of which only 30%-50% (US$60bn-US$100bn) remained unused. However, in the subsequent pages, it claimed unused points to be above US$200bn.
Most recently, in June 2026, an investigation by GPTZero identified various hallucinations in a report published by KPMG International, Redefining excellence in the age of agentic AI, where the case studies provided turned out to be inaccurate.
Such errors are not limited to just the accounting profession. There have been numerous instances in the legal sector where the case references and the facts provided were not factual; in customer services, where incorrect terms and policies were communicated to the customers; and in the media, where in one case tourists were recommended a foodbank as a good place to enjoy a meal.
Lacking rigour
It’s surely almost impossible that the content was published without being subject to review. A more reasonable assumption would be that a vague review was performed but that it was not sufficient to identify and detect hallucinations. Traditional reviews are failing to catch mistakes made by AI, which highlights the need for a higher level of review rigour to preserve quality and public trust.
Conducting analytical reviews would have highlighted EY’s US$200bn error
The emerging research highlights a critical gap: traditional reviews were conducted to identify and correct human errors, such as misjudgment, incorrect extraction of data or incorrect calculations; they were not designed to address the risks posed by AI.
The underlying problem is not that AI is taking the role of preparer and not doing a perfect job at it; rather, that the reviewer and the users of the information are blindly trusting it to be accurate. This is widely due to automation bias, where the information coming from the system is perceived to always be precise and reliable.
More controls
A recent publication by the Committee of Sponsoring Organizations of the Treadway Commission also highlights that generative AI (GenAI) requires additional controls to address the risks of transparency, validation and ethical safeguards.
To address this issue, all processes where GenAI is integrated must be re-engineered to include reviews stringent enough to identify and prevent errors. The reviews required will vary based on the scope, capability and application of GenAI. For instance, task automation, where routine or continuous tasks such as transaction-level reconciliations are automated, will require both detective and preventive reviews – the former to identify errors in the transactions that are matched with a high confidence threshold; the latter to review transactions that are not matched and kept on hold for a human to review.
Human touch
By contrast, reports or insight generation will be subject to a human review, more detective in nature. In this use case, the information vetted is significant and vital, as it will be used for decision-making purposes. The human review should verify the following at a minimum to ensure validation and transparency:
- use of standardised prompt libraries
- use of approved AI tools
- clear tagging of AI and human-generated information
- inclusion and assessment of performance KPIs and confidence levels in the output
- manual verification of inputs and sources
- analytical review of figures in the output.
The scandals discussed above could have been prevented with the use of adequate measures during their review processes. For example, conducting analytical reviews would have highlighted EY Canada’s US$200bn error, while manual verification of inputs would have identified the hallucinated citations in the Deloitte and KPMG reports.
In conclusion, it is critical that the reviewer is mindful of the probabilistic nature of the model generating the information and takes additional measures to address heightened transparency and validation risks to uphold the quality of the information produced for decision-making.
In the ever-evolving age of AI, organisational success will entail robust AI governance structures alongside stringent human reviewers who are aware of the risks posed by AI and work collaboratively to achieve organisational goals.