AI financial model risk

AI Financial Model Risk and Governance: Statistics (2026)

AI Financial Model Risk and Governance: Statistics

Last updated: July 2026

Putting generative AI (Claude, ChatGPT, Excel copilots) on top of financial modeling creates a new class of risk: confidently fabricated numbers, sensitive data leaving the building through unsanctioned tools, and a widening gap between AI usage and the governance that should accompany it. This page collects verified figures with their primary sources, so the numbers can be cited and traced. Each stat links to a primary or authoritative source with its scope and year.

Confident fabrication: where generative AI breaks on financial numbers

81% of financial questions were answered incorrectly or refused by GPT-4-Turbo paired with a retrieval system on the FinanceBench benchmark. With long-context prompts, the failure rate fell to 21% for GPT-4-Turbo and 24% for Claude-2. Sample of 150 cases from FinanceBench (10-K, 10-Q, 8-K and earnings filings from 40 US public companies); 16 model configurations tested, n=2,400 answers manually reviewed; November 2023. (Islam et al., Patronus AI, Contextual AI & Stanford, 2023) + (Patronus AI, 2023)

82.4% accuracy is the best score any model reaches on FinSheet-Bench (Gemini 3.1 Pro), roughly one error every six questions. The authors conclude that "no standalone model achieves error rates low enough for unsupervised use in professional finance applications." 24 evaluation files of varying complexity and layout; 10 model configurations from OpenAI, Google and Anthropic; synthetic portfolio data modeled on real private-equity fund structures; March 2026. (Ravnik et al., Qubera AG / University of Zurich, 2026)

19.6% average accuracy on complex aggregation tasks, versus 89.1% on simple lookups, across all 10 models tested on FinSheet-Bench. Accuracy degrades steadily as task complexity rises. Pooled across all model configurations; ~500 questions per model; 24 evaluation files; synthetic data modeled on private-equity fund structures; 2026. (Ravnik et al., arXiv, 2026)

10 to 20% error rates persist even for frontier models on multi-step numerical reasoning over financial tables, despite high overall accuracy. On the hardest category (multivariate calculation), Claude-Sonnet-4 reaches 80.0% (95.6% overall); the authors caution this subset holds only ~10 cases and should not be over-read. FAITH benchmark; 14 LLMs; 2,406 answerable spans drawn from 2024 10-K reports of 453 S&P 500 companies; context-aware masked-span prediction; ACM ICAIF'25. (Zhang et al., arXiv / ACM ICAIF'25, 2025) + (Cognaptus, 2025)

Hallucinations reach the ledger

86% of CFOs said their finance team had encountered at least one instance of inaccurate or "hallucinated" data while using AI. Survey of 100 CFOs at mid-market US companies, conducted late September 2025, published January 2026; vendor-sponsored (Maximor AI), small sample, fielded by Wakefield Research. (Wakefield Research for Maximor AI, via CFO Dive, 2026) + (Journal of Accountancy, 2026)

14% of CFOs say they completely trust AI to deliver accurate accounting data on its own. Same survey: 100 CFOs at mid-market US companies ($50M-$500M revenue), late September 2025, published 28 January 2026; vendor-sponsored. (Wakefield Research for Maximor AI, 2026) + (CFO Dive, 2026)

Nearly one-third (~33%) of all respondents say their organization has experienced negative consequences from generative AI inaccuracy, the most commonly cited risk to cause negative consequences; 51% of respondents at AI-using organizations report at least one negative consequence. McKinsey Global Survey on AI; 1,491 participants across regions, industries and sizes; data collected 16-31 July 2024; published March 2025. (McKinsey & Company (QuantumBlack), 2025)

93% of UK accountants and bookkeepers who encounter AI-related mistakes estimate they spend up to ten hours per month correcting errors caused by AI-generated advice (44% up to three hours, 39% four to ten hours). Subset of 500 UK accountants and bookkeepers who report encountering public-AI errors; fielded by Censuswide for Dext over the first two weeks of December 2025. (Dext (Censuswide), 2025) + (CFOtech UK, 2025)

What expert modellers will not delegate

0 out of 63 senior financial modellers said they would feel confident relying on an AI-generated model for a high-stakes business decision without independent human review; 75% strongly disagreed and 22% disagreed. It was the single strongest consensus in a 30-question survey. Financial Modeling Global Leaders Council survey; 63 senior modellers across 26 countries, 92% with 10+ years' experience; 100% response rate; expert panel, not a representative sample; fielded May 2026, published July 2026. (Financial Modeling Institute, 2026) + (PR Newswire, 2026)

90% of senior modellers say signing off on model outputs should never be fully delegated to AI, and 86% say the same of ethical judgment calls, the two tasks ranked most firmly human. Only one member of 63 believed all modelling tasks could eventually be fully delegated. Financial Modeling Global Leaders Council survey; 63 senior modellers, 26 countries; multi-select; 100% response rate; expert panel, not a representative sample; published July 2026. (Financial Modeling Institute, 2026) + (PR Newswire, 2026)

59% of modellers named reduced modeller understanding of AI-assisted models their top risk, ahead of hidden or opaque logic (51%), over-reliance / automation bias (48%) and skill atrophy (46%). Subtle numerical errors (32%) and lack of audit trail (17%) ranked lower: the panel fears losing the ability to catch errors more than the errors themselves. Financial Modeling Global Leaders Council survey; 63 senior modellers; members selected up to three risks; expert panel, not a representative sample; published July 2026. (Financial Modeling Institute, 2026) + (PR Newswire, 2026)

Shadow AI in finance

49% of employees report using AI tools not sanctioned by their employer; 58% of those rely on free versions, and 23% admit sharing financial statements or sales data with unsanctioned tools. Sapio Research survey of 2,000 employees (1,000 UK, 1,000 US) at organizations with 500+ staff; fielded November 2025, published 27 January 2026. (BlackFog (Sapio Research), 2026) + (CIO.com, 2026)

16.6% of all detected sensitive-data exposures in enterprise AI prompts were financial: financial projections, investment analysis and sales pipeline data accounted for 95,852 instances, the single largest category. Anonymized telemetry from US and UK enterprises via Harmonic Protect; 22,458,240 prompts and uploads analyzed, 579,113 sensitive-data instances detected, across 665 GenAI tools; 1 Jan-31 Dec 2025. Vendor dataset, not a representative panel. (Harmonic Security, 2025) + (SecurityBrief UK, 2025)

2.6% of 22.4 million enterprise AI prompts and uploads contained company-sensitive data (579,113 detected exposures). Anonymized data from US and UK enterprises monitored via Harmonic Protect, across 665 generative and embedded AI tools; 1 Jan-31 Dec 2025. (Harmonic Security, 2025) + (SecurityBrief UK, 2025)

About one-third of senior finance professionals report that shadow AI is already a noticeable issue in their organization, and 40% say their organization either has no rules governing AI use or they are unsure what those rules are. Survey of 311 senior finance professionals across 22 sectors, global, fielded March-April 2026. (insightsoftware, 2026) + (GlobeNewswire / Yahoo Finance, 2026)

More than 40% of global organizations are predicted to suffer security and compliance incidents from the use of unauthorized AI tools by 2030; 69% of organizations already have evidence or suspect employees are using public generative AI at work. Gartner prediction; the 69% figure is from a survey of 302 cybersecurity leaders worldwide, conducted March-May 2025. (Gartner, Inc., 2025) + (Infosecurity Magazine, 2025)

Contractual and professional limits on AI use

The figures above measure behaviour. This section records the rules and vendor commitments that behaviour is measured against: what a professional confidentiality obligation requires before client data reaches a third party, and what Anthropic and OpenAI state they do with the data they receive. These are documentary sources, not survey results.

Interpretation 1.700.040 of the AICPA Code of Professional Conduct provides that, before disclosing confidential client information to a third-party service provider, a member should either enter into a contractual agreement with that provider to maintain confidentiality and give reasonable assurance it has procedures preventing unauthorized release, or obtain specific consent from the client. It sits under the Confidential Client Information Rule (1.700.001), under which a member in public practice shall not disclose confidential client information without the client's specific consent. AICPA Code of Professional Conduct, applicable to members in public practice in the United States; it covers any third-party service provider used to help deliver professional services. The Code's wording is "should", not "must", and the interpretation predates generative AI and does not name it. Sources are AICPA publications restating the interpretation, not the Code text itself. (Blatch, Journal of Accountancy (AICPA), 2015) + (Journal of Accountancy (AICPA), 2024)

30 days. Anthropic states that for API users it "automatically delete[s] inputs and outputs on our backend within 30 days of receipt or generation", and that retained data "is never used for model training without your express permission." The stated exceptions are features whose retention the customer controls (such as the Files API), agreements to the contrary, enforcement of the Usage Policy, and legal requirements. Anthropic commercial data retention policy and developer documentation, covering the Claude API (api.anthropic.com), Claude Platform on AWS and Claude in Microsoft Foundry, where Anthropic is the data processor. On Amazon Bedrock and Google Cloud's Agent Platform the cloud provider is the processor and its own policy applies. Retrieved July 2026. (Anthropic Privacy Center, commercial retention) + (Anthropic, API and data retention)

Zero data retention is scoped, not global. Under a ZDR arrangement Anthropic "does not store customer prompts or responses at rest after the API response is returned", but the documentation lists what the arrangement does not cover: Console and Workbench; Claude Managed Agents; the Claude Free, Pro and Max consumer plans; the Claude Teams and Claude Enterprise product interfaces; Claude for Excel; the Claude Fable 5 and Claude Mythos 5 models, which are designated Covered Models and require 30-day retention; third-party integrations; and flagged content or legal holds. A separate eligibility table marks further features ineligible, among them code execution, the Files API, batch processing, the MCP connector and Agent Skills. ZDR is enabled per organization and requires an arrangement with Anthropic's sales team. Anthropic developer documentation, "What ZDR does not cover" and the feature eligibility table; retrieved July 2026. An ineligible feature is not blocked under ZDR: the documentation states that using one "is a choice to step outside your ZDR arrangement for that specific data", and the feature's own retention policy then applies. The exclusion list changes over time. (Anthropic, API and data retention)

Up to 2 years. Even with ZDR or HIPAA arrangements in place, Anthropic states it may retain inputs and outputs for up to two years where content is flagged by its automated trust and safety systems, or where retention is required by law. Anthropic developer documentation, "Retention regardless of arrangement"; retrieved July 2026. (Anthropic, API and data retention)

Five years versus 30 days. Under the consumer terms announced on 28 August 2025, Claude Free, Pro and Max users choose whether their chats and coding sessions are used to improve the model. Anthropic states retention of five years "if you allow us to use your data for model training", against the existing 30-day retention period if not. Existing users had until 8 October 2025 to make the selection. The update explicitly does not apply to Claude for Work (Team and Enterprise plans), Claude Gov, Claude for Education, or use via the API, Amazon Bedrock and Google Cloud's Vertex API. Anthropic announcement, consumer plans only. The announcement does not state a default setting; claims that the setting defaults to on are not established by this source. (Anthropic, 2025)

Since 1 March 2023, data sent to the OpenAI API "is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)". Abuse-monitoring logs are retained for up to 30 days, "unless longer retention is required by law, or is reasonably necessary to protect our services or any third party from harm". Zero data retention covers a named list of endpoints and is "subject to prior approval by OpenAI". OpenAI platform documentation, API data controls; retrieved July 2026. ZDR eligibility is granted per customer, not by default. (OpenAI, Data controls in the OpenAI platform)

59% of in-house legal professionals say they do not know whether their outside counsel is using generative AI on their matters, and 80% report that they are neither requiring nor encouraging its use. The second figure measures the absence of a stated position, not the presence of a restriction. ACC / Everlaw survey behind the report "Generative AI's Growing Strategic Value for Corporate Law Departments"; 657 in-house legal professionals across 30 countries; report published October 2025, figures published by ACC in November 2025. Fielding dates are not stated by either source, and the article reporting the two figures does not restate the sample. (Garcia & Whiteman, ACC Corporate Counsel Now, 2025) + (ACC & Everlaw, 2025)

A standard template exists, but no measured adoption rate. The Association of Corporate Counsel publishes sample artificial intelligence guidelines for outside counsel, covering disclosure of AI use, data security, and accuracy and performance. It is a model document, not evidence of how widely such clauses are used. Association of Corporate Counsel resource library; full text is member-gated. No primary survey quantifying the share of client contracts or engagement letters containing an AI restriction was found at the time of writing. Retrieved July 2026. (Association of Corporate Counsel)

The audit and traceability gap

Article 12 of the EU AI Act requires high-risk AI systems to technically enable automatic recording of events (logs) over the system's lifetime, ensuring a level of traceability appropriate to the system's intended purpose. Regulation (EU) 2024/1689; applies to providers and deployers of high-risk AI systems in the EU; obligations apply from 2 August 2026. (European Parliament & Council, 2024) + (European Commission, AI Act Service Desk)

Article 50 of the EU AI Act requires providers of AI systems generating synthetic audio, image, video or text to mark outputs in a machine-readable format, detectable as artificially generated or manipulated, as far as technically feasible. Regulation (EU) 2024/1689; transparency obligation applying from 2 August 2026. (European Parliament & Council, 2024) + (European Commission, AI Act Service Desk)

21% of organizations say they have a mature governance model in place for agentic AI. Among the listed missing capabilities are "audit trails that capture the full chain of agent actions to help ensure accountability." 3,235 IT and business leaders directly involved in AI programs, across 24 countries and 6 industries; data collected August-September 2025. (Deloitte, State of AI in the Enterprise, 2026)

13% of organizations reported breaches of AI models or applications, and of those compromised, 97% reported not having proper AI access controls in place. IBM Cost of a Data Breach Report 2025, research by Ponemon Institute; 600 organizations worldwide that suffered a breach, data collected March 2024-February 2025. (IBM / Ponemon Institute, 2025) + (Kiteworks, 2025)

40% of senior finance professionals are concerned about the lack of formal sign-off processes for AI-generated outputs. Survey of 311 senior finance professionals across 22 sectors, global, fielded March-April 2026. (insightsoftware, 2026) + (GlobeNewswire, 2026)

Regulation is arriving

Up to €35,000,000 or 7% of total worldwide annual turnover (whichever is higher) is the maximum administrative fine under the EU AI Act for breaching the prohibited-practices rules of Article 5. Lower tiers apply for other obligations (€15M or 3%) and for incorrect or misleading information (€7.5M or 1%). Regulation (EU) 2024/1689, Article 99; applies to providers and deployers under EU jurisdiction; for SMEs and startups the lower of the two amounts applies. (European Union (Parliament & Council), 2024) + (European Commission, AI Act Service Desk)

30% increase in legal disputes for technology companies by 2028 is predicted to result from AI regulatory violations. Gartner prediction, tied to a survey of 360 IT leaders involved in deploying GenAI tools, conducted May-June 2025. (Gartner, Inc., 2025) + (Analytics India Magazine, 2025)

92% of UK accountants and bookkeepers believe public AI tools should be regulated and/or restricted when providing financial or tax advice, including 70% who call for formal regulation. Censuswide survey of 500 UK accountants and bookkeepers across firm sizes, regions and sectors; fielded the first two weeks of December 2025; commissioned by an accounting-software vendor. (Dext (Censuswide), 2025) + (FinTech Global, 2025)

The governance maturity gap

63% of breached organizations either have no AI governance policy or are still developing one; only 37% have policies to manage AI or detect shadow AI. IBM Cost of a Data Breach Report 2025, research by Ponemon Institute; 600 breached organizations worldwide, March 2024-February 2025. (IBM / Ponemon Institute, 2025) + (The Actuary (IFoA), 2025)

31% of European organizations have a formal, comprehensive AI policy in place, while 83% of IT and business professionals in Europe believe employees in their organization are using AI (a perception, not a measured usage rate). 561 IT and business professionals in Europe, the European cut of a global ISACA survey of 3,200+ respondents; fieldwork 28 March-14 April 2025. (ISACA, 2025) + (Infosecurity Magazine, 2025)

12% of AI-using financial-services firms have an AI risk management framework, and only 18% have a formal testing program for their AI tools. ACA/NSCP 2024 AI Benchmarking Survey; 200+ compliance and risk leaders at financial-services firms (mostly asset managers), online survey June-July 2024. (ACA Group & NSCP, 2024) + (FinTech Global, 2024)

78% of senior business leaders lack full confidence that their organization could pass an independent AI governance audit within 90 days, a gap Grant Thornton calls the "AI proof gap." AI Impact survey of nearly 1,000 senior leaders across multiple US industries; early 2026; self-reported perceptions, not an attestation. (Grant Thornton, 2026) + (Journal of Accountancy (AICPA & CIMA), 2026)

One-third (33%) of AI use cases deployed by UK financial-services firms are third-party implementations, up from 17% in 2022; the survey flags third-party dependencies as the risk expected to grow most over three years, and notes 46% of firms have only a "partial" understanding of the AI they use. Third joint Bank of England / FCA survey on AI in UK financial services; 118 respondents across 6 sectors; conducted 2024, published 21 November 2024. (Bank of England & FCA, 2024) + (Stephenson Harwood, 2024)

48% of senior financial modellers report their organization has no formal policy governing AI use in financial model development; only 16% describe their policy as clear and documented. Financial Modeling Global Leaders Council survey; 63 senior modellers, 26 countries; 100% response rate; expert panel, not a representative sample; published July 2026. (Financial Modeling Institute, 2026) + (PR Newswire, 2026)

No option above 25%. Asked who should bear primary accountability when an AI-assisted model causes a material error, senior modellers split evenly: the human modeller (25%), shared accountability (22%), the organization (21%), the model owner / sponsor (19%) and the reviewer / approver (13%). The profession has not resolved accountability in an AI-assisted environment. Financial Modeling Global Leaders Council survey; 63 senior modellers; single-select; 100% response rate; expert panel, not a representative sample; published July 2026. (Financial Modeling Institute, 2026) + (PR Newswire, 2026)

The cost of failure

$670,000 in higher breach costs, on average, was observed at organizations with high levels of shadow AI compared with those having low or no shadow AI. IBM Cost of a Data Breach Report 2025, research by Ponemon Institute; 600 breached organizations worldwide, March 2024-February 2025. (IBM / Ponemon Institute, 2025) + (Cybersecurity Dive, 2025)

1 in 5 (20%) organizations reported a breach due to shadow AI, that is, unsanctioned AI tools adopted by employees without IT or security oversight. IBM Cost of a Data Breach Report 2025, research by Ponemon Institute; 600 breached organizations worldwide, March 2024-February 2025. (IBM / Ponemon Institute, 2025) + (Nudge Security, 2025)

37% of the time employees save using AI is lost to rework, including correcting errors, verifying outputs and rewriting low-quality content; only 14% of employees consistently get clear, positive net outcomes from AI. 3,200 full-time employees at organizations with $100M+ revenue, active AI users, split half leaders / half employees, across North America, APAC and EMEA; fielded November 2025. The official release rounds the rework figure to "nearly 40%." (Workday (fieldwork by Hanover Research), 2026) + (CFO.com, 2026)

A$97,000+ (about US$63,000) was refunded by Deloitte Australia on a roughly A$440,000 government contract after a report was found to contain references to nonexistent academic research and a fabricated quote from a federal court judgment; the revised version disclosed use of a generative AI system (Azure OpenAI / GPT-4o). Single documented case (October 2025); report commissioned by Australia's Department of Employment and Workplace Relations, errors identified by a University of Sydney researcher. (CFO Dive, 2025) + (Fortune, 2025)

50% of UK accountants and bookkeepers are aware of businesses that have suffered direct financial losses (overpayments, missed allowances, penalties, fines or compliance issues) after acting on incorrect or misleading AI-generated advice. Censuswide survey of 500 UK accountants and bookkeepers across firm sizes, regions and sectors; fielded the first two weeks of December 2025. The figure measures awareness of affected businesses, not losses suffered by respondents themselves. (Dext (Censuswide), 2025) + (FinTech Global, 2025)

Sources

Changelog

  • 2026-07: Source re-check of the new section against every underlying document. Corrected the AICPA block ("should", not "requires"; removed an unsupported notification claim) and re-sourced it to a second AICPA publication after the PwC Viewpoint link broke. Rewrote the Anthropic 30-day block, which contradicted itself, around the retention policy's own wording. Rebuilt the ZDR block: added the exclusions that were missing (Managed Agents, Covered Models, third-party integrations, flagged content, batch processing) and the fact that an ineligible feature is not blocked. Corrected "Claude for Government" to "Claude Gov" and completed the excluded-services list. Replaced the unsupported inference in the ACC block and added the ACC / Everlaw figures on outside-counsel AI visibility.
  • 2026-07: Added the section "Contractual and professional limits on AI use" (AICPA 1.700.001 / 1.700.040; Anthropic commercial retention, ZDR scope and exclusions, 2-year flagged-content retention, August 2025 consumer terms; OpenAI API training and retention controls; ACC sample outside-counsel AI guidelines). Documentary sources, not survey data.
  • 2026-07: Added Financial Modeling Institute / GLC "The Human Financial Modeller" figures (0-of-63 human-review consensus; 90% sign-off / 86% ethics stay human; top risk = reduced understanding 59%; 48% no formal AI policy for model development; accountability split with no option above 25%).
  • 2026-06: Initial publication.

Further reading


Compiled by Layerz.

Related reading