AI in finance fails on data, not models.

Almost every finance team I speak to is experimenting with AI right now.
ChatGPT for journal entries. Claude for analysing transaction exports. Copilot for summarising reports.
And the experiments mostly work. In isolation. On a clean file. With someone experienced enough to sense when the output feels off.
The problem comes when teams try to move from experimenting to relying on it.
That is where things break. Quietly. Confidently. In a way that is hard to catch.
The confidence problem
There is a specific failure mode with AI in finance that I keep coming back to.
AI does not say “I’m not sure.” It gives you an answer. A complete, well-structured, plausible answer.
Deloitte, December 2025: In high-stakes sectors like finance, where precision is paramount, a difference of only 0.5% could, in certain situations, amount to millions of dollars.
And the compounding risk is this: hallucinations do not sound like errors.
They are convincing answers delivered with complete confidence. If you are not already close enough to the numbers to feel the discrepancy, you will not catch it.
That is not a theoretical risk. That is what happens when a finance team runs a close on data pulled from a PSP, an OMS, a billing system, and the ERP (none of which fully agree) and the AI fills in the gaps with something that looks right but is not.
In a low-volume environment with an experienced team, someone usually catches it.
In a high-volume environment no one does, simply because people aren’t able to keep up and often (unwilfully) agree to estimations rather than facts.
Where AI adoption in finance actually stands
Here is what the research shows.
McKinsey’s November 2025 survey of finance leaders found that nearly two-thirds of organisations have not yet begun scaling AI across the enterprise. Pilots break down under real-world conditions, fail to adapt as new data emerges, and remain poorly integrated into core processes.
The tools are available. The adoption is not following.
Gartner is direct about why. In their February 2025 report on AI readiness, they predict that through 2026, organisations will abandon 60% of AI projects that are unsupported by AI-ready data. The same report found that 63% of organisations either do not have or are not sure they have the right data management practices for AI.
Most organisations investing in AI are building on a foundation they are not confident in.
McKinsey’s framing of what it actually takes is unambiguous: teams must rewire core processes, talent, and technology. Not just add new tools on top of old ways of working.
That word – rewire – matters. They are not talking about a better prompt or a different model. They are talking about what has to exist before the AI tool is even opened.
In a high-volume environment, rewiring is not abstract. It means the PSP, the OMS, the billing system, and the ERP are no longer four separate truths that finance stitches together at month-end. It means the relationships between them are maintained continuously, not reconstructed after the fact.
That is the foundation the research is pointing at. And it is the part that gets skipped.
The export problem
Most finance teams experimenting with AI are working from exports.
Pull the data from your ERP. Export to CSV. Load into the tool. Ask your questions.
This feels practical because it is how most finance work already happens. And for isolated questions, it produces something that looks like an answer.
But a CSV export is a snapshot. A copy of data that has already lost something in transit. It carries no memory of where each record came from. It has no understanding of which order matches which payment. It cannot tell you whether the settlement from last week has reconciled or is still open.
When you ask an AI tool to close the books on that data, you are asking it to reason over something structurally incomplete.
The answer will be confident. It may look right and it may be missing transactions that simply were not in the export. Not because anyone made a mistake, but because the data was never connected in the first place.
This is not a prompt engineering problem. You cannot fix it by asking the question more carefully. It is a data infrastructure problem.
The data lake illusion
Over the last few years, most organisations with serious data ambitions have invested in a data lake.
The promise was clear: centralise everything, and you always have access to what you need. For analytics, dashboards, and business intelligence, this often delivers.
For finance (specifically for high-volume transactional accounting) it does not.
A data lake stores data. It does not maintain the relationships between data. It cannot tell you whether a payment and an order belong together. It does not hold the chain from transaction to settlement to booking in a way that is live, reconciled, and traceable.
McKinsey’s analysis of the path toward the data-driven enterprise flagged this precisely. Seventy percent of organisations report difficulties integrating data into AI models quickly. Not because the data does not exist, but because it is not structured in a way that AI can actually use.
When an AI tool queries a data lake for your month-end position, it gets access to a lot of data. But it does not get the structured truth underneath it.
In a high-volume environment, with tens of thousands of transactions across multiple payment methods, currencies, and channels per day, that gap is not small.
What infrastructure for high-volume accounting actually means
I want to be specific here, because data infrastructure is easy to say and harder to define.
For high-volume accounting, it means three things.
One maintained source of truth.
Not an export. Not a warehouse snapshot. A system that continuously holds the relationship between orders, payments, settlements, and bookings. It also means a system that knows, at any point in time, what has reconciled and what has not.
Reconciliation that happens before anyone asks for it.
The problem with running reconciliation inside an AI tool is that you are doing it downstream, after data has moved through multiple systems and lost its connective tissue. Reconciliation needs to happen at the point of transaction, continuously and automatically. By the time you ask a question, the answer should already exist.
Traceability that starts before the ERP.
Most ERP systems assume they receive clean, correct data. In a high-volume environment, that assumption breaks regularly. Traceability has to exist upstream, from the moment a transaction occurs, or you cannot verify what the AI is working with.
When this infrastructure exists, AI becomes genuinely useful for finance. Not because the model changed. Because what you give the model changed.
You are no longer handing AI a snapshot of what might have happened. You are giving it a live, reconciled, structured dataset. The questions you can ask (and the answers you can trust) are completely different.
The human backstop that does not scale
There is a reason AI experiments hold up in experienced finance teams.
Senior finance managers catch things. They know their numbers well enough to feel when something is off. They push back. They verify.
That instinct is not a process. It is not something you can document or delegate. It is the product of years of closings and deep familiarity with this business, these numbers, this month.
In a low-volume environment with an experienced team, that backstop works.
But in a high-volume environment, it does not scale. The volume is too large, the data too distributed, the cycle too compressed for any individual’s instinct to cover it.
The risk is not that AI will obviously fail. The risk is that it will silently fail in a way no one is positioned to catch.
What this means for finance teams right now
Keep experimenting. Seriously.
The teams building understanding of AI in their context right now are ahead of the ones waiting for the technology to mature. McKinsey’s data is clear: the organisations capturing real value from AI are not the ones with the best tools. They are the ones that rewired how they work.
But know what experiments can teach you and what they cannot fix.
Experiments show you which questions AI answers well. They help you understand where AI speeds things up and where you still need human judgment.
What they do not reveal: whether your data foundation is strong enough to support the answers AI gives you.
That is a separate question. And it is the one that determines whether AI in your finance function ever moves from pilot to production.
The 60% abandonment rate Gartner is projecting is not a failure of ambition. It is a failure of sequence. Teams are investing in AI before investing in the infrastructure that makes AI trustworthy.
The teams that will get this right in finance are the ones that solve the data problem first. Not with a better model. With a better foundation.
The question worth asking
If you are experimenting with AI in your finance function (or trying to understand why the experiments keep stalling) the first question to answer is not which tool should I use. It is: can the AI actually see what happened?
If the honest answer is exports, snapshots, or whatever can be pulled together for the run, there is upstream work to do.
That work is less visible than an AI rollout. It does not make for a good board demo.
But it is the work that separates the teams who trust their numbers from the ones still manually checking them every month.
Sources
McKinsey — How Finance Teams Are Putting AI to Work Today (November 2025)
McKinsey — The State of AI in 2025: Agents, Innovation, and Transformation (November 2025)
McKinsey — Charting a Path to the Data- and AI-Driven Enterprise of 2030 (September 2024)
Gartner — Lack of AI-Ready Data Puts AI Projects at Risk (February 2025)
Deloitte Switzerland — AI Doesn’t Lie, It Hallucinates (December 2025)
