Why financial services AI pilots are still stalling, three years into the hype
The gap between AI investment and AI returns in financial services has become almost uncomfortable to acknowledge. Experts at Alpha FMC explain where pilots are failing to scale, and how firms can finally start seeing returns on their huge levels of AI spending.
Firms have spent tens of billions on artificial intelligence capabilities over the past three years. Yet according to research published by Valliance in mid-2026, 48% of mature organisations with established AI programmes still have AI initiatives stuck in pilot stage. A year earlier, two-thirds of financial services CEOs expected to move past the piloting phase; when surveyed again, 60% were still in it. Three recurring barriers lie behind this gap: fragmented data, governance shortcomings and organisational readiness.
The usual explanation is that pilots fail because of bad execution, or because firms overestimate AI’s capabilities, or because the business case was never solid to begin with. But in regulated financial services, the problem runs deeper. Pilots stall not because the technology doesn’t work, but because the conditions required to deploy it safely and auditably are rarely present when consultants arrive to scope the work.
The difference matters. In a manufacturing setting or a marketing analytics context, an AI system that is 90% accurate might be acceptable. In financial services compliance, it is not. A document review system that misses errors 10% of the time cannot be used in decisions affecting regulatory exposure. A model that cannot explain why it recommended a particular finding cannot defend that finding to a regulator. And a governance framework that has no clear rules for monitoring and updating the system will be rejected by any compliance function worth its salt, no matter how much the business wants the efficiency gain.
Financial services firms are not wrong to be sceptical about AI. They are doing what a regulated business should do: they are holding AI to a standard that matches the stakes of the work it is intended to do. The problem is that most consulting engagements are not designed around that standard. They are designed around speed and demonstrable progress in the early phases. What looks like successful delivery at the end of phase one tends to fall apart once it meets the reality of how a regulated business actually operates.
Three reasons AI pilots fail to scale
Research into AI pilots across European enterprises identifies three barriers to scaling: data infrastructure, governance and organisational readiness. In regulated financial services, each manifests in a way that makes it harder to overcome.
The first is data. When customer, operational and financial data sit in separate systems without a shared structure, no model can build a reliable picture of how a business operates. This is a general problem. In financial services, it is a specific liability. A compliance model that fills data gaps with assumptions will produce findings that quietly contradict what the compliance team sees in their day-to-day work. Once that happens, the team stops trusting the model. And once they stop trusting it, they stop using it. The difference between a pilot that scales and a pilot that dies is often whether the compliance team believes the findings.
The second is governance. Without clear rules on how models are trained, monitored and updated, outputs cannot be used in decisions that carry real risk. That affects accountability and compliance in obvious ways. It also affects something more basic: the ability to explain an outcome when it is questioned. A model that recommends a regulatory finding but cannot tell you which rule it was checking against, or which version of the model made the decision, or when that version was last validated, is a problem. In most regulated industries, that is not acceptable. An AI system that produces results nobody can account for is not a useful system, whatever the prototype looked like.
The third is organisational readiness. Introducing AI changes how decisions are made and who is responsible for them. If the compliance team has not been prepared to act on the model’s output, and if there is no clear ownership of where AI fits into existing workflows, the output stays separate from the work it was built to support. It becomes a report that nobody reads rather than a tool that changes how work happens. In financial services, this is compounded by the fact that the people using the system are often not the people who asked for it. The business case is made to the CFO or CRO, but the system will be used by compliance officers and product specialists who may have had a different set of expectations, or no expectations at all.
What consultants should be doing differently
None of this is especially complicated in theory. The issue is the sequence of work and where the difficult stuff happens to be positioned.
Many AI consulting engagements begin with a phase of rapid model development. A proof of concept is scoped, data is gathered from wherever it is available, a model is trained, and a demo is built. This work is visible and straightforward to bill. A model trained and configured within weeks is easy to present as progress. The harder work, getting data governance right, designing audit trails into the model, preparing the business to use it, is positioned for later phases. Phases that have a habit of not arriving.
A more effective approach to AI consulting in financial services should reverse that sequence. It should start with the conditions that actually allow AI to be deployed safely. That means understanding how data is structured across a client’s systems before any model is configured. It means setting governance rules before outputs go anywhere near a real decision. It means preparing teams to act on what they are getting, and creating clear ownership of where AI fits into existing processes.
This is slower work. It is harder to demonstrate as progress. It does not produce a working prototype that the business can see and touch in week six. But it produces an engagement that actually completes. And it changes the likelihood that a pilot becomes business as usual rather than an expensive experiment.
Ilia Prokashev, Senior Manager at Alpha Financial Markets Consulting, who led the build of the Alpha Concord compliance engine, which launched in June 2026, puts it this way: "The process mattered more than the model. We built Concord around the sequence a compliance or assurance task actually requires: define the rule first, then decide how a finding gets validated, and only then let the system loose on real documents. That ordering is what removes the room for a mistake to slip through unnoticed. If you build it the other way round, speed to production, you inherit exactly the risk you were trying to reduce."
A differently structured starting point
One example of this approach is Concord, which was built specifically for fund compliance, a context where every finding needs to be traced back to a source rule, where the governance framework needs to be auditable and versioned, and where the compliance team using the system needs to understand exactly what it is checking for. The system was designed around auditability from the start, not retrofitted.
The approach began with governance. What rules should every fund document comply with? Which of those can be checked automatically? Which require human judgment? Which require reference to regulatory guidance that changes frequently? Only once those questions were answered was model development scoped. The result is a system whose findings are both accurate and defensible in a regulatory context. Every finding is linked directly to the rule it was checking against. Every decision is traceable to a specific version of the checking logic. Every check is mapped to the relevant regulatory requirement or internal standard.
That structure makes it possible for compliance teams to trust the output. In testing, Concord achieved 99.8% accuracy on compliance checks, compared with 92.5% on manual review of the same population. It reviewed 100% of documents rather than the typical 2-10% sample coverage. And it did so at 8x the speed of a human review team. Those are not incidental benefits. They are the outcome of building the system around auditability and governance from the start.
Part of how that accuracy holds up is a layered review process that happens before any business stakeholder sees a result at all. Isaac Unsworth explains: "We don’t rely on a single check. Every finding goes through an automated pattern check first, looking for the kind of error signature we’ve seen before. Then we use a separate AI system as a judge, essentially asking a second model to mark the first model’s work against the same rule set. Only after that does the output go to a manual review by someone on the implementation team. Three different lenses on the same output. It sounds slow, but it’s what lets us stand behind the accuracy numbers once something reaches a client."
That layered structure, automated pattern checks, an AI acting as judge, and a final manual check, is a small operational decision with large consequences. It means no single point of failure is deciding what a compliance officer eventually sees. It also means that when a finding does reach a stakeholder, it has already survived three different ways of being wrong.
The case for a different approach
Financial services firms are facing a steady rise in the volume of documents they need to review. The transition from KI(I)Ds to Consumer Composite Investment disclosures is one inflection point. The expansion of ESG reporting and climate risk disclosure is another. The growth of fund populations and the complexity of fund structures is a third.
At the same time, regulatory scrutiny of investor communications and disclosure governance is intensifying. Firms that cannot demonstrate they have reviewed documents carefully and consistently are facing real regulatory risk. That combination, rising volume, static or shrinking compliance headcount, rising regulatory expectations, creates a mandate for transformation. But that mandate is not satisfied by a model that works in a sandbox. It is satisfied only by a model that integrates into how the business actually operates.
Consultants that design AI programmes around these realities are more likely to help clients move beyond perpetual pilots and into production. Firms know they need AI in their compliance workflows. They just need it built in a way that a regulated business can use. In regulated financial services, auditability cannot be treated as a feature to add later. It has to shape the engagement from the outset. The pilots that scale are the ones that tackle the hardest problems first.


