Sector
AI consultancy for financial services
Model risk expectations, Consumer Duty evidence and a supervisor who will ask how the decision was reached. AI in financial services is not blocked by capability. It is blocked by explainability and data control.
The constraint
Why financial services AI projects stall
Financial services firms were doing model governance long before generative AI arrived, and the existing framework does not bend for it. A model that influences a customer outcome needs an owner, documented validation, ongoing monitoring and an explanation of its behaviour that a second line function can challenge.
Generative models are awkward against every one of those requirements. Output varies between identical runs, the reasoning is not directly inspectable, and the vendor changes the model underneath you without notice. That last point is the one that causes the most trouble: a validated model that silently becomes a different model has invalidated its own validation.
The data side compounds it. Customer records, transaction histories and unpublished positions are precisely the material that makes an internal assistant useful, and precisely the material nobody wants leaving a controlled environment.
Where it works
What actually earns its place here
Ordered roughly by how quickly they get approved. The first item on this list is usually the right first project, precisely because it is the least contentious.
Policy, procedure and regulatory retrieval
Staff answering questions against the handbook, internal policy and regulatory material, with citations to the governing source. Low risk, high volume, and the accuracy is measurable: a good first build.
Complaint and correspondence triage
Classification, routing and summarisation of inbound complaints, with the vulnerability and Consumer Duty flags surfaced rather than decided. The human keeps the judgement; the machine removes the reading.
Document and KYC extraction
Structured field extraction from statements, certificates, applications and correspondence. Accuracy is measurable against a ground truth, which makes it defensible to validate.
Suitability and file review support
Surfacing files that warrant human review against defined criteria, as a prioritisation aid rather than an assessment. The distinction matters enormously to your second line.
The gate
What your assurance function will ask for
We build so that this evidence is a by-product of delivery rather than a document assembled under pressure afterwards. It is markedly cheaper that way, and considerably more likely to be accurate.
- Model documentation covering purpose, data, limitations and monitoring
- Version pinning, so a validated model does not silently change underneath you
- Logged inputs and outputs sufficient to reconstruct any individual decision
- Explicit statement of where the model does and does not influence customer outcomes
- Human oversight design that is resourced realistically rather than claimed
- Data residency and retention position, written for second line and audit
Illustrative scenario: a composite example of how an engagement typically runs, not a specific client.
Illustrative: a mid-size lender
A lender wants to reduce the time advisers spend searching policy documents for the correct current criteria. The material is internal, non-personal, and changes frequently, meaning the existing process fails not because people are slow but because superseded versions circulate.
The classification work establishes that the corpus itself carries no personal data, so a private UK deployment is sufficient and self-hosting is unnecessary. The build is retrieval with mandatory citation, an effective-date field on every document, and an owner accountable for retiring superseded versions, which turns out to be the change that does most of the work.
The evidence pack covers model version pinning, the evaluation set and measured accuracy, and a written position on why this system does not influence customer outcomes directly. That last document is what allows second line to approve it quickly.
Questions from this sector
Can generative AI ever be used where it affects customer outcomes?
It can, but the bar rises steeply and the architecture has to change: version pinning, comprehensive logging, documented validation against a held-out set, monitoring for drift, and a genuine human decision point. Most firms are better served by starting where the model supports a human rather than influences a customer, building the governance muscle there, and moving closer to the customer once the controls are proven.
How do we validate a model whose outputs vary between runs?
By validating the system rather than treating the model as a deterministic function. That means a held-out evaluation set with reviewed correct answers, measured performance with variance reported rather than hidden, pinned model versions so the thing you validated stays the thing you run, and monitoring that would detect degradation. It is a different validation approach, not an impossible one.
Does our data leave the UK?
Not if the architecture is designed so it cannot. That is the point of establishing the data boundary before choosing components, and of producing a written residency and processing position that your second line and your auditors can rely on rather than infer.
Start with a straight answer
A 30-minute call, no pitch deck. Tell us what you are trying to do and we will tell you whether AI is the right tool, what it would take, and what it would cost, or that you should not bother.