TL;DR: In July 2026 the FCA published a review of how 56 firms practice Consumer Duty outcomes monitoring. The report is mostly qualitative, but four measured results stand out, each from a firm that changed something and tracked the effect. One firm cut first response time from 22 hours to under 2 minutes, and resolution time from 4 days to under 3 hours, over six months, after adding AI-powered in-app chat. One firm raised transaction categorisation accuracy by 12.8% after refining its rules, giving it more reliable monitoring data. One firm improved anti-money laundering pass rates by 5% and bank verification pass rates by 20% after piloting a multi-bureau verification approach. The common thread: the FCA treats the full loop as the standard. Find a problem in the data, act on it, and measure whether the outcome improved.
Introduction
The FCA has published a review of how firms monitor customer outcomes under the Consumer Duty. It surveyed 56 firms and set out what good and poor monitoring looks like in practice. You can read the full publication here: Outcomes monitoring: good practice and areas for improvement.
Most of the report is qualitative. But it contains a handful of specific figures from firms that changed something and measured the result. These numbers are useful for two reasons. They show the scale of improvement the FCA treats as good practice, and they give compliance and operations teams a benchmark they can point to when they make the case for a change internally.
Here are the verified figures from the report, what each firm actually did, and what the numbers tell you.
What figures did the FCA report on Consumer Duty outcomes monitoring?
The report includes four measured results:
- First response time cut from 22 hours to under 2 minutes over six months, after a firm added AI-powered in-app chat.
- Resolution time cut from 4 days to under 3 hours over the same period, at the same firm.
- A 12.8% rise in transaction categorisation accuracy, after a firm refined its categorisation rules.
- A 5% improvement in anti-money laundering pass rates and a 20% improvement in bank verification pass rates, after a firm piloted a multi-bureau verification approach.
Each one comes from a firm that identified a problem in its data, made a specific change, and tracked whether the change worked. The measurement is the point. The FCA was clear that buying a tool or running a process is not evidence on its own. A firm has to show the outcome improved. This is the same principle behind effective Consumer Duty monitoring: the data has to lead to a decision, and the decision has to be shown to work.
What results did firms report to the FCA?
| What the firm changed | Metric | Before | After |
|---|---|---|---|
| Added AI-powered in-app chat | First response time | 22 hours | Under 2 minutes |
| Added AI-powered in-app chat | Resolution time | 4 days | Under 3 hours |
| Refined transaction categorisation rules | Categorisation accuracy | Baseline | +12.8% |
| Piloted multi-bureau verification | AML pass rate | Baseline | +5% |
| Piloted multi-bureau verification | Bank verification pass rate | Baseline | +20% |
How did one firm cut response time from 22 hours to 2 minutes?
The firm used customer feedback and operational data to find that customers wanted a faster way to raise queries outside the formal complaints process. It introduced in-app chat so customers could get an immediate response, then expanded the chat with AI to route customers to the right support team. The same AI used keyword recognition to spot possible signs of vulnerability and escalate those cases for prioritised handling.
It then measured the result. Average first response time dropped from 22 hours to under 2 minutes over a six-month period. Average resolution time dropped from 4 days to under 3 hours.
The lesson the FCA draws is about the full loop, not the chatbot. The firm used contact data to find a problem, made a targeted change, and monitored the data again to confirm the change worked and to identify customers who needed more support earlier. Spotting vulnerability indicators at scale is exactly the kind of task AI monitoring handles well, and it sits at the centre of how firms are expected to support customers in vulnerable circumstances under Consumer Duty.
What does the 12.8% categorisation figure mean?
One firm used transaction-level data and customer testing to find that some payments were being sorted into the wrong categories. HMRC payments were landing in the wrong place, and gambling-related transactions were not always recognised consistently. Miscategorised transactions meant the firm’s own monitoring data was unreliable.
The firm tested sample transaction descriptions against external reference data, refined its categorisation rules, and removed false positives. Categorisation accuracy rose by 12.8%.
The value here is upstream of any single customer outcome. Better categorisation gave the firm a more reliable view of what customers were doing and where it needed to look closer. If the data feeding your monitoring is wrong, every judgement built on it is weaker. Our practical playbook on AI compliance monitoring covers why clean, connected data is the foundation of a Consumer Duty outcomes monitoring programme the regulator will trust.
What did the verification pilot improve?
One firm’s complaints data showed that customers hit delays when they posted original or certified identity documents to complete withdrawals. The firm piloted a multi-bureau verification approach before rolling it out in full, and tracked the result through committee and board reporting.
Anti-money laundering pass rates improved by 5%. Bank verification pass rates improved by 20%. The firm used complaints data to find friction in a specific customer journey, tested a change, and confirmed it worked before committing to it.
Why do these numbers matter for Consumer Duty outcomes monitoring?
Every figure in the report follows the same shape. A firm found a problem in its data, made a change, and measured whether outcomes improved. The FCA treats that loop as the standard, and it is the part firms most often leave incomplete.
The report is direct on this point. Collecting data, listing metrics or reporting management information does not, by itself, show whether customers are getting good outcomes. A firm has to be able to explain what its information tells it, how it uses that to spot risk, what action it takes, and whether that action improved things.
The numbers are worth citing because they show what a completed loop looks like. Not a dashboard, but a measured before-and-after that a firm can put in front of its board or the regulator.
What does the FCA expect from outcomes monitoring?
| The FCA expects firms to explain | In practice |
|---|---|
| What their information tells them | The data is read, not just collected |
| How they use it to identify risk | Findings are turned into risk signals |
| What action they take in response | A decision or change follows |
| How they check outcomes improved | The change is measured after the fact |
How Detect supports outcomes monitoring
Aveni Detect was built to close that loop across every customer interaction. It assesses every call rather than a small sample, shows the reasoning behind each judgement, splits risk by cause including vulnerability drivers, and keeps a record a compliance officer can retrieve months later. That gives firms the before-and-after evidence the FCA asks for, at the scale the Consumer Duty now demands.
See how Detect assesses 100% of your interactions and produces regulator-ready evidence. Book a demo.
All figures are drawn from the FCA’s “Outcomes monitoring: good practice and areas for improvement,” published 27 July 2026, based on a survey of 56 firms. The firm examples are the regulator’s own, anonymised in the source report.