manual qa coverage

The 182,408 high-risk conversations nobody reviewed

Detect has assessed 1.38 million customer conversations across Aveni’s customer base. If those firms had reviewed calls the standard way, sampling 2% of them manually, 182,408 high-risk calls would have gone unreviewed. Those calls contained vulnerability signals, complaints and dissatisfaction. Nobody would have heard them.

Sampling was built for one channel

QA teams at lenders, banks and collections firms review a small share of their interactions each month. Most firms sample fewer than 5%. The common benchmark is 2%. Reviewers pull a selection, score it against the firm’s framework, log the results and share findings with team leads.

That approach was designed when customer contact meant a phone call, and reviewing one meant a person with a headset and a spreadsheet. Customer contact has since spread across channels. A customer might open a complaint by webchat, follow up by email, get a text about a payment plan and only reach the phone at the point of escalation. The complaint, the vulnerability disclosure and the conduct issue can each sit in a different channel.

Most QA programmes still centre on voice. Check what your own firm reviews across webchat, email and SMS, and how those reviews connect back to the same customer. If voice is sampled at 2% and the other channels are reviewed separately or not at all, coverage across the full customer journey is lower than the headline sampling rate suggests.

Risk also refuses to spread evenly. A customer disclosing financial difficulty, an agent mishandling a complaint, a collections conversation that treats a customer unfairly: each sits in a specific interaction. A random 2% sample is as likely to pull a routine balance enquiry as any of them. The riskiest interactions and the most routine ones have the same odds of being picked.

What the unreviewed conversations contained

Detect’s data shows what firms miss. Across the 1.38 million conversations it assessed, Detect surfaced 17.6 million compliance risks. These included 126,000 vulnerability indicators and 165,100 complaints, 59,000 of which were dissatisfaction signals.

At 2% coverage, 182,408 of the high-risk calls in that population would never have reached a reviewer. Each one is a customer whose vulnerability went unrecorded, a complaint that surfaced weeks later through a formal channel, or a conduct problem that continued because nobody saw it.

These misses matter under FCA rules. FG21/1 requires firms to identify vulnerable customers and respond appropriately. Consumer Duty (FG22/5) requires firms to monitor whether customers get fair outcomes and to evidence it. Both apply to every interaction on every channel. A firm that reviews 2% of its calls holds evidence for 2% of its obligations.

What you can show the regulator

When the FCA asks a firm to evidence fair outcomes across its customer journeys, a sampled QA programme supports one answer: here is what we found in the calls we reviewed. It says nothing about the other 98%.

Under SM&CR, senior managers attest that the controls they oversee work. When QA covers 2% of the interactions those controls apply to, the attestation rests on a thin evidence base.

What Detect does instead

Detect assesses every customer interaction automatically, across voice, email, chat, SMS and documents. Each interaction is scored against the firm’s own framework: the assessment criteria, scoring rules and standards the QA team defines. FinLLM, Aveni’s financial-services-specific AI, powers the analysis. It is built for regulated environments and recognises what matters in a conversation, including the vulnerability indicators in FG21/1 and the outcomes Consumer Duty requires firms to monitor.

Human review stays an important part of the process. Detect flags the cases that need review and routes them to the right person with the evidence attached. Reviewers check, override and sign off. While AI handles the volume, people make the final call on flagged cases.

Detect has completed 1.24 million AI assessments across the customer base, roughly 77 automated assessments for every one human review. That has saved an estimated 118,000 hours of review time, equivalent to 66 full-time staff, worth around ÂŁ2.6 million. Firms using Detect run QA up to 6x faster with the same headcount, and reviewers spend their time on the highest-risk cases first.

Run the numbers on your own QA programme

Take your monthly interaction volume and multiply it by your sampling rate. That is the number of interactions your oversight touches. Everything above it goes unreviewed.

The Detect data shows what that costs. At 2% coverage, roughly one high-risk call in every eight goes unseen. If your compliance function is accountable for Consumer Duty outcomes, vulnerability identification and conduct risk on every channel, the answer to “did anyone review this interaction?” should always be yes.

Frequently asked questions

What percentage of customer calls do compliance teams review? Most FCA-regulated firms sample fewer than 5% of customer interactions, and 2% is the common industry benchmark for manual QA coverage. The remaining interactions go unreviewed, including the vulnerability signals, complaints and conduct issues they contain.

Does the FCA require firms to review every customer interaction? The FCA sets no fixed coverage percentage. Consumer Duty (FG22/5) requires firms to monitor and evidence fair customer outcomes, and FG21/1 requires firms to identify and respond to vulnerable customers. Firms decide how much oversight satisfies those obligations, and sampled evidence covers only the interactions reviewed.

Why does sample-based QA miss high-risk interactions? Risk sits in specific interactions, such as vulnerability disclosures, complaints and poor conduct, while random sampling selects with no regard to risk. Across 1.38 million conversations assessed by Detect, 2% sampling would have missed 182,408 high-risk conversations.

How does AI compliance monitoring detect vulnerable customers? Detect assesses every interaction against FCA FG21/1 vulnerability guidance, flagging indicators such as financial difficulty disclosures across voice, email, chat, SMS and documents. Flagged cases route to a human reviewer with the supporting evidence attached. Across its customer base, Detect has surfaced 126,000 vulnerability indicators.

Can compliance teams achieve 100% QA coverage without hiring more staff? Yes. Detect completes roughly 77 automated assessments for every one human review, so full coverage runs alongside existing headcount. Across the customer base this has saved an estimated 118,000 review hours, equivalent to 66 full-time staff, with QA cycles up to 6x faster.

Share with your community!

In this article

Related Articles

Join our newsletter

Be the first to hear about new features, releases, and best-practice guides.

Aveni AI Logo