For years, banking QA has relied on a familiar process: select a small sample of customer interactions, review them manually and use the findings to understand how the wider operation is performing.
That model becomes harder to scale as customer journeys spread across calls, email, webchat, case notes and documents.
For Heads of QA evaluating quality assurance software, the buying decision therefore needs to go further than asking how many calls a platform can score.
The more important test is whether the technology can identify the interactions that matter, assess customer outcomes consistently, connect evidence across channels and give human reviewers enough context to make the final decision.
TL;DR: what should banks look for in quality assurance software?
Quality assurance software for banking should help QA teams move from limited random sampling towards broader, risk-based oversight while keeping humans responsible for material decisions.
Look for software that can:
- assess interactions across voice, digital channels, documents and case notes
- apply your existing QA and compliance frameworks consistently
- prioritise interactions according to risk
- assess customer outcomes alongside employee performance
- monitor Consumer Duty and vulnerability indicators
- connect multiple interactions into one customer case
- explain every automated assessment using supporting evidence
- allow human review, override and sign-off
- create retrievable audit trails
- integrate with existing telephony, CRM and case-management infrastructure
- meet the data governance requirements of a regulated financial-services firm
The FCA does not prescribe a fixed percentage of interactions that firms must review. Its Consumer Duty framework does, however, require firms to understand and evidence the outcomes their customers receive, with outcomes monitoring remaining an active FCA focus in 2026.
That changes what banks should expect from QA technology.
Why are banks reconsidering manual QA sampling?
Manual sampling made sense when every assessment required a person to listen to a complete call.
Its limitation is mathematical.
If a bank reviews 2% of its interactions, 98% sit outside that manual QA sample.
Risk is not evenly distributed across those interactions. A routine balance enquiry and an interaction involving financial difficulty have the same chance of appearing in a random sample.
Across 1.38 million customer conversations assessed by Detect, Aveni found that 182,408 high-risk conversations would have gone unreviewed at 2% manual coverage. Those conversations contained signals including complaints, dissatisfaction and vulnerability. See what 2% QA coverage misses across five channels.
The same banking data shows approximately 77 automated assessments for every manual review, while average assessment time has fallen from around 90 minutes to 15 minutes.
Banking interactions are becoming harder to sample effectively
The Financial Ombudsman Service received 214,600 new complaints in 2025/26. Current accounts alone generated 32,900 complaints, including 18,900 relating to fraud and scams.
UK Finance reported ÂŁ1.17 billion in authorised and unauthorised fraud losses during 2024, including ÂŁ450.7 million in APP fraud losses.
That creates very different QA requirements across:
- fraud disputes
- account restrictions
- collections
- financial difficulty
- complaints
- vulnerable customer interactions
- everyday customer service
A random sample cannot guarantee that the interactions with the greatest potential customer or regulatory impact will reach a reviewer.
This is why selection and prioritisation should be one of the first capabilities assessed when comparing quality assurance software.
Step 1: decide what your quality assurance software needs to measure
Start with the purpose of the QA programme before evaluating the technology.
Traditional QA scorecards tend to ask:
- Did the agent follow the process?
- Was mandatory wording used?
- Was the interaction handled professionally?
- Was the correct procedure followed?
Those remain useful measures.
But they do not necessarily establish what happened to the customer.
Performance scoring and customer outcome assessment are different
The banking Detect material separates these disciplines clearly.
Performance QA measures how an employee behaved.
Outcome-focused QA asks whether the customer received a fair and compliant outcome and whether the firm can evidence that outcome afterwards.
| QA requirement | Performance-focused QA | Outcome-focused QA |
|---|---|---|
| Employee followed process | Primary measure | Supporting context |
| Communication quality | Commonly assessed | Assessed against customer understanding |
| Customer received appropriate outcome | May be inferred | Directly assessed |
| Vulnerability handled appropriately | Depends on sampled interaction | Evaluated as part of outcome |
| Decision can be evidenced | Depends on reviewer notes | Supporting evidence retained |
| Patterns can be analysed across the population | Limited by sample | Possible with broader assessment |
This distinction should influence the vendor shortlist.
A platform can automate employee scoring very effectively without necessarily giving compliance teams stronger evidence of customer outcomes.
Step 2: establish what the software means by “coverage”
Claims of high or complete QA coverage need to be examined carefully.
A platform might analyse every telephone call while leaving email, webchat, complaint correspondence and case documents outside the same QA workflow.
That still creates blind spots.
Consider a customer who:
- reports suspicious activity through webchat
- speaks to the fraud team by phone
- receives an account restriction notice
- provides documents
- raises a complaint by email
Assessing the telephone interaction alone does not establish what happened across the customer journey.
Look for case-level quality assurance software
Banking QA increasingly needs to connect evidence across different interaction types.
Detect can assess voice recordings and written evidence together, including account restriction letters, collections documents and complaint correspondence. The resulting customer journey can then be presented to the reviewer as one case.
This also reduces preparation time.
A reviewer working manually may have to locate the recording, case notes and customer correspondence before the substantive assessment begins. In the banking Detect workflow, a review that previously took 60 to 90 minutes can take around 15 minutes once that evidence is assembled.
Step 3: assess how quality assurance software prioritises human review
Automation should help QA teams spend human attention more deliberately.
That requires more than automatically choosing another sample.
A risk-based QA system first assesses the interaction, then uses what it finds to determine whether human investigation is required.
For a bank, potential prioritisation criteria might include:
| Signal | Example |
|---|---|
| Complaint risk | Customer repeatedly expresses dissatisfaction |
| Vulnerability | Financial difficulty or another FG21/1 indicator |
| Conduct risk | Required process or explanation appears incomplete |
| Fraud | APP fraud dispute with unclear customer outcome |
| Consumer Duty | Customer understanding is not evidenced |
| Collections | Support pathway does not appear to have been followed |
| Documentation | Written evidence conflicts with what was communicated |
Detect pre-assesses interactions and routes higher-risk cases into the QA workflow, while analysts retain responsibility for the final decision.
This is fundamentally different from using automation simply to increase the quantity of randomly selected interactions.
Aveni explores this workflow further in How to prioritise high-risk calls for QA review.
Step 4: make sure the software can use your QA framework
A bank should not have to abandon a well-established QA methodology because a software provider has created its own generic scorecard.
Quality assurance software should be configurable around the firm’s existing controls, including:
- Consumer Duty outcomes
- vulnerability
- complaints
- conduct
- journey-specific requirements
- escalation criteria
- product-specific controls
- internal QA standards
Detect configures a firm’s own QA logic into its scorecards rather than replacing it with a standardised framework. Those criteria can then be applied consistently across the interaction population.
Questions to ask vendors about QA configuration
| Question | What to look for |
|---|---|
| Can our current QA framework be replicated? | Configurable criteria |
| Can different journeys use different scorecards? | Journey-level configuration |
| How are framework changes governed? | Version control and permissions |
| Can analysts see why a result was returned? | Explainability |
| Can a human override the result? | Reviewer control |
| Are assessments comparable over time? | Consistent definitions and reporting |
Consistency matters because manual QA introduces inevitable reviewer variation.
Automation should reduce that variation without removing the judgement of experienced QA professionals.
Step 5: require an explanation and evidence for every assessment
A QA score without supporting evidence has limited value in a regulated environment.
For an automated assessment to be useful, the reviewer should be able to see:
- the result
- why that result was reached
- the underlying evidence
Detect structures checks in this way. Each assessment can return an outcome, an explanation and supporting evidence from the customer interaction.
That becomes particularly important when QA results are used for:
- Consumer Duty reporting
- board MI
- compliance monitoring
- complaint investigations
- FOS referrals
- coaching
- root-cause analysis
The FCA’s Consumer Duty resources now explicitly include outcomes monitoring, board reporting, consumer understanding and vulnerable customers within its published good and poor practice materials.
For the latest monitoring data, see Aveni’s Consumer Duty Outcomes Monitoring: 4 FCA Figures to Cite.
Step 6: evaluate vulnerability monitoring as its own capability
Vulnerability should not be buried inside a generic compliance feature list.
The FCA’s FG21/1 guidance expects firms to understand customer needs, enable staff to recognise vulnerability, respond appropriately and monitor the outcomes vulnerable customers receive. Its guidance page was updated again in July 2026 to point firms towards the latest Consumer Duty expectations for management information and outcomes.
For banking QA, signs of vulnerability may emerge during:
- collections interactions
- account restrictions
- fraud investigations
- bereavement
- affordability discussions
- complaints
- ordinary customer service conversations
The banking Detect configuration identifies vulnerability and financial-difficulty signals across fraud, account restriction and collections journeys, with evidence attached to flags and escalation pathways built into the workflow.
When comparing quality assurance software, assess whether the system can do all three of the following:
identify → evidence → route
A vulnerability alert with no evidence simply gives a QA analyst another case to reconstruct manually.
Step 7: evaluate the audit trail before the dashboard
Almost every QA platform can produce a dashboard.
The more revealing vendor test is to click on one number and ask what sits behind it.
For example:
Complaint risk increased by 18%.
A QA team should then be able to establish:
- which interactions caused the increase
- which teams or journeys were involved
- what evidence supported each classification
- whether a human reviewed the cases
- whether results were overridden
- which QA criteria were applied
Detect creates structured evidence as assessments are completed. For an FOS case, FCA request or board report, the evidence can be retrieved without manually reconstructing the complete interaction trail across multiple systems.
That matters when complaint volumes remain substantial.
The Financial Ombudsman Service resolved more than 224,000 complaints during 2025/26, and 30% of complaints resolved across all financial products were upheld in favour of consumers.
The practical test for quality assurance software is therefore:
How quickly can the QA team move from a trend in the dashboard to the individual customer interaction and the evidence supporting the assessment?
Step 8: understand what AI the quality assurance software actually uses
“AI-powered” is too broad to be a meaningful evaluation criterion on its own.
Banks should examine the model beneath the QA workflow.
Is it built for financial services?
Detect is powered by FinLLM, Aveni’s financial-services-specific language model layer.
Its banking use cases include language and context associated with:
- APP fraud disputes
- account restrictions
- financial difficulty
- collections
- Consumer Duty outcomes
That is materially different from beginning with a general-purpose language model and expecting the QA team to configure every regulatory concept around it.
Can the output be inspected?
Reviewers should be able to see the evidence behind automated conclusions.
Does the human retain authority?
Automated first-stage assessment should support human decision-making rather than obscure it.
How is customer data governed?
Banks should establish:
- where data is hosted
- whether customer information is used to train models
- how permissions work
- how data crosses system boundaries
- what is recorded about automated assessments
Detect’s banking deployment is UK-hosted, with customer data not used for model training.
What does the Mills Review mean for banks evaluating QA software?
AI governance should now form part of the QA technology evaluation itself.
The FCA published the Mills Review in July 2026. It identifies four major AI-driven changes expected to reshape retail financial services, including the transformation of firm operations and the amplification of fraud and cyber risks.
The FCA also reiterated that its approach remains principles-based and outcomes-focused, with the Consumer Duty and Senior Managers Regime central to accountability as AI adoption increases.
For Heads of QA, this adds several questions to the vendor process:
- Who remains accountable for automated QA?
- Can individual assessments be reconstructed?
- Are human interventions recorded?
- Can the firm explain how AI reached a result?
- How is the model monitored after deployment?
- Can the control operate at the same scale as the AI system?
Aveni covers these questions in more detail in:
Step 9: test quality assurance software on real banking journeys
A polished product demo can demonstrate navigation.
It cannot prove whether the technology understands your QA environment.
A proper evaluation should include representative interactions from the journeys your QA team actually handles.
For example:
- APP fraud disputes
- account restrictions
- vulnerable customers
- complaints
- collections
- financial difficulty
- customer understanding
Compare the software’s results with validated assessments already completed by experienced reviewers.
Quality assurance software evaluation scorecard
| Evaluation area | Evidence to request |
|---|---|
| Coverage | Interaction types successfully assessed |
| Accuracy | Agreement against validated assessments |
| Risk detection | Higher-risk cases found outside existing samples |
| Evidence | Supporting information behind every assessment |
| Efficiency | Reviewer time per case |
| Consistency | Same criteria applied across interactions |
| Case reconstruction | Voice and written evidence connected |
| Human oversight | Review, override and escalation capability |
| Auditability | Historic evidence retrievable |
| Integration | Fit with existing systems |
| Security | Hosting and data handling controls |
| Governance | Model and configuration controls |
Do not reduce a pilot to one accuracy percentage.
The important measure is how the technology changes the complete QA workflow from interaction ingestion through to evidence and action.
Step 10: treat QA automation as an operating-model decision
Moving away from manual sampling does not mean removing QA professionals.
It changes where their expertise is used.
A modern QA model can operate as:
Ingest → Assess → Prioritise → Human review → Evidence
Detect follows this approach across voice, documents and case notes. FinLLM applies the configured assessment, higher-risk interactions are prioritised and analysts make the final call before structured evidence is retained.
This allows QA teams to spend less time finding interactions and reconstructing cases and more time investigating the cases where judgement matters.
Quality assurance software checklist for banking QA teams
Before selecting a platform, confirm that it can deliver the following.
| Requirement | What good looks like |
|---|---|
| Multi-channel coverage | Relevant voice, digital and written interactions assessed |
| Risk prioritisation | Higher-risk interactions surface first |
| Configurable QA | Your own framework can be applied |
| Consumer Duty | Customer outcomes assessed and evidenced |
| Vulnerability | Signals identified with supporting evidence |
| Case assessment | Multiple touchpoints treated as one case |
| Explainability | Every result can be inspected |
| Human control | Reviewer override and sign-off supported |
| Audit trail | Evidence can be retrieved later |
| Pattern analysis | Risk compared across teams and journeys |
| Integration | Existing infrastructure can remain in place |
| Data governance | Deployment meets regulated-firm requirements |
The difference between traditional sampling and automated QA is therefore larger than efficiency.
Manual QA starts by selecting a fraction of the interaction population and asks what can be learned from it.
Automated, risk-based QA can assess the wider population first and then direct human expertise towards the interactions most likely to warrant attention.
How Aveni Detect supports banking QA teams
Aveni Detect is quality assurance software built specifically for regulated financial services.
It helps QA and compliance teams:
- assess interactions across voice and written evidence
- apply firm-specific scorecards consistently
- identify vulnerability, complaints, conduct and Consumer Duty risks
- prioritise higher-risk interactions
- review customer journeys at case level
- inspect evidence behind automated assessments
- retain human review and override
- create structured regulatory evidence
Detect has assessed 1.38 million customer conversations across Aveni’s customer base. At 2% manual coverage, 182,408 high-risk conversations would have gone unreviewed.
For Heads of QA comparing quality assurance software, that is the practical benchmark:
How much of what is happening to customers can your current QA process actually see?
FAQs about quality assurance software in banking
What is quality assurance software?
Quality assurance software helps organisations assess customer interactions against defined performance, quality and compliance criteria.
In financial services, it can support interaction assessment, QA scorecards, risk detection, customer outcome monitoring, human review and regulatory evidence.
What is banking quality assurance software?
Banking quality assurance software is QA technology designed for the processes, risks and regulatory requirements of banking operations.
It may assess interactions involving complaints, fraud, collections, vulnerability, customer understanding and other regulated customer journeys.
What is automated quality assurance?
Automated quality assurance uses technology to perform part or all of the initial QA assessment automatically.
Rather than requiring an analyst to manually review every selected interaction, software can analyse a much larger population, apply defined criteria and route relevant cases to human reviewers.
Does the FCA require banks to review 100% of calls?
No. The FCA does not set a fixed percentage of customer calls that firms must review.
Consumer Duty does require firms to monitor and evidence customer outcomes. The appropriate monitoring approach depends on the firm’s business, risks and customer population.
Automated assessment can increase coverage without requiring a human to listen manually to every interaction.
What should banks look for in quality assurance software?
Banks should assess:
- interaction coverage
- QA framework configuration
- risk prioritisation
- customer outcome assessment
- Consumer Duty monitoring
- vulnerability detection
- case-level analysis
- explainability
- human review
- audit trails
- integrations
- data governance
The technology should be tested on representative customer journeys before procurement.
How does quality assurance software reduce manual sampling?
Quality assurance software can automatically assess interactions before human review takes place.
Instead of randomly selecting a small proportion of interactions first, the technology can analyse a broader population and prioritise the cases most likely to require human attention.
Can AI replace banking QA teams?
AI can automate repetitive assessment and prioritisation, but experienced QA professionals remain important for judgement, investigation, escalation and sign-off.
The objective is to use automation to handle scale while directing human expertise towards higher-value decisions.
Can quality assurance software identify vulnerable customers?
Yes, if the system is designed to recognise relevant contextual signals rather than relying solely on simple keyword matching.
The FCA expects firms to understand and monitor the outcomes vulnerable customers receive.
Detect assesses vulnerability signals across interaction types and provides the supporting evidence to human reviewers.
How should banks evaluate AI quality assurance software after the Mills Review?
Banks should evaluate governance alongside performance.
The FCA’s Mills Review reinforces the importance of accountability, oversight and evidence as AI becomes more involved in financial-services operations.
A vendor evaluation should therefore examine human oversight, explainability, model monitoring, data governance and the audit trail surrounding automated assessments.