quality assurance software for banking

Banking QA software: what to look for when replacing manual call sampling

For years, banking QA has relied on a familiar process: select a small sample of customer interactions, review them manually and use the findings to understand how the wider operation is performing.

That model becomes harder to scale as customer journeys spread across calls, email, webchat, case notes and documents.

For Heads of QA evaluating quality assurance software, the buying decision therefore needs to go further than asking how many calls a platform can score.

The more important test is whether the technology can identify the interactions that matter, assess customer outcomes consistently, connect evidence across channels and give human reviewers enough context to make the final decision.

TL;DR: what should banks look for in quality assurance software?

Quality assurance software for banking should help QA teams move from limited random sampling towards broader, risk-based oversight while keeping humans responsible for material decisions.

Look for software that can:

  • assess interactions across voice, digital channels, documents and case notes
  • apply your existing QA and compliance frameworks consistently
  • prioritise interactions according to risk
  • assess customer outcomes alongside employee performance
  • monitor Consumer Duty and vulnerability indicators
  • connect multiple interactions into one customer case
  • explain every automated assessment using supporting evidence
  • allow human review, override and sign-off
  • create retrievable audit trails
  • integrate with existing telephony, CRM and case-management infrastructure
  • meet the data governance requirements of a regulated financial-services firm

The FCA does not prescribe a fixed percentage of interactions that firms must review. Its Consumer Duty framework does, however, require firms to understand and evidence the outcomes their customers receive, with outcomes monitoring remaining an active FCA focus in 2026.

That changes what banks should expect from QA technology.

Why are banks reconsidering manual QA sampling?

Manual sampling made sense when every assessment required a person to listen to a complete call.

Its limitation is mathematical.

If a bank reviews 2% of its interactions, 98% sit outside that manual QA sample.

Risk is not evenly distributed across those interactions. A routine balance enquiry and an interaction involving financial difficulty have the same chance of appearing in a random sample.

Across 1.38 million customer conversations assessed by Detect, Aveni found that 182,408 high-risk conversations would have gone unreviewed at 2% manual coverage. Those conversations contained signals including complaints, dissatisfaction and vulnerability. See what 2% QA coverage misses across five channels.

The same banking data shows approximately 77 automated assessments for every manual review, while average assessment time has fallen from around 90 minutes to 15 minutes.

Banking interactions are becoming harder to sample effectively

The Financial Ombudsman Service received 214,600 new complaints in 2025/26. Current accounts alone generated 32,900 complaints, including 18,900 relating to fraud and scams.

UK Finance reported ÂŁ1.17 billion in authorised and unauthorised fraud losses during 2024, including ÂŁ450.7 million in APP fraud losses.

That creates very different QA requirements across:

  • fraud disputes
  • account restrictions
  • collections
  • financial difficulty
  • complaints
  • vulnerable customer interactions
  • everyday customer service

A random sample cannot guarantee that the interactions with the greatest potential customer or regulatory impact will reach a reviewer.

This is why selection and prioritisation should be one of the first capabilities assessed when comparing quality assurance software.

Step 1: decide what your quality assurance software needs to measure

Start with the purpose of the QA programme before evaluating the technology.

Traditional QA scorecards tend to ask:

  • Did the agent follow the process?
  • Was mandatory wording used?
  • Was the interaction handled professionally?
  • Was the correct procedure followed?

Those remain useful measures.

But they do not necessarily establish what happened to the customer.

Performance scoring and customer outcome assessment are different

The banking Detect material separates these disciplines clearly.

Performance QA measures how an employee behaved.

Outcome-focused QA asks whether the customer received a fair and compliant outcome and whether the firm can evidence that outcome afterwards.

QA requirementPerformance-focused QAOutcome-focused QA
Employee followed processPrimary measureSupporting context
Communication qualityCommonly assessedAssessed against customer understanding
Customer received appropriate outcomeMay be inferredDirectly assessed
Vulnerability handled appropriatelyDepends on sampled interactionEvaluated as part of outcome
Decision can be evidencedDepends on reviewer notesSupporting evidence retained
Patterns can be analysed across the populationLimited by samplePossible with broader assessment

This distinction should influence the vendor shortlist.

A platform can automate employee scoring very effectively without necessarily giving compliance teams stronger evidence of customer outcomes.

Step 2: establish what the software means by “coverage”

Claims of high or complete QA coverage need to be examined carefully.

A platform might analyse every telephone call while leaving email, webchat, complaint correspondence and case documents outside the same QA workflow.

That still creates blind spots.

Consider a customer who:

  1. reports suspicious activity through webchat
  2. speaks to the fraud team by phone
  3. receives an account restriction notice
  4. provides documents
  5. raises a complaint by email

Assessing the telephone interaction alone does not establish what happened across the customer journey.

Look for case-level quality assurance software

Banking QA increasingly needs to connect evidence across different interaction types.

Detect can assess voice recordings and written evidence together, including account restriction letters, collections documents and complaint correspondence. The resulting customer journey can then be presented to the reviewer as one case.

This also reduces preparation time.

A reviewer working manually may have to locate the recording, case notes and customer correspondence before the substantive assessment begins. In the banking Detect workflow, a review that previously took 60 to 90 minutes can take around 15 minutes once that evidence is assembled.

Step 3: assess how quality assurance software prioritises human review

Automation should help QA teams spend human attention more deliberately.

That requires more than automatically choosing another sample.

A risk-based QA system first assesses the interaction, then uses what it finds to determine whether human investigation is required.

For a bank, potential prioritisation criteria might include:

SignalExample
Complaint riskCustomer repeatedly expresses dissatisfaction
VulnerabilityFinancial difficulty or another FG21/1 indicator
Conduct riskRequired process or explanation appears incomplete
FraudAPP fraud dispute with unclear customer outcome
Consumer DutyCustomer understanding is not evidenced
CollectionsSupport pathway does not appear to have been followed
DocumentationWritten evidence conflicts with what was communicated

Detect pre-assesses interactions and routes higher-risk cases into the QA workflow, while analysts retain responsibility for the final decision.

This is fundamentally different from using automation simply to increase the quantity of randomly selected interactions.

Aveni explores this workflow further in How to prioritise high-risk calls for QA review.

Step 4: make sure the software can use your QA framework

A bank should not have to abandon a well-established QA methodology because a software provider has created its own generic scorecard.

Quality assurance software should be configurable around the firm’s existing controls, including:

  • Consumer Duty outcomes
  • vulnerability
  • complaints
  • conduct
  • journey-specific requirements
  • escalation criteria
  • product-specific controls
  • internal QA standards

Detect configures a firm’s own QA logic into its scorecards rather than replacing it with a standardised framework. Those criteria can then be applied consistently across the interaction population.

Questions to ask vendors about QA configuration

QuestionWhat to look for
Can our current QA framework be replicated?Configurable criteria
Can different journeys use different scorecards?Journey-level configuration
How are framework changes governed?Version control and permissions
Can analysts see why a result was returned?Explainability
Can a human override the result?Reviewer control
Are assessments comparable over time?Consistent definitions and reporting

Consistency matters because manual QA introduces inevitable reviewer variation.

Automation should reduce that variation without removing the judgement of experienced QA professionals.

Step 5: require an explanation and evidence for every assessment

A QA score without supporting evidence has limited value in a regulated environment.

For an automated assessment to be useful, the reviewer should be able to see:

  1. the result
  2. why that result was reached
  3. the underlying evidence

Detect structures checks in this way. Each assessment can return an outcome, an explanation and supporting evidence from the customer interaction.

That becomes particularly important when QA results are used for:

  • Consumer Duty reporting
  • board MI
  • compliance monitoring
  • complaint investigations
  • FOS referrals
  • coaching
  • root-cause analysis

The FCA’s Consumer Duty resources now explicitly include outcomes monitoring, board reporting, consumer understanding and vulnerable customers within its published good and poor practice materials.

For the latest monitoring data, see Aveni’s Consumer Duty Outcomes Monitoring: 4 FCA Figures to Cite.

Step 6: evaluate vulnerability monitoring as its own capability

Vulnerability should not be buried inside a generic compliance feature list.

The FCA’s FG21/1 guidance expects firms to understand customer needs, enable staff to recognise vulnerability, respond appropriately and monitor the outcomes vulnerable customers receive. Its guidance page was updated again in July 2026 to point firms towards the latest Consumer Duty expectations for management information and outcomes.

For banking QA, signs of vulnerability may emerge during:

  • collections interactions
  • account restrictions
  • fraud investigations
  • bereavement
  • affordability discussions
  • complaints
  • ordinary customer service conversations

The banking Detect configuration identifies vulnerability and financial-difficulty signals across fraud, account restriction and collections journeys, with evidence attached to flags and escalation pathways built into the workflow.

When comparing quality assurance software, assess whether the system can do all three of the following:

identify → evidence → route

A vulnerability alert with no evidence simply gives a QA analyst another case to reconstruct manually.

Step 7: evaluate the audit trail before the dashboard

Almost every QA platform can produce a dashboard.

The more revealing vendor test is to click on one number and ask what sits behind it.

For example:

Complaint risk increased by 18%.

A QA team should then be able to establish:

  • which interactions caused the increase
  • which teams or journeys were involved
  • what evidence supported each classification
  • whether a human reviewed the cases
  • whether results were overridden
  • which QA criteria were applied

Detect creates structured evidence as assessments are completed. For an FOS case, FCA request or board report, the evidence can be retrieved without manually reconstructing the complete interaction trail across multiple systems.

That matters when complaint volumes remain substantial.

The Financial Ombudsman Service resolved more than 224,000 complaints during 2025/26, and 30% of complaints resolved across all financial products were upheld in favour of consumers.

The practical test for quality assurance software is therefore:

How quickly can the QA team move from a trend in the dashboard to the individual customer interaction and the evidence supporting the assessment?

Step 8: understand what AI the quality assurance software actually uses

“AI-powered” is too broad to be a meaningful evaluation criterion on its own.

Banks should examine the model beneath the QA workflow.

Is it built for financial services?

Detect is powered by FinLLM, Aveni’s financial-services-specific language model layer.

Its banking use cases include language and context associated with:

  • APP fraud disputes
  • account restrictions
  • financial difficulty
  • collections
  • Consumer Duty outcomes

That is materially different from beginning with a general-purpose language model and expecting the QA team to configure every regulatory concept around it.

Can the output be inspected?

Reviewers should be able to see the evidence behind automated conclusions.

Does the human retain authority?

Automated first-stage assessment should support human decision-making rather than obscure it.

How is customer data governed?

Banks should establish:

  • where data is hosted
  • whether customer information is used to train models
  • how permissions work
  • how data crosses system boundaries
  • what is recorded about automated assessments

Detect’s banking deployment is UK-hosted, with customer data not used for model training.

What does the Mills Review mean for banks evaluating QA software?

AI governance should now form part of the QA technology evaluation itself.

The FCA published the Mills Review in July 2026. It identifies four major AI-driven changes expected to reshape retail financial services, including the transformation of firm operations and the amplification of fraud and cyber risks.

The FCA also reiterated that its approach remains principles-based and outcomes-focused, with the Consumer Duty and Senior Managers Regime central to accountability as AI adoption increases.

For Heads of QA, this adds several questions to the vendor process:

  • Who remains accountable for automated QA?
  • Can individual assessments be reconstructed?
  • Are human interventions recorded?
  • Can the firm explain how AI reached a result?
  • How is the model monitored after deployment?
  • Can the control operate at the same scale as the AI system?

Aveni covers these questions in more detail in:

Step 9: test quality assurance software on real banking journeys

A polished product demo can demonstrate navigation.

It cannot prove whether the technology understands your QA environment.

A proper evaluation should include representative interactions from the journeys your QA team actually handles.

For example:

  • APP fraud disputes
  • account restrictions
  • vulnerable customers
  • complaints
  • collections
  • financial difficulty
  • customer understanding

Compare the software’s results with validated assessments already completed by experienced reviewers.

Quality assurance software evaluation scorecard

Evaluation areaEvidence to request
CoverageInteraction types successfully assessed
AccuracyAgreement against validated assessments
Risk detectionHigher-risk cases found outside existing samples
EvidenceSupporting information behind every assessment
EfficiencyReviewer time per case
ConsistencySame criteria applied across interactions
Case reconstructionVoice and written evidence connected
Human oversightReview, override and escalation capability
AuditabilityHistoric evidence retrievable
IntegrationFit with existing systems
SecurityHosting and data handling controls
GovernanceModel and configuration controls

Do not reduce a pilot to one accuracy percentage.

The important measure is how the technology changes the complete QA workflow from interaction ingestion through to evidence and action.

Step 10: treat QA automation as an operating-model decision

Moving away from manual sampling does not mean removing QA professionals.

It changes where their expertise is used.

A modern QA model can operate as:

Ingest → Assess → Prioritise → Human review → Evidence

Detect follows this approach across voice, documents and case notes. FinLLM applies the configured assessment, higher-risk interactions are prioritised and analysts make the final call before structured evidence is retained.

This allows QA teams to spend less time finding interactions and reconstructing cases and more time investigating the cases where judgement matters.

Quality assurance software checklist for banking QA teams

Before selecting a platform, confirm that it can deliver the following.

RequirementWhat good looks like
Multi-channel coverageRelevant voice, digital and written interactions assessed
Risk prioritisationHigher-risk interactions surface first
Configurable QAYour own framework can be applied
Consumer DutyCustomer outcomes assessed and evidenced
VulnerabilitySignals identified with supporting evidence
Case assessmentMultiple touchpoints treated as one case
ExplainabilityEvery result can be inspected
Human controlReviewer override and sign-off supported
Audit trailEvidence can be retrieved later
Pattern analysisRisk compared across teams and journeys
IntegrationExisting infrastructure can remain in place
Data governanceDeployment meets regulated-firm requirements

The difference between traditional sampling and automated QA is therefore larger than efficiency.

Manual QA starts by selecting a fraction of the interaction population and asks what can be learned from it.

Automated, risk-based QA can assess the wider population first and then direct human expertise towards the interactions most likely to warrant attention.

How Aveni Detect supports banking QA teams

Aveni Detect is quality assurance software built specifically for regulated financial services.

It helps QA and compliance teams:

  • assess interactions across voice and written evidence
  • apply firm-specific scorecards consistently
  • identify vulnerability, complaints, conduct and Consumer Duty risks
  • prioritise higher-risk interactions
  • review customer journeys at case level
  • inspect evidence behind automated assessments
  • retain human review and override
  • create structured regulatory evidence

Detect has assessed 1.38 million customer conversations across Aveni’s customer base. At 2% manual coverage, 182,408 high-risk conversations would have gone unreviewed.

For Heads of QA comparing quality assurance software, that is the practical benchmark:

How much of what is happening to customers can your current QA process actually see?

FAQs about quality assurance software in banking

What is quality assurance software?

Quality assurance software helps organisations assess customer interactions against defined performance, quality and compliance criteria.

In financial services, it can support interaction assessment, QA scorecards, risk detection, customer outcome monitoring, human review and regulatory evidence.

What is banking quality assurance software?

Banking quality assurance software is QA technology designed for the processes, risks and regulatory requirements of banking operations.

It may assess interactions involving complaints, fraud, collections, vulnerability, customer understanding and other regulated customer journeys.

What is automated quality assurance?

Automated quality assurance uses technology to perform part or all of the initial QA assessment automatically.

Rather than requiring an analyst to manually review every selected interaction, software can analyse a much larger population, apply defined criteria and route relevant cases to human reviewers.

Does the FCA require banks to review 100% of calls?

No. The FCA does not set a fixed percentage of customer calls that firms must review.

Consumer Duty does require firms to monitor and evidence customer outcomes. The appropriate monitoring approach depends on the firm’s business, risks and customer population.

Automated assessment can increase coverage without requiring a human to listen manually to every interaction.

What should banks look for in quality assurance software?

Banks should assess:

  • interaction coverage
  • QA framework configuration
  • risk prioritisation
  • customer outcome assessment
  • Consumer Duty monitoring
  • vulnerability detection
  • case-level analysis
  • explainability
  • human review
  • audit trails
  • integrations
  • data governance

The technology should be tested on representative customer journeys before procurement.

How does quality assurance software reduce manual sampling?

Quality assurance software can automatically assess interactions before human review takes place.

Instead of randomly selecting a small proportion of interactions first, the technology can analyse a broader population and prioritise the cases most likely to require human attention.

Can AI replace banking QA teams?

AI can automate repetitive assessment and prioritisation, but experienced QA professionals remain important for judgement, investigation, escalation and sign-off.

The objective is to use automation to handle scale while directing human expertise towards higher-value decisions.

Can quality assurance software identify vulnerable customers?

Yes, if the system is designed to recognise relevant contextual signals rather than relying solely on simple keyword matching.

The FCA expects firms to understand and monitor the outcomes vulnerable customers receive.

Detect assesses vulnerability signals across interaction types and provides the supporting evidence to human reviewers.

How should banks evaluate AI quality assurance software after the Mills Review?

Banks should evaluate governance alongside performance.

The FCA’s Mills Review reinforces the importance of accountability, oversight and evidence as AI becomes more involved in financial-services operations.

A vendor evaluation should therefore examine human oversight, explainability, model monitoring, data governance and the audit trail surrounding automated assessments.

Share with your community!

In this article

Related Articles

Join our newsletter

Be the first to hear about new features, releases, and best-practice guides.

Aveni AI Logo