Skip to content
Written by: Aveni Team
Written on: 15 Sep 2026
Reading Time: 14 min
Responsible AI

How to build an AI risk taxonomy for financial services compliance

September 15, 2026

Share Article:

Financial services firms have detailed rules, policies and risk frameworks. The harder task is turning them into definitions that people and technology can apply consistently across thousands of customer interactions.

An AI risk taxonomy provides that structure. It converts broad regulatory and conduct requirements into defined risk categories, observable evidence and clear criteria for assessment.

For compliance and QA teams, this matters because automated monitoring can only be as consistent as the definitions behind it. A model cannot reliably identify poor customer understanding, vulnerability or potential harm if those concepts mean different things to different reviewers.

The FCA’s approach makes this especially relevant. Rather than creating a separate regulatory regime for AI, the FCA has said firms should apply existing frameworks including Consumer Duty, governance requirements and SM&CR to their use of AI.

For firms using AI within compliance monitoring, the practical work starts with defining what the system should look for.

TL;DR: what is an AI risk taxonomy?

An AI risk taxonomy is a structured way of defining, organising and assessing the risks an AI system needs to recognise or manage.

In financial services, a useful taxonomy should:

  • map regulatory requirements and internal policies to clearly defined risk areas
  • describe the behaviour, evidence or customer outcome associated with each risk
  • give human reviewers and AI systems consistent assessment criteria
  • distinguish different levels of severity and the actions they require
  • support testing and evaluation against known examples
  • change when regulation, policies, products or evidence of customer harm changes

The purpose is not to turn regulation into a checklist. It is to provide a consistent language between regulation, compliance specialists, QA teams, model developers and the systems performing the assessment.

What is an AI risk taxonomy in financial services?

A risk taxonomy organises risks into a common structure so that they can be identified, measured and managed consistently.

An AI risk taxonomy takes that concept further by defining risks clearly enough for AI systems to evaluate them.

For a financial services compliance use case, that can mean translating a requirement such as identifying foreseeable harm into more specific questions:

  • What type of harm are we looking for?
  • What evidence could indicate that it has occurred or could occur?
  • What would distinguish a weak signal from a material concern?
  • What surrounding context needs to be considered?
  • When should a human reviewer investigate?
  • What evidence should be retained?

These definitions matter when firms use AI to support . Without them, different models, reviewers and business units can reach different conclusions about the same behaviour.

This principle extends beyond financial services. The US National Institute of Standards and Technology’s AI Risk Management Framework also places structured identification, measurement and management of AI risk at the centre of responsible AI governance.

Why financial services firms need an FCA-aligned AI risk taxonomy

The FCA sets outcomes and standards rather than prescribing a single monitoring methodology.

Under Consumer Duty, firms must monitor the outcomes retail customers receive and identify risks that they are failing to meet the cross-cutting obligations or retail customer outcomes. The nature and frequency of that monitoring depends on the firm, its products, target market and role in the distribution chain.

That flexibility creates an operational challenge.

Terms such as foreseeable harm, customer understanding, vulnerability, fair treatment and good outcomes carry regulatory meaning, but they still need to be translated into something a monitoring process can identify consistently.

A taxonomy provides the bridge between those levels.

Regulatory or conduct areaQuestion the monitoring framework needs to answerExamples of observable evidence
VulnerabilityDid the interaction identify and respond appropriately to relevant customer needs?Customer disclosure, signs of distress, accessibility needs, support offered, escalation
Complaints and dissatisfactionDid the customer express dissatisfaction that required recognition, investigation or escalation?Explicit complaint, repeated dissatisfaction, unresolved issue, request for escalation
Customer understandingWas information communicated in a way that allowed the customer to make an informed decision?Confusion, incorrect interpretation, unclear explanation, missing understanding checks
Conduct and foreseeable harmDid the interaction create or increase a material risk of a poor customer outcome?Pressure, inappropriate steering, failure to respond to known risks, missed support
Case-level outcomeTaken together, did the customer’s interactions produce an appropriate outcome?Issues repeated across calls, emails or webchat, unresolved concerns, actions inconsistent with earlier commitments

These are illustrative examples rather than Aveni’s proprietary taxonomy. The important point is the structure: regulation becomes a defined risk, the risk becomes observable evidence, and that evidence can then support assessment and human review.

How to translate FCA requirements into measurable risk signals

A strong taxonomy starts with the regulation and works towards the evidence. Starting with whatever a model happens to detect risks creating categories around technical capability rather than regulatory purpose.

A practical process looks like this:

StepWhat happensOutput
1. Identify the obligationStart with FCA rules, guidance and relevant internal policyRegulatory requirement
2. Define the riskDescribe the specific failure or customer harm the firm needs to identifyClear risk definition
3. Identify observable evidenceEstablish what behaviour, language or case information could indicate the riskRisk signals
4. Define contextRecord factors that can change the interpretation of the evidenceAssessment criteria
5. Set severity and actionDetermine which findings require review, escalation or thematic analysisOperational response
6. Test the definitionCompare assessments with specialist human judgement across varied examplesEvaluation evidence
7. Refine and maintainUpdate definitions as regulation, policies and observed risks changeCurrent taxonomy

The distinction between a risk and a signal is important.

A customer saying they have lost their job, for example, may indicate financial difficulty or vulnerability. It does not automatically establish that the firm treated them poorly. The assessment also needs to consider what happened next.

The same applies to complaints and customer understanding. Individual words or phrases provide evidence. They rarely provide the whole conclusion.

This is one reason keyword-based compliance monitoring has significant limitations. Context determines whether an interaction represents a genuine risk and what response it warrants.

Clear definitions help humans and models apply the same standard

Consistency is one of the most useful outcomes of a well-designed taxonomy.

Consider a QA framework that asks whether the customer demonstrated sufficient understanding.

One reviewer might treat any customer acknowledgement as a pass. Another might expect the agent to explain a key risk and confirm that the customer understood it. A third may judge the entire conversation more holistically.

Automating the same loosely defined question does not resolve the disagreement. It reproduces it at a larger scale.

The FCA’s March 2026 review of the Consumer Duty’s consumer understanding outcome highlighted similar problems in monitoring practice. Some firms collected MI but could not explain how it informed their assessment of consumer understanding, while others relied on measures such as sales data or the absence of complaints that provided limited assurance. Read the FCA’s consumer understanding findings

A stronger taxonomy sets out what the assessment means before a model or reviewer applies it.

“A model can only assess risk consistently when the risk itself has been defined clearly. That means starting with the regulatory requirement, working with risk and compliance specialists to identify the behaviour and evidence that matters, and testing those definitions across realistic customer interactions. The taxonomy gives the models and the people reviewing their outputs a common language to work from.”

Nicole Nisbett, Technical Product Manager, Models Team, Aveni

An AI risk taxonomy should define context as well as categories

Simple categories can create false confidence if the taxonomy ignores context.

Take vulnerability. The FCA’s guidance expects firms to understand and respond to customer needs, and its review of outcomes for customers in vulnerable circumstances found weaknesses where firms could not clearly define good outcomes or measure whether they were delivering them. Read the FCA’s findings on vulnerable customers

A vulnerability assessment therefore needs more than a binary flag.

The surrounding questions may include:

  • What characteristic or circumstance was identified?
  • Did it affect the customer’s ability to engage with the service?
  • Did the firm recognise it?
  • Was appropriate support offered?
  • Did later interactions follow through on that support?
  • Did the customer ultimately receive an appropriate outcome?

The same principle applies to complaints, dissatisfaction, financial difficulty and customer understanding.

This becomes particularly important when customer journeys span several channels.

Risk taxonomy design needs to work at case level

Many customer risks cannot be understood from one call.

A customer might mention financial difficulty during a webchat, raise dissatisfaction in an email and discuss a repayment arrangement by phone several days later. Looking at each interaction separately can miss how the events relate to one another.

That is why Aveni recently expanded . The same development is particularly relevant for .

A taxonomy built for this type of monitoring needs to account for evidence that accumulates across the customer journey.

An individual interaction may contain a weak risk signal. Several related signals across a case may justify a much higher level of concern.

The taxonomy therefore needs rules for:

  • evidence from different channels
  • repeated or unresolved issues
  • combinations of risk signals
  • changes in customer circumstances
  • actions taken after an issue was identified
  • the final customer outcome

This gives compliance teams a more useful basis for case-level assessment than treating every interaction as an independent event.

How an AI risk taxonomy improves QA prioritisation

Taxonomy design also determines how effectively firms can prioritise human review.

QA teams rarely have the capacity to investigate every interaction manually. Risk-based monitoring changes the order of operations: interactions can first be assessed against defined risk criteria, with higher-risk cases routed to human reviewers.

Aveni’s guide to sets out this process in more detail, including the importance of severity, overlapping risks and clear escalation rules.

A taxonomy can support that process by defining:

Risk type: What has potentially happened?

Severity: How serious could the impact be?

Confidence and evidence: What supports the assessment?

Priority: How quickly does it require review?

Ownership: Which team should investigate?

Action: What should happen after confirmation?

This prevents a risk score from becoming a dead end. The taxonomy connects identification to the operational response.

For banks replacing manual sampling with automated monitoring, these criteria should form part of the technology evaluation itself. Our guide to covers the wider questions firms should ask around coverage, evidence, prioritisation and human oversight.

A taxonomy should support outcomes monitoring, not simply risk detection

Identifying a potential risk is one part of the process. Firms also need to understand what happened afterwards.

The FCA’s July 2026 outcomes-monitoring work emphasised using information to spot potential harm, take action and establish whether that action improved the customer outcome.

Aveni has looked at the same FCA review in more detail in our analysis of .

For taxonomy design, this means the structure should allow a firm to distinguish between:

  1. an initial risk signal
  2. a confirmed issue
  3. the action taken
  4. the eventual customer outcome
  5. recurring patterns across customers, products or journeys

That structure gives compliance teams better information for thematic analysis.

A repeated customer-understanding issue, for example, may point to a wider weakness in a script, communication or customer journey. A series of isolated flags only becomes useful management information when the organisation can group and interpret them consistently.

How to keep an AI risk taxonomy current

A taxonomy cannot be treated as a one-off implementation exercise.

Regulatory guidance develops. Internal policies change. Products and customer journeys evolve. Monitoring also produces new evidence about where customers experience problems.

A sustainable governance process should cover four areas.

1. Regulatory review

Track relevant FCA rules, guidance, thematic reviews and examples of good and poor practice.

The FCA maintains a regularly updated Consumer Duty publications and resources hub containing its latest findings across areas such as outcomes monitoring, consumer understanding, vulnerability, support, products and services.

2. SME ownership

Risk and compliance specialists should remain responsible for the meaning of the risk.

Models teams can translate those definitions into technical assessments, but regulatory interpretation should not sit with engineering alone.

3. Evaluation against reviewed examples

When a risk definition changes, the firm needs evidence that the resulting assessments still behave as intended.

That requires examples with clear expected outcomes, human review and analysis of where assessments disagree.

4. Version control and traceability

Firms should be able to establish which risk definition and assessment criteria applied at a particular point in time.

This becomes increasingly important when monitoring results feed into governance, escalation and board reporting.

How Aveni Detect uses structured risk assessment in compliance monitoring

Aveni Detect applies financial-services-specific AI to QA and compliance monitoring across customer interactions.

It can assess interactions against configurable frameworks, identify potential risks and prioritise cases for human review. Current use cases include complaints, vulnerability, conduct risk, customer understanding and Consumer Duty monitoring.

The role of the taxonomy underneath that process is to make the assessment criteria explicit and consistent.

For compliance teams, that supports three practical goals:

  • broader monitoring coverage
  • more consistent identification and categorisation of risk
  • better evidence for human review and subsequent investigation

The human remains responsible for material judgement and action.

The Models Team’s work on taxonomy and model evaluation supports that underlying capability while keeping the regulatory definition of risk connected to the technical assessment.

AI risk taxonomy checklist for financial services firms

Before using a taxonomy to support automated compliance monitoring, check whether you can answer each of these questions:

  • Does every risk have a clear definition?
  • Can each risk be traced to a regulatory requirement, policy or business control?
  • Have you defined the evidence that could indicate the risk?
  • Does the assessment account for customer and journey context?
  • Can related signals across several interactions be considered together?
  • Have severity and escalation criteria been defined?
  • Can a human reviewer see why an interaction was flagged?
  • Have risk and compliance SMEs validated the definitions?
  • Have the criteria been tested against reviewed examples?
  • Can the taxonomy be updated without losing the history of previous assessments?

If several answers are unclear, the problem sits earlier than model selection.

The risk first needs to be defined well enough to assess.

FAQs about AI risk taxonomy in financial services

What is an AI risk taxonomy?

An AI risk taxonomy is a structured classification of the risks an AI system needs to identify, assess or manage. It defines risk categories, relevant evidence and assessment criteria so that risks can be evaluated consistently.

What is a conduct risk taxonomy?

A conduct risk taxonomy organises the behaviours, events and outcomes that can indicate customer harm or inappropriate conduct. In financial services, it can cover areas such as complaints, vulnerability, customer understanding and other risks relevant to how customers are treated.

Why does an AI model need a risk taxonomy?

A model needs clear definitions of what it is expected to assess. A taxonomy gives those definitions a consistent structure and provides a basis for training, testing, evaluation and human review.

What does FCA-aligned mean in an AI risk taxonomy?

FCA-aligned means the taxonomy has been designed with reference to applicable FCA rules, guidance and regulatory expectations. It does not mean the taxonomy has been approved or endorsed by the FCA.

Does the FCA require firms to use an AI risk taxonomy?

The FCA does not prescribe a specific AI risk taxonomy. Firms remain responsible for meeting applicable regulatory requirements and demonstrating effective governance, controls and outcomes monitoring. A taxonomy is one method for translating those obligations into consistent monitoring criteria.

How does an AI risk taxonomy support Consumer Duty?

A taxonomy can turn Consumer Duty requirements into structured monitoring criteria. This can help firms identify potential harm, vulnerability issues, poor customer understanding and other indicators that warrant investigation, while keeping evidence linked to the relevant customer interaction or case.

Can an AI risk taxonomy replace human compliance judgement?

No. A taxonomy helps AI systems and reviewers apply common criteria, but material findings still require appropriate human judgement, investigation and action.

How often should an AI risk taxonomy be updated?

There is no single required schedule. Firms should review the taxonomy when regulation or FCA guidance changes, when internal policies or products change, when monitoring identifies new patterns of risk, or when evaluation shows that existing definitions are producing inconsistent results.

Give AI a clear definition of the risk it is being asked to find

Compliance monitoring becomes difficult to scale when the same risk means different things to different people, teams or systems.

A structured AI risk taxonomy gives firms a shared starting point. Regulation defines the obligation. Risk and compliance specialists define what that means operationally. Models assess evidence against those criteria. Human reviewers investigate the cases that require judgement.

That structure becomes more important as monitoring expands across larger interaction volumes and complete customer cases.

Sign up for our newsletter