Skip to content
Built by Aveni Labs

FinLLM is the first LLM family built for UK financial services

A suite of small language models trained, evaluated and safety-tested on UK regulatory, advisory and customer-conversation data. It powers every Aveni product.

FinLLM 24B-M · AveniBench +3.6 avg
SUITABILITY REGULATORY VULNERABILITY SAFETY MULTILINGUAL ADVICE
FinLLM 24B-M Frontier avg
Why FinLLM exists

AI that speaks the language of financial services

Generic AI models are trained on the open internet. FinLLM is trained on the data financial services works with: regulation, advice files and real customer interactions. Every model is designed for the work firms need to do, evaluated against real financial services tasks, and deployed in regulated environments.

The model family

One family. Every financial services job.

Benchmarks

Evaluated where it matters

FinLLM is evaluated on AveniBench Finance, a suite built for UK financial services tasks, and on data from live deployments.

AveniBench Finance

Average score across UK financial services tasks.

Model Score
FinLLM-24B-M 57.54
Gemini 2.5 Flash 53.90
GPT-4.1 Mini 53.26
FinLLM-14B-M 51.55
Mistral Small 3.1 24B 43.01

Higher is better. Source: AveniBench Finance, Year of FinLLM report, Nov 2025.

Vulnerability detection

Identifying vulnerable customers in recorded adviser conversations. Missing a vulnerable customer is a regulatory failure, so catching every real case matters most.

Model Macro F1
FinLLM 8B-Compliance 0.957
GPT-4o 0.67
Aveni Vulnerability Detector (prior) 0.67

Higher is better.

Safety

AveniBench Safety score: resistance to toxic, biased and misaligned outputs on finance-specific prompts.

Model Safety
FinLLM 7B Q-Tab 69.37
FinLLM 7B Q (safety-prompted) 67.69
Qwen 2.5 7B 63.73
FinLLM 7B Q (baseline) 62.33

Higher is better.

On the tasks UK financial services runs, FinLLM outperforms frontier models at a fraction of the parameters and cost.

What makes FinLLM different

Designed for regulated financial services

Trained on UK financial services data

Built on AveniVault, Aveni’s corpus of 91B+ tokens of UK regulatory, advisory and customer-conversation data.

Evaluated on financial services tasks

AveniBench measures what matters: suitability, conduct risk, vulnerability, complaint detection, and safety.

Small, specialised, deployable

Models from 3B to 24B parameters. Run on your infrastructure or ours.

Safety-tested for regulated use

Red-teamed (adversarial testing) across toxicity, bias, IP infringement, privacy, misinformation, misalignment, hallucination and sustainability, aligned to FCA, PRA and EU AI Act expectations.
Portrait of Ranil Boteju, Group Chief Data and Analytics Officer at Lloyds Banking Group
Ranil Boteju Group Chief Data & Analytics Officer, Lloyds Banking Group

Aveni’s FinLLM will be a game changer for UK financial services. Recognising its potential, Lloyds Banking Group invested in Aveni in 2024, and since then have worked closely with Aveni to co-create FinLLM and test it on our live AI use cases. Having seen the FinLLM roadmap and integration with the broader Aveni product suite, I’ve been blown away with the progress and ambition. I am excited to see the transformative impact the Aveni FinLLM will have when deployed at scale across Lloyds Banking Group and the industry.

Sri Kanisapakkam Chief Data Officer, Nationwide

Since investing in Aveni and working closely together on co-creating the FinLLM, we are delighted to see its first iteration being released. We’re excited by the performance of FinLLM and the potential benefits it will bring both Nationwide and our customers.