FinLLM is the first LLM family built for UK financial services
A suite of small language models trained, evaluated and safety-tested on UK regulatory, advisory and customer-conversation data. It powers every Aveni product.
AI that speaks the language of financial services
Generic AI models are trained on the open internet. FinLLM is trained on the data financial services works with: regulation, advice files and real customer interactions. Every model is designed for the work firms need to do, evaluated against real financial services tasks, and deployed in regulated environments.
One family. Every financial services job.
A compact model sized for real-time, widget-scale inference. Runs on a single GPU or on-device, with the speed to sit inline inside customer-facing journeys without adding latency.
Evaluated where it matters
FinLLM is evaluated on AveniBench Finance, a suite built for UK financial services tasks, and on data from live deployments.
AveniBench Finance
Average score across UK financial services tasks.
| Model | Score |
|---|---|
| FinLLM-24B-M | 57.54 |
| Gemini 2.5 Flash | 53.90 |
| GPT-4.1 Mini | 53.26 |
| FinLLM-14B-M | 51.55 |
| Mistral Small 3.1 24B | 43.01 |
Higher is better. Source: AveniBench Finance, Year of FinLLM report, Nov 2025.
Vulnerability detection
Identifying vulnerable customers in recorded adviser conversations. Missing a vulnerable customer is a regulatory failure, so catching every real case matters most.
| Model | Macro F1 |
|---|---|
| FinLLM 8B-Compliance | 0.957 |
| GPT-4o | 0.67 |
| Aveni Vulnerability Detector (prior) | 0.67 |
Higher is better.
Safety
AveniBench Safety score: resistance to toxic, biased and misaligned outputs on finance-specific prompts.
| Model | Safety |
|---|---|
| FinLLM 7B Q-Tab | 69.37 |
| FinLLM 7B Q (safety-prompted) | 67.69 |
| Qwen 2.5 7B | 63.73 |
| FinLLM 7B Q (baseline) | 62.33 |
Higher is better.
On the tasks UK financial services runs, FinLLM outperforms frontier models at a fraction of the parameters and cost.