AveniBench: Accessible and Versatile Evaluation of Finance Intelligence
January 31, 2025
Over the past few years, the application of large language models (LLMs) in the finance industry has gained significant attention. However, current benchmarks for financial LLMs often include simple tasks that don’t reflect real-world use cases, and their test sets come with licensing restrictions that limit commercial use.
To address this, we introduce AVENIBENCH, a permissively licensed benchmark designed to evaluate six essential finance-related skills: tabular reasoning, numerical reasoning, question answering, long context modeling, summarization, and dialogue. We’ve refactored the test sets to ensure comparability, providing a unified framework for evaluation. Additionally, AVENIBENCH offers two task difficulty levels—easy and hard—allowing scalable assessments based on real-world deployment needs.
Using AVENIBENCH, we evaluated 20 widely used LLMs, ranging from small open-weight models to proprietary systems like GPT-4. This evaluation launches our public leaderboard, offering valuable insights for both academic research and commercial development.