Conference
Open access
Jun 2026
Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents
This paper introduces FinED-Bench, the first publicly public benchmark for FinED-Bench, which covers nine real-world financial scenarios, and includes over 900 documents reported in 2025 that are unseen by existing language models.
Ying He, Zhouhong Gu, Zhecheng Hu et al.
· Annual Meeting of the Associ... · 2 citations