04:00
2026-08-14
arxiv.org
large-language-models
Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents
Researchers introduced FinED-Bench, the first public benchmark for financial error detection, covering nine real-world scenarios and over 900 documents from 2025. Testing advanced LLMs like GPT-4o andβ¦