{"slug": "are-the-financial-reasoning-from-llms-credible-a-real-world-test-over-long", "title": "Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements", "summary": "A new benchmark called FinIndices, introduced by researchers in a preprint on arXiv (2607.28661v1), tests large language models on financial reasoning over long-horizon statements up to 32K tokens and finds severe vulnerabilities: removing explicit formula hints causes Gemini-3.1-Pro to drop from 70.70% to 38.22% on table tasks, and models regress to shallow heuristics under structural pressure. Supervised fine-tuning yields gains of +8.54% on single-index and +3.82% on table tasks, suggesting structured logic can be partially restored via data-centric alignment.", "body_md": "arXiv:2607.28661v1 Announce Type: new\nAbstract: Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal testbed. Existing benchmarks fail to capture real-world industrial complexity, predominantly relying on multiple-choice questions or single-hop QA over cropped tables while ignoring intricate cross-statement dynamics and temporal de-cumulation.\nTo bridge this gap, we introduce FinIndices, a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens). Utilizing an automated synthesis pipeline with adversarial traps, FinIndices encompasses Single-Index computation and Table-Index tabulation to test complex domain, temporal, and caliber reasoning.\nOur evaluation reveals two severe LLM vulnerabilities. First, a \"Knowledge Bottleneck\": despite memorizing formulas during pre-training, models demonstrate fragile pattern matching. Removing explicit formula hints causes performance to collapse (e.g., Gemini-3.1-Pro drops from 70.70% to 38.22% on table tasks), exposing fatal flaws in temporal de-cumulation and stock-flow caliber mismatch. Second, a \"Structural Bottleneck\": the intense cognitive load of generating multi-metric, multi-period tables actively drains reasoning capacity. Under structural pressure, LLMs that flawlessly execute isolated derivations regress to shallow heuristics, such as fetching incorrect adjacent columns or substituting deep accounting adjustments with lazy literal arithmetic. Finally, Supervised Fine-Tuning (SFT) yields substantial zero-hint gains (+8.54% Single, +3.82% Table), validating that structured logic can be partially restored via data-centric alignment.", "url": "https://wpnews.pro/news/are-the-financial-reasoning-from-llms-credible-a-real-world-test-over-long", "canonical_source": "https://arxiv.org/abs/2607.28661", "published_at": "2026-08-03 04:00:00+00:00", "updated_at": "2026-08-03 04:04:19.026010+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-safety"], "entities": ["FinIndices", "Gemini-3.1-Pro", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/are-the-financial-reasoning-from-llms-credible-a-real-world-test-over-long", "markdown": "https://wpnews.pro/news/are-the-financial-reasoning-from-llms-credible-a-real-world-test-over-long.md", "text": "https://wpnews.pro/news/are-the-financial-reasoning-from-llms-credible-a-real-world-test-over-long.txt", "jsonld": "https://wpnews.pro/news/are-the-financial-reasoning-from-llms-credible-a-real-world-test-over-long.jsonld"}}