Skip to content

Author

Bo-Yu Yang

We have 4 of 11 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal testbed. Existing benchmarks fail to capture real-world industrial complexity, predominantly relying on multiple-choice questions or single-hop QA over cropped tables while ignoring intricate cross-statement dynamics and temporal de-cumulation. To bridge this gap, we introduce FinIndices, a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens). Utilizing an automated synthesis pipeline with adversarial traps, FinIndices encompasses Single-Index computation and Table-Index tabulation to test complex domain, temporal, and caliber reasoning. Our evaluation reveals two severe LLM vulnerabilities. First, a"Knowledge Bottleneck": despite memorizing formulas during pre-training, models demonstrate fragile pattern matching. Removing explicit formula hints causes performance to collapse (e.g., Gemini-3.1-Pro drops from 70.70% to 38.22% on table tasks), exposing fatal flaws in temporal de-cumulation and stock-flow caliber mismatch. Second, a"Structural Bottleneck": the intense cognitive load of generating multi-metric, multi-period tables actively drains reasoning capacity. Under structural pressure, LLMs that flawlessly execute isolated derivations regress to shallow heuristics, such as fetching incorrect adjacent columns or substituting deep accounting adjustments with lazy literal arithmetic. Finally, Supervised Fine-Tuning (SFT) yields substantial zero-hint gains (+8.54% Single, +3.82% Table), validating that structured logic can be partially restored via data-centric alignment.

Xin Tong, Xuanming Zhang, Tianyi Tang et al. · 0 citations
Jul 2026

Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

Results support treating PHD as a distinct target for evaluating how LLMs formulate investigative directions when final conclusions are withheld, with gains for several lower-performing models on HypoArena but regressions for other systems, including a top-performing model.

Tianyun Zhong, Wangyi Jiang, Wei Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs

A diagnostic benchmark for open-domain, open-form, long-horizon counterfactual causal reasoning, containing 220 what-if questions across STEM, HSS, and Hybrid scenarios, and finds that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score.

Yucheng Wang, Yuetian Du, Zheng Liu et al. · 0 citations
Jul 2026

Living-Harness Is an Interactive-Agent Evolver

Living-Harness is proposed, a self-evolving agent harness that converts each completed trajectory and its evaluator signals into posterior evidence for bounded harness updates, and supports retrieval-only reuse of the evolved harness state across model backbones.

Yuetian Du, Yucheng Wang, Helsing Xu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.