Skip to content

FinanceHarness: Autonomous Financial Deep Research Framework

Jul 2026 · arXiv.org · Vol abs/2607.27853 · 1 citation
Computer Science Economics

TL;DR

FinanceHarness is presented, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling, and FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria.

Abstract

Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%, demonstrating the effectiveness of our specialized harness design. However, even pairing FinanceHarness with the most cutting edge LLM (e.g. Opus-5), the FinanceGym score is below 45%, showing that it is a challenging benchmark for financial deep research. Leaderboard is available at: https://financegym.github.io/ and FinanceHarness code is available at: https://github.com/Yijia-Xiao/FinanceHarness.

View source

Similar papers

Preprint Aug 2026

Mint-Agent: Introducing Finance-Native Agentic Foundation Models

Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence, is presented, establishing a path toward trustworthy financial intelligence in which domain expertise, long-horizon execution, and auditable evidence are jointly engineered as a unified foundation for front...

Mint-Agent Team, Kun Wang, Gavin Zhang et al. · 0 citations
Book Open access Aug 2026

Generative AI and Large Language Models for Financial Markets: From Behavioral Prediction to Autonomous Trading

It is shown how recent progress in generative AI becomes genuinely useful in markets when it helps model participant behavior, ground reasoning in live documents and order-flow data, and support research and execution workflows that can survive contact with production.

Z. Iklassov, Hachem Madmoun, J. Duhot et al. · 0 citations
Review Open access Aug 2026

A survey on LLM-enhanced reinforcement learning in financial markets

A three-paradigm taxonomy (feature-based, auxiliary-based, and policy-based) based on the functional role of LLMs within the RL pipeline is proposed, which provides superior scalability and stability, though often at the expense of representational depth.

Ghusoon Hadi al-Aldaffaie, Alireza Taheri, Amirfarhad Farhadi et al. · 0 citations
Preprint Aug 2026

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

Evaluating FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow, finds that the tool harness, not the model alone, strongly shapes quality and efficiency.

Yuhao Zhang, O. Koyluoglu, Thejas Venkatesh et al. · 3 citations · ⚡2
#artificial intelligence Review Sep 2026

LongCat-DeepResearch Technical Report

We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents f...

Mei Zhu, Yue-Ya Xu, Wan-Li Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining

Large language model (LLM)-based multi-agent systems can automate alpha factor mining, but their reliance on external APIs limits control over cost, availability, and confidentiality. Long research loops also tend to revisit a few successful economic mechanisms that lead to research path collapse. To address these limi...

Qing-Zhuo Wang, Zi-Kun Wei, Zhi-Hua Wei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.