IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2607.27853.html

FinanceHarness: Autonomous Financial Deep Research Framework

Author

Listed:
  • Yijia Xiao
  • Rujun Han
  • Yanfei Chen
  • Zifeng Wang
  • Ke Jiang
  • Zhongying CuiZhu
  • Vishy Tirumalashetty
  • Wei Wang
  • Burak Gokturk
  • Tomas Pfister
  • Chen-Yu Lee

Abstract

Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%, demonstrating the effectiveness of our specialized harness design. However, even pairing FinanceHarness with the most cutting edge LLM (e.g. Opus-5), the FinanceGym score is below 45%, showing that it is a challenging benchmark for financial deep research. Leaderboard is available at: https://financegym.github.io/ and FinanceHarness code is available at: https://github.com/Yijia-Xiao/FinanceHarness.

Suggested Citation

  • Yijia Xiao & Rujun Han & Yanfei Chen & Zifeng Wang & Ke Jiang & Zhongying CuiZhu & Vishy Tirumalashetty & Wei Wang & Burak Gokturk & Tomas Pfister & Chen-Yu Lee, 2026. "FinanceHarness: Autonomous Financial Deep Research Framework," Papers 2607.27853, arXiv.org, revised Aug 2026.
  • Handle: RePEc:arx:papers:2607.27853
    as

    Download full text from publisher

    File URL: https://arxiv.org/pdf/2607.27853
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Yijia Xiao & Edward Sun & Tong Chen & Fang Wu & Di Luo & Wei Wang, 2025. "Trading-R1: Financial Trading with LLM Reasoning via Reinforcement Learning," Papers 2509.11420, arXiv.org.
    2. Shijie Wu & Ozan Irsoy & Steven Lu & Vadim Dabravolski & Mark Dredze & Sebastian Gehrmann & Prabhanjan Kambadur & David Rosenberg & Gideon Mann, 2023. "BloombergGPT: A Large Language Model for Finance," Papers 2303.17564, arXiv.org, revised Dec 2023.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Zheng Li, 2026. "Design and Empirical Study of a Large Language Model-Based Multi-Agent Investment System for Chinese Public REITs," Papers 2602.00082, arXiv.org.
    2. Zuoyou Jiang & Li Zhao & Rui Sun & Ruohan Sun & Zhongjian Li & Jing Li & Daxin Jiang & Zuo Bai & Cheng Hua, 2025. "Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning," Papers 2512.23515, arXiv.org, revised Sep 2026.
    3. Mohammed-Khalil Ghali & Cecil Pang & Oscar Molina & Carlos Gershenson-Garcia & Daehan Won, 2025. "Forecasting Commodity Price Shocks Using Temporal and Semantic Fusion of Prices Signals and Agentic Generative AI Extracted Economic News," Papers 2508.06497, arXiv.org.
    4. Ching-Nam Hang & Pei-Duo Yu & Roberto Morabito & Chee-Wei Tan, 2024. "Large Language Models Meet Next-Generation Networking Technologies: A Review," Future Internet, MDPI, vol. 16(10), pages 1-29, October.
    5. Dangxing Chen & Pengzhan Guo, 2026. "Shapley in Context: Explaining Financial Language with Domain Expertise," Papers 2607.00856, arXiv.org.
    6. Xia Li & Hanghang Zheng & Xiwei Zhuang & Zhong Wang & Xiao Chen & Hong Liu & Jasmine Bai & Mao Mao, 2025. "Class-Imbalanced-Aware Adaptive Dataset Distillation for Scalable Pretrained Model on Credit Scoring," Papers 2501.10677, arXiv.org, revised Mar 2026.
    7. Lezhi Li & Ting-Yu Chang & Hai Wang, 2023. "Multimodal Gen-AI for Fundamental Investment Research," Papers 2401.06164, arXiv.org.
    8. Cristina Angelico & Enrico Bernardini, 2026. "Can GenAI fill banks' emissions data gaps?," Questioni di Economia e Finanza (Occasional Papers) 1003, Bank of Italy, Economic Research and International Relations Area.
    9. Thanos Konstantinidis & Giorgos Iacovides & Mingxue Xu & Tony G. Constantinides & Danilo Mandic, 2024. "FinLlama: Financial Sentiment Classification for Algorithmic Trading Applications," Papers 2403.12285, arXiv.org.
    10. Yijia Xiao & Edward Sun & Tong Chen & Fang Wu & Di Luo & Wei Wang, 2025. "Trading-R1: Financial Trading with LLM Reasoning via Reinforcement Learning," Papers 2509.11420, arXiv.org.
    11. Meyer, Julian Anton, 2025. "Success factors and development areas for the implementation of Generative AI in companies," Junior Management Science (JUMS), Junior Management Science e. V., vol. 10(1), pages 1-23.
    12. Hui Gong, 2026. "AI Agents in Financial Markets: Architecture, Applications, and Systemic Implications," Papers 2603.13942, arXiv.org, revised Apr 2026.
    13. Frank Xing, 2024. "Designing Heterogeneous LLM Agents for Financial Sentiment Analysis," Papers 2401.05799, arXiv.org.
    14. Maher Hamid, 2026. "Implementing domain-specific LLMs for strategic investment decisions: a retrospective case study comparing AI and human expertise," Digital Finance, Springer, vol. 8(1), pages 1-134, March.
    15. Yikuan Huang & Zheqi Fan & Kaiqi Hu & Yifan Ye, 2026. "From Hypotheses to Factors: Constrained LLM Agents in Cryptocurrency Markets," Papers 2604.26747, arXiv.org.
    16. Haofei Yu & Fenghai Li & Jiaxuan You, 2025. "LiveTradeBench: Seeking Real-World Alpha with Large Language Models," Papers 2511.03628, arXiv.org.
    17. Geofrey Ntale, 2026. "AI Trading: Evaluating Large Language Models for Technical Market Analysis," Papers 2607.15414, arXiv.org.
    18. Ankur Sinha & Chaitanya Agarwal & Pekka Malo, 2025. "FinBloom: Knowledge Grounding Large Language Model with Real-time Financial Data," Papers 2502.18471, arXiv.org, revised Feb 2026.
    19. Hoyoung Lee & Youngsoo Choi & Yuhee Kwon, 2024. "Quantifying Qualitative Insights: Leveraging LLMs to Market Predict," Papers 2411.08404, arXiv.org.
    20. Seppälä, Timo & Mucha, Tomasz & Mattila, Juri, 2023. "Beyond AI, Blockchain Systems, and Digital Platforms: Digitalization Unlocks Mass Hyper-Personalization and Mass Servitization," ETLA Working Papers 106, The Research Institute of the Finnish Economy.

    More about this item

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2607.27853. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: https://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.