IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2607.00856.html

Shapley in Context: Explaining Financial Language with Domain Expertise

Author

Listed:
  • Dangxing Chen
  • Pengzhan Guo

Abstract

In recent years, large language models have achieved remarkable success and have seen growing adoption in financial applications. At the same time, explainability remains critical in finance, a domain characterized by high stakes and strict regulatory requirements. Although numerous methods have been proposed to explain black box machine learning models, the majority of these approaches are designed for general purpose tasks and do not incorporate domain specific knowledge. In this work, we study the explainability of financial textual data modeled by large language models through the lens of the Shapley value. Specifically, we investigate whether Shapley based attributions align with established financial domain knowledge. Through rigorous theoretical analysis and extensive empirical evaluations, we demonstrate that Shapley values can yield explanations that are consistent with financial reasoning and can offer meaningful insights into the model's behavior in text based financial applications.

Suggested Citation

  • Dangxing Chen & Pengzhan Guo, 2026. "Shapley in Context: Explaining Financial Language with Domain Expertise," Papers 2607.00856, arXiv.org.
  • Handle: RePEc:arx:papers:2607.00856
    as

    Download full text from publisher

    File URL: https://arxiv.org/pdf/2607.00856
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Haim Shalit, 2021. "The Shapley value decomposition of optimal portfolios," Annals of Finance, Springer, vol. 17(1), pages 1-25, March.
    2. Dangxing Chen & Weicheng Ye, 2023. "How to address monotonicity for model risk management?," Papers 2305.00799, arXiv.org, revised Sep 2023.
    3. Nikola Tarashev & Kostas Tsatsaronis & Claudio Borio, 2016. "Risk Attribution Using the Shapley Value: Methodology and Policy Applications," Review of Finance, European Finance Association, vol. 20(3), pages 1189-1213.
    4. Rosen, Dan & Saunders, David, 2010. "Risk factor contributions in portfolio credit risk models," Journal of Banking & Finance, Elsevier, vol. 34(2), pages 336-349, February.
    5. Attilio Meucci, 2005. "Risk and Asset Allocation," Springer Finance, Springer, number 978-3-540-27904-4, January.
    6. Julian Junyan Wang & Victor Xiaoqi Wang, 2025. "Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks," Papers 2503.16974, arXiv.org, revised Sep 2025.
    7. Guy Stephane Waffo Dzuyo & Gaël Guibon & Christophe Cerisara & Luis Belmar-Letelier, 2025. "Linking Industry Sectors and Financial Statements: A Hybrid Approach for Company Classification," Post-Print hal-05031499, HAL.
    8. Mukund Sundararajan & Amir Najmi, 2019. "The many Shapley values for model explanation," Papers 1908.08474, arXiv.org, revised Feb 2020.
    9. Federico Siano, 2025. "The News in Earnings Announcement Disclosures: Capturing Word Context Using LLM Methods," Management Science, INFORMS, vol. 71(11), pages 9831-9855, November.
    10. Dangxing Chen & Weicheng Ye, 2022. "Monotonic Neural Additive Models: Pursuing Regulated Machine Learning Models for Credit Scoring," Papers 2209.10070, arXiv.org.
    11. Dangxing Chen & Jingfeng Chen & Weicheng Ye, 2024. "Why Groups Matter: Necessity of Group Structures in Attributions," Papers 2408.05701, arXiv.org.
    12. Dangxing Chen, 2025. "Explaining Risks: Axiomatic Risk Attributions for Financial Models," Papers 2506.06653, arXiv.org.
    13. Steve Y. Yang & Sheung Yin Kevin Mo & Anqi Liu, 2015. "Twitter financial community sentiment and its predictive relationship to stock market movement," Quantitative Finance, Taylor & Francis Journals, vol. 15(10), pages 1637-1656, October.
    14. Dangxing Chen, 2025. "Explaining risks: axiomatic risk attributions for financial models," Quantitative Finance, Taylor & Francis Journals, vol. 25(6), pages 1007-1014, June.
    15. Shijie Wu & Ozan Irsoy & Steven Lu & Vadim Dabravolski & Mark Dredze & Sebastian Gehrmann & Prabhanjan Kambadur & David Rosenberg & Gideon Mann, 2023. "BloombergGPT: A Large Language Model for Finance," Papers 2303.17564, arXiv.org, revised Dec 2023.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Giacomo Morelli, 2024. "Responsible investing and portfolio selection: a shapley - CVaR approach," Annals of Operations Research, Springer, vol. 342(3), pages 1991-2019, November.
    2. Dangxing Chen, 2025. "Explaining Risks: Axiomatic Risk Attributions for Financial Models," Papers 2506.06653, arXiv.org.
    3. Mohammed-Khalil Ghali & Cecil Pang & Oscar Molina & Carlos Gershenson-Garcia & Daehan Won, 2025. "Forecasting Commodity Price Shocks Using Temporal and Semantic Fusion of Prices Signals and Agentic Generative AI Extracted Economic News," Papers 2508.06497, arXiv.org.
    4. Nicola Jean & Giacomo Le Pera & Lorenzo Giada & Claudio Nordio, 2025. "A Framework for Waterfall Pricing Using Simulation-Based Uncertainty Modeling," Papers 2507.13324, arXiv.org.
    5. Jang, Junkyu, 2025. "Selective news selection model for explainable stock prediction via cross-attention integration," Finance Research Letters, Elsevier, vol. 85(PD).
    6. Ching-Nam Hang & Pei-Duo Yu & Roberto Morabito & Chee-Wei Tan, 2024. "Large Language Models Meet Next-Generation Networking Technologies: A Review," Future Internet, MDPI, vol. 16(10), pages 1-29, October.
    7. Nan Zhang & Heng Xu, 2024. "Fairness of Ratemaking for Catastrophe Insurance: Lessons from Machine Learning," Information Systems Research, INFORMS, vol. 35(2), pages 469-488, June.
    8. Algieri, Bernardina & Leccadito, Arturo, 2017. "Assessing contagion risk from energy and non-energy commodity markets," Energy Economics, Elsevier, vol. 62(C), pages 312-322.
    9. Dangxing Chen & Luyao Zhang, 2023. "Monotonicity for AI ethics and society: An empirical study of the monotonic neural additive model in criminology, education, health care, and finance," Papers 2301.07060, arXiv.org.
    10. Xia Li & Hanghang Zheng & Xiwei Zhuang & Zhong Wang & Xiao Chen & Hong Liu & Jasmine Bai & Mao Mao, 2025. "Class-Imbalanced-Aware Adaptive Dataset Distillation for Scalable Pretrained Model on Credit Scoring," Papers 2501.10677, arXiv.org, revised Mar 2026.
    11. Meng-Jou Lu & Cathy Yi-Hsuan Chen & Wolfgang Karl Härdle, 2017. "Copula-based factor model for credit risk analysis," Review of Quantitative Finance and Accounting, Springer, vol. 49(4), pages 949-971, November.
    12. Marco Gregnanin & Johannes De Smedt & Giorgio Gnecco & Maurizio Parton, 2026. "A Generative Adversarial Graph Neural Network for Synthetic Time Series Data," Papers 2605.22215, arXiv.org.
    13. Alireza Rezazadeh & Yasamin Jafarian & Ali Kord, 2022. "Explainable Ensemble Machine Learning for Breast Cancer Diagnosis Based on Ultrasound Image Texture Features," Forecasting, MDPI, vol. 4(1), pages 1-13, February.
    14. Lezhi Li & Ting-Yu Chang & Hai Wang, 2023. "Multimodal Gen-AI for Fundamental Investment Research," Papers 2401.06164, arXiv.org.
    15. Cristina Angelico & Enrico Bernardini, 2026. "Can GenAI fill banks' emissions data gaps?," Questioni di Economia e Finanza (Occasional Papers) 1003, Bank of Italy, Economic Research and International Relations Area.
    16. Arnold Polanski & Evarist Stoja & Ching‐Wai (Jeremy) Chiu, 2021. "Tail risk interdependence," International Journal of Finance & Economics, John Wiley & Sons, Ltd., vol. 26(4), pages 5499-5511, October.
    17. Hu'e Sullivan & Hurlin Christophe & P'erignon Christophe & Saurin S'ebastien, 2022. "Measuring the Driving Forces of Predictive Performance: Application to Credit Scoring," Papers 2212.05866, arXiv.org, revised Jan 2025.
    18. Thanos Konstantinidis & Giorgos Iacovides & Mingxue Xu & Tony G. Constantinides & Danilo Mandic, 2024. "FinLlama: Financial Sentiment Classification for Algorithmic Trading Applications," Papers 2403.12285, arXiv.org.
    19. Hugh Chen & Scott M. Lundberg & Su-In Lee, 2022. "Explaining a series of models by propagating Shapley values," Nature Communications, Nature, vol. 13(1), pages 1-15, December.
    20. John R. Graham & Campbell R. Harvey & Manish Jha, 2026. "CFOs Meet LLMs," Papers 2606.13812, arXiv.org.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2607.00856. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: https://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.