IDEAS home Printed from https://ideas.repec.org/a/das/njaigs/v10y2026i01p1-23id488.html

Evidence-Calibrated Financial Language Models for Macro-Policy Stance Classification and Decision Cards

Author

Listed:
  • Hugo Yan

Abstract

Central-bank language often conveys policy direction through qualifying clauses rather than isolated sentiment terms. This study evaluates hawkish, dovish, and neutral stance classification on all 496 FinBen-FOMC excerpts using leakage-controlled five-fold stratified-group cross-validation. Seven systems span a class-prior baseline, an n-gram language model, word-and-character TF-IDF, frozen FinBERT transfer scores, frozen MiniLM sentence embeddings, and two fused models. Four-fold inner cross-validation fits temperature and isotonic calibrators without access to outer-test labels. Evaluation combines macro-F1, Matthews correlation coefficient, expected calibration error, Brier score, confusion analysis, clustered bootstrap intervals, selective risk, and contrastive lexical evidence. TFIDF-LR achieved the highest pooled point estimates for macro-F1 (0.493) and MCC (0.246), followed closely by EvidenceFusion at 0.489 and 0.233. None of the macro-F1 differences between TFIDF-LR and another learned model was significant after Holm correction. Temperature scaling sharply reduced overconfidence in the n-gram and embedding systems; TFIDF+MiniLM attained the lowest temperature-scaled ECE (0.014), while TFIDF-LR achieved a Brier score of 0.587. Isotonic calibration lowered probability loss further for several models but reduced directional recall. Removing TFIDF-LR’s selected evidence terms reduced predicted-class confidence by 0.170 and changed 69.4% of labels, whereas matched random removal changed confidence by −0.019. At a 0.70 acceptance threshold, TFIDF-LR covered 6.45% of records with 21.9% risk, including a high-confidence error driven by tightening vocabulary despite explicit negation. Evidence-calibrated decision cards are therefore most appropriate for conservative analyst triage rather than autonomous policy interpretation.

Suggested Citation

  • Hugo Yan, 2026. "Evidence-Calibrated Financial Language Models for Macro-Policy Stance Classification and Decision Cards," Journal of Artificial Intelligence General science (JAIGS) ISSN:3006-4023, Open Knowledge, vol. 10(01), pages 1-23.
  • Handle: RePEc:das:njaigs:v:10:y:2026:i:01:p:1-23:id:488
    as

    Download full text from publisher

    File URL: https://newjaigs.com/index.php/JAIGS/article/view/488
    Download Restriction: no
    ---><---

    More about this item

    Keywords

    ;
    ;
    ;
    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:das:njaigs:v:10:y:2026:i:01:p:1-23:id:488. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Open Knowledge (email available below). General contact details of provider: https://newjaigs.com/index.php/JAIGS/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.