Author
Listed:
- KM Niharika
- Dr
- Jeetendra Singh Yadav
Abstract
Enterprises routinely process large volumes of semi-structured business documents such as invoices, receipts, purchase orders and forms, and reliably extracting key fields from them remains a challenging information extraction (IE) problem. Purely rule-based systems are precise but brittle to layout variation, while purely learned models generalize well but behave as opaque black boxes and typically require substantial labelled data. This paper proposes an intelligent hybrid framework for automated document processing that integrates a deterministic rule/pattern layer, a machine-learning confidence classifier, and a Fuzzy Inference System (FIS) that fuses the two to resolve ambiguous or conflicting extractions. The framework follows a five-stage pipeline: document preprocessing, candidate field generation, feature extraction, fuzzy-based decision fusion, and final field output. The system was implemented end-to-end and evaluated on a labelled invoice document corpus (600 documents, three document-quality conditions) covering five field types: invoice number, invoice date, vendor name, total amount, and tax identification number. Experimental results show that the proposed hybrid NLP+ML+Fuzzy framework achieves an overall F1-score of 0.951 and field-level accuracy of 94.6%, compared to 0.878 F1 for a rule-based baseline and 0.883 F1 for the same hybrid architecture without the fuzzy fusion stage, confirming the specific contribution of fuzzy-based decision fusion, particularly for semantically ambiguous fields such as vendor names (F1 improvement from 0.444 to 0.806). The framework maintains a practical average processing latency of under 200 ms per document. An ablation study and per-field, per-condition breakdowns further validate the robustness and interpretability benefits of the proposed approach for real-world document automation pipelines.
Suggested Citation
KM Niharika & Dr & Jeetendra Singh Yadav, 2026.
"An Intelligent Framework for Automated Document Processing and Information Extraction Using Hybrid Natural Language Processing and Machine Learning,"
International Journal of Scientific Research in Artificial Intelligence and Machine Learning, International Journal of Scientific Research in Artificial Intelligence and Machine Learning, vol. 2(4), pages 53-61, July.
Handle:
RePEc:jbo:ijsrml:v2:y2026:i4:id:92
Note: Article URL: https://ijsraiml.com/home/article/view/IJSRAIML262418
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:jbo:ijsrml:v2:y2026:i4:id:92. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Pankaj Sharma (email available below). General contact details of provider: https://ijsraiml.com/home .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.