IDEAS home Printed from https://ideas.repec.org/a/gam/jftint/v18y2026i5p251-d1938505.html

AI-Based Framework for Arabic Language Proficiency Assessment: A Deep Learning ASR Model with Enhanced Similarity Measures

Author

Listed:
  • Sufian A. Badawi

    (Department of Autonomous Systems, Faculty of Artificial Intelligence, Al-Balqa Applied University, Al-Salt 19117, Jordan)

  • Maen Takruri

    (College of Engineering and Technology, American University of the Middle East, Egaila 54200, Kuwait)

  • Khouloud Salameh

    (Advanced Technology and Artificial Intelligence Center (ATAIC), American University of Ras Al Khaimah, Ras Al Khaimah 72603, United Arab Emirates)

  • Mohammad Al-Badawi

    (English Language and Literature Department, Faculty of Arts, Al-Zaytoonah University of Jordan, Amman 11733, Jordan)

  • Nowar Alani

    (Department of General Education, RAK Medical and Health Sciences University, Ras Al Khaimah 11172, United Arab Emirates)

  • Isam ElBadawi

    (Industrial Engineering Department, College of Engineering, University of Ha’il, Ha’il 81481, Saudi Arabia)

  • Aws Al-Qaisi

    (College of Engineering and Technology, American University of the Middle East, Egaila 54200, Kuwait)

  • Ghaleb Aldoboni

    (Advanced Technology and Artificial Intelligence Center (ATAIC), American University of Ras Al Khaimah, Ras Al Khaimah 72603, United Arab Emirates)

Abstract

This work presents an innovative approach to test the Arabic language proficiency assessment via Automatic Speech Recognition (ASR) by enhancing the proficiency of the Whisper model in transcribing Arabic speech. The core of our research involved fine-tuning the Whisper model using a substantial, large-scale Arabic speech corpus, with a specific focus on Modern Standard Arabic. This process used a 2000-h Arabic-labeled speech corpus, the QASR dataset, and improved the model’s Word Error Rate (WER). After optimization, the fine-tuned Whisper model’s WER was reduced from 35% to 7% on the QASR dataset, corresponding to an absolute reduction of 28 percentage points (approximately 80% relative reduction). These results demonstrate the strong generalization ability of the fine-tuned model across multiple Arabic ASR benchmarks. A key component of our methodology was the development of a sophisticated scoring system. This system integrates various similarity metrics, such as cosine similarity, the Jaccard index, and the Levenshtein distance, with a machine learning regression model. This multifaceted system provides a comprehensive assessment of reading proficiency, proposing a practical automated assessment method that contributes to the field of AI language transcription and to its application in the assessment of students’ reading. Our research also introduces the ICONET dataset, an augmented Arabic speech corpus comprising 3160 h of diverse and tailored audio–text pairs designed for fine-tuning ASR models. This study demonstrates the potential of fine-tuning pretrained models for specific linguistic contexts (Arabic), establishing a foundation for future research in ASR and language technology.

Suggested Citation

  • Sufian A. Badawi & Maen Takruri & Khouloud Salameh & Mohammad Al-Badawi & Nowar Alani & Isam ElBadawi & Aws Al-Qaisi & Ghaleb Aldoboni, 2026. "AI-Based Framework for Arabic Language Proficiency Assessment: A Deep Learning ASR Model with Enhanced Similarity Measures," Future Internet, MDPI, vol. 18(5), pages 1-29, May.
  • Handle: RePEc:gam:jftint:v:18:y:2026:i:5:p:251-:d:1938505
    as

    Download full text from publisher

    File URL: https://www.mdpi.com/1999-5903/18/5/251/pdf
    Download Restriction: no

    File URL: https://www.mdpi.com/1999-5903/18/5/251/
    Download Restriction: no
    ---><---

    More about this item

    Keywords

    ;
    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:gam:jftint:v:18:y:2026:i:5:p:251-:d:1938505. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: MDPI Indexing Manager The email address of this maintainer does not seem to be valid anymore. Please ask MDPI Indexing Manager to update the entry or send us the correct address (email available below). General contact details of provider: https://www.mdpi.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.