IDEAS home Printed from https://ideas.repec.org/a/gam/jftint/v18y2026i8p421-d2011936.html

Bimodal Speech Emotion Recognition Using a Hybrid CNN-LSTM Architecture with Sentiment Fusion

Author

Listed:
  • Tze-Syn Yap

    (Centre for Advanced Analytics for Artificial Intelligence, Faculty of Information Science and Technology (FIST), Multimedia University, Jalan Ayer Keroh Lama, Melaka 75450, Malaysia)

  • Lee-Yeng Ong

    (Centre for Advanced Analytics for Artificial Intelligence, Faculty of Information Science and Technology (FIST), Multimedia University, Jalan Ayer Keroh Lama, Melaka 75450, Malaysia)

Abstract

Speech emotion recognition (SER) is a fundamental task in affective computing; however, traditional unimodal approaches often struggle to capture the complex emotional cues present in spontaneous conversational speech. Bimodal frameworks that integrate acoustic and textual information have therefore emerged to provide complementary semantic and acoustic representations. This study proposes a bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture. Using the Multimodal EmotionLines Dataset (MELD), the framework combines temporal acoustic features, statistical acoustic features, and predicted textual sentiment. Experimental results indicate that the proposed model achieves reliable recognition of majority emotion classes but exhibits limited performance on underrepresented minority classes due to severe class imbalance. To better understand the contribution of each modality, feature sufficiency and feature necessity analyses were conducted. Furthermore, an evaluation of alternative fusion strategies showed that the expressive attention networks did not provide meaningful performance improvements over simple feature concatenation. These findings suggest that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER, highlighting the importance of addressing data imbalance before pursuing more sophisticated multimodal architectures.

Suggested Citation

  • Tze-Syn Yap & Lee-Yeng Ong, 2026. "Bimodal Speech Emotion Recognition Using a Hybrid CNN-LSTM Architecture with Sentiment Fusion," Future Internet, MDPI, vol. 18(8), pages 1-18, August.
  • Handle: RePEc:gam:jftint:v:18:y:2026:i:8:p:421-:d:2011936
    as

    Download full text from publisher

    File URL: https://www.mdpi.com/1999-5903/18/8/421/pdf
    Download Restriction: no

    File URL: https://www.mdpi.com/1999-5903/18/8/421/
    Download Restriction: no
    ---><---

    More about this item

    Keywords

    ;
    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:gam:jftint:v:18:y:2026:i:8:p:421-:d:2011936. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: MDPI Indexing Manager The email address of this maintainer does not seem to be valid anymore. Please ask MDPI Indexing Manager to update the entry or send us the correct address (email available below). General contact details of provider: https://www.mdpi.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.