IDEAS home Printed from https://ideas.repec.org/a/ijs/ijsrse/v11y2024i2id28.html

Survey on Silentinterpreter : Analysis of Lip Movement and Extracting Speech using Deep Learning

Author

Listed:
  • Ameen Hafeez
  • Rohith M K
  • Sakshi Prashant
  • Sinchana Hegde
  • Shwetha K S

Abstract

Lip reading is a complex but interesting path for the growth of speech recognition algorithms. It is the ability of deciphering spoken words by evaluating visual cues from lip movements. In this study, we suggest a unique method for lip reading that converts lip motions into textual representations by using deep neural networks. Convolutional neural networks are used in the methodology to extract visual features, recurrent neural networks are used to simulate temporal context, and the Connectionist Temporal Classification loss function is used to align lip features with corresponding phonemes. The study starts with a thorough investigation of data loading methods, which include alignment extraction and video preparation. A well selected dataset with video clips and matching phonetic alignments is presented. We select relevant face regions, convert frames to grayscale, then standardize the resulting data so that it can be fed into a neural network. The neural network architecture is presented in depth, displaying a series of bidirectional LSTM layers for temporal context understanding after 3D convolutional layers for spatial feature extraction. Careful consideration of input shapes, layer combinations, and parameter selections forms the foundation of the model's design. To train the model, we align predicted phoneme sequences with ground truth alignments using the CTC loss. Dynamic learning rate scheduling and a unique callback mechanism for training visualization of predictions are integrated into the training process. After training on a sizable dataset, the model exhibits remarkable convergence and proves its capacity to understand intricate temporal correlations. Through the use of both quantitative and qualitative evaluations, the results are thoroughly assessed. We visually check the model's lip reading abilities and assess its performance using common speech recognition criteria. It is explored how different model topologies and hyperparameters affect performance, offering guidance for future research. The trained model is tested on external video samples to show off its practical application. Its accuracy and resilience in lip-reading spoken phrases are demonstrated. By providing a deep learning framework for precise and effective speech recognition, this research adds to the rapidly changing field of lip reading devices. The results offer opportunities for additional development and implementation in various fields, such as assistive technologies, audio-visual communication systems, and human-computer interaction.

Suggested Citation

  • Ameen Hafeez & Rohith M K & Sakshi Prashant & Sinchana Hegde & Shwetha K S, 2024. "Survey on Silentinterpreter : Analysis of Lip Movement and Extracting Speech using Deep Learning," International Journal of Scientific Research in Science, Engineering and Technology, Technoscience Academy, vol. 11(2), pages 183-191, April.
  • Handle: RePEc:ijs:ijsrse:v11:y2024:i2:id:28
    DOI: 10.32628/IJSRSET2411219
    as

    Download full text from publisher

    File URL: https://ijsrset.com/home/article/view/IJSRSET2411219
    File Function: Abstract page
    Download Restriction: no

    File URL: https://ijsrset.com/home/article/download/IJSRSET2411219/IJSRSET2411223
    File Function: Full text
    Download Restriction: no

    File URL: https://libkey.io/10.32628/IJSRSET2411219?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Keywords

    ;
    ;
    ;
    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:ijs:ijsrse:v11:y2024:i2:id:28. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Pankaj Sharma (email available below). General contact details of provider: https://ijsrset.com/home .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.