z-logo
open-access-imgOpen Access
A hybrid CNN-LiGRU acoustic modeling using raw waveform sincnet for Hindi ASR
Author(s) -
Ankit Kumar,
Rajesh Kumar Aggarwal
Publication year - 2020
Publication title -
computer science
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.145
H-Index - 5
eISSN - 2300-7036
pISSN - 1508-2806
DOI - 10.7494/csci.2020.21.4.3748
Subject(s) - computer science , recurrent neural network , speech recognition , convolution (computer science) , convolutional neural network , waveform , hindi , artificial intelligence , signal (programming language) , pattern recognition (psychology) , computation , artificial neural network , algorithm , telecommunications , radar , programming language
Deep Neural Network (DNN) is currently playing the most vital role in Automatic Speech Recognition (ASR). Convolution Neural Network (CNN) and Recurrent Neural Network (RNN) are the advanced versions of DNN. CNN and RNN are right to deal with spatial and temporal properties of the speech signal, respectively, and both properties have a higher impact on accuracy. In today’s scenario, many acoustic modeling techniques often switches due to the battle of CNNs and RNNs. In the last few years, CNN, with raw speech signal, shows their superiority over precomputed acoustic features. Recently, a novel first convolution layer named SincNet was proposed to produce the interpretable filters with better accuracy. In this work, we proposed a hybrid SincNet-CNN-RNN architecture with low computation cost and high accuracy. Different configurations of the hybrid model were extensively examined to achieve this goal. All experiments were performed on the Hindi speech dataset.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom