Artificial speech detection using image-based features and random forest classifier | Zendy

Choon Beng Tan | Zendy; Mohd Hanafi Ahmad Hijazi | Zendy; Frazier Kok | Zendy; Mohd Saberi Mohamad | Zendy; Puteri N. E. Nohuddin | Zendy

AI Assistant Blog Pricing

Home ZAIA Blog

Open Access

Artificial speech detection using image-based features and random forest classifier

Author(s) -

Choon Beng Tan,

Mohd Hanafi Ahmad Hijazi,

Frazier Kok,

Mohd Saberi Mohamad,

Puteri N. E. Nohuddin

Publication year - 2022

Publication title -

iaes international journal of artificial intelligence

Language(s) - English

Resource type - Journals

SCImago Journal Rank - 0.341

H-Index - 7

eISSN - 2252-8938

pISSN - 2089-4872

DOI - 10.11591/ijai.v11.i1.pp161-172

Subject(s) - computer science , random forest , spoofing attack , classifier (uml) , artificial intelligence , pattern recognition (psychology) , mel frequency cepstrum , speech recognition , word error rate , feature extraction , computer network

The ASVspoof 2015 Challenge was one of the efforts of the research community in the field of speech processing to foster the development of generalized countermeasures against spoofing attacks. However, most countermeasures submitted to the ASVspoof 2015 Challenge failed to detect the S10 attack effectively, the only attack that was generated using the waveform concatenation approach. Hence, more informative features are needed to detect previously unseen spoofing attacks. This paper presents an approach that uses data transformation techniques to engineer image-based features together with random forest classifier to detect artificial speech. The objectives are two-fold: (i) to extract image-based features from the melfrequency cepstral coefficients representation of the speech signal and (ii) to compare the performance of using the extracted features and Random Forest to determine the authenticity of voices with the existing approaches. An audio-to-image transformation technique was used to engineer new features in classifying genuine and spoof voices. An experiment was conducted to find the appropriate combination of the engineered features and classifier. Experimental results showed that the proposed approach was able to detect speech synthesis and voice conversion attacks effectively, with an equal error rate of 0.10% and accuracy of 99.93%.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.

Having issues? You can contact us here

Empowering knowledge with every search

About

About Careers Publisher Partners Contact Us

Learn

FAQs Blog Terms of Use Privacy Policy

About

Learn

Discover

Explore