Advanced acoustic modelling techniques in MP3 speech recognition

Borsky Michal<sup>*</sup>; Pollak Petr; Mizera Petr

doi:10.1186/s13636-015-0064-7

摘要

The automatic recognition of MP3 compressed speech presents a challenge to the current systems due to the lossy nature of compression which causes irreversible degradation of the speech wave. This article evaluates the performance of a recognition system optimized for MP3 compressed speech with current state-of-the-art acoustic modelling techniques and one specific front-end compensation method. The article concentrates on acoustic model adaptation, discriminative training, and additional dithering as prominent means of compensating for the described distortion in the task of phoneme and large vocabulary continuous speech recognition (LVCSR). The experiments presented on the phoneme task show a dramatic increase of the recognition error for unvoiced speech units as a direct result of compression. The application of acoustic model adaptation has proved to yield the highest relative contribution while the gain of discriminative training diminished with decreasing bit-rate. The application of additional dithering yielded a consistent improvement only for the MFCC features, but the overall results were still worse than those for the PLP features.

出版日期2015-7-28

全文

访问全文

收藏分享被引(2) 浏览

更新时间：2024-04-28 05:37

Advanced acoustic modelling techniques in MP3 speech recognition

摘要

全文

产品服务

站内浏览

服务支持

联系方式

科研之友