Computational auditory induction as a missing-data model-fitting problem with Bregman divergence

Le Roux Jonathan<sup>*</sup>; Kameoka Hirokazu; Ono Nobutaka; de Cheveigne Alain; Sagayama Shigeki

doi:10.1016/j.specom.2010.08.009

摘要

The human auditory system has the ability, known as auditory induction, to estimate the missing parts of a continuous auditory stream briefly covered by noise and perceptually resynthesize them. In this article, we formulate this ability as a model-based spectrogram analysis and clustering problem with missing data, show how to solve it using an auxiliary function method, and explain how this method is generally related to the expectation-maximization (EM) algorithm for a certain type of divergence measures called Bregman divergences, thus enabling the use of prior distributions on the parameters. We illustrate how our method can be used to simultaneously analyze a scene and estimate missing information with two algorithms: the first, based on non-negative matrix factorization (NMF), performs analysis of polyphonic multi-instrumental musical pieces. Our method allows this algorithm to cope with gaps within the audio data, estimating the timbre of the instruments and their pitch, and reconstructing the missing parts. The second, based on a recently introduced technique for the analysis of complex acoustical scenes called harmonic-temporal clustering (FITC), enables us to perform robust fundamental frequency estimation from incomplete speech data.

出版日期2011-6

全文

访问全文

收藏分享被引(4) 浏览

更新时间：2018-02-09 10:55

Computational auditory induction as a missing-data model-fitting problem with Bregman divergence

摘要

全文

产品服务

站内浏览

服务支持

联系方式

科研之友