logo Idiap Research Institute        
 [BibTeX] [Marc21]
A Sector-Based, Frequency-Domain Approach to Detection and Localization of Multiple Speakers
Type of publication: Idiap-RR
Citation: lathoud-rr-04-54
Number: Idiap-RR-54-2004
Year: 2004
Institution: IDIAP
Address: Martigny, Switzerland
Note: To appear in ``Proceedings of the ICASSP 2005''
Abstract: Detection and localization of speakers with microphone arrays is a difficult task due to the wideband nature of speech signals, the large amount of overlaps between speakers in spontaneous conversations, and the presence of noise sources. Many existing audio multi-source localization methods rely on prior knowledge of the sectors containing active sources and/or the number of active sources. This paper proposes sector-based, frequency-domain approaches that address both detection and localization problems by measuring relative phases between microphones. The first approach is similar to delay-sum beamforming. The second approach is novel: it relies on systematic optimization of a centroid in phase space, for each sector. It provides major, systematic improvement over the first approach as well as over previous work. Very good results are obtained on more than one hour of recordings in real meeting room conditions, including cases with up to 3 concurrent speakers.
Userfields: ipdinar={2004}, ipdmembership={speech}, language={English},
Keywords:
Projects Idiap
Authors Lathoud, Guillaume
Magimai.-Doss, Mathew
Crossref by lathoud05a
Added by: [UNK]
Total mark: 0
Attachments
  • rr-04-54.pdf
  • rr-04-54.ps.gz
Notes