Toggle Main Menu Toggle Search

Open Access padlockePrints

The Newcastle University research output collection, currently available on ePrints, will shortly be moving to a new open repository platform, Figshare. To prepare for the data migration we have paused adding new content to ePrints, and will resume once the new repository is launched. During this time you will continue to have access to ePrints (but no new content will appear). We will share updates here when available.

Experiments with the Log Filter-Bank

Lookup NU author(s): Dr Derek Wilson, Emeritus Professor Lindsay MarshallORCiD

Downloads


Licence

This is the final published version of a report that has been published in its final definitive form by School of Computing Science, University of Newcastle upon Tyne, 2016.

For re-use rights please refer to the publisher's terms and conditions.


Abstract

The Log Filter-bank is comprised of a set of overlapping passband filters— arranged at linear frequency intervals up to the frequency ft, and at logarithmic intervals above ft—which operate on the log magnitude of the Fourier transform of the speech signal; the purpose being to model to some extent the psychophysical aspects of human speech and hearing. Historically the Log Filter-Bank has featured as one of the key processes in the generation of the Mel Frequency Cepstral Coefficients(MFCC) representation of speech for Automatic Speech Recognition (ASR)—and although the MFCC technique has latterly been surpassed in performance by techniques using Deep Neural Networks there remains a role in ASR for the Log Filter-bank.From the literature on the characteristics of human hearing we know that the audible frequency spectrum can be described by a set of frequency bands—known as Critical Bands or Bark Bands; and that within each of these bands, that the human auditory system somehow amplifies the most powerful harmonic whilst simultaneously masking nearby harmonics. This knowledge has not had a great impact on speech processing, as typically neither the most powerful harmonic per band, nor simultaneous masking feature in current log filter-bank design. That this is the case, together with our belief that the hearing model of speech provides the most compact and accurate speech representation, provides the motivation for our work on a qualitative and quantitative comparison of reconstructed audio for various configurations of the log filter-bank, and in particular, an assessment of the strategy of arranging simultaneous masking around the dynamically derived most powerful harmonic for each of the Bark bands. With this strategy we find that the spectra of the reconstructed speech more closely resembles the spectrum of the speech master waveform. We also find that we can reconstruct a fully understandable (though far less than perfect) speech waveform using only the most powerful harmonic for each of the Bark bands; but if we reconstruct the speech waveform using only the centre harmonic for each of the Bark bands, the arbitrarily chosen harmonics dominate and the resultant ‘speech’ is incomprehensible. We can conclude from this that the contribution to speech of the most powerful harmonics is disproportionate to their number—thus providing further evidence for Bark Bands and simultaneous masking. That is, to achieve a more accurate sparse model of speech, the most powerful harmonic in each of the Bark bands should be included.


Publication metadata

Author(s): Wilson D, Marshall L

Publication type: Report

Publication status: Published

Series Title: School of Computing Science Technical Report Series

Year: 2016

Pages: 9

Print publication date: 01/07/2016

Acceptance date: 01/07/2016

Report Number: 1497

Institution: School of Computing Science, University of Newcastle upon Tyne

Place Published: Newcastle upon Tyne

URL: http://www.cs.ncl.ac.uk/publications/trs/papers/1497.pdf


Share