Inclusion of temporal information into features for speech recognition

B. Milner
Proceeding of Fourth International Conference on Spoken Language Processing. ICSLP '96  
Conventional methods for incorporating temporal information into speech features apply regression to a series of successive cepstral vectors to generate differential cepstra, or apply a cosine transform to generate cepstral-time matrices. This paper aims to generalise these techniques such that a series of stacked cepstral vectors is multiplied by a temporal transform matrix to produce the final speech feature. This can made to incorporate both static and dynamic speech information. Using this
more » ... ethod, the coding of temporal information is not restricted to regression or cosine coefficients -any suitable transform may used. Results are presented for a variety of transforms, such as Legendre, Karhunen-Loeve, Cosine, Rectangle, where it is shown that the transform based techniques offer higher performance than conventional differential cepstrum.
doi:10.1109/icslp.1996.607093 fatcat:b35obs25vzdnng276tq5ofejoa