dc.contributor.author |
Pitsikalis, V |
en |
dc.contributor.author |
Maragos, P |
en |
dc.date.accessioned |
2014-03-01T01:29:51Z |
|
dc.date.available |
2014-03-01T01:29:51Z |
|
dc.date.issued |
2009 |
en |
dc.identifier.issn |
0167-6393 |
en |
dc.identifier.uri |
https://dspace.lib.ntua.gr/xmlui/handle/123456789/19371 |
|
dc.subject |
Broad class phoneme classification |
en |
dc.subject |
Feature extraction |
en |
dc.subject |
Generalized fractal dimensions |
en |
dc.subject.classification |
Acoustics |
en |
dc.subject.classification |
Communication |
en |
dc.subject.classification |
Computer Science, Interdisciplinary Applications |
en |
dc.subject.classification |
Language & Linguistics |
en |
dc.subject.other |
Broad class phoneme classification |
en |
dc.subject.other |
Classification of speech |
en |
dc.subject.other |
Feature vectors |
en |
dc.subject.other |
Fractal feature |
en |
dc.subject.other |
Fractal theory |
en |
dc.subject.other |
Generalized fractal dimensions |
en |
dc.subject.other |
Mel-frequency cepstral coefficients |
en |
dc.subject.other |
Non-linear signal processing |
en |
dc.subject.other |
Phase spaces |
en |
dc.subject.other |
Phoneme classification |
en |
dc.subject.other |
Qualitative aspects |
en |
dc.subject.other |
Raw measurements |
en |
dc.subject.other |
Spectral content |
en |
dc.subject.other |
Speech signals |
en |
dc.subject.other |
Speech sounds |
en |
dc.subject.other |
Statistical parameters |
en |
dc.subject.other |
Dynamical systems |
en |
dc.subject.other |
Fractal dimension |
en |
dc.subject.other |
Initiators (chemical) |
en |
dc.subject.other |
Linguistics |
en |
dc.subject.other |
Signal processing |
en |
dc.subject.other |
Feature extraction |
en |
dc.title |
Analysis and classification of speech signals by generalized fractal dimension features |
en |
heal.type |
journalArticle |
en |
heal.identifier.primary |
10.1016/j.specom.2009.06.005 |
en |
heal.identifier.secondary |
http://dx.doi.org/10.1016/j.specom.2009.06.005 |
en |
heal.language |
English |
en |
heal.publicationDate |
2009 |
en |
heal.abstract |
We explore nonlinear signal processing methods inspired by dynamical systems and fractal theory in order to analyze and characterize speech sounds. A speech signal is at first embedded in a multidimensional phase-space and further employed for the estimation of measurements related to the fractal dimensions. Our goals are to compute these raw measurements in the practical cases of speech signals, to further utilize them for the extraction of simple descriptive features and to address issues on the efficacy of the proposed features to characterize speech sounds. We observe that distinct feature vector elements obtain values or show statistical trends that on average depend on general characteristics such as the voicing, the manner and the place of articulation of broad phoneme classes. Moreover the way that the statistical parameters of the features are altered as an effect of the variation of phonetic characteristics seem to follow some roughly formed patterns. We also discuss some qualitative aspects concerning the linear phoneme-wise correlation between the fractal features and the commonly employed mel-frequency cepstral coefficients (MFCCs) demonstrating phonetic cases of maximal and minimal correlation. In the same context we also investigate the fractal features' spectral content, in terms of the most and least correlated components with the MFCC. Further the proposed methods are examined under the light of indicative phoneme classification experiments. These quantify the efficacy of the features to characterize broad classes of speech sounds. The results are shown to be comparable for some classification scenarios with the corresponding ones of the MFCC features. (C) 2009 Elsevier B.V. All rights reserved. |
en |
heal.publisher |
ELSEVIER SCIENCE BV |
en |
heal.journalName |
Speech Communication |
en |
dc.identifier.doi |
10.1016/j.specom.2009.06.005 |
en |
dc.identifier.isi |
ISI:000274888800005 |
en |
dc.identifier.volume |
51 |
en |
dc.identifier.issue |
12 |
en |
dc.identifier.spage |
1206 |
en |
dc.identifier.epage |
1223 |
en |