Skip to main content

LUP Student Papers

LUND UNIVERSITY LIBRARIES

Continuous-Speech Parkinson's Disease Detection Using Acoustic and Inharmonicity Features

Li, Rujia LU (2026) In Master's Theses in Mathematical Sciences MASM02 20261
Mathematical Statistics
Abstract
Speech has become an increasingly important source of non-invasive information for Parkinson's disease detection, but much existing work still relies on sustained vowels. This thesis investigates whether carefully represented continuous speech can provide added predictive value over a strong sustained-vowel benchmark. Using the NeuroVoz corpus, sustained-vowel recordings were first used to establish a controlled single-vowel reference condition.

The continuous-speech framework was built around selected vowel-centred regions rather than whole-recording summaries. These regions were used in two ways. First, matched short-time acoustic descriptors were summarised into a recording-level acoustic representation. Second, the same extracted... (More)
Speech has become an increasingly important source of non-invasive information for Parkinson's disease detection, but much existing work still relies on sustained vowels. This thesis investigates whether carefully represented continuous speech can provide added predictive value over a strong sustained-vowel benchmark. Using the NeuroVoz corpus, sustained-vowel recordings were first used to establish a controlled single-vowel reference condition.

The continuous-speech framework was built around selected vowel-centred regions rather than whole-recording summaries. These regions were used in two ways. First, matched short-time acoustic descriptors were summarised into a recording-level acoustic representation. Second, the same extracted frames were analysed through signed harmonic-offset patterns and pooled into a person-level inharmonicity representation based on covariance and geometric structure. The acoustic and inharmonicity models were evaluated separately and then combined through late probability fusion.

Among the sustained vowels, /u/ gave the strongest single-vowel benchmark. The continuous-speech acoustic model exceeded this benchmark, and the best observed performance was obtained when the acoustic and inharmonicity scores were combined using simple weighted fusion. The inharmonicity model was weaker as a standalone classifier, but the fusion results suggest that it may provide complementary information to the acoustic representation.

These results support the use of continuous speech for Parkinson's disease (PD) detection when the representation is restricted to reliable vowel-centred regions and evaluated at speaker level. The inharmonicity features did not outperform the acoustic model on their own, but the highest observed performance was obtained when their probability scores were combined with the acoustic-model scores. (Less)
Popular Abstract (Swedish)
Parkinsons sjukdom förknippas ofta med rörelsesymtom, som skakningar,
stelhet och långsammare rörelser. Sjukdomen kan dock också påverka talet.
Rösten kan bli svagare, mindre varierad eller mindre stabil, och uttalet kan
bli svårare att kontrollera. Eftersom tal är enkelt att spela in och inte kräver
några invasiva undersökningar är det intressant att undersöka om tal kan
användas som en källa till information om Parkinsons sjukdom.

I detta examensarbete studeras om talinspelningar kan användas för att skilja
personer med Parkinsons sjukdom från friska kontrollpersoner. En viktig fråga
är vilken typ av tal som är mest användbar. Mycket tidigare forskning har
använt hållna vokaler, till exempel när en person säger ett långt... (More)
Parkinsons sjukdom förknippas ofta med rörelsesymtom, som skakningar,
stelhet och långsammare rörelser. Sjukdomen kan dock också påverka talet.
Rösten kan bli svagare, mindre varierad eller mindre stabil, och uttalet kan
bli svårare att kontrollera. Eftersom tal är enkelt att spela in och inte kräver
några invasiva undersökningar är det intressant att undersöka om tal kan
användas som en källa till information om Parkinsons sjukdom.

I detta examensarbete studeras om talinspelningar kan användas för att skilja
personer med Parkinsons sjukdom från friska kontrollpersoner. En viktig fråga
är vilken typ av tal som är mest användbar. Mycket tidigare forskning har
använt hållna vokaler, till exempel när en person säger ett långt ``aaa''. En sådan
inspelning är enkel att analysera, men den beskriver bara en liten del av hur vi
faktiskt talar. I vanligt sammanhängande tal måste talaren däremot hela tiden
växla mellan olika ljud, rytmer och artikulatoriska rörelser. Därför kan
sammanhängande tal innehålla mer information, men det är också svårare att
analysera.

Arbetet använder talmaterial från NeuroVoz-korpusen, som innehåller spanska
inspelningar från personer med Parkinsons sjukdom och friska kontrollpersoner.
Först byggdes en jämförelsemodell baserad på hållna vokaler. Därefter
undersöktes korta lyssna-och-upprepa-inspelningar som en form av
sammanhängande tal. I stället för att analysera hela yttranden direkt valdes
korta vokalcentrerade delar av talet ut. Dessa delar bedömdes vara mer stabila
och mer lämpliga för akustisk analys.

Två typer av information användes från det sammanhängande talet. Den första
var vanliga akustiska egenskaper, till exempel egenskaper kopplade till tonhöjd,
ljudstyrka och spektral struktur. Den andra var ett mer specialiserat mått på hurväl talets spektrum följer en ideal harmonisk struktur. Detta kallas här
inharmonicitet och kan ses som ett sätt att beskriva små avvikelser i röstens
periodiska struktur.

Resultaten visar att den akustiska modellen baserad på sammanhängande tal
gav bättre resultat än den bästa modellen baserad på en enskild hållen vokal.
Inharmonicitetsmodellen var svagare när den användes ensam, men den verkade
bidra med kompletterande information när den kombinerades med den akustiska
modellen. Det bästa resultatet erhölls därför när de två modellerna kombinerades.

Sammanfattningsvis tyder arbetet på att sammanhängande tal kan vara en
värdefull källa till information för talbaserad analys av Parkinsons sjukdom,
särskilt när analysen fokuserar på stabila och noggrant utvalda delar av talet.
Resultaten bör inte tolkas som ett färdigt diagnostiskt verktyg, men de visar atttal kan innehålla användbara mönster som är värda att studera vidare. (Less)
Please use this url to cite or link to this publication:
author
Li, Rujia LU
supervisor
organization
course
MASM02 20261
year
type
H2 - Master's Degree (Two Years)
subject
keywords
Parkinson's disease detection, voice anomaly detection, continuous speech, inharmonicity estimation, speech analysis
publication/series
Master's Theses in Mathematical Sciences
report number
LUNFMS-3139-2026
ISSN
1404-6342
other publication id
2026:E21
language
English
id
9226135
date added to LUP
2026-05-13 10:42:33
date last changed
2026-05-13 10:42:33
@misc{9226135,
  abstract     = {{Speech has become an increasingly important source of non-invasive information for Parkinson's disease detection, but much existing work still relies on sustained vowels. This thesis investigates whether carefully represented continuous speech can provide added predictive value over a strong sustained-vowel benchmark. Using the NeuroVoz corpus, sustained-vowel recordings were first used to establish a controlled single-vowel reference condition.

The continuous-speech framework was built around selected vowel-centred regions rather than whole-recording summaries. These regions were used in two ways. First, matched short-time acoustic descriptors were summarised into a recording-level acoustic representation. Second, the same extracted frames were analysed through signed harmonic-offset patterns and pooled into a person-level inharmonicity representation based on covariance and geometric structure. The acoustic and inharmonicity models were evaluated separately and then combined through late probability fusion.

Among the sustained vowels, /u/ gave the strongest single-vowel benchmark. The continuous-speech acoustic model exceeded this benchmark, and the best observed performance was obtained when the acoustic and inharmonicity scores were combined using simple weighted fusion. The inharmonicity model was weaker as a standalone classifier, but the fusion results suggest that it may provide complementary information to the acoustic representation.

These results support the use of continuous speech for Parkinson's disease (PD) detection when the representation is restricted to reliable vowel-centred regions and evaluated at speaker level. The inharmonicity features did not outperform the acoustic model on their own, but the highest observed performance was obtained when their probability scores were combined with the acoustic-model scores.}},
  author       = {{Li, Rujia}},
  issn         = {{1404-6342}},
  language     = {{eng}},
  note         = {{Student Paper}},
  series       = {{Master's Theses in Mathematical Sciences}},
  title        = {{Continuous-Speech Parkinson's Disease Detection Using Acoustic and Inharmonicity Features}},
  year         = {{2026}},
}