Personalization for robust voice pathology detection in sound waves
(2023) p.1708-1712- Abstract
- Automatic voice pathology detection is promising for non-invasive screening and early intervention using sound signals. Nevertheless, existing methods are susceptible to covariate shifts due to background noises, human voice variations, and data selection biases leading to severe performance degradation in real-world scenarios. Hence, we propose a non-invasive framework that contrastively learns personalization from sound waves as a pre-train and predicts latent-spaced profile features through semi-supervised learning. It allows all subjects from various distributions (e.g., regionality, gender, age) to benefit from personalized predictions for robust voice pathology in a privacy-fulfilled manner. We extensively evaluate the framework on... (More)
- Automatic voice pathology detection is promising for non-invasive screening and early intervention using sound signals. Nevertheless, existing methods are susceptible to covariate shifts due to background noises, human voice variations, and data selection biases leading to severe performance degradation in real-world scenarios. Hence, we propose a non-invasive framework that contrastively learns personalization from sound waves as a pre-train and predicts latent-spaced profile features through semi-supervised learning. It allows all subjects from various distributions (e.g., regionality, gender, age) to benefit from personalized predictions for robust voice pathology in a privacy-fulfilled manner. We extensively evaluate the framework on four real-world respiratory illnesses datasets, including Coswara, COUGHVID, ICBHI and our private dataset - ASound, under multiple covariate shift settings (i.e., cross-dataset), improving up to 4.12% in overall performance. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/record/79193861-63cf-4350-835a-fab176503436
- author
- Tran, Khanh-Tung ; Hoang, Truong ; Nguyen, Duy Khuong ; Nguyen, Hoang D and Vu, Xuan-Son LU
- publishing date
- 2023
- type
- Chapter in Book/Report/Conference proceeding
- publication status
- published
- subject
- host publication
- Proceedings of the annual conference of the international speech communication association, INTERSPEECH,
- pages
- 1708 - 1712
- publisher
- International Speech Communication Association
- external identifiers
-
- scopus:85171525230
- DOI
- 10.21437/Interspeech.2023-1332
- language
- English
- LU publication?
- no
- id
- 79193861-63cf-4350-835a-fab176503436
- date added to LUP
- 2026-02-11 00:06:28
- date last changed
- 2026-03-10 11:43:35
@inproceedings{79193861-63cf-4350-835a-fab176503436,
abstract = {{Automatic voice pathology detection is promising for non-invasive screening and early intervention using sound signals. Nevertheless, existing methods are susceptible to covariate shifts due to background noises, human voice variations, and data selection biases leading to severe performance degradation in real-world scenarios. Hence, we propose a non-invasive framework that contrastively learns personalization from sound waves as a pre-train and predicts latent-spaced profile features through semi-supervised learning. It allows all subjects from various distributions (e.g., regionality, gender, age) to benefit from personalized predictions for robust voice pathology in a privacy-fulfilled manner. We extensively evaluate the framework on four real-world respiratory illnesses datasets, including Coswara, COUGHVID, ICBHI and our private dataset - ASound, under multiple covariate shift settings (i.e., cross-dataset), improving up to 4.12% in overall performance.}},
author = {{Tran, Khanh-Tung and Hoang, Truong and Nguyen, Duy Khuong and Nguyen, Hoang D and Vu, Xuan-Son}},
booktitle = {{Proceedings of the annual conference of the international speech communication association, INTERSPEECH,}},
language = {{eng}},
pages = {{1708--1712}},
publisher = {{International Speech Communication Association}},
title = {{Personalization for robust voice pathology detection in sound waves}},
url = {{http://dx.doi.org/10.21437/Interspeech.2023-1332}},
doi = {{10.21437/Interspeech.2023-1332}},
year = {{2023}},
}