Comparative diagnostic performance of machine learning models and traditional scores for HFpEF in older adults
(2026) In European Journal of Heart Failure- Abstract
AIMS: Diagnosing heart failure with preserved ejection fraction (HFpEF) remains challenging, particularly in older individuals. We hypothesized that machine learning (ML) approaches could improve diagnostic accuracy compared with HFpEF scores.
METHODS: We evaluated the diagnostic performance of four supervised ML algorithms (random forest [RF], extreme gradient boosting [XGBoost], support vector machines, and decision trees) to identify HFpEF in individuals aged 60 to 80 years. The models were trained on three derivation cohorts (N = 1474; HFpEF: KaRen, MEDIA cohorts; community-based without HF: Malmö Preventive Project) and validated in two independent cohorts (N = 542; HFpEF: HF-Nancy cohort; community-based without HF:... (More)
AIMS: Diagnosing heart failure with preserved ejection fraction (HFpEF) remains challenging, particularly in older individuals. We hypothesized that machine learning (ML) approaches could improve diagnostic accuracy compared with HFpEF scores.
METHODS: We evaluated the diagnostic performance of four supervised ML algorithms (random forest [RF], extreme gradient boosting [XGBoost], support vector machines, and decision trees) to identify HFpEF in individuals aged 60 to 80 years. The models were trained on three derivation cohorts (N = 1474; HFpEF: KaRen, MEDIA cohorts; community-based without HF: Malmö Preventive Project) and validated in two independent cohorts (N = 542; HFpEF: HF-Nancy cohort; community-based without HF: STANISLAS cohort). Performance metrics included accuracy, F-measure, area under the receiver operating characteristic curve (AUC), and C-index. ML models were also compared with HFA-PEFF, H2FPEF, and HFpEF-ABA scores.
RESULTS: Among 2017 participants, RF and XGBoost demonstrated the highest diagnostic value, outperforming traditional HFpEF scores (AUC: RF, 0.98; XGBoost, 0.96; HFA-PEFF, 0.86; H2FPEF, 0.79). RF and XGBoost also showed the greatest gain in discriminative capacity among ML algorithms when compared with H2FPEF (ΔC-index: RF +0.20, XGBoost +0.18), HFA-PEFF (ΔC-index: RF +0.12, XGBoost +0.10), and HFpEF-ABA score (ΔC-index: RF +0.17, XGBoost +0.15). Elevated natriuretic peptides were by far the most influential feature in both RF and XGBoost models (36% of model explainability).
CONCLUSIONS: Machine learning algorithms, particularly RF and XGBoost, demonstrated superior diagnostic accuracy compared to established HFpEF scoring systems. These findings support the potential integration of ML-based tools into clinical workflows to facilitate earlier identification of HFpEF and prompt initiation of guideline-recommended therapies.
(Less)
- author
- organization
- publishing date
- 2026-03-20
- type
- Contribution to journal
- publication status
- epub
- subject
- in
- European Journal of Heart Failure
- publisher
- Elsevier
- external identifiers
-
- pmid:41859834
- ISSN
- 1879-0844
- DOI
- 10.1093/ejhf/xuag039
- language
- English
- LU publication?
- yes
- additional info
- © The Author(s) 2026. Published by Oxford University Press on behalf of the European Society of Cardiology. All rights reserved. For commercial re-use, please contact reprints@oup.com for reprints and translation rights for reprints. All other permissions can be obtained through our RightsLink service via the Permissions link on the article page on our site—for further information please contact journals.permissions@oup.com.
- id
- 4f4df661-fd47-4862-b0e3-e0e1176fbde4
- date added to LUP
- 2026-03-23 10:16:22
- date last changed
- 2026-03-23 12:40:20
@article{4f4df661-fd47-4862-b0e3-e0e1176fbde4,
abstract = {{<p>AIMS: Diagnosing heart failure with preserved ejection fraction (HFpEF) remains challenging, particularly in older individuals. We hypothesized that machine learning (ML) approaches could improve diagnostic accuracy compared with HFpEF scores.</p><p>METHODS: We evaluated the diagnostic performance of four supervised ML algorithms (random forest [RF], extreme gradient boosting [XGBoost], support vector machines, and decision trees) to identify HFpEF in individuals aged 60 to 80 years. The models were trained on three derivation cohorts (N = 1474; HFpEF: KaRen, MEDIA cohorts; community-based without HF: Malmö Preventive Project) and validated in two independent cohorts (N = 542; HFpEF: HF-Nancy cohort; community-based without HF: STANISLAS cohort). Performance metrics included accuracy, F-measure, area under the receiver operating characteristic curve (AUC), and C-index. ML models were also compared with HFA-PEFF, H2FPEF, and HFpEF-ABA scores.</p><p>RESULTS: Among 2017 participants, RF and XGBoost demonstrated the highest diagnostic value, outperforming traditional HFpEF scores (AUC: RF, 0.98; XGBoost, 0.96; HFA-PEFF, 0.86; H2FPEF, 0.79). RF and XGBoost also showed the greatest gain in discriminative capacity among ML algorithms when compared with H2FPEF (ΔC-index: RF +0.20, XGBoost +0.18), HFA-PEFF (ΔC-index: RF +0.12, XGBoost +0.10), and HFpEF-ABA score (ΔC-index: RF +0.17, XGBoost +0.15). Elevated natriuretic peptides were by far the most influential feature in both RF and XGBoost models (36% of model explainability).</p><p>CONCLUSIONS: Machine learning algorithms, particularly RF and XGBoost, demonstrated superior diagnostic accuracy compared to established HFpEF scoring systems. These findings support the potential integration of ML-based tools into clinical workflows to facilitate earlier identification of HFpEF and prompt initiation of guideline-recommended therapies.</p>}},
author = {{Monzo, Luca and Huttin, Olivier and Bresso, Emmanuel and Duarte, Kevin and Linde, Cecilia and Lund, Lars H and Hage, Camilla and Donal, Erwan and Magnusson, Martin and Nilsson, Peter and Leosdottir, Margret and Bozec, Erwan and Baudry, Guillaume and Zannad, Faiez and Girerd, Nicolas}},
issn = {{1879-0844}},
language = {{eng}},
month = {{03}},
publisher = {{Elsevier}},
series = {{European Journal of Heart Failure}},
title = {{Comparative diagnostic performance of machine learning models and traditional scores for HFpEF in older adults}},
url = {{http://dx.doi.org/10.1093/ejhf/xuag039}},
doi = {{10.1093/ejhf/xuag039}},
year = {{2026}},
}
