Xpectrass : an evaluation-driven preprocessing and interpretable machine learning platform for large-scale FTIR polymer classification
(2026) In Digital Discovery- Abstract
Fourier transform infrared (FTIR) spectroscopy is widely used for polymer identification in microplastics research. However, raw spectra are often affected by noise, baseline drift, and other acquisition-related artifacts that are not chemically meaningful. Since preprocessing steps can strongly influence multivariate analysis and machine learning (ML) classification, we developed Xpectrass, an open-source framework designed for systematic evaluation of FTIR preprocessing, exploratory analysis, and interpretable ML. A dataset covering multiple polymer classes was used to compare different denoising, baseline correction, and normalization strategies. A preprocessing workflow based on wavelet denoising, adaptive smoothness parameter... (More)
Fourier transform infrared (FTIR) spectroscopy is widely used for polymer identification in microplastics research. However, raw spectra are often affected by noise, baseline drift, and other acquisition-related artifacts that are not chemically meaningful. Since preprocessing steps can strongly influence multivariate analysis and machine learning (ML) classification, we developed Xpectrass, an open-source framework designed for systematic evaluation of FTIR preprocessing, exploratory analysis, and interpretable ML. A dataset covering multiple polymer classes was used to compare different denoising, baseline correction, and normalization strategies. A preprocessing workflow based on wavelet denoising, adaptive smoothness parameter penalized least squares (asPLS) baseline correction, and either standard normal variate (SNV) or spectral moments normalization provided a harmonized feature space for downstream analyses. Starting from a harmonized collection of 12 189 spectra, we selected a labeled, non-redundant subset of 4018 spectra across eight major polymer classes for the main exploratory and machine-learning analyses. Unsupervised methods such as PCA, t-SNE, and UMAP moderately separated polymer classes with partial overlap among chemically similar polyolefins. We further evaluated 41 ML model configurations across 8 algorithmic families and present detailed results for a representative model based on XGBoost. A reduced leave-one-dataset-out validation for PP versus PE+ showed mixed dataset-level transfer, with high performance for three hold-out datasets but poor performance for the Frond hold-out fold. Model interpretation using SHapley Additive exPlanations (SHAP) associated predictions with wavenumber regions, facilitating chemical interpretability. In summary, Xpectrass provides a structured and scalable workflow that links heterogeneous FTIR spectra to reproducible polymer classification in microplastics research.
(Less)
- author
- Khanam, M. Maksuda
; Younus, Saleena
LU
; Mousafi Alasal, Laila
LU
; Uddin, M. Khabir
and Kazi, Julhash U.
LU
- organization
- publishing date
- 2026
- type
- Contribution to journal
- publication status
- in press
- subject
- in
- Digital Discovery
- article number
- d6dd00137h
- publisher
- Royal Society of Chemistry
- external identifiers
-
- scopus:105044713459
- ISSN
- 2635-098X
- DOI
- 10.1039/d6dd00137h
- language
- English
- LU publication?
- yes
- additional info
- Publisher Copyright: This journal is © The Royal Society of Chemistry, 2026.
- id
- 4fb885b9-99fd-4f76-9e90-e6b48b427b27
- date added to LUP
- 2026-08-01 19:36:24
- date last changed
- 2026-08-03 08:43:18
@article{4fb885b9-99fd-4f76-9e90-e6b48b427b27,
abstract = {{<p>Fourier transform infrared (FTIR) spectroscopy is widely used for polymer identification in microplastics research. However, raw spectra are often affected by noise, baseline drift, and other acquisition-related artifacts that are not chemically meaningful. Since preprocessing steps can strongly influence multivariate analysis and machine learning (ML) classification, we developed Xpectrass, an open-source framework designed for systematic evaluation of FTIR preprocessing, exploratory analysis, and interpretable ML. A dataset covering multiple polymer classes was used to compare different denoising, baseline correction, and normalization strategies. A preprocessing workflow based on wavelet denoising, adaptive smoothness parameter penalized least squares (asPLS) baseline correction, and either standard normal variate (SNV) or spectral moments normalization provided a harmonized feature space for downstream analyses. Starting from a harmonized collection of 12 189 spectra, we selected a labeled, non-redundant subset of 4018 spectra across eight major polymer classes for the main exploratory and machine-learning analyses. Unsupervised methods such as PCA, t-SNE, and UMAP moderately separated polymer classes with partial overlap among chemically similar polyolefins. We further evaluated 41 ML model configurations across 8 algorithmic families and present detailed results for a representative model based on XGBoost. A reduced leave-one-dataset-out validation for PP versus PE+ showed mixed dataset-level transfer, with high performance for three hold-out datasets but poor performance for the Frond hold-out fold. Model interpretation using SHapley Additive exPlanations (SHAP) associated predictions with wavenumber regions, facilitating chemical interpretability. In summary, Xpectrass provides a structured and scalable workflow that links heterogeneous FTIR spectra to reproducible polymer classification in microplastics research.</p>}},
author = {{Khanam, M. Maksuda and Younus, Saleena and Mousafi Alasal, Laila and Uddin, M. Khabir and Kazi, Julhash U.}},
issn = {{2635-098X}},
language = {{eng}},
publisher = {{Royal Society of Chemistry}},
series = {{Digital Discovery}},
title = {{Xpectrass : an evaluation-driven preprocessing and interpretable machine learning platform for large-scale FTIR polymer classification}},
url = {{http://dx.doi.org/10.1039/d6dd00137h}},
doi = {{10.1039/d6dd00137h}},
year = {{2026}},
}