Longitudinal analysis is crucial in different areas of medicine. Finding similarities between patients’ data over time (longitudinal data) allows the definition of subgroups of similar patients, unveiling new aspects of disease progression and fostering new therapies specifically tailored for each group. Longitudinal clustering has become an interesting tool for grouping subjects and analyzing data trajectories over time. These trajectories are often represented as time series. In the healthcare domain, the analysis of longitudinal data poses some challenges due to their complex representation and inherent nature, such as high dimensionality and variability. To address these issues, we propose a LongShort-Term Memory (LSTM) based autoencoder to enhance the k-means longitudinal clustering by encoding high-dimensional and high-variable longitudinal data into a compressed representation. The proposed approach was evaluated on synthetic datasets, showing an improvement in performance compared to using only longitudinal k-means clustering. Moreover, it was applied to a real longitudinal dataset of Alzheimer’s Disease (AD) patients to also assess the efficacy of the proposed approach in real longitudinal health data.
Enhancing k-Means Longitudinal Clustering with an LSTM-Based Autoencoder for High-Dimensional and High-Variable Data
Patrizia Ribino
;Maria Mannone;Claudia Di Napoli;Giovanni Paragliola;
2025
Abstract
Longitudinal analysis is crucial in different areas of medicine. Finding similarities between patients’ data over time (longitudinal data) allows the definition of subgroups of similar patients, unveiling new aspects of disease progression and fostering new therapies specifically tailored for each group. Longitudinal clustering has become an interesting tool for grouping subjects and analyzing data trajectories over time. These trajectories are often represented as time series. In the healthcare domain, the analysis of longitudinal data poses some challenges due to their complex representation and inherent nature, such as high dimensionality and variability. To address these issues, we propose a LongShort-Term Memory (LSTM) based autoencoder to enhance the k-means longitudinal clustering by encoding high-dimensional and high-variable longitudinal data into a compressed representation. The proposed approach was evaluated on synthetic datasets, showing an improvement in performance compared to using only longitudinal k-means clustering. Moreover, it was applied to a real longitudinal dataset of Alzheimer’s Disease (AD) patients to also assess the efficacy of the proposed approach in real longitudinal health data.| File | Dimensione | Formato | |
|---|---|---|---|
|
KES_k-means.pdf
solo utenti autorizzati
Tipologia:
Documento in Post-print
Licenza:
Altro tipo di licenza
Dimensione
3.36 MB
Formato
Adobe PDF
|
3.36 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


