Longitudinal analysis is crucial in different areas of medicine. Finding similarities between patients’ data over time (longitudinal data) allows the definition of subgroups of similar patients, unveiling new aspects of disease progression and fostering new therapies specifically tailored for each group. Longitudinal clustering has become an interesting tool for grouping subjects and analyzing data trajectories over time. These trajectories are often represented as time series. In the healthcare domain, the analysis of longitudinal data poses some challenges due to their complex representation and inherent nature, such as high dimensionality and variability. To address these issues, we propose a LongShort-Term Memory (LSTM) based autoencoder to enhance the k-means longitudinal clustering by encoding high-dimensional and high-variable longitudinal data into a compressed representation. The proposed approach was evaluated on synthetic datasets, showing an improvement in performance compared to using only longitudinal k-means clustering. Moreover, it was applied to a real longitudinal dataset of Alzheimer’s Disease (AD) patients to also assess the efficacy of the proposed approach in real longitudinal health data.

Enhancing k-Means Longitudinal Clustering with an LSTM-Based Autoencoder for High-Dimensional and High-Variable Data

Patrizia Ribino
;
Maria Mannone;Claudia Di Napoli;Giovanni Paragliola;
2025

Abstract

Longitudinal analysis is crucial in different areas of medicine. Finding similarities between patients’ data over time (longitudinal data) allows the definition of subgroups of similar patients, unveiling new aspects of disease progression and fostering new therapies specifically tailored for each group. Longitudinal clustering has become an interesting tool for grouping subjects and analyzing data trajectories over time. These trajectories are often represented as time series. In the healthcare domain, the analysis of longitudinal data poses some challenges due to their complex representation and inherent nature, such as high dimensionality and variability. To address these issues, we propose a LongShort-Term Memory (LSTM) based autoencoder to enhance the k-means longitudinal clustering by encoding high-dimensional and high-variable longitudinal data into a compressed representation. The proposed approach was evaluated on synthetic datasets, showing an improvement in performance compared to using only longitudinal k-means clustering. Moreover, it was applied to a real longitudinal dataset of Alzheimer’s Disease (AD) patients to also assess the efficacy of the proposed approach in real longitudinal health data.
2025
Istituto di Calcolo e Reti ad Alte Prestazioni - ICAR
978-3-032-21375-4
Longitudinal clustering
k-means
LSTM
Autoencoder
File in questo prodotto:
File Dimensione Formato  
KES_k-means.pdf

solo utenti autorizzati

Tipologia: Documento in Post-print
Licenza: Altro tipo di licenza
Dimensione 3.36 MB
Formato Adobe PDF
3.36 MB Adobe PDF   Visualizza/Apri   Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14243/598509
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact