NADA is a synthetic dataset of 1 500 000 (128 × 128 pixels) images of geometric shapes with non-elementary multivariate parameter distributions designed to benchmark and test novel probabilistic deep learning models. Benchmarking uncertainty-aware techniques is critical in real-world scenarios, especially in high-stakes domains such as automated driving or health data analysis. This is because the most robust and reliable AI methods are rooted in Bayesian reasoning and uncertainty analysis, yet there is no consensus on how to test or compare them. In particular, the public, synthetic, and real datasets currently available in the literature are inadequate for benchmarking most probabilistic methodologies: first, due to a lack of control over population variability; and second, due to an oversimplification of the distributional properties of latent variables. From this perspective, NADA is organized into three main From this perspective, NADA is organized into three main repositories, specifically developed to challenge uncertainty-aware methods and provide a unified benchmark reference dataset across three key areas: 1) characterization of complex latent space and evaluation of disentangling ability (NADA_Dis: 300 000 images); 2) identification of different types of aleatoric and epistemic uncertainties (NADA_AlEp: 500 000 images); and 3) detection of out-of-distribution elements for reliable AI (NADA_OOD: 700 000 images). Each repository includes the dataframe describing the image parameters (e.g., rotation, position, shape, color, deformation, and noise), the dataset-generating hyperparameters (e.g., marginal distributions and correlation matrix), and evaluation plots to assess data quality. In addition, the dataset is coupled with an open-source Python synthetic generator, allowing easy modification and adaptation to specific research questions.

Descriptor: Not-A-DAtabase of synthetic shapes benchmarking dataset (NADA-SynShapes)

Del Corso G.
Conceptualization
;
Volpini F.
Software
;
Caudai C.
Conceptualization
;
Moroni D.
Funding Acquisition
;
Colantonio S.
Funding Acquisition
2025

Abstract

NADA is a synthetic dataset of 1 500 000 (128 × 128 pixels) images of geometric shapes with non-elementary multivariate parameter distributions designed to benchmark and test novel probabilistic deep learning models. Benchmarking uncertainty-aware techniques is critical in real-world scenarios, especially in high-stakes domains such as automated driving or health data analysis. This is because the most robust and reliable AI methods are rooted in Bayesian reasoning and uncertainty analysis, yet there is no consensus on how to test or compare them. In particular, the public, synthetic, and real datasets currently available in the literature are inadequate for benchmarking most probabilistic methodologies: first, due to a lack of control over population variability; and second, due to an oversimplification of the distributional properties of latent variables. From this perspective, NADA is organized into three main From this perspective, NADA is organized into three main repositories, specifically developed to challenge uncertainty-aware methods and provide a unified benchmark reference dataset across three key areas: 1) characterization of complex latent space and evaluation of disentangling ability (NADA_Dis: 300 000 images); 2) identification of different types of aleatoric and epistemic uncertainties (NADA_AlEp: 500 000 images); and 3) detection of out-of-distribution elements for reliable AI (NADA_OOD: 700 000 images). Each repository includes the dataframe describing the image parameters (e.g., rotation, position, shape, color, deformation, and noise), the dataset-generating hyperparameters (e.g., marginal distributions and correlation matrix), and evaluation plots to assess data quality. In addition, the dataset is coupled with an open-source Python synthetic generator, allowing easy modification and adaptation to specific research questions.
2025
Istituto di Scienza e Tecnologie dell'Informazione "Alessandro Faedo" - ISTI
Synthetic data generator
Uncertainty quantification benchmarking
Feature disentanglement
Out-of-distribution identification
File in questo prodotto:
File Dimensione Formato  
Descriptor_Not-A-DAtabase_of_Synthetic_Shapes_Benchmarking_Dataset_NADA-SynShapes.pdf

accesso aperto

Descrizione: Descriptor: Not-A-DAtabase of Synthetic Shapes Benchmarking Dataset (NADA-SynShapes)
Tipologia: Versione Editoriale (PDF)
Licenza: Creative commons
Dimensione 3.58 MB
Formato Adobe PDF
3.58 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14243/593182
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact