We present ArchiClip, a method for learning multi-modal joint representations of text and freeform architectural surfaces. Building on the widely adopted multi-modal contrastive learning paradigm, our approach focuses specifically on distilling knowledge from the architectural domain and enabling robust understanding of freeform 3D geometry. To support this, we introduce a strategy for constructing ArchiShape, a curated 3D dataset enriched with domain-specific, rich textual descriptions. ArchiShape is generated automatically by combining procedural modeling and geometric shape descriptors with pre-trained large language models, eliminating the need for manual annotation. We then design a bi-modal pre-training framework that learns aligned embeddings for both textual descriptions and architectural freeform surfaces. The framework incorporates a 3D backbone network tailored to architectural geometry and a custom batch-sampling scheme to ensure efficient training. We evaluate ArchiClip on cross-modal 3D retrieval of architectural freeforms, demonstrating its ability to encode rich geometric and domain-specific concepts (e.g., curvature, spatial organization), and highlighting the benefits of domain-aware multi-modal representation learning.

ArchiClip: learning joint text–geometry representations for 3D architectural freeform surfaces

Favilli Andrea;Laccone Francesco;Malomo Luigi;Messina Nicola;Carrara Fabio;Cignoni Paolo;Giorgi Daniela
2026

Abstract

We present ArchiClip, a method for learning multi-modal joint representations of text and freeform architectural surfaces. Building on the widely adopted multi-modal contrastive learning paradigm, our approach focuses specifically on distilling knowledge from the architectural domain and enabling robust understanding of freeform 3D geometry. To support this, we introduce a strategy for constructing ArchiShape, a curated 3D dataset enriched with domain-specific, rich textual descriptions. ArchiShape is generated automatically by combining procedural modeling and geometric shape descriptors with pre-trained large language models, eliminating the need for manual annotation. We then design a bi-modal pre-training framework that learns aligned embeddings for both textual descriptions and architectural freeform surfaces. The framework incorporates a 3D backbone network tailored to architectural geometry and a custom batch-sampling scheme to ensure efficient training. We evaluate ArchiClip on cross-modal 3D retrieval of architectural freeforms, demonstrating its ability to encode rich geometric and domain-specific concepts (e.g., curvature, spatial organization), and highlighting the benefits of domain-aware multi-modal representation learning.
2026
Istituto di Scienza e Tecnologie dell'Informazione "Alessandro Faedo" - ISTI
Freeform architectural geometry, AI-ready annotated architectural dataset, 3D captioning for architectural design, Multimodal representation learning, 3D contrastive learning, Cross-modal 3D architectural shape retrieval
File in questo prodotto:
File Dimensione Formato  
1-s2.0-S0010448526001260-main.pdf

accesso aperto

Descrizione: ArchiClip: Learning joint text–geometry representations for 3D architectural freeform surfaces
Tipologia: Versione Editoriale (PDF)
Licenza: Creative commons
Dimensione 8.86 MB
Formato Adobe PDF
8.86 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14243/597663
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact