We present ArchiClip, a method for learning multi-modal joint representations of text and freeform architectural surfaces. Building on the widely adopted multi-modal contrastive learning paradigm, our approach focuses specifically on distilling knowledge from the architectural domain and enabling robust understanding of freeform 3D geometry. To support this, we introduce a strategy for constructing ArchiShape, a curated 3D dataset enriched with domain-specific, rich textual descriptions. ArchiShape is generated automatically by combining procedural modeling and geometric shape descriptors with pre-trained large language models, eliminating the need for manual annotation. We then design a bi-modal pre-training framework that learns aligned embeddings for both textual descriptions and architectural freeform surfaces. The framework incorporates a 3D backbone network tailored to architectural geometry and a custom batch-sampling scheme to ensure efficient training. We evaluate ArchiClip on cross-modal 3D retrieval of architectural freeforms, demonstrating its ability to encode rich geometric and domain-specific concepts (e.g., curvature, spatial organization), and highlighting the benefits of domain-aware multi-modal representation learning.
ArchiClip: learning joint text–geometry representations for 3D architectural freeform surfaces
Favilli Andrea;Laccone Francesco;Malomo Luigi;Messina Nicola;Carrara Fabio;Cignoni Paolo;Giorgi Daniela
2026
Abstract
We present ArchiClip, a method for learning multi-modal joint representations of text and freeform architectural surfaces. Building on the widely adopted multi-modal contrastive learning paradigm, our approach focuses specifically on distilling knowledge from the architectural domain and enabling robust understanding of freeform 3D geometry. To support this, we introduce a strategy for constructing ArchiShape, a curated 3D dataset enriched with domain-specific, rich textual descriptions. ArchiShape is generated automatically by combining procedural modeling and geometric shape descriptors with pre-trained large language models, eliminating the need for manual annotation. We then design a bi-modal pre-training framework that learns aligned embeddings for both textual descriptions and architectural freeform surfaces. The framework incorporates a 3D backbone network tailored to architectural geometry and a custom batch-sampling scheme to ensure efficient training. We evaluate ArchiClip on cross-modal 3D retrieval of architectural freeforms, demonstrating its ability to encode rich geometric and domain-specific concepts (e.g., curvature, spatial organization), and highlighting the benefits of domain-aware multi-modal representation learning.| File | Dimensione | Formato | |
|---|---|---|---|
|
1-s2.0-S0010448526001260-main.pdf
accesso aperto
Descrizione: ArchiClip: Learning joint text–geometry representations for 3D architectural freeform surfaces
Tipologia:
Versione Editoriale (PDF)
Licenza:
Creative commons
Dimensione
8.86 MB
Formato
Adobe PDF
|
8.86 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


