<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/CINECAstyle.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-07-22T23:31:39Z</responseDate><request verb="GetRecord" identifier="oai:iris.cnr.it:20.500.14243/590143" metadataPrefix="oai_dc">https://iris.cnr.it/oai/request</request><GetRecord><record><header><identifier>oai:iris.cnr.it:20.500.14243/590143</identifier><datestamp>2026-07-09T00:17:38Z</datestamp><setSpec>com_20.500.14243_46</setSpec><setSpec>com_20.500.14243_21</setSpec><setSpec>col_20.500.14243_47</setSpec><setSpec>ou_ou239</setSpec></header><metadata><oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:doc="http://www.lyncode.com/xoai" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
<dc:title>Linguistic Profiling of Transformer Embedding Geometry</dc:title>
<dc:creator>Lucia Domenichelli</dc:creator>
<dc:creator>Dominique Brunato</dc:creator>
<dc:creator>Felice Dell'Orletta</dc:creator>
<dc:contributor>Domenichelli, Lucia</dc:contributor>
<dc:contributor> Brunato, Dominique</dc:contributor>
<dc:contributor> Dell'Orletta, Felice</dc:contributor>
<dc:subject>embedding space, neural language models, linguistic profiling, isotropy, non linear intrinsic dimensionality</dc:subject>
<dc:description>Transformer language models embed tokens in high-dimensional spaces, but whether geometry reflects linguistic structure remains unclear. We analyse token representations in BERT and GPT-2, selected as canonical encoder-only and decoder-only Transformer architectures, through a linguistically-grounded geometric lens. We partition tokens from the Universal Dependencies English Web treebank by surface and syntactic features (position, length, POS, head distance and arity) and examine how their representational geometry evolves across layers. We employ complementary diagnostic metrics, including isotropy, linear and nonlinear intrinsic dimensionality, to capture distinct aspects of embedding structure. Our findings reveal that BERT maintains more isotropic and higher-dimensional subspaces, whereas GPT-2 exhibits stronger anisotropy driven by a compact cluster of sentence-initial tokens. Across models, open-class words, longer tokens, and predicates with several dependents occupy more isotropic, higher-dimensional manifolds than short function words and pre-head modifiers, indicating that semantic richness and syntactic centrality play a key role in structuring embedding space. Our analysis provides a reusable framework for profiling how linguistic abstractions organize the geometry of Transformer embeddings.</dc:description>
<dc:date>2026</dc:date>
<dc:type>info:eu-repo/semantics/conferenceObject</dc:type>
<dc:identifier>https://hdl.handle.net/20.500.14243/590143</dc:identifier>
<dc:identifier>10.18653/v1/2026.conll-main.0</dc:identifier>
<dc:language>eng</dc:language>
<dc:relation>ispartofbook:Proceedings of the 30th Conference on Computational Natural Language Learning</dc:relation>
<dc:relation>30th Conference on Computational Natural Language Learning (CoNLL 2026)</dc:relation>
<dc:relation>firstpage:145</dc:relation>
<dc:relation>lastpage:164</dc:relation>
<dc:relation>numberofpages:20</dc:relation>
<dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
<dc:publisher>Association for Computational Linguistics</dc:publisher>
<dc:rights>license:Creative commons</dc:rights>
<dc:rights>license uri:http://creativecommons.org/licenses/by-nc-nd/4.0/</dc:rights>
</oai_dc:dc></metadata></record></GetRecord></OAI-PMH>