CNR Institutional Research Information System

Obtaining high-quality labelled data for training a classifier in a new application domain is often costly. Transfer Learning(a.k.a. "Inductive Transfer") tries to alleviate these costs by transferring, to the "target"domain of interest, knowledge available from a different "source"domain. In transfer learning the lack of labelled information from the target domain is compensated by the availability at training time of a set of unlabelled examples from the target distribution. Transductive Transfer Learning denotes the transfer learning setting in which the only set of target documents that we are interested in classifying is known and available at training time. Although this definition is indeed in line with Vapnik's original definition of "transduction", current terminology in the field is confused. In this article, we discuss how the term "transduction"has been misused in the transfer learning literature, and propose a clarification consistent with the original characterization of this term given by Vapnik. We go on to observe that the above terminology misuse has brought about misleading experimental comparisons, with inductive transfer learning methods that have been incorrectly compared with transductive transfer learning methods. We then, give empirical evidence that the difference in performance between the inductive version and the transductive version of a transfer learning method can indeed be statistically significant (i.e., that knowing at training time the only data one needs to classify indeed gives an advantage). Our clarification allows a reassessment of the field, and of the relative merits of the major, state-of-The-Art algorithms for transfer learning in text classification.

Lost in transduction: transductive transfer learning in text classification

Moreo Fernandez A.;Esuli A.;Sebastiani F.

2021

Abstract

Obtaining high-quality labelled data for training a classifier in a new application domain is often costly. Transfer Learning(a.k.a. "Inductive Transfer") tries to alleviate these costs by transferring, to the "target"domain of interest, knowledge available from a different "source"domain. In transfer learning the lack of labelled information from the target domain is compensated by the availability at training time of a set of unlabelled examples from the target distribution. Transductive Transfer Learning denotes the transfer learning setting in which the only set of target documents that we are interested in classifying is known and available at training time. Although this definition is indeed in line with Vapnik's original definition of "transduction", current terminology in the field is confused. In this article, we discuss how the term "transduction"has been misused in the transfer learning literature, and propose a clarification consistent with the original characterization of this term given by Vapnik. We go on to observe that the above terminology misuse has brought about misleading experimental comparisons, with inductive transfer learning methods that have been incorrectly compared with transductive transfer learning methods. We then, give empirical evidence that the difference in performance between the inductive version and the transductive version of a transfer learning method can indeed be statistically significant (i.e., that knowing at training time the only data one needs to classify indeed gives an advantage). Our clarification allows a reassessment of the field, and of the relative merits of the major, state-of-The-Art algorithms for transfer learning in text classification.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
				2021
			
	Strutture organizzative
	
				Istituto di Scienza e Tecnologie dell'Informazione "Alessandro Faedo" - ISTI
			
	Parole chiave
	
				Transduction
Induction
Transfer learning
Text classification
Distributional hypothesis
			
	Appare nelle tipologie:
	
				01.01 Articolo in rivista

File in questo prodotto:

File	Dimensione	Formato
prod_456427-doc_176636.pdf accesso aperto Descrizione: Postprint - Lost in transduction: transductive transfer learning in text classification Tipologia: Documento in Post-print Licenza: Nessuna licenza dichiarata (non attribuibile a prodotti successivi al 2023) Dimensione 747.17 kB Formato Adobe PDF Visualizza/Apri	747.17 kB	Adobe PDF	Visualizza/Apri
prod_456427-doc_176654.pdf solo utenti autorizzati Descrizione: Lost in transduction: transductive transfer learning in text classification Tipologia: Versione Editoriale (PDF) Licenza: NON PUBBLICO - Accesso privato/ristretto Dimensione 334.12 kB Formato Adobe PDF Visualizza/Apri Richiedi una copia	334.12 kB	Adobe PDF	Visualizza/Apri Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14243/399577

Citazioni

ND

15

ND

social impact