Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection

Weijia Xu; Sweta Agrawal; Eleftheria Briakou; Marianna J. Martindale; Marine Carpuat

doi:10.1162/tacl_a_00563

Transactions of the Association for Computational Linguistics (Jan 2023)

Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection

Weijia Xu,
Sweta Agrawal,
Eleftheria Briakou,
Marianna J. Martindale,
Marine Carpuat

Affiliations

Weijia Xu: Microsoft Research, Redmond, USA. [email protected]
Sweta Agrawal: University of Maryland, USA. [email protected]
Eleftheria Briakou: University of Maryland, USA. [email protected]
Marianna J. Martindale: University of Maryland, USA. [email protected]
Marine Carpuat: University of Maryland, USA. [email protected]

DOI: https://doi.org/10.1162/tacl_a_00563
Journal volume & issue: Vol. 11
pp. 546 – 564

Abstract

Read online

AbstractNeural sequence generation models are known to “hallucinate”, by producing outputs that are unrelated to the source text. These hallucinations are potentially harmful, yet it remains unclear in what conditions they arise and how to mitigate their impact. In this work, we first identify internal model symptoms of hallucinations by analyzing the relative token contributions to the generation in contrastive hallucinated vs. non-hallucinated outputs generated via source perturbations. We then show that these symptoms are reliable indicators of natural hallucinations, by using them to design a lightweight hallucination detector which outperforms both model-free baselines and strong classifiers based on quality estimation or large pre-trained models on manually annotated English-Chinese and German-English translation test beds.

Published in Transactions of the Association for Computational Linguistics

ISSN: 2307-387X (Online)
Publisher: The MIT Press
Country of publisher: United States
LCC subjects: Language and Literature: Philology. Linguistics: Computational linguistics. Natural language processing
Website: https://direct.mit.edu/tacl

About the journal