The AMR-PT corpus and the semantic annotation of challenging sentences from journalistic and opinion texts

Marcio Lima Inácio; Marco Antonio Sobrevilla Cabezudo; Renata Ramisch; Ariani Di Felippo; Thiago Alexandre Salgueiro Pardo

doi:10.1590/1678-460x202339355159

DELTA: Documentação de Estudos em Lingüística Teórica e Aplicada (Sep 2023)

The AMR-PT corpus and the semantic annotation of challenging sentences from journalistic and opinion texts

Marcio Lima Inácio,
Marco Antonio Sobrevilla Cabezudo,
Renata Ramisch,
Ariani Di Felippo,
Thiago Alexandre Salgueiro Pardo

Affiliations

Marcio Lima Inácio: ORCiD
Marco Antonio Sobrevilla Cabezudo: ORCiD
Renata Ramisch: ORCiD
Ariani Di Felippo: ORCiD
Thiago Alexandre Salgueiro Pardo: ORCiD

DOI: https://doi.org/10.1590/1678-460x202339355159
Journal volume & issue: Vol. 39, no. 3

Abstract

Read online Read online

ABSTRACT One of the most popular semantic representation languages in Natural Language Processing (NLP) is Abstract Meaning Representation (AMR). This formalism encodes the meaning of single sentences in directed rooted graphs. For English, there is a large annotated corpus that provides qualitative and reusable data for building or improving existing NLP methods and applications. For building AMR corpora for non-English languages, including Brazilian Portuguese, automatic and manual strategies have been conducted. The automatic annotation methods are essentially based on the cross-linguistic alignment of parallel corpora and the inheritance of the AMR annotation. The manual strategies focus on adapting the AMR English guidelines to a target language. Both annotation strategies have to deal with some phenomena that are challenging. This paper explores in detail some characteristics of Portuguese for which the AMR model had to be adapted and introduces two annotated corpora: AMRNews, a corpus of 870 annotated sentences from journalistic texts, and OpiSums-PT-AMR, comprising 404 opinionated sentences in AMR.

Published in DELTA: Documentação de Estudos em Lingüística Teórica e Aplicada

ISSN: 0102-4450 (Print); 1678-460X (Online)
Publisher: Pontifícia Universidade Católica de São Paulo
Country of publisher: Brazil
LCC subjects: Language and Literature: Philology. Linguistics
Website: http://www.scielo.br/scielo.php?pid=0102-4450&script=sci_serial

About the journal

Abstract

Keywords