Representation of structured data of the text genre as a technique for automatic text processing

Claudia Aparecida Fonseca; Marcus Vinícius Carvalho  Guelpeli; Rafael Santiago de Souza  Netto

doi:10.35699/1983-3652.2022.35445

Texto Livre: Linguagem e Tecnologia (Jan 2022)

Representation of structured data of the text genre as a technique for automatic text processing

Claudia Aparecida Fonseca,
Marcus Vinícius Carvalho Guelpeli,
Rafael Santiago de Souza Netto

Affiliations

Claudia Aparecida Fonseca: Universidade Federal dos Vales do Jequitinhonha e Mucuri, Departamento de Letras, Diamantina, Minas Gerais, Brasil
Marcus Vinícius Carvalho Guelpeli: Universidade Federal dos Vales do Jequitinhonha e Mucuri, Departamento de sistema de informação, Diamantina, Minas Gerais, Brasil
Rafael Santiago de Souza Netto: Centro Universitário de Barra Mansa, Departamento de Ciência da Computação, Barra Mansa, Rio de Janeiro, Brasil

DOI: https://doi.org/10.35699/1983-3652.2022.35445
Journal volume & issue: Vol. 15

Abstract

Read online

The present article was developed in the field of Natural Language Processing and Language Studies based on a corpus compiled by computational tools. This study is based on the assumption that it is helpful to trace a close relationship between corpus generation/annotation and the assessment of the constitutive elements of the text genre source. It aims to demonstrate, through specific studies of structured data from the text genre ‘scientific article’, alternatives to automatic text processing techniques. In order to reach the intended goal, the authors created a computational model for the compilation of a linguistic, specialized Corpus, representative of the genre Scientific Article - CorpACE. The object of study includes the constitutive elements of scientific articles, marked in XML, extracted and collected from the SciELO-Scientific Electronic Library On-line database. The final product was a database obtained with information extracted and structured in XML format, which designates and identifies the markups of the genre being analyzed and is available for many tools and applications. The results demonstrate how the representation of constitutive elements of the genre can condense available information with hierarchical and dynamic processes built during the compilation. At the end of the study, it is believed that more research will be required for bringing Language Science and Computer Science closer with emphasis on NLP in the attempt to represent and manipulate linguistic knowledge in its many levels – morphological, syntactic, semantic and discursive – in order to improve implementation and manipulation of automatic text processing.

Published in Texto Livre: Linguagem e Tecnologia

ISSN: 1983-3652 (Online)
Publisher: Universidade Federal de Minas Gerais
Country of publisher: Brazil
LCC subjects: Technology; Language and Literature
Website: https://periodicos.ufmg.br/index.php/textolivre/

About the journal

Abstract

Keywords