CustFRE: An annotated dataset for extraction of family relations from English text

Raabia Mumtaz; Muhammad Abdul Qadir; Asif Saeed

Data in Brief (Apr 2022)

CustFRE: An annotated dataset for extraction of family relations from English text

Raabia Mumtaz,
Muhammad Abdul Qadir,
Asif Saeed

Affiliations

Raabia Mumtaz: Capital University of Science & Technology, Islamabad, Pakistan; Corresponding author.
Muhammad Abdul Qadir: Capital University of Science & Technology, Islamabad, Pakistan
Asif Saeed: COMSATS University Islamabad, Attock Campus, Pakistan

Journal volume & issue: Vol. 41
p. 107980

Abstract

Read online

Meaningful Information extraction is an extremely important and challenging task due to the ever growing size of data. Training and evaluating automated systems for the task requires annotated datasets which are rarely available because of the great amount of human effort and time required for annotating data. The dataset described in this manuscript, CustFRE, is meant for systems that learn extracting family relations from text. Sentences having at least two persons have been collected from the internet. The texts are first processed using Stanford's NLP pipeline for basic NLP tagging. Next, a team of natural language processing experts annotated the dataset. All family relations among persons in the texts have been annotated, or a no_relation is annotated if no family relation between two persons can be inferred from the text. After annotation, the dataset was verified by an NLP expert for completeness and correctness. CustFRE contains in total 2,716 annotations. The dataset can be used by information extraction researchers as a benchmark for evaluating their systems, and can also be used for training and evaluating family relation extraction systems.

Published in Data in Brief

ISSN: 2352-3409 (Online)
Publisher: Elsevier
Country of publisher: United States
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Science: Science (General)
Website: http://www.journals.elsevier.com/data-in-brief/

About the journal

Abstract

Keywords