Improved Machine Reading Comprehension Using Data Validation for Weakly Labeled Data

Yunyeong Yang; Sangwoo Kang; Jungyun Seo

doi:10.1109/access.2019.2963569

IEEE Access (Jan 2020)

Improved Machine Reading Comprehension Using Data Validation for Weakly Labeled Data

Yunyeong Yang,
Sangwoo Kang,
Jungyun Seo

Affiliations

Yunyeong Yang: ORCiD; Department of Computer Science and Engineering, Sogang University, Seoul, South Korea
Sangwoo Kang: ORCiD; Department of Software, Gachon University, Seongnam, South Korea
Jungyun Seo: ORCiD; Department of Computer Science and Engineering, Sogang University, Seoul, South Korea

DOI: https://doi.org/10.1109/access.2019.2963569
Journal volume & issue: Vol. 8
pp. 5667 – 5677

Abstract

Read online

Machine reading comprehension (MRC) is a natural language processing task wherein a given question is answered according to a holistic understanding of a given context. Recently, many researchers have shown interest in MRC, for which a considerable number of datasets are being released. Datasets for MRC, which are composed of the context-query-answer triple, are designed to answer a given query by referencing and understanding a readily-available, relevant context text. The TriviaQA dataset is a weakly labeled dataset, because it contains irrelevant context that forms no basis for answering the query. The existing syntactic data cleaning method struggles to deal with the contextual noise this irrelevancy creates. Therefore, a semantic data cleaning method using reasoning processes is necessary. To address this, we propose a new MRC model in which the TriviaQA dataset is validated and trained using a high-quality dataset. The data validation method in our MRC model improves the quality of the training dataset, and the answer extraction model learns with the validated training data, because of our validation method. Our proposed method showed a 4.33% improvement in performance for the TriviaQA Wiki, compared to the existing baseline model. Accordingly, our proposed method can address the limitation of irrelevant context in MRC better than the human supervision.

Published in IEEE Access

ISSN: 2169-3536 (Online)
Publisher: IEEE
Country of publisher: United States
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering
Website: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=6287639

About the journal

Abstract

Keywords