A comparison of three methods to determine the subject matter in textual data

George A. Barnett; Christopher Calabrese; Jeanette B. Ruiz

doi:10.3389/frma.2023.1104691

Frontiers in Research Metrics and Analytics (Jun 2023)

A comparison of three methods to determine the subject matter in textual data

George A. Barnett,
Christopher Calabrese,
Jeanette B. Ruiz

Affiliations

George A. Barnett: Department of Communication, University of California, Davis, Davis, CA, United States
Christopher Calabrese: Department of Communication, Clemson University, Clemson, SC, United States
Jeanette B. Ruiz: Department of Communication, University of California, Davis, Davis, CA, United States

DOI: https://doi.org/10.3389/frma.2023.1104691
Journal volume & issue: Vol. 8

Abstract

Read online

This study compares three different methods commonly employed for the determination and interpretation of the subject matter of large corpuses of textual data. The methods reviewed are: (1) topic modeling, (2) community or group detection, and (3) cluster analysis of semantic networks. Two different datasets related to health topics were gathered from Twitter posts to compare the methods. The first dataset includes 16,138 original tweets concerning HIV pre-exposure prophylaxis (PrEP) from April 3, 2019 to April 3, 2020. The second dataset is comprised of 12,613 tweets about childhood vaccination from July 1, 2018 to October 15, 2018. Our findings suggest that the separate “topics” suggested by semantic networks (community detection) and/or cluster analysis (Ward's method) are more clearly identified than the topic modeling results. Topic modeling produced more subjects, but these tended to overlap. This study offers a better understanding of how results may vary based on method to determine subject matter chosen.

Published in Frontiers in Research Metrics and Analytics

ISSN: 2504-0537 (Online)
Publisher: Frontiers Media S.A.
Country of publisher: Switzerland
LCC subjects: Bibliography. Library science. Information resources
Website: http://journal.frontiersin.org/journal/research-metrics-and-analytics

About the journal

Abstract

Keywords