A guide on extracting and tidying tweets with R

Julia Bahia Adams; Carlos Augusto Jardim Chiarelli

doi:10.25189/2675-4916.2021.v2.n4.id410

Cadernos de Linguística (Dec 2021)

A guide on extracting and tidying tweets with R

Julia Bahia Adams,
Carlos Augusto Jardim Chiarelli

Affiliations

Julia Bahia Adams: Universidade Estadual de Campinas (UNICAMP)
Carlos Augusto Jardim Chiarelli: Universidade Estadual de Campinas (NICAMP)

DOI: https://doi.org/10.25189/2675-4916.2021.v2.n4.id410
Journal volume & issue: Vol. 2, no. 4

Abstract

Read online

Social media platforms represent a deep resource for academic research and a wide range of untapped possibilities for linguists (D'ARCY; YOUNG, 2012). This rapidly developing field presents various ethical issues and unique challenges regarding methods to retrieve and analyze data. This tutorial provides a straightforward guide to harvesting and tidying Twitter data, focused mainly on the Tweets' text, by using the R programming language (R CORE TEAM, 2020) via Twitter's APIs. The R code was developed in Adams (2020), based on the rtweet package (KEARNEY, 2018), and successfully resulted in a script for corpora compilation. In this tutorial, we discuss limitations, problems, and solutions in our framework for conducting ethical research on this social networking site. Our ethical concerns go beyond what we "agree to" in terms of use and privacy policies, that is, we argue that their content does not contemplate all the concerns researchers need to attend to. Additionally, our aim is to show that using Twitter as a data source does not require advanced computational skills.

Published in Cadernos de Linguística

ISSN: 2675-4916 (Online)
Publisher: Associação Brasileira de Linguística
Country of publisher: Brazil
LCC subjects: Language and Literature: Philology. Linguistics
Website: https://cadernos.abralin.org

About the journal

Abstract

Keywords