From categories to gradience: Auto-coding sociophonetic variation with random forests

Dan Villarreal; Jennifer Hay; Kevin Watson; Lynn Clark

doi:10.5334/labphon.216

Laboratory Phonology (Jun 2020)

From categories to gradience: Auto-coding sociophonetic variation with random forests

Dan Villarreal,
Jennifer Hay,
Kevin Watson,
Lynn Clark

Affiliations

Dan Villarreal: ORCiD; Department of Linguistics, University of Pittsburgh, Pittsburgh, PA
Jennifer Hay: University of Canterbury
Kevin Watson
Lynn Clark: ORCiD; New Zealand Institute of Language, Brain and Behaviour, University of Canterbury, Christchurch; Department of Linguistics, University of Canterbury, Christchurch

DOI: https://doi.org/10.5334/labphon.216
Journal volume & issue: Vol. 11, no. 1

Abstract

Read online Read online

The time-consuming nature of coding sociophonetic variables that are typically treated as categorical represents an impediment to addressing research questions around these variables that require large volumes of data. In this paper, we apply a machine learning method, random forest classification (Breiman, 2001), to automate coding (categorical prediction) of two English sociophonetic variables traditionally treated as categorical, non-prevocalic /r/ and word-medial intervocalic /t/, based on tokens’ acoustic signatures. We found good performance for binary classifiers of non-prevocalic /r/ (Absent versus Present) and medial /t/ (Voiced versus Voiceless), but not for medial /t/ with a six-way coding distinction (largely due to some codes being sparsely represented in the training data). This method also yields rankings of acoustic measures in terms of importance in classification. Beyond any individual measures, this method generates probabilistic predictions of variation (classifier probabilities) that represent a composite of the acoustic cues fed into the model. In a listening experiment, we found that not only did classifier probabilities significantly capture gradience in trained listeners’ perceptions of rhoticity, they better predicted listeners’ perceptions than individual acoustic measures. This method thus represents a new approach to reconciling the categorical and continuous dimensions of sociophonetic variation.

Published in Laboratory Phonology

ISSN: 1868-6354 (Online)
Publisher: Open Library of Humanities
Country of publisher: United Kingdom
LCC subjects: Language and Literature: Philology. Linguistics: Language. Linguistic theory. Comparative grammar
Website: https://www.journal-labphon.org

About the journal

Abstract

Keywords