Glossa (Nov 2020)
Using social-media data to investigate morphosyntactic variation and dialect syntax in a lesser-used language: Two case studies from Welsh
Abstract
Data gathered from social media have been used extensively to examine lexical dialect variation in widely used languages such as English and Spanish, but their use to date in morphosyntax and for lesser-used languages has been more limited. This paper tests the usefulness of using data derived from Twitter to address traditional questions in dialect syntax and sociolinguistics. It uses two cases studies from Welsh – the form of the second-person singular pronoun in various syntactic contexts, and the availability of auxiliary deletion – to assess whether datasets based on Twitter data can successfully replicate and enhance results derived by traditional means. The results of the case studies coincide to a large extent with distributions established in existing studies, even ones using entirely different methods, such as dialect questionnaires or acceptability judgment tests. Twitter data also show considerable success in establishing implicational hierarchies and conditioning factors comparable to those typical of the field. Where the results differ from existing studies, the differences may be due to the younger demographics of Twitter users, or to differences in the quantity of data provided by different methodologies. The results produce patterns closer to spoken data than to written data, giving us reasonable confidence in such data as a relatively good proxy for spoken usage of large numbers of language users.
Keywords