Supervised versus unsupervised approaches to classification of accelerometry data

Maitreyi Sur; Jonathan C. Hall; Joseph Brandt; Molly Astell; Sharon A. Poessel; Todd E. Katzner

doi:10.1002/ece3.10035

Ecology and Evolution (May 2023)

Supervised versus unsupervised approaches to classification of accelerometry data

Maitreyi Sur,
Jonathan C. Hall,
Joseph Brandt,
Molly Astell,
Sharon A. Poessel,
Todd E. Katzner

Affiliations

Maitreyi Sur: Conservation Science Global, Inc. West Cape May New Jersey USA
Jonathan C. Hall: Department of Biology Eastern Michigan University Ypsilanti Michigan USA
Joseph Brandt: U.S. Fish and Wildlife Service, Hopper Mountain National Wildlife Refuge Complex Ventura California USA
Molly Astell: U.S. Fish and Wildlife Service, Hopper Mountain National Wildlife Refuge Complex Ventura California USA
Sharon A. Poessel: U.S. Geological Survey, Forest and Rangeland Ecosystem Science Center Boise Idaho USA
Todd E. Katzner: U.S. Geological Survey, Forest and Rangeland Ecosystem Science Center Boise Idaho USA

DOI: https://doi.org/10.1002/ece3.10035
Journal volume & issue: Vol. 13, no. 5
pp. n/a – n/a

Abstract

Read online

Abstract Sophisticated animal‐borne sensor systems are increasingly providing novel insight into how animals behave and move. Despite their widespread use in ecology, the diversity and expanding quality and quantity of data they produce have created a need for robust analytical methods for biological interpretation. Machine learning tools are often used to meet this need. However, their relative effectiveness is not well known and, in the case of unsupervised tools, given that they do not use validation data, their accuracy can be difficult to assess. We evaluated the effectiveness of supervised (n = 6), semi‐supervised (n = 1), and unsupervised (n = 2) approaches to analyzing accelerometry data collected from critically endangered California condors (Gymnogyps californianus). Unsupervised K‐means and EM (expectation–maximization) clustering approaches performed poorly, with adequate classification accuracies of 0.81. Kappa statistics were also highest for RF and kNN, in most cases substantially greater than for other modeling approaches. Unsupervised modeling, which is commonly used for the classification of a priori‐defined behaviors in telemetry data, can provide useful information but likely is instead better suited to post hoc definition of generalized behavioral states. This work also shows the potential for substantial variation in classification accuracy among different machine learning approaches and among different metrics of accuracy. As such, when analyzing biotelemetry data, best practices appear to call for the evaluation of several machine learning techniques and several measures of accuracy for each dataset under consideration.

Published in Ecology and Evolution

ISSN: 2045-7758 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Science: Biology (General): Ecology
Website: https://onlinelibrary.wiley.com/journal/20457758

About the journal

Abstract

Keywords