From Street Photos to Fashion Trends: Leveraging User-Provided Noisy Labels for Fashion Understanding

Fu-Hsien Huang; Hsin-Min Lu; Yao-Wen Hsu

doi:10.1109/ACCESS.2021.3069245

IEEE Access (Jan 2021)

From Street Photos to Fashion Trends: Leveraging User-Provided Noisy Labels for Fashion Understanding

Fu-Hsien Huang,
Hsin-Min Lu,
Yao-Wen Hsu

Affiliations

Fu-Hsien Huang: Master Program in Statistics, National Taiwan University, Taipei, Taiwan
Hsin-Min Lu: ORCiD; Department of Information Management, National Taiwan University, Taipei, Taiwan
Yao-Wen Hsu: Department of International Business, National Taiwan University, Taipei, Taiwan

DOI: https://doi.org/10.1109/ACCESS.2021.3069245
Journal volume & issue: Vol. 9
pp. 49189 – 49205

Abstract

Read online

There is increased interest in using street photos to understand fashion trends. Though street photos usually contain rich clothing information, there are several technical challenges to their analysis. First, street photos collected from social media sites often contain user-provided noisy labels, and training models using these labels may deteriorate prediction performance. Second, most existing methods predict multiple clothing attributes individually and do not consider the potential to share knowledge between related tasks. In addition to these technical challenges, most fashion image datasets created by previous studies focus on American and European fashion styles. To address these technical challenges and understand fashion trends in Asia, we created RichWear, a new street fashion dataset containing 322,198 images with various text labels for fashion analysis. This dataset, collected from an Asian social network site, focuses on street styles in Japan and other Asian areas. RichWear provides a subset of expert-verified labels in addition to user-provided noisy labels for model training and evaluation. We propose the Fashion Attributes Recognition Network (FARNet) based on the multi-task learning framework to improve fashion recognition. Instead of predicting each clothing attribute individually, FARNet predicts three types of attributes simultaneously, and, once trained, this network leverages the noisy labels and generates corrected labels based on the input images. Experimental results show that this approach significantly outperforms existing methods. Applying the trained model to the RichWear dataset, we report Asian fashion trends and street styles based on predicted labels and image clusters from latent feature vectors.

Published in IEEE Access

ISSN: 2169-3536 (Online)
Publisher: IEEE
Country of publisher: United States
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering
Website: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=6287639

About the journal

Abstract

Keywords