Y-Rank: A Multi-Feature-Based Keyphrase Extraction Method for Short Text

Qiang Liu; Yan Hui; Shangdong Liu; Yimu Ji

doi:10.3390/app14062510

Applied Sciences (Mar 2024)

Y-Rank: A Multi-Feature-Based Keyphrase Extraction Method for Short Text

Qiang Liu,
Yan Hui,
Shangdong Liu,
Yimu Ji

Affiliations

Qiang Liu: School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing 210023, China
Yan Hui: School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing 210023, China
Shangdong Liu: School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing 210023, China
Yimu Ji: School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing 210023, China

DOI: https://doi.org/10.3390/app14062510
Journal volume & issue: Vol. 14, no. 6
p. 2510

Abstract

Read online

Keyphrase extraction is a critical task in text information retrieval, which traditionally employs both supervised and unsupervised approaches. Supervised methods generally rely on large corpora, which introduce the problems of availability, while unsupervised methods are independent of out-sources but also lead to defects like imperfect statistical features or low accuracy. Particularly in short-text scenarios, limited text features often result in low-quality candidate ranking. To address this issue, this paper proposes Y-Rank, a lightweight unsupervised keyphrase extraction method that extracts the average information content of candidate sentences as the key statistical features from a single document, and follows a graph construction approach based on similarity to obtain the semantic features of keyphrase with high-quality and ranking accuracy. Finally, the top-ranked keyphrases are acquired by the fusion of these features. The experimental results on five datasets illustrate that Y-Rank outperforms the other nine unsupervised methods, achieves enhancements on six accuracy metrics, including Precision, Recall, F-Measure, MRR, MAP, and Bpref, and performs the highest improvement in short text scenarios.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords