IEEE Access (Jan 2024)
Morpho-Phraseological Based Classification of CEFR Italian L2 Learner Writing Proficiency
Abstract
This paper presents a model for the automatic classification of writing proficiency in Italian as a second language (L2) according to the Common European Framework of Reference for languages (CEFR). The proposed method integrates lexical and morphosyntactic quantitative analysis with phraseological dimensions. Phraseological aspects include the ability to use and understand fixed expressions, idioms, and other multi-word units that are common in a language and reflect the depth of language comprehension typically manifested by native speakers. Specific techniques for encoding phraseological features have been introduced, and basic phraseological statistics, previously unavailable for Italy, have been extracted from an Italian corpus. The proposed model was experimentally compared with widely used machine learning models on a dataset of written texts produced by non-native speakers for the official Italian CEFR certification exams. The experimental results outperformed previous work on the CEFR classification of Italian L2 proficiency in terms of accuracy and all relevant prediction metrics, demonstrating the effectiveness of the proposed approach, which integrates morphosyntactic and phraseological features.
Keywords