Asia-Pacific Journal of Information Technology and Multimedia (Dec 2020)

QUANTIFYING SEMANTIC SHIFT VISUALLY ON A MALAY DOMAIN SPECIFIC CORPUS USING TEMPORAL WORD EMBEDDING APPROACH

  • Sabrina Tiun,
  • Saidah Saad,
  • Nor Fariza Mohd Noor,
  • Azhar Jalaludin,
  • Anis Nadiah Che Abdul Rahman

DOI
https://doi.org/10.17576/apjitm-2020-0902-01
Journal volume & issue
Vol. 09, no. 02
pp. 1 – 10

Abstract

Read online

In this study, we propose an alternative approach to analyzing a domain-specific time series corpus for detecting word evolution. The method trains a target corpus in time series into a temporal word embedding (TWE) model. The advantage of TWE is that one can see how the meaning of a word changes over time. We have chosen the TWEC approach to model a Malay domain-specific time-series corpus, the Malaysian Hansard Corpus (MHC), to a TWE model and called the model as MHC-TWEC. Two primary analyses, i.e., self-similarity analysis and user-defined method analysis, were performed to validate the effectiveness of the MHC-TWEC model in quantifying semantic shift on MHC visually. From those analyses, we visually find out that the TWE model can capture the semantic shift in the temporal corpus (the MHC).

Keywords