| dc.description.abstract |
Language mixing and informal usage in code-mixed text make sentiment analysis a challenging task. Although a large number of linguistic resources are available for monolingual languages, lexicon and dataset resources are scarce for most code-mixed language pairs. In this study, we introduce a Sinhala-English code-mixed dataset which comprise over more than 13,000 manually annotated YouTube comments. The dataset achieved a Krippendorff's alpha of 0.885 in nominal metric and 0.890 in ordinal metric. Several machine learning algorithms, such as Logistic Regression, Support Vector Machine, Multinomial Naive Bayes, K-Nearest Neighbor, Decision Tree, and Random Forest, were used on the dataset to create a baseline. In addition to that, transformer-based models, such as BERT, DistilBERT, ALBERT, and XLM-R also fine-tuned on the dataset. The results provide a strong baseline for sentiment classification of Sinhala-English code-mixed text and facilitate future research in multilingual and code-mixed natural language processing. |
en_US |