Lexicon-Based and Machine Learning Approaches for Sentiment Classification of Telemedicine Reviews in Indonesia
DOI:
https://doi.org/10.21009/JKOMA.091.05Keywords:
Sentiment Analysis, Telemedicine, lexicon-based, machine learning, cross-validationAbstract
The rapid growth of telemedicine services in Indonesia has generated a substantial volume of user reviews, an important source for evaluating service quality. This study compares the effectiveness of lexicon-based and machine learning approaches in classifying review sentiment, using the three largest telemedicine applications in Indonesia, Alodokter, Halodoc, and KlikDokter as a case study. A publicly available dataset of 19,389 reviews, manually labelled by two annotators under the guidance of a psychologist, served as the gold standard and was preprocessed through case folding, normalisation, stopword removal, and stemming. The lexicon-based approach employs the InSet dictionary, while the machine learning approach applies Support Vector Machine (SVM), Naïve Bayes (NB), and Random Forest (RF) with TF-IDF features, evaluated via stratified 10-fold cross-validation. Given the highly imbalanced class distribution (positive 74.77%, negative 21.69%, neutral 3.54%), macro-F1 is adopted as the primary metric alongside accuracy. Machine learning approaches substantially outperformed the lexicon-based approach: SVM achieved the highest macro-F1 of 70.43% (88.68% accuracy), far exceeding the InSet dictionary's macro-F1 of 27.67% (34.92% accuracy). A further key finding is an evaluation paradox: although Naïve Bayes attained the highest accuracy (89.91%), its macro-F1 was comparatively low (60.11%) due to bias toward the majority class, demonstrating how accuracy alone can mislead on imbalanced data. The InSet dictionary also performed poorly in the telemedicine domain, frequently misclassifying positive reviews as negative. Finally, comparing three imbalance-handling strategies (no handling, class-weight, and random oversampling) showed that cost-sensitive learning (class-weight) was most effective, raising the neutral-class F1-score from 26.56% to 35.8% without compromising majority-class performance.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Ahsanun Naseh Khudori

This work is licensed under a Creative Commons Attribution 4.0 International License.