Improving Sentiment Classification of Indonesian Skincare Reviews through Fine-Tuned IndoBERT and Data Augmentation
Abstract
This study analyzes sentiment in customer reviews of local skincare serum products on the Tokopedia e-commerce platform using a fine-tuned IndoBERT model enhanced with data augmentation techniques. A total of 5,000 reviews were collected from ten local skincare brands through web scraping and labeled according to star ratings into three sentiment classes: Positive, Neutral, and Negative. The dataset exhibited extreme class imbalance, with the Positive class representing 94.66% of all observations, creating substantial challenges for minority-class recognition. The data were divided through stratified sampling into 70% training, 15% validation, and 15% test sets to preserve class distributions. To mitigate imbalance, back-translation from Indonesian to English and back to Indonesian, together with synonym replacement, was applied exclusively to minority classes within the training set. The IndoBERT-base-p1 model was subsequently fine-tuned using focal loss combined with class weighting and compared against a baseline model trained without augmentation. Experimental results show that the proposed model achieved 94.40% accuracy, a Macro F1-score of 61.43%, and a Weighted F1-score of 95.14%. Although the baseline model obtained higher overall accuracy of 97.47%, it completely failed to identify the Neutral class, producing an F1-score of 0.00%. In contrast, the proposed approach increased the Neutral F1-score to 23.53% and improved the Macro F1-score by 2.30 percentage points, demonstrating more balanced performance across sentiment classes. The resulting model was deployed as SerumSense, a web-based application developed using Streamlit and SQLite, supporting both single-review and batch sentiment analysis. Black-box testing across 20 functional scenarios confirmed that all application features operated successfully as intended. These findings demonstrate that combining IndoBERT fine-tuning, targeted data augmentation, focal loss, and class weighting offers a practical approach for improving minority-class recognition in highly imbalanced Indonesian e-commerce review datasets.
Keywords
References
D. N. Aryani and H. Wandebori, “The Impact of Online Reviews on Consumer Purchase Intention: An Empirical Study in Indonesian E-Commerce,” J. Distrib. Sci., vol. 21, no. 1, pp. 89–98, 2023, doi: 10.15722/jds.21.01.202301.89.
T. Hidayat, Y. Ruldeviyani, B. R. Aditya, G. R. Madya, and A. W. Nugraha, “Sentiment Analysis of Twitter Data Related to Rinca Island Development Using Doc2Vec and SVM and Logistic Regression as Classifier,” Procedia Comput. Sci., vol. 197, pp. 660–667, 2021, doi: 10.1016/j.procs.2021.12.187.
A. P. Ramadhani and A. Purnomo, “The Rise of Local Beauty Brands in Indonesia: A Case Study of Millennial Consumer Behavior,” J. Bus. Manag. Rev., vol. 3, no. 8, pp. 564–579, 2022, doi: 10.47153/jbmr38.5202022.
N. E. Putri, T. Suryani, and D. Wulandari, “Consumer Trust and Purchase Intention of Local Skincare Products in Indonesia,” Sustainability, vol. 15, no. 3, Art. no. 2456, 2023, doi: 10.3390/su15032456.
W. Zhang, X. Li, Y. Deng, L. Bing, and W. Lam, “A Survey on Aspect-Based Sentiment Analysis: Tasks, Methods, and Challenges,” IEEE Trans. Knowl. Data Eng., vol. 35, no. 11, pp. 11019–11038, 2021, doi: 10.1109/TKDE.2021.3126820.
M. Birjali, M. Kasri, and A. Beni-Hssane, “A Comprehensive Survey on Sentiment Analysis: Approaches, Challenges and Trends,” Knowl.-Based Syst., vol. 226, Art. no. 107134, 2021, doi: 10.1016/j.knosys.2021.107134.
H. Ahmadian, T. F. Abidin, H. Riza, and K. Muchtar, “Hybrid Models for Emotion Classification and Sentiment Analysis in Indonesian Language,” Appl. Comput. Intell. Soft Comput., vol. 2024, Art. no. 2826773, 2024, doi: 10.1155/2024/2826773.
K. Kurniawan, S. Louvan, Y. Wibisono, and R. E. Prasojo, “Leveraging Pre-Trained Language Model for Speech Sentiment Analysis,” in Proc. Int. Conf. Recent Adv. Nat. Lang. Process. (RANLP), 2021, pp. 711–718, doi: 10.26615/978-954-452-072-4_082.
M. Bayer, M.-A. Kaufhold, and C. Reuter, “A Survey on Data Augmentation for Text Classification,” ACM Comput. Surv., vol. 55, no. 7, Art. no. 146, 2021, doi: 10.1145/3544558.
A. Taheri, A. Zamanifar, and A. Farhadi, “Enhancing Aspect-Based Sentiment Analysis Using Data Augmentation Based on Back-Translation,” Int. J. Data Sci. Anal., 2024, doi: 10.1007/s41060-024-00622-w.
Y. Susanti, T. Tokunaga, H. Nishikawa, and H. Obari, “IDENTIC: An Indonesian News Article Dataset for Subjectivity Classification,” Lang. Resour. Eval., vol. 56, pp. 1177–1203, 2022, doi: 10.1007/s10579-021-09570-x.
R. Sutoyo and A. Chowanda, “Indonesian News Classification Using Naïve Bayes and Long Short-Term Memory,” Procedia Comput. Sci., vol. 179, pp. 52–61, 2021, doi: 10.1016/j.procs.2020.12.007.
B. Min et al., “Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey,” ACM Comput. Surv., vol. 56, no. 2, Art. no. 30, 2023, doi: 10.1145/3605943.
A. Palanivinayagam, C. Z. El-Bayeh, and R. Damaševi?ius, “Twenty Years of Machine-Learning-Based Text Classification: A Systematic Review,” Algorithms, vol. 16, no. 5, Art. no. 236, 2023, doi: 10.3390/a16050236.
A. Vaswani et al., “Attention is All You Need,” in Proc. 31st Conf. Neural Inf. Process. Syst. (NeurIPS), 2017, pp. 5998–6008.
G. M. de Santana Correia and E. L. Colombini, “Attention, Please! A Survey of Neural Attention Models in Deep Learning,” Artif. Intell., vol. 311, Art. no. 103788, 2022, doi: 10.1016/j.artint.2022.103788.
G. Brauwers and F. Frasincar, “A Survey on Aspect-Based Sentiment Analysis: Datasets, Methods, and Trends,” Comput. Sci. Rev., vol. 49, Art. no. 100575, 2023, doi: 10.1016/j.cosrev.2023.100575.
T. Lin, Y. Wang, X. Liu, and X. Qiu, “BERT Has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model,” in Findings Assoc. Comput. Linguist.: NAACL, 2022, pp. 191–205, doi: 10.18653/v1/2022.findings-naacl.16.
M. A. Khder, “Web Scraping or Web Crawling: State of Art, Techniques, Approaches and Application,” Int. J. Adv. Soft Comput. Appl., vol. 13, no. 3, pp. 145–168, 2021, doi: 10.15849/IJASCA.211128.11.
A. Satriajati, A. D. Cahyani, and R. Chandra, “Implementation of Web Scraping for E-Commerce Product Price Monitoring” [in Indonesian], J. Teknol. dan Komput., vol. 7, no. 1, pp. 58–67, 2021, doi: 10.31294/jtk.v7i1.9651.
N. M. D. P. Suari, A. A. K. A. C. Sudana, and D. G. H. Divayana, “Sentiment Analysis on Twitter Using Machine Learning Approach,” Int. J. Comput. Inf. Technol., vol. 9, no. 1, pp. 45–52, 2023, doi: 10.31294/ijcit.v9i1.14249.
A. Rendragraha, M. S. Mubarok, and Adiwijaya, “Sentiment Classification of Tokopedia User Reviews Using the Naive Bayes Algorithm” [in Indonesian], J. Tek. Elektro, vol. 1, no. 2, pp. 89–96, 2021, doi: 10.47134/jte.v1i2.21.
B. Wilie, Vincentio, Y. Xu, S. W. Lim, H. Lovenia, and A. Purwarianti, “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proc. 1st Conf. Asia-Pac. Chapter Assoc. Comput. Linguist. and 10th Int. Joint Conf. Nat. Lang. Process. (AACL-IJCNLP), 2021, pp. 843–857.
M. Rizwan, L. Cruz, and M. Fernandez, “Effectiveness of Data Augmentation for Sentiment Analysis,” Nat. Lang. Process. Res., vol. 1, no. 3–4, pp. 1–9, 2022, doi: 10.2991/nlpr.d.200522.001.
F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-Trained Language Model for Indonesian NLP,” in Proc. 28th Int. Conf. Comput. Linguist. (COLING), 2022, pp. 757–770, doi: 10.18653/v1/2020.coling-main.66.
R. R. Ula and D. S. Rusdianto, “Indonesian Twitter Sentiment Analysis Using Convolutional Neural Network and Long Short-Term Memory,” J. Ilm. Teknol. Inf. Asia, vol. 15, no. 2, pp. 129–138, 2021, doi: 10.32815/jitika.v15i2.543.
A. P. Nugraha and I. Budi, “Sentiment Analysis of Product Reviews in Indonesian E-Commerce Using BERT,” J. Big Data, vol. 9, no. 1, Art. no. 104, 2022, doi: 10.1186/s40537-022-00656-4.
J. Wei and K. Zou, “EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks,” 2023.
S. Nurmaini, D. Stiawan, and B. Y. Suprapto, “Handling Class Imbalance in Sentiment Analysis Using SMOTE and Deep Learning,” Int. J. Adv. Comput. Sci. Appl., vol. 12, no. 5, pp. 678–685, 2021, doi: 10.14569/IJACSA.2021.0120577.
DOI: https://doi.org/10.52088/ijesty.v6i3.1835
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Nadia Thahira, Ar Razi




























