Comparative Analysis of Machine Learning Models for Identifying Cybercrimes in Social Media Comments
DOI:
https://doi.org/10.33474/infotron.v4i2.23069Keywords:
Comparative Analysis, Cybercrimes, Machine Learning, Social Media CommentsAbstract
The rapid growth of social media has created opportunities for digital interaction but has also introduced challenges, particularly in addressing cybercrimes such as defamation, threats, and SARA-related content. Cybercrime detection on social media is critical as it helps mitigate the spread of harmful behavior, safeguard users, and support law enforcement in addressing violations like Indonesia's Information and Electronic Transactions Law (UU ITE). This study conducts a comparative analysis of machine learning algorithms—Naive Bayes, Support Vector Machines (SVM), and Random Forests—to identify cybercrimes in social media comments. Using a sentiment-labeled dataset obtained from Kaggle, consisting of Indonesian social media comments from Twitter (X), the comments are categorized into seven specific classes: Neutral Sentiment, Positive Sentiment, Negative Sentiment, Insulting Government, Insulting or Defaming Others, Threatening Others, and SARA-Based Content. The results show that Random Forest achieved the highest overall accuracy (91%) and performed best in detecting moderately represented classes such as Insulting Government. SVM demonstrated robust performance with 88% accuracy, particularly excelling in identifying dominant classes like Negative Sentiment, while Naive Bayes, though computationally efficient, struggled with minority classes, achieving an accuracy of 73%. However, the dataset's imbalance posed challenges for all algorithms, particularly with underrepresented categories. This limitation underscores the need for more diverse and representative datasets to improve model performance and ensure broader applicability of the findings.
References
A. Sharma and U. Ghose, “Sentimental Analysis of Twitter Data with respect to General Elections in India,” in Procedia Computer Science, Elsevier B.V., 2020, pp. 325–334. doi: 10.1016/j.procs.2020.06.038.
S. Zervoudakis, E. Marakakis, H. Kondylakis, and S. Goumas, “OpinionMine: A Bayesian-based framework for opinion mining using Twitter Data,” Mach. Learn. with Appl., vol. 3, p. 100018, Mar. 2021, doi: 10.1016/j.mlwa.2020.100018.
M. T. Hasan, M. A. E. Hossain, M. S. H. Mukta, A. Akter, M. Ahmed, and S. Islam, “A Review on Deep-Learning-Based Cyberbullying Detection,” Futur. Internet, vol. 15, no. 5, pp. 1–47, 2023, doi: 10.3390/fi15050179.
S. J. Lesmana, I. Sofia, and F. Felina, “Law Enforcement in Efforts to Combat Cyber Crime in Indonesia Building Future Digital Security,” Int. J. Law Rev. …, vol. 1, no. 3, pp. 120–128, 2023, [Online]. Available: https://www.ijems.id/index.php/ijlrsa/article/view/90%0Ahttps://www.ijems.id/index.php/ijlrsa/article/download/90/77
Waluyadi, “Law Enforcement Against Cyber Crimes in Indonesia : Analysis of The Role of The Ite Law in Handling Cyber Crimes,” Indones. Cyber Law Rev., vol. 1, no. 1, pp. 38–46, 2024.
D. Sukardi, “Legal Protection Against Cyber Crime Threats for Business Economics 4.0,” Leg. Br., vol. 11, no. 4, pp. 2227–2235, 2022, doi: 10.35335/legal.
B. G. Bokolo and Q. Liu, “Artificial Intelligence in Social Media Forensics: A Comprehensive Survey and Analysis,” Electron., vol. 13, no. 9, 2024, doi: 10.3390/electronics13091671.
F. R. S. Nascimento, G. D. C. Cavalcanti, and M. Da Costa-Abreu, “Exploring Automatic Hate Speech Detection on Social Media: A Focus on Content-Based Analysis,” SAGE Open, vol. 13, no. 2, pp. 1–19, 2023, doi: 10.1177/21582440231181311.
G. Kovács, P. Alonso, and R. Saini, “Challenges of Hate Speech Detection in Social Media: Data Scarcity, and Leveraging External Resources,” SN Comput. Sci., vol. 2, no. 2, pp. 1–15, 2021, doi: 10.1007/s42979-021-00457-3.
A. Muneer and S. M. Fati, “A comparative analysis of machine learning techniques for cyberbullying detection on twitter,” Futur. Internet, vol. 12, no. 11, pp. 1–21, 2020, doi: 10.3390/fi12110187.
I. S. Arfan, S. Fauziah, and I. Nawangsih, “Sentiment Analyst of Cyber Bullying in X Using Naïve Bayes Algorithm Analisa Sentimen Terhadap Cyber Bullying di X Menggunakan Algoritma Naïve Bayes,” MALCOM Indones. J. Mach. Learn. Comput. Sci., vol. 4, no. October, pp. 1411–1419, 2024.
A. Muzakir, H. Syaputra, and F. Panjaitan, “A Comparative Analysis of Classification Algorithms for Cyberbullying Crime Detection: An Experimental Study of Twitter Social Media in Indonesia,” Sci. J. Informatics, vol. 9, no. 2, pp. 133–138, 2022, doi: 10.15294/sji.v9i2.35149.
N. Novalita, A. Herdiani, I. Lukmana, and D. Puspandari, “Cyberbullying identification on twitter using random forest classifier,” J. Phys. Conf. Ser., vol. 1192, no. 1, 2019, doi: 10.1088/1742-6596/1192/1/012029.
D. Hidayat, H. Firmanda, and M. H. Wafi, “Analysis of Hate Speech in the Perspective of Changes to the Electronic Information and Transaction Law,” Fiat Justisia J. Ilmu Huk., vol. 18, no. 1, pp. 31–48, 2024, doi: 10.25041/fiatjustisia.v18no1.3146.
F. Isnawan, “TINJAUAN HUKUM PIDANA TENTANG FENOMENA CYBERBULLYING YANG DILAKUKAN OLEH REMAJA,” J. Interpret. Huk., vol. 4, no. 1, pp. 145–163, 2023.
J. Guo et al., “A Deep Look into neural ranking models for information retrieval,” Inf. Process. Manag., vol. 57, no. 6, p. 102067, 2020, doi: 10.1016/j.ipm.2019.102067.
M. A. Rofiqi, A. C. Fauzan, A. P. Agustin, and A. A. Saputra, “Implementasi Term-Frequency Inverse Document Frequency ( TF- IDF ) Untuk Mencari Relevansi Dokumen Berdasarkan Query,” J. Comput. Sci. Appl. Informatics, vol. 1, no. 2, pp. 58–64, 2019.
J. Pardede and D. P. Pamungkas, “The Impact of Balanced Data Techniques on Classification Model Performance,” Sci. J. Informatics, vol. 11, no. 2, pp. 401–412, 2023, doi: 10.15294/sji.v11i2.3649.
A. M. Halim, M. Dwifebri, and F. Nhita, “Handling Imbalanced Data Sets Using SMOTE and ADASYN to Improve Classification Performance of Ecoli Data Sets,” Build. Informatics, Technol. Sci., vol. 5, no. 1, pp. 246–253, 2023, doi: 10.47065/bits.v5i1.3647.
K. Hikmah and A. C. Fauzan, “Sentiment Analysis of Vaccine Booster during Covid-19: Indonesian Netizen Perspective Based on Twitter Dataset,” J. Teknol. Komput. dan Sist. Inf., vol. 5, no. 2, pp. 102–106, 2022.
A. C. Fauzan and K. Hikmah, “Implementasi Algoritma Naive Bayes Dalam Analisis Polarisasi Opini Masyarakat Terkait Vaksin Covid-19,” Rabit J. Teknol. dan Sist. Inf. Univrab, vol. 7, no. 2, pp. 122–128, 2022, doi: 10.36341/rabit.v7i2.2403.
P. Handayani and A. Charis Fauzan, “Machine Learning Klasifikasi Status Gizi Balita Menggunakan Algoritma Random Forest,” KLIK Kaji. Ilm. Inform. dan Komput., vol. 4, no. 6, pp. 3064–3072, 2024, doi: 10.30865/klik.v4i6.1909.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Abd. Charis Fauzan, Mochammad Arifin, Veradella Yuelisa Mafula Veradella

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.


