Klasifikasi sentimen teks code-mixed Indonesia–Inggris non-formal pada X menggunakan model Fine-Tuned DistilBERT

Nurbayani, Syifa Arifah and Maylawati, Dian Sa’adillah and Atmadja, Aldy Rialdy (2026) Klasifikasi sentimen teks code-mixed Indonesia–Inggris non-formal pada X menggunakan model Fine-Tuned DistilBERT. EXPLORER Journal of Computer Science and Information Technology, 6 (2). pp. 54-65. ISSN 2774-4647

[img]
Preview
Text
2856-Article Text-11276-2-10-20260725.pdf - Published Version

Download (683kB) | Preview
Official URL: https://journal.fkpt.org/index.php/Explorer/articl...

Abstract

INDONESIA: Perkembangan media sosial mendorong meningkatnya penggunaan bahasa campuran (code-mixed) Indonesia–Inggris dalam komunikasi digital, khususnya pada media sosial X. Karakteristik teks media sosial yang non-formal, seperti penggunaan slang, singkatan, emoji, dan perpindahan bahasa dalam satu kalimat, menyebabkan proses analisis sentimen menjadi lebih kompleks dibandingkan teks monolingual. Penelitian ini bertujuan melakukan klasifikasi sentimen teks code-mixed dengan melakukan pengukuran performa model transformer ringan DistilBERTdan membandingkannya dengan model IndoBERTweet. Penelitian menggunakan metodologi CRISP-DM yang meliputi tahapan Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, dan Deployment. Dataset yang digunakan terdiri dari data primer hasil crawling media sosial X periode 2022–2026sebanyak 1.108 data dan dataset sekunder dari penelitian sebelumnya sebanyak 5.048 data. Penelitian ini menerapkan tiga skenario pengujian, yaitu translasi ke Bahasa Indonesia, translasi ke Bahasa Inggris, dan raw data tanpa translasi. Hasil penelitian menunjukkan bahwa DistilBERT memperoleh performa terbaik pada skenario translasi Bahasa Inggris dengan akurasi sebesar 73,66% pada dataset primer dan 72,00% pada dataset sekunder. Sementara itu, IndoBERTweet menunjukkan performa terbaik pada skenario translasi Bahasa Indonesia dengan akurasi mencapai 79,01%. Hasil tersebut menunjukkan bahwa kesesuaian bahasa data dengan karakteristik pretraining model berpengaruh terhadap performa klasifikasi sentimen pada teks code-mixed non-formal ENGLISH: The rapid growth of social media has increased the use of Indonesian–English code-mixed language in digital communication, particularly on social media X. The non-formal characteristics of social media text, such as slang, abbreviations, emojis, and language switching within a single sentence, make sentiment analysis more challenging than monolingual text. This study aims to perform sentiment classification on code-mixed text by evaluating the performance of the lightweight Transformer model DistilBERT and comparing it with IndoBERTweet. The study adopts the CRISP-DM methodology, which consists of Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, and Deployment stages. The dataset comprises 1,108 primary data collected from social media X between 2022 and 2026 and 5,048 secondary data obtained from previous research. Three experimental scenarios were applied: translation into Indonesian, translation into English, and raw data without translation. The results show that DistilBERT achieved its best performance on the English translation scenario, with accuracies of 73.66% on the primary dataset and 72.00% on the secondary dataset. Meanwhile, IndoBERTweet obtained the highest performance on the Indonesian translation scenario, achieving an accuracy of 79.01%. These findings indicate that the alignment between the language of the input data and the pre-training characteristics of the model significantly affects sentiment classification performance on non-formal code-mixed text.

Item Type: Article
Uncontrolled Keywords: Klasifikasi Sentimen; Code-Mixed; DistilBERT; Media Sosial; Transformer Ringan
Subjects: Data Processing, Computer Science > Systems Analysis and Computer Design
Data Processing, Computer Science > Computer Performance Evaluation
Special Computer Methods
Special Computer Methods > Artificial Intelligence
Technology, Applied Sciences
Divisions: Fakultas Sains dan Teknologi > Program Studi Teknik Informatika
Depositing User: Syifa Arifah Nurbayani
Date Deposited: 03 Aug 2026 08:08
Last Modified: 03 Aug 2026 08:08
URI: https://digilib.uinsgd.ac.id/id/eprint/137259

Actions (login required)

View Item View Item