Trahestiawan, Akbar Bimantara (2026) Deteksi berita palsu berbahasa Indonesia pada data tidak seimbang menggunakan IndoBERT Dengan Focal Loss. Undergraduate thesis, Universitas Islam Negeri Maulana Malik Ibrahim.
|
Text (Fulltext)
220605110080.pdf - Accepted Version Available under License Creative Commons Attribution Non-commercial No Derivatives. (1MB) | Preview |
Abstract
INDONESIA:
Penyebaran berita palsu (hoaks) di ruang digital Indonesia menjadi permasalahan serius yang mengancam stabilitas sosial dan kepercayaan publik. Klasifikasi berita hoaks berbahasa Indonesia berperan penting dalam menekan penyebaran informasi palsu, namun masih menghadapi tantangan berupa ketidakseimbangan distribusi data (class imbalance). Penelitian ini mengusulkan pendekatan klasifikasi berita hoaks menggunakan teknologi Natural Language Processing (NLP) melalui fine-tuning model IndoBERT-base-p1. Dataset yang digunakan adalah Indonesia False News Dataset 2025 – Turnbackhoax.id yang terdiri dari 970 sampel, dengan distribusi 698 data kategori SALAH (71,96%) dan 272 data kategori PENIPUAN (28,04%). Penelitian dilakukan menggunakan dua skenario klasifikasi biner, yaitu Cross Entropy Loss sebagai baseline dan Focal Loss (α=0,75; γ=2,0) untuk mengatasi ketidakseimbangan data. Focal Loss dirancang untuk meningkatkan pembelajaran pada sampel kelas minoritas dengan mengurangi kontribusi sampel yang mudah diklasifikasikan. Hasil evaluasi menunjukkan bahwa kombinasi IndoBERT dan Focal Loss menghasilkan akurasi 98,63% dan macro F1-score 98,33%, lebih baik dibandingkan Cross Entropy Loss yang memperoleh akurasi 97,26% dan macro F1-score 96,56%. Peningkatan terbesar terjadi pada kelas PENIPUAN dengan F1-score meningkat dari 95,00% menjadi 97,62%, sementara nilai *loss* menurun dari 0,0923 menjadi 0,0059 (93,61%). Hasil penelitian menunjukkan bahwa penerapan Focal Loss pada IndoBERT efektif meningkatkan kinerja klasifikasi berita hoaks, pada kelas minoritas.
ENGLISH:
The spread of fake news (hoaxes) in Indonesia's digital space has become a serious problem that threatens social stability and public trust. The classification of Indonesian-language hoax news plays an important role in suppressing the dissemination of false information, yet it still faces challenges in the form of imbalanced data distribution (class imbalance). This study proposes a hoax news classification approach using Natural Language Processing (NLP) technology through fine-tuning of the IndoBERT-base-p1 model. The dataset used is the Indonesia False News Dataset 2025 – Turnbackhoax.id, consisting of 970 samples, with a distribution of 698 data points in the SALAH (FALSE) category (71.96%) and 272 data points in the PENIPUAN (DECEPTION) category (28.04%). The research was conducted using two binary classification scenarios: Cross Entropy Loss as a baseline and Focal Loss (α=0.75; γ=2.0) to address data imbalance. Focal Loss is designed to enhance learning on minority class samples by reducing the contribution of easily classified samples. The evaluation results show that the combination of IndoBERT and Focal Loss achieved an accuracy of 98.63% and a macro F1-score of 98.33%, outperforming Cross Entropy Loss, which obtained an accuracy of 97.26% and a macro F1-score of 96.56%. The most significant improvement occurred in the PENIPUAN (DECEPTION) class, with the F1-score increasing from 95.00% to 97.62%, while the loss value decreased from 0.0923 to 0.0059 (93.61%). The findings indicate that the application of Focal Loss to IndoBERT is effective in improving the performance of hoax news classification, particularly for the minority class.
ARABIC:
ي ُ عد انتشار األخبار الكاذب ةومات املضللة، إال أنه ال يزال يواجه حتدًي ً تصنيف األخبار الكاذبة املكتوبة ابللغة اإلندونيسية دور ً ا مهم ً ا يف احلد من انتشار املعل هدف هذه الدراسة إىل اقرتاح منهج لتصنيف األخبار الكاذبة ت .)Data Imbalance( يتمثل يف عدم توازن توزيع البيانت Fine-Tuning من خالل )Natural Language Processing - NLP( ابستخدام تقنيات معاجلة اللغات الطبيعية Indonesia False News Dataset است ُخدمت يف هذه الدراسة جمموعة بيانت IndoBERT-base-p1 لنموذج 2025 – Turnbackhoax.id ، عينة ضمن فئة 698 عينة، تتوزع إىل 970 واليت تضم SALAH (FALSE) بنسبة خدام ون ُفذت الدراسة ابست %28.04 بنسبة PENIPUAN (DECEPTION) عينة ضمن فئة 272 ، و 71.96% Focal Loss واستخدام ، )Baseline( كنموذج مرجعي Cross Entropy Loss صنيف الثنائي، ومها استخدام سيناريوهني للت لتعزيز عملية التعلم على عينات Focal Loss ملعاجلة مشكلة عدم توازن البيانت. وقد صُ ممت γ = 2.0 و α = 0.75 بقي م IndoBERT أظهرت نتائج التقييم أن اجلمع بني منوذج . أتثري العينات اليت يسهل تصنيفها الفئة األقل متثيالً من خالل تقليل متفوق ًا على منوذج ، %98.33 بلغت Macro F1-score ودرجة ، %98.63 بلغت )Accuracy( حقق دقة Focal Loss و Cross Entropy Loss ودرجة %97.26 الذي حقق دقة Macro F1-score حتسن يف ر أكرب وقد ظه .%96.56 بلغت يف حني اخنفضت ، %97.62 إىل %95.00 من F1-score حيث ارتفعت قيمة ، PENIPUAN (DECEPTION) فئة Focal وتشري نتائج هذه الدراسة إىل أن تطبيق .%93.61 بنسبة اخنفاض بلغت ، 0.0059 إىل 0.0923 من Loss قيمة Loss على منوذج IndoBERT وال سيما يف تصنيف عينات الفئة األقل ء تصنيف األخبار الكاذبة، ي ُ عد فعاالً يف حتسني أدا .
| Item Type: | Thesis (Undergraduate) |
|---|---|
| Supervisor: | Abidin, Zainal and Sari, Nur Fitriyah Ayu Tunjung |
| Keywords: | Deteksi Berita Palsu; Focal Loss; IndoBERT; Klasifikasi Teks; Ketidakseimbangan Data; Fake News Detection;Text Classification, Data Imbalance ; تصنيف النصوص; عدم توازن البيانتالكاذبة : ا |
| Subjects: | 08 INFORMATION AND COMPUTING SCIENCES > 0801 Artificial Intelligence and Image Processing > 080107 Natural Language Processing |
| Departement: | Fakultas Sains dan Teknologi > Jurusan Teknik Informatika |
| Depositing User: | Akbar Bimantara Trahestiawan |
| Date Deposited: | 28 Jul 2026 13:46 |
| Last Modified: | 28 Jul 2026 13:46 |
| URI: | http://etheses.uin-malang.ac.id/id/eprint/88185 |
Downloads
Downloads per month over past year
Actions (login required)
![]() |
View Item |
