Responsive Banner

Klasifikasi cyberbullying menggunakan metode desicion tree dengan menggunakan ekstraksi fitur term frequency-inverce document (tf-idf)

Lestari, Sari Indah (2026) Klasifikasi cyberbullying menggunakan metode desicion tree dengan menggunakan ekstraksi fitur term frequency-inverce document (tf-idf). Undergraduate thesis, Universitas Islam Negeri maulana Malik Ibrahim.

[img]
Preview
Text (Fulltext)
19650045.pdf - Accepted Version
Available under License Creative Commons Attribution Non-commercial No Derivatives.

(6MB) | Preview

Abstract

INDONESIA

Seiring meningkatnya penggunaan jejaring media sosial oleh berbagai macam kalangan, cyberbullying bisa menjadi hal yang tidak bisa dihindari akan tetapi masih bisa disaring, meski demikian setiap individu juga harus selalu waspada terhadap cyberbullying dan diimbau untuk selalu melindungi diri ketika melakukan kegiatan di media sosial, sebab apabila dibiarkan saja tanpa penanganan cyberbullying dapat mengganggu kegiatan positif di media sosial. Dalam penelitan ini, penelitia akan memanfaatkan machine learning dalam mengklasifikasikan teks cyberbullying yang beredar pada teks tweet di twitter. Metode yang digunakan adalah Decision Tree untuk mengklasifikasikan data teks tweet sebanyak 47.692 data yang terbagi menjadi enam jenis kelas yakni, religon, gender, age, ethnicity, other cyberbullying, dan not cyberbullying. Setelah melalui tahapan pra proses data yakni pembersihan dan penyaringan teks, data nantinya akan dihitung bobotnya dengan TF-IDF, selanjutnya data akan diuji dengan tiga jenis skenario pembagian data, lalu diimplementasikan pada metode Decision Tree dengan kedalaman maksimum terbaik berdasarkan skenario data. Hasil yang diperoleh menujukkan skenario data 90:10 memiliki rata-rata akurasi dengan tinggi 83,63% dimana model mampu mengklasifikasikan data tesk cyberbullying berdasarkan kelas gender, age, ethnicity, religion, sedangkan model kurang dalam mengklasifikasi jenis kelas other cyberbullying dan not cyberbullying.

INGGRIS

Along with the increasing use of social media networks by various groups, cyberbullying can become unavoidable but can still be filtered. Nevertheless, every individual must remain vigilant against cyberbullying and is urged to always protect themselves when engaging in social media activities, as if left unhandled, cyberbullying can disrupt positive activities on social media. In this study, the researcher utilizes machine learning to classify cyberbullying texts found in tweets on Twitter. The method used is the Decision Tree to classify 47,692 tweet text data, which are divided into six classes: religion, gender, age, ethnicity, other cyberbullying, and not cyberbullying. After going through the data preprocessing stages, namely text cleaning and filtering, the data weights are calculated using TF-IDF. Subsequently, the data is tested using three types of data splitting scenarios and then implemented into the Decision Tree method with the best maximum depth based on the data scenarios. The results obtained show that the 90:10 data scenario yields the highest mean accuracy of 83.63%, where the model is capable of classifying cyberbullying text data based on gender, age, ethnicity, and religion classes, while the model is less effective in classifying the other cyberbullying and not cyberbullying classes.

ARABIK

مع زيادة استخدام شبكات التواصل الاجتماعي من قبل مختلف الفئات، يمكن أن يصبح التنمر السيبران أمرًا لا مفر منه ولكنه لا يزال قابلًا للتصفية .ومع ذلك، يجب على كل فرد أن يظل يقظًا دائمًا ضد التنمر السيبران ويُنصح بحماية نفسه دائمًا عند القيام .بالأنشطة على وسائل التواصل الاجتماعي، لأنه إذا تُرك دون معالجة، فقد يعيق الأنشطة الإيجابية على وسائل التواصل الاجتماعي ف هذا البحث، ستقوم الباحثة بالاستفادة من التعلم الآلي ف تصنيف نصوص التنمر السيبران المنتشرة ف نصوص التغريدات على منصة تويتر .الطريقة المستخدمة هي شجرة القرار لتصنيف ٤٧,٦٩٢ من بيانات نصوص التغريدات المقسمة إلى ستة أنواع من الفئات وهي :الدين، والجنس، والعمر، والعرق، وتنمر سيبران آخر، وليس تنمرًا سيبرانيًا .بعد المرور بمراحل المعالجة المسبقة للبيانات وهي ث سيتم اختبار البيانات بثلاثة أنواع من سيناريوهات ،TF-IDF تنظيف النصوص وتصفيتها، سيتم حساب وزن البيانات باستخدام تقسيم البيانات، ث تطبيقها على طريقة شجرة القرار بأفضل عمق أقصى بناءً على سيناريوهات البيانات .وأظهرت النتائج المحققة أن سيناريو البيانات ٩٠:١٠ حقق أعلى دقة بنسبة ٦٣٫٨٣٪ حيث كان النموذج قادرًا على تصنيف بيانات نصوص التنمر السيبران .بناءً على فئات الجنس، والعمر، والعرق، والدين، بينما كان النموذج أقل كفاءة ف تصنيف فئتي تنمر سيبران آخر وليس تنمرًا سيبرانيًا

Item Type: Thesis (Undergraduate)
Supervisor: Aziz, Okta Qomarudin and Wahyu Prakasa, Johan Ericka
Keywords: Klasifikasi teks cyberbullying; Decision Tree; Cyberbullying; Cyberbullying text classification; Decision Tree; Cyberbullying; تصنيف نصوص التنمر السيبراني; شجرة القرار; التنمر السيبراني
Subjects: 08 INFORMATION AND COMPUTING SCIENCES > 0801 Artificial Intelligence and Image Processing > 080199 Artificial Intelligence and Image Processing not elsewhere classified
Departement: Fakultas Sains dan Teknologi > Jurusan Teknik Informatika
Depositing User: Sari Indah Lestari
Date Deposited: 31 Jul 2026 08:54
Last Modified: 31 Jul 2026 08:54
URI: http://etheses.uin-malang.ac.id/id/eprint/88893

Downloads

Downloads per month over past year

Actions (login required)

View Item View Item