Ghinafikar, Zulham (2026) Deteksi kemiripan judul skripsi berbasis Semantic Similarity menggunakan Sentence BERT. Undergraduate thesis, Universitas Islam Negeri Maulana Malik Ibrahim.
|
Text (Fulltext)
220605110184.pdf - Accepted Version Available under License Creative Commons Attribution Non-commercial No Derivatives. (2MB) | Preview |
Abstract
INDONESIA:
Peningkatan volume skripsi di perguruan tinggi memunculkan tantangan besar dalam memverifikasi kebaruan judul guna menghindari duplikasi penelitian. Sistem deteksi kemiripan leksikal konvensional seringkali gagal mengenali kesamaan makna pada dua judul yang menggunakan susunan kata yang berbeda. Penelitian ini bertujuan mengevaluasi kinerja pipeline deteksi kemiripan judul skripsi berbasis semantic similarity menggunakan model Sentence-BERT dan Cosine Similarity. Guna memastikan validitas evaluasi, penelitian ini menyusun dataset dari 1.887 judul skripsi melalui teknik pairwise combination, yang kemudian disaring menggunakan Stratified Under-sampling dan divalidasi oleh domain expert hingga menghasilkan 700 pasangan ground truth. Pengujian performa sistem dilakukan melalui skema Nested K-Fold Cross Validation dengan membandingkan metode pencarian threshold berbasis Grid Search dari penelitain Ghinassi dan Holis. Hasil uji coba menunjukkan bahwa skenario pencarian Ghinassi pada rentang 0.05 - 0.95 menggunakan skema 5-Fold Cross Validation memberikan performa paling optimal, dengan mencetak tingkat Akurasi 82,00%, Presisi 83,20%, Recall 95,44%, F1-Score 88,76%, serta specificity 42,39%. Melalui pengujian lanjutan Flat K-Fold CV, ditetapkan bahwa threshold 0,50 merupakan titik keseimbangan tertinggi dengan tingkat Akurasi 82,00% dan F1-Score 89,03%. Penelitian ini menyimpulkan bahwa model SBERT terbukti sangat efektif dan peka dalam menangkap konteks semantik, sementara konfigurasi threshold 0,50 mampu memberikan titik deteksi yang stabil.
ENGLISH:
The increasing volume of theses in universities poses a major challenge in verifying the novelty of titles to avoid research duplication. Conventional lexical similarity detection systems often fail to recognize the similarity of meaning in two titles that use different wording. This study aims to evaluate the performance of a semantic similarity-based thesis title similarity detection pipeline using the Sentence-BERT and Cosine Similarity models. To ensure the validity of the evaluation, this study compiled a dataset of 1,887 thesis titles through a pairwise combination technique, which was then filtered using Stratified Under-sampling and validated by domain experts to produce 700 ground truth pairs. System performance testing was carried out using the Nested K-Fold Cross Validation scheme by comparing the Grid Search-based threshold search method from the Ghinassi and Holis research. The test results show that the Ghinassi search scenario in the range of 0.05 - 0.95 using the 5-Fold Cross Validation scheme provides the most optimal performance, with an Accuracy level of 82.00%, Precision of 83.20%, Recall of 95.44%, F1-Score of 88.76%, and specificity 42.39%. Through further Flat K-Fold CV testing, it was determined that the threshold of 0.50 was the highest balance point with an Accuracy level of 82.00% and an F1-Score of 89.03%. This study concludes that the SBERT model is proven to be very effective and sensitive in capturing semantic context, while the threshold configuration of 0.50 is able to provide stable detection points.
ARABIC:
يشكل حجم الرسائل المتزايدة في الجامعات تحديا كبيرا في التحقق من حداثة العناوين لتجنب تكرار الأبحاث. غالبا ما تفشل أنظمة اكتشاف التشابه المعجمي التقليدية في التعرف على تشابه المعنى في عنوانين يستخدمان صياغة مختلفة. تهدف هذه الدراسة إلى تقييم أداء خط كشف تشابه عنوان أطروحة يعتمد على التشابه الدلالي باستخدام نموذجي (SBERT) و(Cosine Similarity). لضمان صحة التقييم، جمعت هذه الدراسة مجموعة بيانات تضم 1,887 عنوان أطروحة عبر تقنية الدمج الزوجي، والتي تم تصفيتها بعد ذلك باستخدام التخفيض الطبقي والتحقق منها من قبل خبراء المجال لإنتاج 700 زوج حقيقة أرضية. تم إجراء اختبار أداء النظام باستخدام مخطط (CV K-Fold Nested) من خلال مقارنة طريقة البحث العتبي المعتمدة على البحث الشبكي من أبحاث غيناسي وهوليس. تظهر نتائج الاختبار أن سيناريو البحث في غيناسي في النطاق من 0.05 إلى 0.95 باستخدام نظام التحقق المتقاطع (5-Fold) يوفر الأداء الأمثل، مع مستوى (دقة) 82.00٪، (دقة) 83.20٪، استرجاع 95.44٪، و(درجة F1) 88.76٪ و(خصوصية) 42.39٪. ومن خلال اختبارات إضافية (Flat K-Fold CV)، تم تحديد أن عتبة 0.50 هي أعلى نقطة توازن مع مستوى (دقة) 82.00٪ ودرجة F1 تبلغ 89.03٪. خلصت هذه الدراسة إلى أن نموذج (SBERT) ثبت فعاليته وحساسيته جدا في التقاط السياق الدلالي، بينما تكوين العتبة 0.50 قادر على توفير نقاط كشف مستقرة.
| Item Type: | Thesis (Undergraduate) |
|---|---|
| Supervisor: | Abidin, Zainal and Holle, Khadijah Fahmi hayati |
| Keywords: | Semantic Similarity; Sentence-BERT; Cosine Similarity; Deteksi Kemiripan Teks; Skripsi; Semantic Similarity; Sentence-BERT; Cosine Similarity; Text Similarity Detection; Thesis; التشابه الدلالي; SBERT; تشابه جيب التمام; اكتشاف تشابه النص; الأطروحة |
| Subjects: | 08 INFORMATION AND COMPUTING SCIENCES > 0801 Artificial Intelligence and Image Processing > 080107 Natural Language Processing 20 LANGUAGE, COMMUNICATION AND CULTURE > 2003 Language Studies > 200313 Indonesian Languages 20 LANGUAGE, COMMUNICATION AND CULTURE > 2004 Linguistics > 200408 Linguistic Structures (incl. Grammar, Phonology, Lexicon, Semantics) > 20040804 Semantics |
| Departement: | Fakultas Sains dan Teknologi > Jurusan Teknik Informatika |
| Depositing User: | Zulham Ghinafikar |
| Date Deposited: | 27 Jul 2026 08:38 |
| Last Modified: | 27 Jul 2026 08:38 |
| URI: | http://etheses.uin-malang.ac.id/id/eprint/87614 |
Downloads
Downloads per month over past year
Actions (login required)
![]() |
View Item |
