conference-paper

Automated Duplicate Question Analysis with NLP

Research footprint

At a glance

الاستشهادات
0
المراجع
14
Comments
0
Paper overview

Abstract

As the capacity of user-generated content endures to upsurge, question duplication has emerged as a must-have for increasing efficiency in question-answering systems. The aim of this paper is to explore how to look for duplicate questions using machine and deep learning methods, reducing duplication in large sets of question data with great accuracy and thus improving user experience at knowledge platforms. This exploration includes models of Logistic Regression, Naive Bayes, Decision Trees, LSTM networks, BERT embeddings, and advanced ensemble methods, such as XGBoost. From above, it is seen that the performance metrics vary across different models, and its accuracy is shown as follows - accuracy of Logistic Regression at 78.9%, 76.4% accuracy of Naive Bayes, 72.3% of Decision Trees, and 82.6% in the LSTM model; the best accuracy has been obtained with the model of XGBoostoptimised by Optuna at a rate of 91.5%, even higher than some deep learning models such as BERT. The results depicted above indicate the outcome of hyperparameter tuning via Optuna on model performance. Our results highlight better effectiveness from optimized ensemble learning techniques in question deduplication, very useful in the implementation of efficient question-answering and knowledge management solutions.

Record transparency

Publication details

DOI
10.1109/idciot64235.2025.10914861
OpenAlex
W4408399615
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.