conference-paper

Addressing LLM Challenges: A Hybrid Framework for Duplicate Question Detection

Research footprint

At a glance

Citations
0
References
24
Comments
0
Paper overview

Abstract

Detecting similar or duplicate texts or questions is a critical subtask in information retrieval. Despite the remarkable advancements achieved by large language models (LLMs) in the question answering (QA) domain in recent years, challenges such as unreliable responses caused by LLM hallucinations and outdated information have highlighted the continued importance of information retrieval techniques. The integration of retrieval-augmented generation (RAG) with LLMs has mitigated some of these limitations, enhancing the reliability of applying LLMs to domain-specific knowledge. This paper uses Chinese educational questions as a case study and introduces an innovative framework for duplicate question detection, termed the Balanced Duplicate Question Detector (BDQD). The BDQD incorporates multiple extensible detectors and a machine learning-based balancer. Experimental results reveal that detecting duplicate questions involves an inherent trade-off between precision and recall. Logistic regression proves effective in balancing thresholds across different detectors, establishing a more stable and reliable standard for duplicate question determination. Finally, we propose a hybrid QA framework that considers both cost and efficiency, integrating the lightweight question retrieval architecture developed in this study with RAG and LLMs. This framework offers practical recommendations for building robust and efficient QA systems.

Record transparency

Publication details

DOI
10.1109/icairc64177.2024.10900148
OpenAlex
W4408145767
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.