conference-paper

IDSpider: Indonesian Standard Dataset for Text-to-SQL

Research footprint

At a glance

Citations
0
References
21
Comments
0
Paper overview

Öz

This research presents IDSpider, a new standard dataset for Text-to-SQL research in Indonesian, developed by translating the well-established Spider dataset into Indonesian. Text-to-SQL is a critical technology that bridges the gap between natural language processing (NLP) and database management systems, allowing non-technical users to retrieve data from databases using simple, natural language queries. However, most existing datasets are in English, leaving nonEnglish languages, especially Indonesian, underrepresented in the field. This paper describes the challenges encountered in translating both natural language questions and SQL queries, given the need for precision in maintaining SQL syntax and database structure integrity. We implemented a four-stage process consisting of Extraction, Translation, Cleaning, and Conversion to generate the IDSpider dataset. Furthermore, we tested the performance of the BRIDGE model on both the original Spider dataset and the translated IDSpider dataset. Results show that while the BRIDGE model performs well on Spider, its performance drops significantly when applied to IDSpider, primarily due to language differences and translation complexities. This research establishes the first benchmark for Indonesian Text-to-SQL tasks and lays the groundwork for further improvements in cross-lingual natural language interfaces for databases.

Record transparency

Publication details

DOI
10.1109/icic64337.2024.10956918
OpenAlex
W4409474733
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.