Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
At a glance
- Citations
- 1
- References
- 0
- Comments
- 0
Abstract
Code-switching is a common practice for millions of multilingual speakers but remains challenging for Large Language Models (LLMs).This paper investigates LLM capabilities in generating code-switched text, conducting extensive experiments across five diverse language pairs: English paired with Hindi, Tamil, Malayalam, and Indonesian, as well as Indonesian-Javanese.Our analysis, grounded in comprehensive human evaluations by native speakers, uncovers a directional asymmetry: LLMs consistently produce higher-quality (more accurate and fluent) code-switched text when prompted with a lower-resource language (e.g., Hindi, Tamil, Javanese) as the source, compared to when a higher-resource language (English, Indonesian) serves as the source.This asymmetry mirrors sociolinguistic patterns, particularly the Matrix Language Frame model, suggesting LLMs implicitly learn common codeswitching structures from their training data where regional languages often form the grammatical base.Furthermore, we find that explicit linguistic guidance, applied through Equivalence Constraint Theory (ECT) to identify switching points, primarily benefits generation quality only in the less common, higherresource-source direction where LLMs intrinsically struggle.These findings highlight a crucial interplay between the implicit linguistic knowledge captured by LLMs and the targeted utility of explicit linguistic constraints.We also introduce CSPREF, a pairwise preference dataset derived from our human evaluations, to facilitate future research in code-switching generation and evaluation.
Publication details
- DOI
- 10.18653/v1/2026.cdl-1.1
- OpenAlex
- W4404342393
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.