preprint Open access

Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
3
References
0
Comments
0
Paper overview

Öz

To ensure that text generated by large language models (LLMs) is in an expected format, constrained decoding proposes to enforce strict formal language constraints during generation. However, as we show in this work, not only do such methods incur performance overhead during generation, but many of them also significantly impair task accuracy, if they do not correctly align the underlying LLM sub-word vocabularies with external constraints. To address this, we present a novel decoding algorithm, DOMINO, that can enforce constraints in a fully subword-aligned fashion, while leveraging pre-computation and speculative decoding to achieve virtually no overhead and in some cases even almost 2$\times$ speedup over unconstrained decoding -- thereby outperforming existing approaches by a wide margin.

Record transparency

Publication details

DOI
10.48550/arxiv.2403.06988
OpenAlex
W4392781601
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.