Expanding the Context of Large Language Models Via Linear Interpolation of Positional Embeddings
At a glance
- Citations
- 1
- References
- 29
- Comments
- 0
Abstract
This study explores the problem of limited context in Russian large language models (LLMs) and proposes a new method for increasing their context size in an efficient manner. The proposed method is based on the linear interpolation of positional embeddings, allowing for a significant increase in the context size of the LLMs. This, in turn, has great practical significance for processing long documents and developing applications that require extended interaction with the LLMs. During the research, the ruGPT-3.5 model with 13 billion parameters was used, trained with a context of 2048 tokens. Through linear interpolation of positional embeddings, it was possible to successfully expand the model's context up to 6144 tokens. This substantial increase in context opens new possibilities for processing long text data and enhancing the performance of chatbots and other applications that work with Russian LLMs.
Publication details
- DOI
- 10.1109/itnt60778.2024.10582292
- OpenAlex
- W4400449854
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.