conference-paper

Expanding the Context of Large Language Models Via Linear Interpolation of Positional Embeddings

Research footprint

At a glance

Citations
1
References
29
Comments
0
Paper overview

Abstract

This study explores the problem of limited context in Russian large language models (LLMs) and proposes a new method for increasing their context size in an efficient manner. The proposed method is based on the linear interpolation of positional embeddings, allowing for a significant increase in the context size of the LLMs. This, in turn, has great practical significance for processing long documents and developing applications that require extended interaction with the LLMs. During the research, the ruGPT-3.5 model with 13 billion parameters was used, trained with a context of 2048 tokens. Through linear interpolation of positional embeddings, it was possible to successfully expand the model's context up to 6144 tokens. This substantial increase in context opens new possibilities for processing long text data and enhancing the performance of chatbots and other applications that work with Russian LLMs.

Record transparency

Publication details

DOI
10.1109/itnt60778.2024.10582292
OpenAlex
W4400449854
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.