preprint وصول مفتوح

Attention Mechanism, Transformers, BERT, and GPT: Tutorial and Survey

Research footprint

At a glance

الاستشهادات
3
المراجع
0
Comments
0
Paper overview

Abstract

This is a tutorial and survey paper on the attention mechanism, transformers, BERT, and GPT. We first explain attention mechanism, sequence-to-sequence model without and with attention, self-attention, and attention in different areas such as natural language processing and computer vision. Then, we explain transformers which do not use any recurrence. We explain all the parts of encoder and decoder in the transformer, including positional encoding, multihead self-attention and cross-attention, and masked multihead attention. Thereafter, we introduce the Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformer (GPT) as the stacks of encoders and decoders of transformer, respectively. We explain their characteristics and how they work.

Record transparency

Publication details

DOI
10.31219/osf.io/mru2x
OpenAlex
W4400450746
Document type
preprint
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.