BERT, short for Bidirectional Encoder Representations from Transformers, is a language model and text-representation system developed by Google researchers and introduced in October 2018. It uses the encoder of the Transformer architecture to learn from unlabeled text, conditioning representations jointly on preceding and following context. BERT can then be adapted to natural-language processing tasks through fine-tuning, usually with a small task-specific output layer. The original paper was published at NAACL in June 2019. (arxiv.org)
Background and design
BERT belongs to the pretrain-and-adapt approach to transfer learning: a model first learns reusable representations from a large text collection, then learns a particular task using labeled examples. Its central contribution was to combine deep bidirectional Transformer representations with a relatively uniform adaptation procedure across different tasks. (aclanthology.org)
Earlier approaches included static word embeddings and contextual systems such as ELMo. Static embeddings assign a word the same vector regardless of its surroundings. BERT instead produces context-dependent representations: the representation of “bank,” for example, changes with the surrounding sentence. Unlike a left-to-right language model, its encoder can incorporate information on both sides of a token at every layer. This distinction concerns how context is processed, not an ability to understand language in the human sense. (github.com)
Architecture and input representation
BERT consists of stacked Transformer encoder blocks. Each block uses multi-head attention, specifically self-attention, followed by a position-wise feed-forward network, with residual connections and layer normalization. Its feed-forward layers use the Gaussian error linear unit activation. The original principal configurations were BERT-Base, with 12 layers, 768-dimensional hidden representations, 12 attention heads, and approximately 110 million parameters; and BERT-Large, with 24 layers, 1,024-dimensional representations, 16 heads, and approximately 340 million parameters. (aclanthology.org)
Input tokenization uses WordPiece, which can divide words into subword units. Each input position combines a token embedding, a segment embedding indicating membership in one of two input segments, and learned position embeddings. The special token [CLS] begins the sequence, while [SEP] separates segments or marks their ends. For classification, the final representation at [CLS] is commonly passed to an output layer. The original models support sequences of up to 512 tokens, including special tokens. (aclanthology.org)
Pretraining objectives
BERT’s original pretraining combines two objectives derived automatically from text, making it an example of self-supervised learning.
Masked language modeling trains the model to recover selected input tokens. In the original procedure, 15% of token positions are selected. Of those selected positions, 80% are replaced with [MASK], 10% with a random token, and 10% remain unchanged. Prediction is evaluated only at selected positions. This corruption procedure prevents straightforward copying while reducing dependence on a special token absent from ordinary downstream inputs. (aclanthology.org)
Next-sentence prediction trains a binary classifier to distinguish consecutive text segments from pairs in which the second segment is randomly sampled. Half the training pairs are positive examples and half are negative examples. Despite its name, the objective does not generate a next sentence. The original English training data comprised BooksCorpus, reported as 800 million words, and English Wikipedia, reported as 2.5 billion words. (aclanthology.org)
The objectives use cross-entropy losses. Pretraining optimizes the combined loss function using Adam, with a warming-up and decaying learning-rate schedule. BERT learns representations through these prediction tasks rather than through manually assigned linguistic annotations. (aclanthology.org)
Adaptation and evaluation
In downstream supervised learning, BERT’s pretrained parameters and an added output layer are usually updated together. Sequence classification uses a sequence-level representation; sequence labeling assigns predictions to token representations; extractive question answering predicts the beginning and end of an answer span within a supplied passage. These tasks share the encoder but differ in their outputs and training losses. (github.com)
The original paper reported state-of-the-art results on eleven tasks. Its reported results included a GLUE score of 80.5, MultiNLI accuracy of 86.7%, and SQuAD 1.1 test F1 of 93.2. These are historical benchmark results under the paper’s evaluation procedures, rather than fixed performance characteristics of every BERT model or application. (aclanthology.org)
Variants and limitations
Google released English cased and uncased checkpoints, a Chinese model, and multilingual checkpoints. The multilingual cased model was pretrained on Wikipedia text in 104 languages. Research demonstrated cross-lingual transfer in which a model fine-tuned using annotations in one language was evaluated in another, although this does not imply equal performance across languages. (github.com)
RoBERTa, introduced in 2019, investigated BERT’s training recipe and improved results through changes including longer training, additional data, dynamic masking, and removal of next-sentence prediction. DistilBERT used knowledge distillation during pretraining to produce a smaller, faster encoder. These models illustrate that architecture, training objectives, and computational budget are distinct contributors to performance. (arxiv.org)
BERT’s original context length restricts direct processing of long documents, while dense self-attention becomes more costly as sequences grow. Its masked-token objective also differs from the next-token objective commonly associated with GPT models. Unmodified BERT is primarily an encoder for representation and prediction tasks, not a left-to-right text generator. RoBERTa’s replication study further showed that comparisons between pretrained models depend substantially on training data, duration, and hyperparameter choices. (github.com)