BERT (Bidirectional Encoder Representations from Transformers) is a language representation model designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. Unlike prior architectures that used unidirectional language models or shallow concatenations of left-to-right and right-to-left models, BERT uses a Masked Language Model (MLM) objective alongside a Next Sentence Prediction (NSP) task during pre-training. The pre-trained BERT model can be fine-tuned with just one additional output layer to construct state-of-the-art models for a broad range of natural language processing tasks without task-specific architecture modifications. Evaluated on eleven natural language processing benchmarks, BERT achieves new state-of-the-art performance, including pushing the GLUE score to 80.5%, MultiNLI accuracy to 86.7%, and SQuAD v1.1 question answering Test F1 to 93.2. Furthermore, ablation studies demonstrate that deep bidirectionality and extreme model scaling substantially improve fine-tuning performance across both sentence-level and token-level downstream tasks.
Key Takeaways
BERT utilizes a Masked Language Model (MLM) pre-training objective, masking 15% of WordPiece tokens at random to learn deep bidirectional contextual representations.
During pre-training, BERT also incorporates a Next Sentence Prediction (NSP) binary classification task, reaching 97%-98% accuracy in predicting whether sentence B follows sentence A.
BERT obtains an average GLUE benchmark score of 80.5% with BERT-Large, delivering a 7.7% absolute improvement over previous state-of-the-art models.
On SQuAD v1.1, single-model BERT-Large achieves 91.8 Test F1, while an ensemble augmented with TriviaQA achieves 93.2 Test F1, outperforming human performance.
Scaling model size from BERT-Base (110M parameters) to BERT-Large (340M parameters) provides continuous accuracy gains, even on small-scale datasets like MRPC with only 3,600 training examples.
Learning Objectives
Understand how the Masked Language Model (MLM) objective enables deep bidirectional Transformer pre-training.
Differentiate between fine-tuning and feature-based transfer learning approaches using pre-trained representations.
Analyze the impact of model scaling and joint pre-training tasks on downstream NLP benchmark performance.
Evaluate BERT's empirical performance across GLUE, SQuAD, and SWAG benchmarks.
Glossary
BERT
Bidirectional Encoder Representations from Transformers, a language representation model pre-trained on unlabeled text using deep bidirectional representations.
Masked Language Model (MLM)
A pre-training objective where a percentage of input tokens are randomly masked and predicted based on surrounding left and right context.
Next Sentence Prediction (NSP)
A pre-training task where the model predicts whether a second sentence logically follows the first sentence in a document pair.
Transformer
A multi-layer neural network architecture based on self-attention mechanisms, used as the foundational building block of BERT.
Fine-tuning
A transfer learning approach where pre-trained model parameters are adjusted on a downstream task using labeled data and minimal task-specific output layers.
GLUE
General Language Understanding Evaluation benchmark, a collection of diverse natural language understanding tasks used to evaluate model performance.
WordPiece Embeddings
A tokenization method with a 30,000 token vocabulary used by BERT to convert text into sub-word token sequences.
Mind Map
Everything is expanded by default. Use the − buttons to collapse a branch, or the controls below.
BERT Language Model
Pre-training Objectives
Masked Language Model
Next Sentence Prediction
Model Architectures
BERT-Base (110M Params)
BERT-Large (340M Params)
Downstream Applications
GLUE Benchmark
SQuAD Question Answering
BERT Model Breakthroughs and Benchmarks
Deep Bidirectional Transformers Revolutionize Natural Language Processing
award
80.5%
GLUE Benchmark Score
check-circle
93.2
SQuAD v1.1 Test F1 Score
file-text
83.1
SQuAD v2.0 Test F1 Score
cpu
340M
Parameters in BERT-Large Model
percent
15%
Random Token Masking Rate
bar-chart
86.7%
MultiNLI Classification Accuracy
Deep Bidirectionality
Jointly conditions on both left and right context across all Transformer layers using a Masked Language Model objective.
Unified Fine-Tuning
Eliminates task-specific architecture changes by modifying only a single output layer for downstream applications.
Extreme Model Scaling
Scaling from 110M to 340M parameters yields continuous performance gains even on small downstream datasets.
Why is BERT's bidirectional context superior to traditional unidirectional language models?
Traditional language models process text either left-to-right or right-to-left, which limits their ability to integrate future context when understanding words. BERT's Masked Language Model objective allows every token to attend to both left and right context simultaneously across all Transformer layers, resulting in richer contextual representations for downstream NLP tasks.
How does fine-tuning BERT differ from traditional feature-based transfer learning?
In feature-based approaches like ELMo, pre-trained embeddings are kept frozen and fed as input features into complex, task-specific architectures. In fine-tuning with BERT, almost all pre-trained parameters are updated end-to-end on downstream tasks with only a single minimal output layer added, making training faster and architectures simpler.
How does BERT represent two input sentences in a single sequence?
BERT concatenates the two sentences into a single sequence, separated by a special [SEP] token and starting with a [CLS] token. It then adds learned segment embeddings (Segment A and Segment B) along with position embeddings to differentiate tokens belonging to each sentence.
What datasets were used to pre-train BERT?
BERT was pre-trained on two large monolingual text corpora: BooksCorpus (800 million words) and English Wikipedia (2,500 million words, extracting text passages while ignoring tables and headers).
References
Attention is all you need
Deep contextualized word representations
Improving language understanding with unsupervised learning
SQuAD: 100,000+ questions for machine comprehension of text