📋 KEY INSIGHTS
- The Transformer architecture (Vaswani et al., 2017) replaced RNNs for sequence modelling by using self-attention โ every token can directly attend to every other token, eliminating vanishing gradients and enabling full parallelisation during training.
- Self-attention computes three matrices โ Query, Key, Value โ and scores each token pair as Softmax(QK^T / sqrt(d_k)) V, allowing the model to learn which tokens are most relevant to each other.
- BERT is encoder-only, pre-trained with Masked Language Modelling and Next Sentence Prediction โ it excels at classification, NER, and question answering. GPT is decoder-only, trained with next-token prediction โ it excels at generation.
- Positional encodings (sinusoidal or learned) are added to token embeddings because self-attention is permutation-invariant โ without them, the model has no sense of word order.
- Fine-tuning a pre-trained transformer on a domain-specific dataset requires only a small labelled dataset (hundreds to thousands of examples) and typically takes minutes on a GPU.
- The primary bottleneck of Transformers is quadratic memory complexity O(nยฒ) with sequence length n โ efficient attention variants (FlashAttention, Longformer, BigBird) address this for long documents.
The Transformer architecture introduced in “Attention Is All You Need” (Vaswani et al., 2017) did not just improve NLP โ it fundamentally replaced the dominant paradigm. Before Transformers, recurrent networks (LSTM, GRU) processed sequences token by token, creating two crippling limitations: sequential processing prevented parallelisation, making training slow, and the vanishing gradient problem made it hard to learn dependencies across long distances. The Transformer solved both problems at once with self-attention, enabling models to be trained on massive corpora in parallel. Today, Transformers power not just language models (BERT, GPT-4, Gemini) but also computer vision (ViT), protein structure prediction (AlphaFold2), and code generation (Codex, GitHub Copilot). Understanding how they work from first principles is essential for any data scientist working with modern ML.
Self-Attention โ The Core Mechanism
Self-attention allows each position in a sequence to attend to all other positions in the same sequence, computing a weighted representation of the entire context for every token. Given an input sequence of token embeddings X (shape: n ร d_model), self-attention projects X into three matrices using learned weight matrices W_Q, W_K, W_V:
Q = X W_Q (Queries โ “what am I looking for?”)
K = X W_K (Keys โ “what do I contain?”)
V = X W_V (Values โ “what do I provide if attended to?”)
The attention output for each token is: Attention(Q, K, V) = Softmax(QK^T / โd_k) ยท V. The dot product QK^T gives a raw compatibility score between every query-key pair โ how much each token should attend to every other. Dividing by โd_k (the key dimension) stabilises gradients when d_k is large (otherwise softmax saturates and gradients vanish). Softmax normalises scores into a probability distribution. The output is a weighted sum of Value vectors โ for each token, a context-aware representation blending information from all relevant positions.
Multi-Head Attention runs h parallel attention heads, each with its own W_Q, W_K, W_V, then concatenates and projects the results. This allows different heads to specialise โ one might capture syntactic dependencies, another semantic relationships, another coreference. In BERT-base, h=12, d_k=64, d_model=768 (12ร64=768).
import torch
import torch.nn as nn
import torch.nn.functional as F
import math
class MultiHeadSelfAttention(nn.Module):
'''
Multi-head self-attention from scratch.
d_model: total embedding dimension
n_heads: number of parallel attention heads
'''
def __init__(self, d_model=512, n_heads=8, dropout=0.1):
super().__init__()
assert d_model % n_heads == 0, 'd_model must be divisible by n_heads'
self.d_model = d_model
self.n_heads = n_heads
self.d_k = d_model // n_heads # dimension per head
# Projection matrices for Q, K, V, and output
self.W_q = nn.Linear(d_model, d_model, bias=False)
self.W_k = nn.Linear(d_model, d_model, bias=False)
self.W_v = nn.Linear(d_model, d_model, bias=False)
self.W_o = nn.Linear(d_model, d_model)
self.dropout = nn.Dropout(dropout)
def split_heads(self, x, batch_size):
# x: (batch, seq_len, d_model) -> (batch, n_heads, seq_len, d_k)
x = x.view(batch_size, -1, self.n_heads, self.d_k)
return x.transpose(1, 2)
def forward(self, x, mask=None):
batch_size, seq_len, _ = x.size()
# Project to Q, K, V and split into heads
Q = self.split_heads(self.W_q(x), batch_size) # (B, H, T, d_k)
K = self.split_heads(self.W_k(x), batch_size)
V = self.split_heads(self.W_v(x), batch_size)
# Scaled dot-product attention
scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(self.d_k)
# scores: (B, H, T, T) โ every token attends to every token
if mask is not None:
# Causal mask: set future positions to -inf (decoder only)
scores = scores.masked_fill(mask == 0, float('-inf'))
attn_weights = F.softmax(scores, dim=-1) # (B, H, T, T)
attn_weights = self.dropout(attn_weights)
# Weighted sum of values
context = torch.matmul(attn_weights, V) # (B, H, T, d_k)
# Concatenate heads and project
context = context.transpose(1, 2).contiguous().view(batch_size, seq_len, self.d_model)
return self.W_o(context), attn_weights
# Test forward pass
mha = MultiHeadSelfAttention(d_model=512, n_heads=8)
x_in = torch.randn(4, 20, 512) # batch=4, seq_len=20, d_model=512
out, attn = mha(x_in)
print('Input shape: ', x_in.shape) # (4, 20, 512)
print('Output shape: ', out.shape) # (4, 20, 512)
print('Attn weights: ', attn.shape) # (4, 8, 20, 20)
class TransformerEncoderLayer(nn.Module):
'''One Transformer encoder block: self-attention + FFN with residual + LayerNorm.'''
def __init__(self, d_model=512, n_heads=8, d_ff=2048, dropout=0.1):
super().__init__()
self.self_attn = MultiHeadSelfAttention(d_model, n_heads, dropout)
self.ffn = nn.Sequential(
nn.Linear(d_model, d_ff),
nn.GELU(), # BERT uses GELU, original Transformer used ReLU
nn.Dropout(dropout),
nn.Linear(d_ff, d_model),
)
self.norm1 = nn.LayerNorm(d_model)
self.norm2 = nn.LayerNorm(d_model)
self.dropout = nn.Dropout(dropout)
def forward(self, x, mask=None):
# Pre-norm variant (more stable training than original post-norm)
attn_out, _ = self.self_attn(self.norm1(x), mask)
x = x + self.dropout(attn_out) # residual connection
x = x + self.dropout(self.ffn(self.norm2(x)))
return x
layer = TransformerEncoderLayer(d_model=512, n_heads=8, d_ff=2048)
out = layer(x_in)
print('Encoder layer output:', out.shape) # (4, 20, 512)
BERT vs GPT โ Architecture Differences and When to Use Each
Both BERT and GPT are built from Transformer blocks, but their training objectives and architecture choices make them suited to fundamentally different tasks. Understanding the distinction is one of the most common interview topics in applied NLP.
BERT (Bidirectional Encoder Representations from Transformers): Uses only the Transformer encoder. Pre-training tasks: (1) Masked Language Modelling (MLM) โ randomly mask 15% of input tokens and predict them from context; (2) Next Sentence Prediction (NSP) โ predict whether sentence B follows sentence A. Because BERT sees the full context in both directions simultaneously, it builds rich bidirectional representations. Best for: text classification, named entity recognition, question answering, semantic similarity, any task where you need to understand a piece of text.
GPT (Generative Pre-trained Transformer): Uses only the Transformer decoder with a causal (autoregressive) attention mask โ each token can only attend to previous tokens, never future ones. Pre-training task: next-token prediction on a massive text corpus. Because the model must predict what comes next, it learns deep language generation capability. Best for: text generation, summarisation, translation, code generation, few-shot prompting, and instruction-following with RLHF.
| Property | BERT (Encoder-only) | GPT (Decoder-only) | T5 / BART (Encoder-Decoder) |
|---|---|---|---|
| Attention type | Bidirectional (full) | Causal (left-to-right) | Both |
| Pre-training task | MLM + NSP | Next-token prediction | Text-to-text denoising |
| Best for | Classification, NER, QA | Generation, completion | Translation, summarisation |
| Input format | [CLS] text [SEP] | prompt tokens | encoder: input, decoder: target |
| Notable models | BERT, RoBERTa, DistilBERT | GPT-2/3/4, Llama, Mistral | T5, BART, mT5 |
Fine-Tuning BERT for Text Classification
Fine-tuning a pre-trained Transformer takes the model’s learned language representations and adapts them to a specific task with a small labelled dataset. For classification, you add a linear head on top of the [CLS] token’s output representation and train end-to-end with a lower learning rate (2e-5 to 5e-5 โ too high and you catastrophically forget pre-training knowledge). With HuggingFace Transformers, this takes fewer than 20 lines of meaningful code.
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
Trainer, TrainingArguments)
from datasets import Dataset
import numpy as np
from sklearn.metrics import accuracy_score, f1_score
# Example: binary sentiment classification
texts = ['Great product, highly recommend!',
'Terrible experience, waste of money.',
'Average, nothing special.',
'Absolutely loved it, will buy again.',
'Poor quality, broke after a week.']
labels = [1, 0, 0, 1, 0] # 1=positive, 0=negative
model_name = 'distilbert-base-uncased' # lighter than BERT-base, 97% of performance
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(
model_name, num_labels=2)
def tokenize(batch):
return tokenizer(batch['text'], padding='max_length',
truncation=True, max_length=128)
dataset = Dataset.from_dict({'text': texts, 'label': labels})
dataset = dataset.map(tokenize, batched=True)
dataset = dataset.train_test_split(test_size=0.2, seed=42)
def compute_metrics(eval_pred):
logits, labels = eval_pred
preds = np.argmax(logits, axis=-1)
return {'accuracy': accuracy_score(labels, preds),
'f1': f1_score(labels, preds, average='weighted')}
training_args = TrainingArguments(
output_dir = './bert_sentiment',
num_train_epochs = 3,
per_device_train_batch_size = 8,
per_device_eval_batch_size = 16,
learning_rate = 2e-5,
weight_decay = 0.01,
evaluation_strategy = 'epoch',
save_strategy = 'epoch',
load_best_model_at_end = True,
metric_for_best_model = 'f1',
report_to = 'none', # set to 'wandb' for experiment tracking
)
trainer = Trainer(
model = model,
args = training_args,
train_dataset = dataset['train'],
eval_dataset = dataset['test'],
compute_metrics = compute_metrics,
)
trainer.train()
# Inference on new text
inputs = tokenizer('This is an outstanding purchase!',
return_tensors='pt', padding=True, truncation=True)
import torch
with torch.no_grad():
logits = model(**inputs).logits
pred_label = torch.argmax(logits, dim=1).item()
print('Predicted:', 'Positive' if pred_label == 1 else 'Negative')
Interview Q&A โ Transformers and BERT
Q: Why does self-attention have O(nยฒ) complexity and how is it addressed?
Each of the n tokens attends to all other n tokens, producing an nรn attention matrix. Memory is O(nยฒยทd_k) and computation is O(nยฒยทd_k) per layer โ for n=512 this is manageable but for n=100,000 (long documents) it becomes prohibitive. Solutions: FlashAttention (IO-aware recomputation, same complexity but much faster in practice), Longformer/BigBird (sparse local + global attention, O(n)), Linformer (low-rank approximation), and Performer (kernel-based approximation).
Q: What is positional encoding and why is it needed?
Self-attention is permutation-invariant โ if you shuffle the input tokens, the attention scores change but the model architecture has no inherent way to distinguish “word at position 3” from “word at position 7”. Positional encodings add position information to token embeddings. The original Transformer used sinusoidal encodings (fixed formulas using sin/cos at different frequencies for each dimension). BERT uses learned position embeddings. Modern LLMs use RoPE (Rotary Position Embedding) or ALiBi, which encode relative rather than absolute position.
Q: What is the difference between BERT and RoBERTa?
RoBERTa (Robustly Optimised BERT Pretraining Approach) made three key changes: trained for longer with larger batches on more data (160GB vs 16GB), removed the Next Sentence Prediction task (found it hurt downstream performance), and used dynamic masking (new mask pattern each epoch rather than static). The result was consistently better downstream task performance with the same architecture. RoBERTa is generally preferred over BERT as a starting point for fine-tuning.
The broader NLP pipeline context โ tokenisation, preprocessing, word embeddings โ that feeds into Transformers is covered in our NLP Pipeline guide. For interview questions covering BERT, attention, and NLP system design, our NLP Interview Q&A has 40+ questions. The transfer learning principles behind fine-tuning are in our Transfer Learning guide. The neural network architectures (RNN, LSTM) that Transformers replaced are compared in our Neural Network Architectures guide.



