📋 KEY INSIGHTS
- Word embeddings represent words as dense vectors in a continuous space where semantically similar words are geometrically close — enabling arithmetic on meaning: “king” − “man” + “woman” ≈ “queen”. This geometric encoding of semantics is the foundation for every modern NLP system.
- Static embeddings (word2vec, GloVe, FastText) assign one fixed vector per word regardless of context — “bank” has the same vector in “river bank” and “central bank.” Contextual embeddings (BERT, GPT) assign different vectors depending on surrounding context, which is why they outperform static embeddings on most tasks.
- Sentence embeddings — dense vector representations of entire sentences or paragraphs — are the enabling technology for semantic search, RAG (Retrieval-Augmented Generation), and document similarity systems. Models like Sentence-BERT produce sentence embeddings that can be compared by cosine similarity.
- Vector databases (Qdrant, Weaviate, Pinecone, pgvector) are purpose-built for approximate nearest-neighbour (ANN) search in high-dimensional embedding spaces — they can find the most similar vectors among hundreds of millions in milliseconds using algorithms like HNSW.
- The quality of a RAG pipeline is determined primarily by retrieval quality — if the relevant chunks are not retrieved, the generator cannot produce a correct answer regardless of its capability. Evaluating retrieval precision and recall separately from generation quality is essential.
- Fine-tuning embedding models on domain-specific data (using contrastive learning objectives like InfoNCE or triplet loss) substantially improves retrieval performance for specialised domains — a general-purpose embedding model trained on web text will underperform a domain-tuned model for legal, medical, or financial document retrieval.
The ability to represent meaning as geometry — to place words, sentences, and documents at positions in a vector space such that similar meanings are close together — is one of the most powerful ideas in modern NLP. Word embeddings made it possible to do arithmetic on meaning and to transfer knowledge learned from large corpora to downstream tasks. Sentence embeddings made it possible to find semantically similar documents, build knowledge-aware chatbots, and power the retrieval component of modern AI systems. Understanding how embeddings work, how they are trained, and how they are deployed in vector search systems is now a core skill for data scientists working on any text-related application. This guide covers static and contextual word embeddings, sentence embeddings, dense retrieval architectures, and the vector database infrastructure that enables embedding-based search at scale.
Static Word Embeddings — word2vec, GloVe and FastText
The central insight behind word embeddings is the distributional hypothesis: words that appear in similar contexts have similar meanings. Word2vec (Mikolov et al., 2013) operationalises this by training a shallow neural network to predict a target word from its context (CBOW — Continuous Bag of Words) or to predict context words from a target word (Skip-gram). The learned weights of the hidden layer become the word vectors. Because the network is trained to capture co-occurrence patterns, the resulting vectors encode semantic and syntactic relationships geometrically: the direction from “man” to “woman” encodes the gender relationship, and the same direction from “king” gives “queen.” These relationships emerge from the training objective — they are not explicitly supervised.
Word2vec’s training is closely related to matrix factorisation: Levy and Goldberg (2014) showed that the Skip-gram model with negative sampling implicitly factorises a shifted pointwise mutual information (PMI) matrix of word co-occurrences. GloVe (Global Vectors, Pennington et al., 2014) makes this factorisation explicit: it directly trains vectors to approximate the logarithm of word co-occurrence probabilities across the entire corpus, achieving comparable performance to word2vec with more interpretable training. The linear algebra of SVD and matrix factorisation is directly relevant to understanding GloVe’s optimisation.
FastText (Bojanowski et al., 2017) extends word2vec by representing each word as a bag of character n-grams and computing the word vector as the sum of its n-gram vectors. This gives FastText two important advantages over word2vec and GloVe: it handles out-of-vocabulary words (any word can be represented from its character n-grams, even if it was not in the training corpus), and it works well for morphologically rich languages where word forms carry systematic meaning (German compound words, Turkish suffixes, Hindi conjugations). FastText embeddings trained on 294 languages are available as a multilingual baseline for cross-lingual NLP tasks — a useful starting point before the full NLP pipeline is set up.
| Model | Training Objective | OOV Handling | Context-Aware? | Best For |
|---|---|---|---|---|
| word2vec (CBOW) | Predict target from context window | No | No (one vector per word) | General semantic tasks, fast training |
| word2vec (Skip-gram) | Predict context from target word | No | No | Rare words, larger vocabulary |
| GloVe | Factorize log co-occurrence matrix | No | No | Interpretable geometry, analogies |
| FastText | Skip-gram + character n-grams | Yes (n-gram decomposition) | No | Morphologically rich languages, OOV |
| BERT (contextual) | Masked language model | Yes (subword tokenisation) | Yes (per-token, context-dependent) | Downstream fine-tuning, NER, QA |
| Sentence-BERT | Siamese network, contrastive loss | Yes | Yes | Sentence similarity, semantic search |
Contextual Embeddings — BERT and Beyond
The fundamental limitation of static embeddings is polysemy: a word like “bank,” “match,” “lead,” or “run” has multiple distinct meanings, but a static embedding assigns it a single vector that is a blurred average across all uses. Contextual embeddings solve this by producing different representations for the same word depending on its surrounding context — the vector for “bank” in “river bank” is different from the vector in “central bank.”
BERT (Devlin et al., 2018) produces contextual word-level embeddings through bidirectional Transformer encoding. Pre-trained on masked language modelling (predict randomly masked tokens) and next-sentence prediction, BERT’s internal representations encode rich semantic and syntactic information that can be extracted and fine-tuned for downstream tasks. BERT-based models dominate NLP applications including sentiment analysis, named entity recognition, and question answering. The architecture is explained in depth in our Transformers and Attention guide. Fine-tuning BERT for a specific task using a small labelled dataset is the standard approach for achieving high performance with limited annotation effort.
For semantic search and retrieval, standard BERT embeddings are not optimal — BERT was trained for understanding, not for producing sentence-level similarity representations. Sentence-BERT (Reimers and Gurevych, 2019) addresses this by fine-tuning BERT with a Siamese network architecture using a natural language inference dataset: pairs of semantically equivalent sentences are trained to produce similar embeddings (high cosine similarity), while dissimilar pairs are pushed apart. The resulting sentence embeddings can be compared efficiently using cosine similarity, enabling fast semantic search. SBERT and its successors (BGE, E5, GTE) are the standard embedding models for RAG pipelines and document retrieval systems.
Vector Databases and Semantic Search at Scale
Once you have sentence embeddings for a corpus of documents, the challenge becomes finding the most similar embeddings to a query embedding quickly at scale. Exact nearest-neighbour search requires computing the distance between the query and every document embedding — O(n) operations per query, which is 10 seconds for 10 million documents at typical embedding similarity computation speeds. This is too slow for interactive applications. Approximate Nearest Neighbour (ANN) algorithms find near-optimal results in O(log n) or even O(1) time, with a controllable trade-off between speed and recall.
HNSW (Hierarchical Navigable Small World graphs, Malkov and Yashunin, 2018) is the dominant ANN algorithm in production vector databases. It builds a multi-layer graph where each node connects to its nearest neighbours in the embedding space, with higher layers providing coarser-grained long-range connections for fast navigation and lower layers providing fine-grained local search. Query time is O(log n) with very high recall (typically 95–99% of true nearest neighbours retrieved). HNSW is the default index in Qdrant, Weaviate, and Milvus. The trade-off is memory: HNSW requires storing the graph structure alongside the vectors, typically 2–4× the raw vector memory.
The recommendation system and generative AI RAG pipeline use cases are the two most important production applications of vector search. In recommendation, user and item embeddings are compared to find the most relevant items for each user. In RAG, document chunk embeddings are searched to find the passages most relevant to a user’s question, which are then passed as context to the LLM generator. The evaluation of retrieval quality uses standard information retrieval metrics: precision@k (fraction of top-k retrieved chunks that are relevant), recall@k (fraction of all relevant chunks that appear in the top-k), and MRR (Mean Reciprocal Rank). The model evaluation framework for retrieval systems differs from classification metrics and should be set up before any optimisation is attempted.
| Vector DB | Index Type | Hosting | Filtering | Best For |
|---|---|---|---|---|
| Qdrant | HNSW | Self-hosted + cloud | Payload filter (fast) | Performance, open-source, production |
| Weaviate | HNSW | Self-hosted + cloud | GraphQL filter | Hybrid search, knowledge graphs |
| Pinecone | Proprietary ANN | Managed only | Metadata filter | Fastest time-to-value, serverless |
| pgvector | IVFFlat / HNSW | PostgreSQL extension | Full SQL | Existing PostgreSQL stack, small scale |
| Milvus | HNSW + IVF + DiskANN | Self-hosted + cloud | Attribute filter | Very large scale (billions of vectors) |
| Chroma | HNSW | Self-hosted / embedded | Metadata filter | Local prototyping, notebooks |
✦ SUMMARIZE THIS ARTICLE WITH AI
The NLP preprocessing pipeline — tokenisation, subword encoding, stopword removal — that feeds into embedding models is covered in our NLP Pipeline guide. The Transformer architecture that produces contextual embeddings is explained in our Transformers and Attention guide. Building complete RAG pipelines using sentence embeddings and vector databases is covered in our Generative AI guide. The NLP applications — sentiment analysis, NER, text classification — that use these embeddings as inputs are in our NLP Applications guide. The dimensionality reduction techniques (PCA, t-SNE, UMAP) used to visualise embedding spaces are in our Dimensionality Reduction guide.



