Thursday, October 8, 2026
HomeData ScienceLinear Algebra for Data Scientists – Matrices, Eigenvalues and SVD Explained

Linear Algebra for Data Scientists – Matrices, Eigenvalues and SVD Explained

Table of Content

📋 KEY INSIGHTS

  • Linear algebra is the mathematical language of machine learning — neural networks are sequences of matrix multiplications and non-linearities; PCA is eigendecomposition of a covariance matrix; linear regression is a matrix equation solved by least squares; and word embeddings are vectors in a high-dimensional space.
  • The matrix multiplication A × B produces a transformation: it rotates, scales, and shears vectors in a way determined by A. Understanding matrix multiplication geometrically — as a composition of linear transformations — is more useful than memorising the element-wise formula.
  • Eigenvalues and eigenvectors reveal the “invariant structure” of a linear transformation: eigenvectors are the directions that are only scaled (not rotated) by the transformation, and eigenvalues are the scale factors. PCA is entirely based on this concept — it finds the directions of maximum variance in data.
  • Singular Value Decomposition (SVD) is the most practically important matrix factorisation in data science — it underlies PCA, latent semantic analysis, collaborative filtering, image compression, and the low-rank approximations used in LoRA fine-tuning of LLMs.
  • The dot product of two vectors measures their similarity: it is maximised when vectors point in the same direction and is zero when they are perpendicular. This is why cosine similarity (the normalised dot product) is the standard metric in recommendation systems and embedding-based retrieval.
  • Matrix rank tells you the true dimensionality of the information in a dataset — a rank-k matrix can be exactly represented as the product of two much smaller matrices. Low-rank structure is ubiquitous in real data (images, text, user-item matrices) and is exploited by every matrix factorisation method.

Linear algebra appears at the foundation of almost every machine learning algorithm, yet it is frequently taught in a way that emphasises computation over intuition — pages of matrix arithmetic without an explanation of what any of it means geometrically or why it matters for machine learning. This guide takes the opposite approach: it builds geometric intuition first, explains the connection to machine learning throughout, and covers the computations only in service of understanding. By the end, you will understand why PCA is eigendecomposition, why word embeddings use dot products for similarity, what SVD has to do with recommendation systems, and why neural networks are fundamentally a composition of matrix operations. This is not a complete linear algebra course — it is a focused guide to the concepts that appear most frequently in ML and that are most commonly tested in data science interviews.

Vectors, Matrices and the Geometry of Transformations

A vector is a directed quantity with both magnitude and direction — in data science, a vector represents an observation (a row in a dataset) or a learned representation (a word embedding, a node embedding, a latent code). In n dimensions, a vector is a list of n real numbers [x₁, x₂, …, xₙ]. The length (magnitude) of a vector is computed as the Euclidean norm: ||v|| = √(x₁² + x₂² + … + xₙ²). In high-dimensional spaces, all vectors of unit length lie on the surface of the unit hypersphere — a geometric fact that underlies the behaviour of cosine similarity in embedding spaces.

The dot product of two vectors u and v is defined as u · v = u₁v₁ + u₂v₂ + … + uₙvₙ. Geometrically, it equals ||u|| × ||v|| × cos(θ), where θ is the angle between the vectors. When the vectors are unit-length (||u|| = ||v|| = 1), the dot product equals cos(θ) — the cosine similarity. This geometric interpretation is why dot products measure similarity: two vectors pointing in the same direction have dot product equal to the product of their magnitudes (maximum similarity); two perpendicular vectors have dot product zero (no similarity); two anti-parallel vectors have negative dot product (opposite). Embedding models (word2vec, GloVe, sentence transformers) are trained to make semantically similar items have high cosine similarity by construction.

A matrix is a rectangular array of numbers that can be interpreted as a linear transformation of space. Multiplying a vector v by a matrix A produces a new vector Av — the vector v transformed by A. The key insight is that matrix multiplication is a composition of geometric operations: rotation, reflection, scaling, shearing, and projection. The identity matrix I leaves vectors unchanged (Iv = v). A diagonal matrix scales each coordinate independently. A rotation matrix rotates vectors without changing their length. Understanding which transformation a matrix represents is more useful than computing its product element by element.

OperationDefinitionGeometric MeaningML Application
Dot product u·vΣ uᵢvᵢ = ||u||||v||cos(θ)Projection of u onto v (scaled)Cosine similarity, attention scores
Matrix multiplication AB(AB)ᵢⱼ = Σₖ AᵢₖBₖⱼComposition of two transformationsNeural network forward pass
Transpose Aᵀ(Aᵀ)ᵢⱼ = AⱼᵢReflection across the main diagonalGram matrix, covariance (XᵀX)
Matrix inverse A⁻¹AA⁻¹ = IUndoing the transformation AOLS: β = (XᵀX)⁻¹Xᵀy
Outer product uvᵀ(uvᵀ)ᵢⱼ = uᵢvⱼRank-1 matrix from two vectorsLoRA: ΔW = AB where A, B are low-rank
Hadamard product A⊙B(A⊙B)ᵢⱼ = AᵢⱼBᵢⱼElement-wise scalingAttention masking, dropout

Eigenvalues and Eigenvectors — The Invariant Structure of Transformations

Diagram illustrates various forces and vectors
Photo by Bozhin Karaivanov on Unsplash

For a square matrix A, an eigenvector v is a non-zero vector that, when transformed by A, only changes in length (not direction): Av = λv. The scalar λ is the corresponding eigenvalue — the factor by which the eigenvector is scaled. Eigenvectors reveal the “natural axes” of a transformation — the directions along which the transformation acts purely as a stretch or compression, without rotation.

The connection to PCA is direct. Given a dataset X (n samples × d features), the covariance matrix C = (1/n)XᵀX is a d×d symmetric matrix. The eigenvectors of C are the principal components — the directions of maximum variance in the data. The eigenvalues are the variances along each principal component — larger eigenvalue means more variance is captured by that direction. PCA consists of: (1) computing the covariance matrix; (2) finding its eigenvectors and eigenvalues; (3) sorting eigenvectors by eigenvalue (descending); (4) projecting the data onto the top-k eigenvectors to obtain a k-dimensional representation. This projection is optimal in the sense that it retains more variance than any other k-dimensional linear projection.

The spectral theorem guarantees that any real symmetric matrix (like a covariance matrix) has real eigenvalues and orthogonal eigenvectors. This is why PCA principal components are always orthogonal to each other — they point in completely independent (uncorrelated) directions in the original feature space. For non-symmetric matrices (like weight matrices in neural networks), eigenvalues can be complex, and eigenvectors need not be orthogonal — which is why SVD, not eigendecomposition, is used for general matrix factorisation.

Matrix TypeEigenvalue PropertiesEigenvector PropertiesExample in ML
Symmetric (Aᵀ = A)All realOrthogonal to each otherCovariance matrix, Gram matrix
Positive Definite (xᵀAx > 0)All positive realOrthogonalValid covariance matrix, Hessian at minimum
Orthogonal (AᵀA = I)Magnitude 1 (complex)OrthonormalRotation matrices, Q in QR decomposition
DiagonalDiagonal entriesStandard basis vectorsScaled space; feature standardisation
General rectangularN/A (not square)N/AUse SVD instead of eigendecomposition

Singular Value Decomposition (SVD) — The Universal Matrix Factorisation

SVD is the generalisation of eigendecomposition to arbitrary (including non-square) matrices. Every real matrix A of shape m×n can be factorised as A = UΣVᵀ, where U is an m×m orthogonal matrix (the left singular vectors), Σ is an m×n diagonal matrix of singular values (non-negative real numbers in descending order), and Vᵀ is an n×n orthogonal matrix (the right singular vectors). The singular values in Σ measure the “importance” of each rank-1 component — the first singular value captures the most structure, the second the second-most, and so on.

The most powerful application of SVD is low-rank approximation. The best rank-k approximation to A (in the sense of minimising the Frobenius norm of the approximation error) is obtained by keeping only the top-k singular values: Aₖ = Uₖ Σₖ Vₖᵀ. This is the Eckart-Young theorem, and it is the mathematical foundation for PCA, latent semantic analysis, collaborative filtering (matrix factorisation for recommendation), image compression, and — most recently — LoRA (Low-Rank Adaptation of LLMs), where the weight update matrix ΔW is constrained to be the product of two low-rank matrices, dramatically reducing the number of trainable parameters.

The relationship between SVD and PCA is exact: the principal components of XᵀX are the right singular vectors of X, and the eigenvalues of XᵀX are the squares of the singular values of X. In practice, PCA is always computed using SVD rather than by explicitly forming and eigen-decomposing the covariance matrix, because SVD is more numerically stable (especially when n < d, the “wide data” regime) and is available for non-square matrices without modification.

ApplicationHow SVD Is UsedWhat the Components Represent
PCASVD of data matrix XV columns = principal components; Σ² = variances
Latent Semantic AnalysisSVD of term-document matrixU = topic-term; V = topic-document; Σ = topic importance
Collaborative FilteringSVD of user-item rating matrixU = user latent factors; V = item latent factors
Image CompressionRank-k SVD of pixel matrixTop k components capture most visual structure
LoRA (LLM Fine-tuning)ΔW ≈ AB (low-rank product)A and B replace full ΔW; only A, B are trained
Noise ReductionTruncate small singular valuesSmall σ = noise; large σ = signal

✦ SUMMARIZE THIS ARTICLE WITH AI

PCA and other dimensionality reduction techniques that apply eigendecomposition and SVD in practice are covered in depth in our Dimensionality Reduction guide. The neural network architectures that implement linear algebra operations as computational graphs are explained in our Neural Network Architectures guide. Statistics foundations — covariance, correlation, and sampling distributions — that complement linear algebra in data science are in our Statistics Fundamentals guide. The SVD-based matrix factorisation methods used in recommendation systems are in our Recommendation Systems guide.

Leave feedback about this

  • Rating

Durgesh Kekare
Durgesh Kekarehttps://www.dataexpertise.in
Durgesh Kekare is a data science educator and founder of DataExpertise.in. With expertise in Python, machine learning, and analytics, he helps 10,000+ learners break into data careers.

Latest Posts

List of Categories