📋 KEY INSIGHTS
- Anomaly detection (also called outlier detection) is an unsupervised or semi-supervised task โ labels for anomalies are rare or nonexistent in most real-world datasets like fraud, sensor faults, and network intrusion.
- Isolation Forest works by randomly partitioning the feature space with trees: anomalies are isolated in fewer splits than normal points because they occupy sparse, low-density regions.
- Statistical methods (Z-score, IQR, Mahalanobis distance) are fast and interpretable but assume specific distributions and fail in high-dimensional feature spaces.
- Autoencoders detect anomalies through reconstruction error โ a model trained only on normal data struggles to reconstruct anomalous inputs, so high reconstruction error signals an outlier.
- The threshold for flagging anomalies is a business decision, not a statistical one โ it is set by balancing precision (avoiding false alarms) and recall (catching real anomalies).
- Multivariate anomaly detection is significantly harder than univariate: a sensor reading that is normal in isolation can be anomalous given the readings of correlated sensors.
Anomaly detection is one of the most practically valuable โ and statistically tricky โ problems in data science. Unlike classification, you almost never have a balanced labelled dataset: fraud events are 0.01% of transactions, equipment failures are rare by definition, and network intrusions are vastly outnumbered by normal traffic. The core challenge is learning a model of “normal” from predominantly normal data, then flagging observations that deviate significantly. This guide covers the full toolkit: statistical methods, tree-based isolation, density-based approaches, and deep learning with autoencoders โ with Python code and guidance on setting thresholds for production deployment.
Statistical Methods โ Z-Score, IQR, and Mahalanobis Distance
Statistical methods are the right starting point for anomaly detection: they are fast, interpretable, and require no training data. They work well for low-dimensional, approximately Gaussian data. Z-score flags points more than k standard deviations from the mean. IQR (Interquartile Range) method flags points below Q1 โ 1.5ยทIQR or above Q3 + 1.5ยทIQR โ more robust than Z-score for skewed distributions. Mahalanobis distance generalises Z-score to multivariate data, accounting for correlations between features โ a point that is unusual in the joint distribution of features is flagged, even if each individual feature looks normal.
import numpy as np
import pandas as pd
from scipy import stats
from scipy.spatial.distance import mahalanobis
# Simulate transaction data: amount, time_of_day, merchant_category_code
np.random.seed(42)
n_normal = 5000
n_fraud = 50
df_normal = pd.DataFrame({
'amount': np.random.lognormal(mean=4.0, sigma=1.2, size=n_normal),
'hour': np.random.randint(6, 23, size=n_normal),
'mcc_code': np.random.randint(100, 200, size=n_normal),
})
df_fraud = pd.DataFrame({
'amount': np.random.lognormal(mean=7.5, sigma=0.5, size=n_fraud), # unusually large
'hour': np.random.randint(0, 4, size=n_fraud), # late night
'mcc_code': np.random.randint(500, 600, size=n_fraud), # different MCCs
})
df = pd.concat([df_normal, df_fraud], ignore_index=True)
df['is_fraud'] = [0]*n_normal + [1]*n_fraud
# --- 1. Z-score (univariate: amount only) ---
df['amount_zscore'] = np.abs(stats.zscore(df['amount']))
zscore_flags = df['amount_zscore'] > 3.0
print('Z-score anomalies detected:', zscore_flags.sum())
print('Fraud recall:', df.loc[zscore_flags, 'is_fraud'].sum() / n_fraud)
# --- 2. IQR method ---
Q1, Q3 = df['amount'].quantile(0.25), df['amount'].quantile(0.75)
IQR = Q3 - Q1
iqr_flags = (df['amount'] < Q1 - 3.0*IQR) | (df['amount'] > Q3 + 3.0*IQR)
print('
IQR anomalies detected:', iqr_flags.sum())
# --- 3. Mahalanobis distance (multivariate) ---
features = df[['amount', 'hour', 'mcc_code']].values
cov_matrix = np.cov(features.T)
cov_inv = np.linalg.inv(cov_matrix)
mean_vector = features.mean(axis=0)
maha_dists = [mahalanobis(row, mean_vector, cov_inv) for row in features]
df['maha_dist'] = maha_dists
# Chi-squared threshold: df=3 features, alpha=0.001
from scipy.stats import chi2
threshold = chi2.ppf(0.999, df=features.shape[1])
maha_flags = df['maha_dist'] > np.sqrt(threshold)
print('
Mahalanobis anomalies detected:', maha_flags.sum())
print('Fraud recall:', df.loc[maha_flags, 'is_fraud'].sum() / n_fraud)
Isolation Forest โ Tree-Based Anomaly Detection
Isolation Forest (Liu et al., 2008) is the most widely used algorithm for anomaly detection in tabular data. The core idea is elegant: anomalies are easier to isolate than normal points. Normal observations cluster together in dense regions and require many random splits to isolate. Anomalous observations sit in sparse regions and are isolated quickly. The algorithm builds an ensemble of random Isolation Trees, each built by: (1) randomly selecting a feature, (2) randomly selecting a split value between the feature’s min and max, (3) recursively splitting until each point is isolated. The anomaly score for a point is based on the average path length across all trees โ shorter paths = more anomalous.
from sklearn.ensemble import IsolationForest
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import classification_report, roc_auc_score
import matplotlib.pyplot as plt
features_cols = ['amount', 'hour', 'mcc_code']
X = df[features_cols].values
scaler = StandardScaler()
X_sc = scaler.fit_transform(X)
# Isolation Forest: contamination = expected fraction of anomalies
# If you don't know the fraction, use 'auto' (uses 0.1)
iso_forest = IsolationForest(
n_estimators = 200,
contamination = 0.01, # we expect ~1% fraud
max_samples = 'auto',
random_state = 42,
n_jobs = -1
)
iso_forest.fit(X_sc)
# Predict: -1 = anomaly, 1 = normal
preds = iso_forest.predict(X_sc)
scores = iso_forest.score_samples(X_sc) # lower = more anomalous
df['if_pred'] = (preds == -1).astype(int) # 1 = flagged as anomaly
df['if_score'] = -scores # negate: higher = more anomalous
print('Isolation Forest โ Results:')
print(classification_report(df['is_fraud'], df['if_pred'],
target_names=['Normal', 'Anomaly']))
print('ROC-AUC:', round(roc_auc_score(df['is_fraud'], df['if_score']), 4))
# --- Visualise anomaly scores ---
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
axes[0].hist(df.loc[df['is_fraud']==0, 'if_score'], bins=50,
alpha=0.7, label='Normal', color='steelblue')
axes[0].hist(df.loc[df['is_fraud']==1, 'if_score'], bins=20,
alpha=0.7, label='Fraud', color='red')
axes[0].set_xlabel('Anomaly Score'); axes[0].set_title('Score Distribution')
axes[0].legend()
# Scatter: colour by anomaly score
sc = axes[1].scatter(df['amount'], df['hour'],
c=df['if_score'], cmap='RdYlBu_r',
alpha=0.4, s=10)
plt.colorbar(sc, ax=axes[1], label='Anomaly Score')
axes[1].set_xlabel('Amount'); axes[1].set_ylabel('Hour of Day')
axes[1].set_title('Anomaly Scores in Feature Space')
plt.tight_layout(); plt.show()
# --- One-Class SVM (alternative: better for non-linear boundaries) ---
from sklearn.svm import OneClassSVM
ocsvm = OneClassSVM(kernel='rbf', nu=0.01, gamma='scale')
ocsvm.fit(X_sc[df['is_fraud']==0]) # train on ONLY normal data
ocsvm_preds = (ocsvm.predict(X_sc) == -1).astype(int)
print('
One-Class SVM:')
print(classification_report(df['is_fraud'], ocsvm_preds,
target_names=['Normal', 'Anomaly']))
Autoencoder-Based Anomaly Detection
Autoencoders are neural networks trained to compress input data into a lower-dimensional latent representation and then reconstruct the original input. When trained exclusively on normal data, an autoencoder learns to reconstruct normal patterns well. When presented with anomalous data at inference time, the reconstruction fails โ yielding a high reconstruction error (MSE or MAE between input and output). The reconstruction error becomes the anomaly score: set a threshold on it, and flag any observation above the threshold as an anomaly.
Autoencoders excel over Isolation Forest when: (1) you have high-dimensional data (many features), (2) you have sufficient normal training data, (3) features have complex non-linear relationships that tree-based methods miss. They are the standard approach for anomaly detection on images (manufacturing defect detection) and time series (predictive maintenance).
import torch
import torch.nn as nn
class TabularAutoencoder(nn.Module):
'''Autoencoder for tabular anomaly detection.'''
def __init__(self, input_dim, latent_dim=8, hidden_dim=32, dropout=0.1):
super().__init__()
self.encoder = nn.Sequential(
nn.Linear(input_dim, hidden_dim),
nn.BatchNorm1d(hidden_dim),
nn.ReLU(),
nn.Dropout(dropout),
nn.Linear(hidden_dim, latent_dim),
nn.ReLU(),
)
self.decoder = nn.Sequential(
nn.Linear(latent_dim, hidden_dim),
nn.BatchNorm1d(hidden_dim),
nn.ReLU(),
nn.Dropout(dropout),
nn.Linear(hidden_dim, input_dim), # no activation: reconstruct raw values
)
def forward(self, x):
z = self.encoder(x)
recon = self.decoder(z)
return recon
def reconstruction_error(self, x):
'''Per-sample MSE reconstruction error (anomaly score).'''
with torch.no_grad():
recon = self.forward(x)
return ((x - recon) ** 2).mean(dim=1)
# Prepare data: train ONLY on normal samples
X_normal = X_sc[df['is_fraud'] == 0]
X_all = torch.FloatTensor(X_sc)
X_tr = torch.FloatTensor(X_normal)
model = TabularAutoencoder(input_dim=3, latent_dim=2, hidden_dim=16)
optimiser = torch.optim.Adam(model.parameters(), lr=1e-3, weight_decay=1e-5)
criterion = nn.MSELoss()
# Training loop
model.train()
for epoch in range(100):
optimiser.zero_grad()
recon = model(X_tr)
loss = criterion(recon, X_tr)
loss.backward()
optimiser.step()
if (epoch + 1) % 20 == 0:
print('Epoch ' + str(epoch+1) + '/100 Loss: ' + str(round(loss.item(), 5)))
# Anomaly scoring on all data
model.eval()
ae_scores = model.reconstruction_error(X_all).numpy()
df['ae_score'] = ae_scores
# Set threshold: 99th percentile of reconstruction errors on training data
ae_train_scores = model.reconstruction_error(X_tr).numpy()
threshold_ae = np.percentile(ae_train_scores, 99)
df['ae_pred'] = (df['ae_score'] > threshold_ae).astype(int)
print('
Autoencoder Results:')
print(classification_report(df['is_fraud'], df['ae_pred'],
target_names=['Normal', 'Anomaly']))
print('ROC-AUC:', round(roc_auc_score(df['is_fraud'], df['ae_score']), 4))
| Method | Best Scenario | Needs Labels? | Interpretable? | Scales to High-Dim? |
|---|---|---|---|---|
| Z-Score / IQR | Low-dim, Gaussian, quick checks | No | Yes | No |
| Mahalanobis Distance | Multivariate, correlated features | No | Partial | No (n < features) |
| Isolation Forest | General tabular data, scalable | No | Partial | Yes |
| One-Class SVM | Non-linear boundaries, small data | No (needs normal) | No | Medium |
| Autoencoder | High-dim, images, time series | No (needs normal) | No | Yes |
✦ SUMMARIZE THIS ARTICLE WITH AI
Anomaly detection in production requires the monitoring infrastructure covered in our MLOps Interview Q&A โ model drift, data drift (which is itself a form of anomaly detection on input distributions), and alerting pipelines. The statistical foundations for Z-score and Mahalanobis distance are covered in our Probability Distributions guide and Statistics Interview Q&A. The autoencoder architecture used here is explained in our Neural Network Architectures guide. Clustering algorithms (DBSCAN, Gaussian Mixture Models) as an alternative anomaly detection approach are covered in our Clustering Algorithms guide.



