📋 KEY INSIGHTS
- Feature selection reduces dimensionality, speeds up training, reduces overfitting, and can improve model generalisation — but the right method depends on the model type and data size.
- Filter methods are model-agnostic and fast (correlation, mutual information, chi-squared) — they rank features independently of each other and miss interaction effects.
- Wrapper methods (RFE, forward/backward selection) search for the best subset by training the model repeatedly — they find optimal subsets but are computationally expensive.
- Embedded methods (L1 regularisation / Lasso, tree feature importance, SHAP) perform selection during training — they are the best balance of accuracy and efficiency in practice.
- Multicollinearity — highly correlated features — inflates variance in linear models and makes coefficients unreliable; VIF (Variance Inflation Factor) is the standard diagnostic.
- For high-dimensional data (text, genomics), use variance thresholding and mutual information first to reduce to a manageable set, then apply wrapper or embedded methods.
Every dataset has features that don’t contribute to prediction — noise columns, redundant duplicates, highly correlated pairs, and low-variance constants. Including them wastes compute, can hurt generalisation through the curse of dimensionality, and makes models harder to interpret. Feature selection is the process of systematically identifying and removing irrelevant or redundant features. There are three paradigms: filter methods (rank features independently, no model), wrapper methods (search for the best subset using the model itself), and embedded methods (selection happens as part of model training). This guide covers all three with Python code and practical guidance on when to use each.
Filter Methods — Fast, Model-Agnostic Screening
Filter methods rank features using statistical measures computed independently of any model. They are the right first step for large datasets — they run in O(n) time and eliminate obvious noise before you invest in more expensive methods.
import numpy as np, pandas as pd
from sklearn.datasets import make_classification
from sklearn.feature_selection import (SelectKBest, chi2, mutual_info_classif,
VarianceThreshold, f_classif)
from scipy.stats import spearmanr
X, y = make_classification(n_samples=5000, n_features=20, n_informative=8,
n_redundant=4, random_state=42)
X_df = pd.DataFrame(X, columns=['f' + str(i) for i in range(20)])
# 1. Variance threshold — remove near-constant features
sel_var = VarianceThreshold(threshold=0.01)
X_var = sel_var.fit_transform(X_df)
print('After variance threshold:', X_var.shape[1], 'features remain')
# 2. Correlation filter — drop one of each highly correlated pair
corr_matrix = X_df.corr().abs()
upper = corr_matrix.where(np.triu(np.ones(corr_matrix.shape), k=1).astype(bool))
drop_cols = [c for c in upper.columns if any(upper[c] > 0.90)]
print('Highly correlated (drop):', drop_cols)
X_uncorr = X_df.drop(columns=drop_cols)
# 3. Mutual information — non-linear dependency with target
mi_scores = mutual_info_classif(X_df, y, random_state=42)
mi_df = pd.Series(mi_scores, index=X_df.columns).sort_values(ascending=False)
print('
Top-10 by Mutual Information:')
print(mi_df.head(10))
# 4. ANOVA F-test (SelectKBest) — fast for continuous features
selector = SelectKBest(score_func=f_classif, k=10)
X_kbest = selector.fit_transform(X_df, y)
selected = X_df.columns[selector.get_support()].tolist()
print('
Top-10 by ANOVA F-test:', selected)
# 5. VIF — detect multicollinearity
from statsmodels.stats.outliers_influence import variance_inflation_factor
vif_df = pd.DataFrame({
'feature': X_df.columns,
'VIF': [variance_inflation_factor(X_df.values, i) for i in range(X_df.shape[1])]
}).sort_values('VIF', ascending=False)
print('
Features with VIF > 10 (high multicollinearity):')
print(vif_df[vif_df['VIF'] > 10])
Wrapper Methods — Search-Based Subset Selection
Wrapper methods evaluate feature subsets by training the target model and measuring performance. They find better subsets than filter methods (because they account for interactions) but are computationally expensive for large feature sets. Recursive Feature Elimination (RFE) starts with all features, fits the model, removes the least important feature (by coefficient or importance), and repeats — it is the most widely used wrapper method in scikit-learn.
from sklearn.feature_selection import RFE, RFECV, SequentialFeatureSelector
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import StratifiedKFold
# --- RFE: eliminate features step by step ---
estimator = LogisticRegression(max_iter=500, random_state=42)
rfe = RFE(estimator=estimator, n_features_to_select=8, step=1)
rfe.fit(X_df, y)
selected_rfe = X_df.columns[rfe.support_].tolist()
print('RFE selected:', selected_rfe)
print('Ranking (1 = selected):')
print(pd.Series(rfe.ranking_, index=X_df.columns).sort_values())
# --- RFECV: automatically selects n_features via cross-validation ---
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
rfecv = RFECV(estimator=estimator, cv=cv, scoring='roc_auc', min_features_to_select=3)
rfecv.fit(X_df, y)
print('
Optimal n_features (RFECV):', rfecv.n_features_)
print('Selected:', X_df.columns[rfecv.support_].tolist())
# --- Sequential Forward Selection (SFS) ---
sfs = SequentialFeatureSelector(
RandomForestClassifier(n_estimators=100, random_state=42),
n_features_to_select=8, direction='forward', cv=3, scoring='roc_auc'
)
sfs.fit(X_df, y)
print('
SFS selected:', X_df.columns[sfs.get_support()].tolist())
Embedded Methods — Selection During Training
Embedded methods perform feature selection as part of the model training process. They are more efficient than wrapper methods (only one training run) and more accurate than filter methods (they account for the model’s inductive bias). The two main embedded techniques are L1 regularisation (Lasso) — which drives uninformative feature coefficients exactly to zero — and tree-based importance (especially SHAP values, which provide a more reliable importance estimate than split count or gain).
from sklearn.linear_model import Lasso, LassoCV
from sklearn.preprocessing import StandardScaler
# --- L1 / Lasso regularisation ---
scaler = StandardScaler()
X_sc = scaler.fit_transform(X_df)
# LassoCV: automatically finds optimal alpha via cross-validation
lasso_cv = LassoCV(cv=5, random_state=42, max_iter=2000)
lasso_cv.fit(X_sc, y)
print('Optimal alpha:', round(lasso_cv.alpha_, 5))
coefs = pd.Series(lasso_cv.coef_, index=X_df.columns)
non_zero = coefs[coefs != 0].sort_values(key=abs, ascending=False)
print('Features selected by Lasso (coef != 0):')
print(non_zero)
zero_features = coefs[coefs == 0].index.tolist()
print('Zeroed out (dropped):', zero_features)
# --- SHAP-based feature selection ---
import shap
rf = RandomForestClassifier(n_estimators=200, random_state=42).fit(X_df, y)
explainer = shap.TreeExplainer(rf)
shap_vals = explainer.shap_values(X_df)[1]
mean_shap = pd.Series(np.abs(shap_vals).mean(axis=0),
index=X_df.columns).sort_values(ascending=False)
print('
Mean |SHAP| importance (top 10):')
print(mean_shap.head(10))
# Select top-K by SHAP
k = 8
shap_selected = mean_shap.head(k).index.tolist()
print('
SHAP-selected features:', shap_selected)
| Method | Type | Pros | Cons |
|---|---|---|---|
| Variance Threshold | Filter | Instant, removes constants | Ignores target |
| Correlation / VIF | Filter | Removes redundancy | Misses non-linear relations |
| Mutual Information | Filter | Captures non-linear dependency | Ignores feature interactions |
| RFE / RFECV | Wrapper | Model-aware, finds good subsets | Slow for large feature sets |
| Lasso (L1) | Embedded | Automatic, regularises simultaneously | Linear models only |
| SHAP Importance | Embedded | Theoretically grounded, handles interactions | Requires tree model |
✦ SUMMARIZE THIS ARTICLE WITH AI
SHAP values used for embedded feature selection are explained in depth in our Model Interpretability guide. The feature engineering context for creating the features you then select is covered in our Feature Engineering guide. Regularisation techniques including L1/L2 and their effect on coefficients are in our Regularisation guide. Dimensionality reduction (PCA, UMAP) as an alternative to selection is covered in our Dimensionality Reduction guide.



