Monday, September 28, 2026
HomeData ScienceFeature Selection Techniques – Filter, Wrapper and Embedded Methods for Machine Learning

Feature Selection Techniques – Filter, Wrapper and Embedded Methods for Machine Learning

Table of Content

📋 KEY INSIGHTS

  • Feature selection reduces dimensionality, speeds up training, reduces overfitting, and can improve model generalisation — but the right method depends on the model type and data size.
  • Filter methods are model-agnostic and fast (correlation, mutual information, chi-squared) — they rank features independently of each other and miss interaction effects.
  • Wrapper methods (RFE, forward/backward selection) search for the best subset by training the model repeatedly — they find optimal subsets but are computationally expensive.
  • Embedded methods (L1 regularisation / Lasso, tree feature importance, SHAP) perform selection during training — they are the best balance of accuracy and efficiency in practice.
  • Multicollinearity — highly correlated features — inflates variance in linear models and makes coefficients unreliable; VIF (Variance Inflation Factor) is the standard diagnostic.
  • For high-dimensional data (text, genomics), use variance thresholding and mutual information first to reduce to a manageable set, then apply wrapper or embedded methods.

Every dataset has features that don’t contribute to prediction — noise columns, redundant duplicates, highly correlated pairs, and low-variance constants. Including them wastes compute, can hurt generalisation through the curse of dimensionality, and makes models harder to interpret. Feature selection is the process of systematically identifying and removing irrelevant or redundant features. There are three paradigms: filter methods (rank features independently, no model), wrapper methods (search for the best subset using the model itself), and embedded methods (selection happens as part of model training). This guide covers all three with Python code and practical guidance on when to use each.

Filter Methods — Fast, Model-Agnostic Screening

Filter methods rank features using statistical measures computed independently of any model. They are the right first step for large datasets — they run in O(n) time and eliminate obvious noise before you invest in more expensive methods.

import numpy as np, pandas as pd
from sklearn.datasets import make_classification
from sklearn.feature_selection import (SelectKBest, chi2, mutual_info_classif,
                                        VarianceThreshold, f_classif)
from scipy.stats import spearmanr

X, y = make_classification(n_samples=5000, n_features=20, n_informative=8,
                            n_redundant=4, random_state=42)
X_df = pd.DataFrame(X, columns=['f' + str(i) for i in range(20)])

# 1. Variance threshold — remove near-constant features
sel_var = VarianceThreshold(threshold=0.01)
X_var   = sel_var.fit_transform(X_df)
print('After variance threshold:', X_var.shape[1], 'features remain')

# 2. Correlation filter — drop one of each highly correlated pair
corr_matrix = X_df.corr().abs()
upper = corr_matrix.where(np.triu(np.ones(corr_matrix.shape), k=1).astype(bool))
drop_cols = [c for c in upper.columns if any(upper[c] > 0.90)]
print('Highly correlated (drop):', drop_cols)
X_uncorr = X_df.drop(columns=drop_cols)

# 3. Mutual information — non-linear dependency with target
mi_scores = mutual_info_classif(X_df, y, random_state=42)
mi_df = pd.Series(mi_scores, index=X_df.columns).sort_values(ascending=False)
print('
Top-10 by Mutual Information:')
print(mi_df.head(10))

# 4. ANOVA F-test (SelectKBest) — fast for continuous features
selector = SelectKBest(score_func=f_classif, k=10)
X_kbest  = selector.fit_transform(X_df, y)
selected = X_df.columns[selector.get_support()].tolist()
print('
Top-10 by ANOVA F-test:', selected)

# 5. VIF — detect multicollinearity
from statsmodels.stats.outliers_influence import variance_inflation_factor
vif_df = pd.DataFrame({
    'feature': X_df.columns,
    'VIF': [variance_inflation_factor(X_df.values, i) for i in range(X_df.shape[1])]
}).sort_values('VIF', ascending=False)
print('
Features with VIF > 10 (high multicollinearity):')
print(vif_df[vif_df['VIF'] > 10])

Wrapper Methods — Search-Based Subset Selection

A book sitting on top of a wooden table
Photo by Thorium on Unsplash

Wrapper methods evaluate feature subsets by training the target model and measuring performance. They find better subsets than filter methods (because they account for interactions) but are computationally expensive for large feature sets. Recursive Feature Elimination (RFE) starts with all features, fits the model, removes the least important feature (by coefficient or importance), and repeats — it is the most widely used wrapper method in scikit-learn.

from sklearn.feature_selection import RFE, RFECV, SequentialFeatureSelector
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import StratifiedKFold

# --- RFE: eliminate features step by step ---
estimator = LogisticRegression(max_iter=500, random_state=42)
rfe = RFE(estimator=estimator, n_features_to_select=8, step=1)
rfe.fit(X_df, y)
selected_rfe = X_df.columns[rfe.support_].tolist()
print('RFE selected:', selected_rfe)
print('Ranking (1 = selected):')
print(pd.Series(rfe.ranking_, index=X_df.columns).sort_values())

# --- RFECV: automatically selects n_features via cross-validation ---
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
rfecv = RFECV(estimator=estimator, cv=cv, scoring='roc_auc', min_features_to_select=3)
rfecv.fit(X_df, y)
print('
Optimal n_features (RFECV):', rfecv.n_features_)
print('Selected:', X_df.columns[rfecv.support_].tolist())

# --- Sequential Forward Selection (SFS) ---
sfs = SequentialFeatureSelector(
    RandomForestClassifier(n_estimators=100, random_state=42),
    n_features_to_select=8, direction='forward', cv=3, scoring='roc_auc'
)
sfs.fit(X_df, y)
print('
SFS selected:', X_df.columns[sfs.get_support()].tolist())

Embedded Methods — Selection During Training

Embedded methods perform feature selection as part of the model training process. They are more efficient than wrapper methods (only one training run) and more accurate than filter methods (they account for the model’s inductive bias). The two main embedded techniques are L1 regularisation (Lasso) — which drives uninformative feature coefficients exactly to zero — and tree-based importance (especially SHAP values, which provide a more reliable importance estimate than split count or gain).

from sklearn.linear_model import Lasso, LassoCV
from sklearn.preprocessing import StandardScaler

# --- L1 / Lasso regularisation ---
scaler = StandardScaler()
X_sc   = scaler.fit_transform(X_df)

# LassoCV: automatically finds optimal alpha via cross-validation
lasso_cv = LassoCV(cv=5, random_state=42, max_iter=2000)
lasso_cv.fit(X_sc, y)
print('Optimal alpha:', round(lasso_cv.alpha_, 5))

coefs = pd.Series(lasso_cv.coef_, index=X_df.columns)
non_zero = coefs[coefs != 0].sort_values(key=abs, ascending=False)
print('Features selected by Lasso (coef != 0):')
print(non_zero)
zero_features = coefs[coefs == 0].index.tolist()
print('Zeroed out (dropped):', zero_features)

# --- SHAP-based feature selection ---
import shap
rf = RandomForestClassifier(n_estimators=200, random_state=42).fit(X_df, y)
explainer  = shap.TreeExplainer(rf)
shap_vals  = explainer.shap_values(X_df)[1]
mean_shap  = pd.Series(np.abs(shap_vals).mean(axis=0),
                       index=X_df.columns).sort_values(ascending=False)
print('
Mean |SHAP| importance (top 10):')
print(mean_shap.head(10))

# Select top-K by SHAP
k = 8
shap_selected = mean_shap.head(k).index.tolist()
print('
SHAP-selected features:', shap_selected)
MethodTypeProsCons
Variance ThresholdFilterInstant, removes constantsIgnores target
Correlation / VIFFilterRemoves redundancyMisses non-linear relations
Mutual InformationFilterCaptures non-linear dependencyIgnores feature interactions
RFE / RFECVWrapperModel-aware, finds good subsetsSlow for large feature sets
Lasso (L1)EmbeddedAutomatic, regularises simultaneouslyLinear models only
SHAP ImportanceEmbeddedTheoretically grounded, handles interactionsRequires tree model

✦ SUMMARIZE THIS ARTICLE WITH AI

SHAP values used for embedded feature selection are explained in depth in our Model Interpretability guide. The feature engineering context for creating the features you then select is covered in our Feature Engineering guide. Regularisation techniques including L1/L2 and their effect on coefficients are in our Regularisation guide. Dimensionality reduction (PCA, UMAP) as an alternative to selection is covered in our Dimensionality Reduction guide.

Leave feedback about this

  • Rating

Durgesh Kekare
Durgesh Kekarehttps://www.dataexpertise.in
Durgesh Kekare is a data science educator and founder of DataExpertise.in. With expertise in Python, machine learning, and analytics, he helps 10,000+ learners break into data careers.

Latest Posts

List of Categories