Wednesday, September 30, 2026
HomeData ScienceCausal Inference for Data Scientists โ€“ Potential Outcomes, DiD and Observational Studies

Causal Inference for Data Scientists โ€“ Potential Outcomes, DiD and Observational Studies

Table of Content

📋 KEY INSIGHTS

  • Correlation is not causation โ€” a model that predicts churn well does not tell you what intervention will reduce churn. Causal inference provides the framework for answering “what happens if we change X?”.
  • The Potential Outcomes Framework (Rubin Causal Model) defines causal effect as the difference between what would happen under treatment vs control for the same unit โ€” the fundamental problem of causal inference is that we observe only one of these.
  • Randomised Controlled Trials (RCTs / A/B tests) are the gold standard for causal inference because random assignment ensures the treatment and control groups are comparable on all confounders, observed and unobserved.
  • Propensity Score Matching and Inverse Probability Weighting (IPW) are observational study techniques that control for measured confounders when running an RCT is impossible.
  • Difference-in-Differences (DiD) estimates causal effects from observational data by comparing the before-after change in a treated group to the before-after change in an untreated control group.
  • The DoWhy library and CausalML package are the standard Python tools for causal inference โ€” they enforce an explicit causal graph, making your assumptions visible and testable.

Every data scientist eventually faces questions where standard predictive modelling is not enough: “Will adding this feature reduce churn?” “Does this pricing change cause fewer cancellations?” “Which users should we target with this intervention for maximum lift?” These questions require causal inference โ€” the discipline of estimating cause-and-effect relationships from data, as opposed to purely correlational associations. Causal inference is what separates a data scientist who can describe what happened from one who can advise what to do. This guide covers the potential outcomes framework, randomised experiments and their limitations, and three key methods for extracting causal effects from observational data when experiments are impossible.

The Potential Outcomes Framework and the Fundamental Problem

The potential outcomes framework (Rubin, 1974) formalises causal inference. For unit i, define: Y_i(1) as the potential outcome if i receives treatment, and Y_i(0) as the potential outcome if i receives control. The Individual Treatment Effect (ITE) for unit i is: ฯ„_i = Y_i(1) โˆ’ Y_i(0). The Average Treatment Effect (ATE) is: E[Y(1) โˆ’ Y(0)] = E[ฯ„]. The Fundamental Problem of Causal Inference: for any unit i, we observe either Y_i(1) or Y_i(0), never both. The counterfactual (what would have happened under the other treatment) is unobservable. This is why causal inference is hard โ€” we must make assumptions to impute the missing counterfactual.

The three core assumptions needed to identify causal effects from observational data are: (1) SUTVA (Stable Unit Treatment Value Assumption) โ€” no interference between units and only one version of treatment; (2) Ignorability / Unconfoundedness โ€” treatment assignment is independent of potential outcomes given observed covariates X: (Y(0), Y(1)) โŠฅ T | X; (3) Overlap / Positivity โ€” every unit has a nonzero probability of receiving either treatment: 0 < P(T=1|X) < 1. Randomisation in an RCT guarantees ignorability by design โ€” it is this guarantee that makes RCTs the gold standard.

import numpy as np
import pandas as pd
from scipy import stats

np.random.seed(42)
n = 2000

# Simulate observational study: email marketing campaign
# Confounder: user_value (high-value users are both more likely to receive email
# AND more likely to purchase โ€” creating spurious positive correlation)
user_value = np.random.normal(50, 15, n)       # 0-100, higher = more valuable
propensity = 1 / (1 + np.exp(-(user_value - 50) / 15))  # S-curve
treatment  = np.random.binomial(1, propensity)           # did they receive the email?

# True causal effect of email: +5 percentage points on purchase probability
# But confounded by user_value
purchase_prob  = (0.02 + 0.005 * user_value / 100 +
                  0.05 * treatment +              # true causal effect
                  np.random.normal(0, 0.02, n))
purchase_prob  = np.clip(purchase_prob, 0, 1)
purchased      = np.random.binomial(1, purchase_prob)

df_obs = pd.DataFrame({
    'user_value': user_value,
    'propensity': propensity,
    'treatment':  treatment,
    'purchased':  purchased,
})

# Naive estimate: compare outcomes between treated and control
# This will be BIASED because treated users have higher user_value
naive_effect = (df_obs.loc[df_obs['treatment']==1, 'purchased'].mean() -
                df_obs.loc[df_obs['treatment']==0, 'purchased'].mean())
print('Naive (confounded) estimate of email effect: ', round(naive_effect, 4))
# This overestimates the true effect of 0.05 because high-value users both receive
# more emails AND purchase more regardless of the email

# True effect for comparison (from simulation)
potential_outcomes_1 = 0.02 + 0.005 * user_value / 100 + 0.05  # all treated
potential_outcomes_0 = 0.02 + 0.005 * user_value / 100          # all control
true_ate = (potential_outcomes_1 - potential_outcomes_0).mean()
print('True ATE (from simulation):                  ', round(true_ate, 4))

Propensity Score Matching and Inverse Probability Weighting

a green tray with dices on top of it
Photo by Colin Davis on Unsplash

Propensity Score Matching (PSM) addresses confounding by matching each treated unit to a control unit with a similar probability of receiving treatment (the propensity score P(T=1|X)). After matching, the matched control units serve as the counterfactual for treated units, and the difference in outcomes is the estimated treatment effect. IPW (Inverse Probability Weighting) takes a different approach: it reweights observations so that the treatment and control groups are balanced on observed confounders. Units that have a low propensity score but were treated get high weight (they are “surprising” treated units and represent the broader population well), and vice versa.

from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler

# Step 1: Estimate propensity scores P(T=1 | X)
X_ps = df_obs[['user_value']].values
scaler = StandardScaler()
X_sc   = scaler.fit_transform(X_ps)

ps_model = LogisticRegression()
ps_model.fit(X_sc, df_obs['treatment'])
df_obs['ps_estimated'] = ps_model.predict_proba(X_sc)[:, 1]

# --- Method A: Inverse Probability Weighting (IPW) ---
# Treated weight: 1/P(T=1|X), Control weight: 1/P(T=0|X)
# ATE-IPW: E[YยทT/e(X)] - E[Yยท(1-T)/(1-e(X))]
eps = 1e-6
w_treated = df_obs['treatment'] / (df_obs['ps_estimated'] + eps)
w_control = (1 - df_obs['treatment']) / (1 - df_obs['ps_estimated'] + eps)

ate_ipw = ((df_obs['purchased'] * w_treated).mean() -
           (df_obs['purchased'] * w_control).mean())
print('IPW estimate of ATE:', round(ate_ipw, 4))

# --- Method B: Propensity Score Matching ---
treated_df = df_obs[df_obs['treatment'] == 1].copy()
control_df = df_obs[df_obs['treatment'] == 0].copy()

matched_controls = []
for _, t_row in treated_df.iterrows():
    # Nearest neighbour match: find control with closest propensity score
    diffs = np.abs(control_df['ps_estimated'] - t_row['ps_estimated'])
    match_idx = diffs.idxmin()
    matched_controls.append(control_df.loc[match_idx, 'purchased'])

psm_ate = treated_df['purchased'].values.mean() - np.mean(matched_controls)
print('PSM estimate of ATE: ', round(psm_ate, 4))

# --- Method C: DoWhy for causal inference with explicit causal graph ---
# pip install dowhy
# import dowhy
# causal_model = dowhy.CausalModel(
#     data=df_obs,
#     treatment='treatment',
#     outcome='purchased',
#     common_causes=['user_value']
# )
# identified_estimand = causal_model.identify_effect()
# estimate = causal_model.estimate_effect(
#     identified_estimand,
#     method_name='backdoor.propensity_score_weighting'
# )
# print('DoWhy ATE:', estimate.value)

Difference-in-Differences and A/B Test Pitfalls

Difference-in-Differences (DiD) is used when you cannot run an experiment but have panel data (multiple units observed before and after a treatment). It estimates the causal effect by comparing the before-after change in the treated group to the before-after change in an untreated control group, under the parallel trends assumption: in the absence of treatment, treated and control groups would have followed parallel trends over time.

# DiD: policy change affected one city but not another
cities = pd.DataFrame({
    'city':    ['City_A']*4 + ['City_B']*4,
    'quarter': ['Q1','Q2','Q3','Q4']*2,
    'post':    [0,0,1,1]*2,               # Q3-Q4 = after policy
    'treated': [1,1,1,1,0,0,0,0],         # City_A received the policy
    'revenue': [100,105,130,135,          # City_A: big jump in Q3
                95, 98, 102, 104],        # City_B: small natural growth
})

# DiD estimate (regression approach โ€” handles controls easily)
import statsmodels.formula.api as smf
did_model = smf.ols('revenue ~ post + treated + post:treated', data=cities).fit()
print(did_model.summary().tables[1])
# Coefficient on post:treated = DiD estimate = causal effect of the policy

# Manual DiD calculation:
before_treated  = cities.loc[(cities['city']=='City_A') & (cities['post']==0), 'revenue'].mean()
after_treated   = cities.loc[(cities['city']=='City_A') & (cities['post']==1), 'revenue'].mean()
before_control  = cities.loc[(cities['city']=='City_B') & (cities['post']==0), 'revenue'].mean()
after_control   = cities.loc[(cities['city']=='City_B') & (cities['post']==1), 'revenue'].mean()
did_manual = (after_treated - before_treated) - (after_control - before_control)
print('
Manual DiD estimate:', round(did_manual, 2))
# = (132.5 - 102.5) - (103 - 96.5) = 30 - 6.5 = 23.5

Common A/B test pitfalls that require causal thinking: (1) Network effects / SUTVA violation โ€” if treating user A affects user B’s behaviour (social networks, marketplaces), the standard A/B framework breaks; use cluster randomisation. (2) Novelty effect โ€” users change behaviour just because something is new; run experiments long enough for novelty to wear off. (3) Simpson’s Paradox โ€” the treatment may appear harmful in aggregate but beneficial within every subgroup (when group composition is confounded with treatment assignment). (4) Multiple testing โ€” running 20 metrics and claiming the significant ones without Bonferroni or FDR correction gives a 64% chance of at least one false positive.

✦ SUMMARIZE THIS ARTICLE WITH AI

The A/B testing framework that applies causal inference principles in practice is covered in depth in our A/B Testing guide. The statistical tests used in DiD and propensity score analysis โ€” t-tests, regression โ€” are in our Hypothesis Testing guide. The Bayesian approach to causal inference, including Bayesian A/B testing, is in our Bayesian Statistics guide. For the ML system design interview questions that test causal thinking (“how would you measure the impact of this feature?”), see our ML System Design guide.

Leave feedback about this

  • Rating

Durgesh Kekare
Durgesh Kekarehttps://www.dataexpertise.in
Durgesh Kekare is a data science educator and founder of DataExpertise.in. With expertise in Python, machine learning, and analytics, he helps 10,000+ learners break into data careers.

Latest Posts

List of Categories