📋 KEY INSIGHTS
- The Central Limit Theorem (CLT) is the single most important theorem in applied statistics: it guarantees that the sampling distribution of the mean approaches normality as sample size grows, regardless of the original distribution — which is why normal-distribution-based inference works in practice even when data is skewed.
- A confidence interval does NOT mean “there is a 95% probability the true parameter lies in this interval.” Once calculated, the interval either contains the true parameter or it does not — probability is a property of the procedure, not the interval. This distinction matters in any role that involves communicating uncertainty to stakeholders.
- Statistical power — the probability of correctly detecting a real effect when one exists — is determined by effect size, sample size, significance level, and variance. Most practitioners focus on significance (p-values) while neglecting power, leading to underpowered studies that produce unreliable results even when they reach significance.
- The p-value is the probability of observing a test statistic at least as extreme as the one observed, assuming the null hypothesis is true — it is not the probability that the null hypothesis is true, and it says nothing about the practical significance or magnitude of an effect.
- Sampling bias systematically skews estimates when the sample is not representative of the population — no amount of sample size increase or statistical correction can overcome a fundamentally biased sampling mechanism. Survey data, observational studies, and convenience samples all carry this risk.
- Bootstrap resampling is the most practically useful statistical technique for data scientists who want to compute confidence intervals for any statistic (median, correlation, model AUC) without relying on distributional assumptions — it works by repeatedly resampling from the observed data with replacement.
Statistics is the mathematical foundation on which all of data science rests. Machine learning algorithms are statistical estimators. A/B tests are hypothesis tests. Feature selection is about statistical dependence. Model evaluation requires understanding sampling distributions. And yet statistics is also the area where even experienced data scientists most frequently hold misconceptions — about what p-values mean, what confidence intervals guarantee, and what the Central Limit Theorem actually says. These misconceptions are not trivial: they lead to wrong conclusions, miscommunicated results, and flawed experimental designs. This guide covers the core concepts of sampling theory and statistical inference with the precision and practical orientation that a working data scientist needs — including the correct and incorrect interpretations that interviewers and collaborators use to probe understanding.
Sampling Theory — Populations, Samples, and Estimators
Statistical inference is the process of drawing conclusions about a population from a sample. The population is the complete set of individuals or observations you want to make claims about. The sample is a subset of the population that you have actually observed. Because we almost never observe the entire population, every statistical estimate carries uncertainty — and quantifying that uncertainty rigorously is what inferential statistics is about.
A statistic is any function of the sample data — the sample mean, median, standard deviation, or correlation coefficient. A parameter is the corresponding quantity for the population — the population mean (μ), standard deviation (σ), or proportion (p). An estimator is a rule (a formula applied to the sample) for estimating a parameter. The sample mean x̄ is an estimator of the population mean μ. A good estimator has three properties: it is unbiased (its expected value equals the true parameter), consistent (it converges to the true parameter as sample size grows), and efficient (it has the minimum variance among all unbiased estimators — the Cramér-Rao lower bound formalises this).
The sampling distribution of an estimator is the distribution of that estimator across all possible samples of the same size from the population. If you drew 10,000 samples of size n=50 from a population, computed the mean of each, and plotted the resulting 10,000 means, you would have an empirical sampling distribution of the mean. The standard deviation of the sampling distribution is called the standard error — it measures how much the estimator varies across samples of size n. For the sample mean, the standard error is σ/√n (where σ is the population standard deviation), which shows that standard error decreases as sample size increases. Doubling the sample size halves the standard error, but quadrupling the sample size halves it again — the law of diminishing returns in sampling precision.
| Concept | Symbol | Definition | Example |
|---|---|---|---|
| Population parameter | μ, σ, p | True value for the whole population | Mean age of all customers |
| Sample statistic | x̄, s, p̂ | Estimate computed from sample data | Mean age of 500 sampled customers |
| Standard error | SE = σ/√n | Std dev of the sampling distribution | How much x̄ varies across samples of size 500 |
| Bias | E[θ̂] − θ | Systematic difference between estimator and parameter | Sample variance with n denominator (biased); n−1 is unbiased |
| Efficiency | Var(θ̂) | Variance of estimator; lower is better | Mean is more efficient than median for normal data |
| Consistency | θ̂ →p θ as n→∞ | Estimator converges to parameter in probability | Sample mean is consistent for population mean |
The Central Limit Theorem — Why Normal Distributions Are Everywhere
The Central Limit Theorem (CLT) is the theorem that makes most of classical statistics work in practice. Informally stated: the sampling distribution of the mean of a large enough sample is approximately normal, regardless of the distribution of the original population. More precisely: if X₁, X₂, …, Xₙ are independent and identically distributed random variables with mean μ and finite variance σ², then the standardised sample mean (x̄ − μ)/(σ/√n) converges in distribution to the standard normal N(0, 1) as n → ∞.
What does this mean in practice? It means that even if your individual data points follow a heavily skewed distribution — customer purchase amounts, which are right-skewed with a long tail — the mean of a large sample follows an approximately normal distribution. This is why t-tests, z-tests, and confidence intervals based on the normal distribution are valid in practice even when the original data is not normally distributed. The convergence is faster (requires smaller n) when the original distribution is symmetric and has light tails, and slower when it is skewed or heavy-tailed. A common rule of thumb is n ≥ 30 for the CLT approximation to be adequate for moderately skewed distributions — but for heavily skewed distributions (like the log-normal distributions common in financial and customer data), n ≥ 100 or more may be needed.
The CLT has an important extension for sums of random variables: the sum of n independent random variables with any distribution also approaches normality. This is why the normal distribution appears so frequently in nature — many natural phenomena are the result of many small, independent random effects summing together (height is affected by hundreds of genetic variants and environmental factors; measurement error is the sum of many small independent sources of noise).
Confidence Intervals — Correct and Incorrect Interpretations
A 95% confidence interval for a population mean is a random interval constructed from sample data such that, if the procedure were repeated many times on different samples from the same population, 95% of the resulting intervals would contain the true population mean. This is the frequentist definition, and it is widely misunderstood even by practising statisticians.
Correct interpretation: The confidence interval is a statement about the procedure, not about any single interval. Before the data is collected, 95% of all intervals that will be constructed by this procedure will contain μ. Once the data is collected and the specific interval [2.3, 4.7] is computed, that specific interval either contains μ or it does not — there is no probability involved. The probability was used to construct the procedure, not to describe the specific realised interval.
Common incorrect interpretation: “There is a 95% probability that the true mean lies between 2.3 and 4.7.” This assigns a probability to the parameter μ, which is a fixed unknown constant in frequentist statistics. Saying μ has a 95% probability of being in a specific interval is the Bayesian interpretation (a credible interval), which requires a prior distribution on μ. The two frameworks give numerically similar results in many practical settings but have fundamentally different interpretations.
The width of a confidence interval is determined by three factors: the sample size n (larger n → narrower interval, all else equal), the population variance σ² (higher variance → wider interval), and the confidence level (higher confidence level → wider interval). A 99% confidence interval is wider than a 95% interval for the same data — achieving higher confidence requires a wider net. This trade-off between confidence level and precision is fundamental: you can always achieve 100% confidence by reporting an infinitely wide interval, but it contains no information.
| Statement | Correct? | Why |
|---|---|---|
| “95% of data points fall within the CI” | No | CIs are about the mean, not individual data points (that’s a prediction interval) |
| “There is 95% probability μ is in [2.3, 4.7]” | No (frequentist) | μ is fixed; probability applies to the procedure, not the realised interval |
| “If we repeated this 100 times, ~95 CIs would contain μ” | Yes | Correct frequentist interpretation of the procedure |
| “Wider CI means more uncertainty about μ” | Yes | Higher variance or smaller n → wider CI → less precise estimate |
| “Non-overlapping 95% CIs implies p < 0.05” | Not always | Two 95% CIs can overlap and still give p < 0.05; non-overlap implies p < ~0.005 |
| “A narrower CI is always better” | No | Narrower CI at 50% confidence is worse than wider CI at 95% confidence |
Statistical Power and Sample Size
Statistical power is the probability of correctly rejecting a false null hypothesis — equivalently, the probability of detecting a real effect when one exists. Power = 1 − β, where β is the Type II error rate (the probability of failing to detect a real effect). Power is determined by four quantities: effect size (how large is the true difference or association?), sample size (more data → more power), significance level α (higher α → more power, at the cost of more false positives), and variance (higher variance → lower power). Power analysis is the process of choosing the sample size needed to achieve a target power (typically 80% or 90%) at a given effect size and significance level.
Neglecting power analysis produces two types of costly mistakes. Underpowered studies fail to detect real effects — they produce inconclusive results that are often misinterpreted as evidence that no effect exists. An A/B test run for one week on a small traffic segment may have only 40% power — meaning even if the intervention truly improves conversion by 10%, there is a 60% chance the test returns a non-significant result and the improvement is missed. Overpowered studies detect statistically significant but practically meaningless effects — with n = 1,000,000, a 0.001% difference in conversion rate will be statistically significant but costs more to investigate than it could ever return. Effect sizes must always be evaluated for practical significance alongside statistical significance.
✦ SUMMARIZE THIS ARTICLE WITH AI
Hypothesis testing — the formal procedures for testing null hypotheses including t-tests, chi-squared tests, and non-parametric alternatives — is covered in depth in our Hypothesis Testing guide. Probability distributions — the building blocks of the statistical models described here — are in our Probability Distributions guide. Bayesian statistics, which provides the alternative interpretation of probability and credible intervals, is in our Bayesian Statistics guide. Applying these statistical concepts to experiment design for product decisions is in our A/B Testing guide.



