Statistics
The maths that lets you say whether a number means anything.
Data Analyst
Junior
Write correct SQL, build trusted dashboards, ask good questions.
- Mean, Median and ModeThree ways of naming a 'typical' value, and why the choice changes the story.
- Descriptive StatisticsSummarising a dataset with counts, averages, spread and shape.
- Distribution ShapesSymmetry, tails and peaks — the shape that tells you which summary is honest.
- Standard DeviationHow far a typical value sits from the average, in the same units as the data.
Mid-level
Own an analysis end to end, from vague question to recommendation.
- Base Rate FallacyForgetting how rare something is before testing for it.
- Benford's LawWhy real numbers start with 1 far more often.
- Binomial DistributionCounting successes in a fixed number of tries.
- Bootstrap ResamplingEstimating uncertainty by resampling your own data.
- Conditional ProbabilityUpdating chances when you learn something new.
- Expected ValueThe long-run average of an uncertain quantity.
- Exponential DistributionWaiting times when events arrive randomly.
- ExtrapolationPredicting beyond the data is where models break.
- Lognormal DistributionWhy money and latency skew right.
- Odds RatioComparing chances between two groups.
- Poisson DistributionCounting events that arrive randomly over time.
- R-squaredHow much variation a model explains — and what it hides.
- Random VariablesQuantities whose values come from chance.
- Relative RiskHow many times more likely an outcome is in one group.
- Sampling MethodsChoosing who or what you measure so the result stands for the whole population.
- WinsorizationTaming outliers by capping instead of dropping.
- Z-testComparing rates and means with large samples.
- Law of Large NumbersAverages settle toward the true value as you gather more observations.
- Hypothesis TestingDeciding, in advance, what evidence would change your mind — then checking it.
- Null HypothesisThe 'nothing is happening' assumption a test measures against.
- Quartiles and IQRSplitting data into quarters to summarise the middle half without letting extremes rule.
- Long-Tail DistributionA few huge values and a very long tail of small ones, where the average misleads.
- OutliersValues far from the rest of the data, and when to keep, cap or exclude them.
- The Chi-Square TestComparing counts and proportions between groups, and testing whether categories are independent.
- Variance and Standard DeviationHow spread out numbers are, and why the average alone misleads.
- Correlation CoefficientOne number from -1 to 1 for how tightly two measures move together.
- Effect SizeHow big the difference is, separately from how sure you are that it is real.
- Normal DistributionThe bell curve many natural and sampled quantities roughly follow, and the results that lean on it.
- Regression to the MeanWhy the worst week of the year is almost always followed by a better one.
- Confounding VariableA third factor driving both things you measured, producing a real but meaningless relationship.
- Linear RegressionThe straight-line fit that answers 'what changes when this number moves'.
- Logistic RegressionPredicting yes-or-no outcomes with a linear model and a cutoff.
- Margin of ErrorThe plus-or-minus around a polled number, and why small samples make it large.
- The t-TestComparing two averages and asking whether the gap is bigger than chance.
Senior
Own experimentation and metrics design; call out bad numbers.
- Central Limit TheoremWhy averages and rates behave predictably even when the underlying data does not.
- Standard ErrorHow much your estimate would wobble if you repeated the measurement on new data.
- ANOVATesting whether several group averages really differ.
- Bayes' TheoremUpdating a belief with new evidence, and why 'how likely is that really' beats 'how surprising'.
- Bias-Variance TradeoffSimple models underfit, complex ones overfit.
- Bonferroni CorrectionRaising the bar when you run many tests.
- Credible IntervalThe Bayesian twin of the confidence interval.
- False Discovery RateControlling the share of false alarms among hits.
- HeteroskedasticityWhen variability itself varies across the range.
- Interaction EffectsWhen an effect depends on another variable.
- Mann-Whitney U TestComparing two groups without assuming bell curves.
- Missing Data MechanismsWhether missingness is random decides what you may do.
- MulticollinearityWhen predictors overlap and coefficients go unstable.
- Permutation TestTesting by shuffling labels and re-measuring.
- Propensity ScoresBalancing groups statistically when you cannot randomize.
- Regression DiscontinuityLearning from a cutoff that assigns treatment.
- Sample Size CalculationHow much data you need before starting a test so its result is worth reading.
- Statistical PowerThe chance a test would have detected an effect of the size you care about.
- Multiple RegressionFitting several drivers at once so each one is read holding the others still.
- Regression AnalysisFitting a line or curve so you can describe relationships and make estimates.
- ResidualsWhat the model failed to explain — the leftover that tells you where the fit breaks.
- Non-Parametric TestsTests that fall back to ranks when the data does not meet the assumptions of a t-test.