You need to enable JavaScript to run this app.
Math Says Yes
Glossary
Look up 51 statistics terms and 32 ideas in plain language.
English
Español
Français
Terms
Quick definitions of the words statistics uses.
Absolute Risk
The actual chance of an event, usually stated as a count or percentage.
absolute increase · risk difference
Base Rate
How common something is before you use new evidence about a specific case.
baseline rate · prior rate
Bayes' Theorem
The rule for updating a probability when new evidence arrives, weighing how well the evidence fits against how common the thing was to begin with.
Bayes' rule · Bayes rule · Bayesian updating
Bias
A systematic tilt that pushes results away from the truth in one direction, no matter how much data you collect.
systematic error · statistical bias
Central Limit Theorem
The result that averages of many independent values follow a normal, bell-shaped distribution — almost regardless of the shape of the raw data.
CLT
Confidence Interval
A range of values that remain plausible for an estimate under the method used.
confidence range · 95% interval
Confounder
A hidden third factor that influences both things you are comparing, creating a link that is not causal.
confounding variable · lurking variable · third variable
Control Group
The group in an experiment that does not get the treatment, providing the baseline the treated group is judged against.
control arm · comparison group
Correlation
A number between -1 and 1 that says how strongly two quantities move together.
correlation coefficient · Pearson correlation · linear association
Distribution
The overall pattern of where values land — which values occur, and how often each one does.
frequency distribution · data distribution
Effect Size
How big a difference or relationship actually is — separate from whether it is statistically significant.
treatment effect · magnitude of effect
Expected Value
The long-run average result when each outcome is weighted by its probability.
long-run average · expected outcome
False Negative
A test result that says something is not there when it is — a true case the test misses.
missed case · type II error
False Positive
A test result that says something is there when it is not — a healthy person flagged as positive.
false alarm · type I error
Histogram
A chart that slices the number line into bins and draws a bar counting how many values land in each — making the data's shape visible.
frequency chart
Incidence
The rate at which new cases of a condition appear in a population over a period of time.
incidence rate · new-case rate
Independence
Two events are independent when knowing one happened tells you nothing about whether the other did.
independent events · statistical independence
Law of Large Numbers
The rule that averages of more and more independent observations settle ever closer to the true underlying value.
law of averages · LLN
Margin of Error
The plus-or-minus range attached to a poll or estimate, showing how far off it could plausibly be from sampling alone.
error margin · plus-or-minus range
Mean
The arithmetic average: add the values and divide by how many values there are.
average · arithmetic average · sample mean
Median
The middle value after sorting the data from smallest to largest.
middle value · 50th percentile
Mutually Exclusive
Two outcomes are mutually exclusive when they cannot both happen at once, so at most one of them occurs.
disjoint events · non-overlapping outcomes
Normal Distribution
The symmetric bell-shaped distribution where values crowd around the mean and become rare evenly on both sides.
bell curve · Gaussian distribution
Null Hypothesis
The default assumption a study tries to challenge — usually that there is no effect or no difference.
H0 · no-effect hypothesis
Odds
A way of expressing chance as the ratio of ways for something to happen against the ways for it not to happen.
odds in favor · betting odds
Odds Ratio
A comparison of two groups made by dividing the odds of an event in one group by the odds in the other.
cross-product ratio · relative odds
Outlier
A value that sits far away from most of the other values.
extreme value · unusual value
P-value
How surprising the observed data would be if a chosen null model were true.
p value
Percentile
The value below which a given percentage of the data falls.
percentile rank · centile
Placebo Effect
Improvement people experience after an inactive treatment, simply because they expect it to help.
placebo response · sugar-pill effect
Population
The entire group you want conclusions about — every case, not just the ones you measured.
target population · statistical population
Prevalence
The share of a population that has a condition at one moment in time.
prevalence rate · point prevalence
Prior
The probability you assign to something before a new piece of evidence arrives.
prior probability · prior belief
Probability
A number between 0 and 1 that measures how likely something is to happen — 0 means impossible, 1 means certain.
chance · likelihood
Quartile
One of the three cut points that split sorted data into four equal-sized groups.
first quartile · upper quartile
Randomized Controlled Trial
A study that randomly assigns people to a treatment group or a control group, so outcome differences can be pinned on the treatment.
RCT · randomized trial · randomised controlled trial
Regression
A method that fits a line through data to describe how one variable changes with another and to make predictions.
linear regression · regression analysis · regression line
Relative Risk
A comparison of risks as a ratio or percentage change.
relative increase
Risk Ratio
The risk of an event in one group divided by the risk in another — how many times more likely the event is.
cumulative incidence ratio
Sample
The smaller group you actually measure, drawn from the larger group you want to learn about.
random sample · study sample
Sample Size
The number of observations in a study or sample — the n behind every estimate.
number of observations · study size
Scatter Plot
A chart that draws one dot per observation, placed by its values on two axes, to reveal how two quantities relate.
scatterplot · scatter diagram · scattergram
Sensitivity
The share of people who truly have a condition that a test correctly flags as positive.
true positive rate · detection rate
Skew
A lopsided distribution: values pile up on one side and stretch out in a long tail on the other.
skewness · skewed data
Specificity
The share of people who truly do not have a condition that a test correctly clears as negative.
true negative rate
Standard Deviation
A measure of how spread out individual values are around their mean.
data spread · spread of values
Standard Error
The typical sampling wobble of an estimate such as a mean.
standard error of the mean · sampling error
Statistical Significance
A label meaning a result would be unusual if pure chance were the only thing going on — typically a p-value below 0.05.
statistically significant · significance testing
Survivorship Bias
The error of judging a group by the members that made it through some selection, while the ones that did not stay invisible.
survivor bias · survival bias
Variance
A measure of spread: the average of the squared distances between each value and the mean.
sample variance · mean squared deviation
Z-Score
How many standard deviations a value sits above or below the mean.
standard score · z value · z score
Ideas
The bigger ideas behind the facts, explained in plain language.
Absolute vs Relative Risk
A relative change ('50% more', 'doubles your risk') hides how big the underlying risk actually is. A large relative increase on a tiny base rate is still tiny. Always ask for the absolute numbers — how many in 100 before and after — not just the percentage change.
Anchoring
An initial number — even an arbitrary or irrelevant one — pulls later estimates toward it. People adjust away from the anchor too little, so the starting point quietly shapes the final judgment.
Averages Mislead
The mean adds everything up and divides, so a few extreme values can drag it far from the typical case. In skewed data the median — the middle value — better represents what is normal. Ask about the shape of the data, not just its average.
Base Rates
The background frequency of something before you consider new evidence. A strong signal can still produce many false alarms when the thing being tested for is rare, because there are many more non-cases than cases. Good reasoning combines the new signal with the base rate.
Central Limit Theorem
The central limit theorem says that averages of many independent, similar-sized random influences tend to form a bell-shaped sampling distribution, even when the original data are not bell-shaped. It explains why averages are often more predictable than individual observations.
Collider Bias
Collider bias appears when you only look at cases selected by a shared outcome, such as hospital patients or admitted applicants. Conditioning on that shared outcome can create a relationship between causes that were not related in the full population.
Comparing Risks
Putting different dangers on one common scale — like the micromort, a one-in-a-million chance of death — lets you compare activities honestly, instead of ranking them by how frightening they feel.
Conditional Probability
The probability of something after new information changes what is still possible. The important move is to update the denominator: cases ruled out by the new information should no longer be counted. Many mistakes happen when people keep using the original odds after the situation has narrowed.
Confidence Intervals
A confidence interval is a range produced by a method that captures the true value a stated share of the time across repeated samples. It describes uncertainty in the estimating method, not a personal guarantee that one already-computed interval contains the truth.
Correlation vs Causation
Two things moving together does not mean one causes the other. A third variable (a confounder) can drive both, or the link can be coincidence. Establishing causation needs more than correlation — usually a controlled comparison.
Error Cancellation
When many independent estimates each carry random error in different directions, averaging them cancels much of the error and leaves the shared signal. Some guesses land too high and some too low, so the over- and under-shoots offset one another in the average. This is why a crowd average — or a repeated measurement — can beat most individual guesses.
Expected Value
The long-run average outcome: each result weighted by how likely it is. It tells you whether a bet is favourable on average, but it ignores how badly a single rare outcome would hurt — which is why people rationally pay to avoid losses they could not absorb.
Exponential Growth
Quantities that grow by a constant percentage multiply rather than add, so each step is bigger than the last. The total stays deceptively flat for a long time, then erupts. Doubling is the clearest case: every step is as large as everything that came before it combined.
Framing Effects
The same true numbers can give very different impressions depending on how they are presented: axis ranges, baselines, absolute versus relative, or wording. Framing changes perception without changing the facts, so check how a number is shown, not just what it is.
Goodhart's Law
Goodhart's law says that when a measure becomes a target, it tends to stop being a good measure. People optimize what is rewarded, so the metric can improve while the thing it was meant to represent gets worse.
Independence
Independent events do not influence each other: the probability of the next outcome is the same no matter what came before. A coin, a die, and a roulette wheel have no memory of their past results.
Law of Large Numbers
As a sample grows, its average gets closer to the true average and random swings shrink, roughly in proportion to 1 over the square root of the sample size. Small samples can land far from the truth purely by chance, which is why a handful of observations is so easy to misread.
Logarithmic Scale
A way of measuring multiplicative growth, where equal distances mean equal ratios instead of equal differences. Moving from 1 to 10 is the same size jump as moving from 10 to 100 because both are times ten. Log scales are useful when quantities span several orders of magnitude.
Margin of Error
A margin of error is the plus-or-minus range around a survey estimate caused by sampling uncertainty. If two estimates have overlapping uncertainty ranges, the apparent leader may not be truly ahead.
Orders of Magnitude
A factor-of-ten step. We handle small additive differences well but badly underestimate how far apart numbers are when they span many tens — which is why a million and a billion feel similar even though one is a thousand times the other.
P-Values
A p-value is the probability of seeing data at least this extreme if a specific null model were true. It is not the probability that the claim is false, not the chance the result was random, and not a measure of how large or important the effect is.
Pairwise Comparisons
When every item can be compared with every other item, the number of possible pairs grows much faster than the number of items. A group of n items has n x (n - 1) / 2 pairs, so each new item connects to everything already there. This is why coincidences become common before any single comparison feels likely.
Publication Bias
Publication bias happens when studies with exciting or statistically significant results are more likely to be published, shared, or noticed than studies with null results. The visible literature can then exaggerate how strong an effect really is.
Randomness and Clustering
Truly random events clump and streak far more than intuition expects. Runs and clusters are a normal feature of randomness, not proof of a pattern, a cause, or a hot streak. An even, tidy spread is the unlikely outcome.
Regression to the Mean
An extreme measurement is usually a mix of a stable, repeatable part and a one-off lucky part. Because the lucky part rarely shows up the same way twice, the next measurement of the same thing tends to land closer to the average. This is not a force that pulls things toward the middle. It is simply what selecting on luck looks like once you measure again.
Sampling Bias
A distorted result caused by looking at data that does not represent the population you care about. More data does not fix the problem if the same kinds of cases are still missing. Before trusting a conclusion, ask who had a chance to appear in the sample and who was left out.
Size-Biased Sampling
A way of sampling where an item's chance of being picked grows with its size or its number of connections, so larger or more-connected items are over-represented. It is why a name occurrence pooled from every friend list looks more popular than a random person: well-connected people appear in many lists, so they get drawn more often. The same effect makes classes feel bigger when you ask students than when you ask the registrar, because more students sit in the large classes.
Standard Deviation
Standard deviation describes how spread out individual observations are around their mean. It is about variation in the data itself. Standard error is different: it describes how much an estimate such as the sample mean would vary across repeated samples.
Standard Error
Standard error is the typical sampling wobble of an estimate such as a mean. For independent observations, the standard error of a mean shrinks with the square root of the sample size, so cutting the error in half usually takes four times as much data.
Statistical Power
Statistical power is the chance that a study will detect an effect of a given size if that effect is real. Low-powered studies can miss real effects, while very large studies can detect effects too small to matter in practice.
Summary Statistics
Summary statistics compress a dataset into numbers such as the mean, variance, correlation, or regression line. They are useful, but they can hide shape, outliers, clusters, and curved relationships that a graph would reveal immediately.
Uncertainty
The honest range left after measurement limits, model choices, and missing information. A single number can be useful as a midpoint, but decisions usually need plausible low and high values too. Good estimates state the assumptions that would move the answer.
Download on the App Store
·
Get the app on Google Play