Every dataset reaches you through a collection process. That process filters who is included, who is excluded, who responds, who survives, who is visible, and who is measured accurately. A can be huge and still biased if the filter is pointed in the wrong direction. The first question is therefore not how many observations there are, but whether the observations represent the you want to describe.
The missing cases carry information
is powerful because the absent cases are often exactly the cases that would change the conclusion. Failed companies are missing from founder advice. People who disliked a product may never leave a review. Damaged aircraft that did not return cannot be inspected. When missingness is related to the outcome, the visible data does not merely have gaps. It points in a systematically misleading direction.
Precision does not fix bias
A biased can produce neat charts, narrow confidence intervals, and repeatable results while still answering the wrong question. More data reduces random noise, but it does not automatically repair a distorted collection method. If a survey only reaches one audience, adding more people from that same audience makes the estimate look more confident, not more representative. Sampling quality is a design issue before it is a calculation issue.
How to inspect a sample
Before trusting a claim, trace the path from to data. Who had a to be included? Who had a reason not to respond? Which failures, quiet users, small groups, or edge cases are invisible? Then ask whether the missing groups would probably have different outcomes from the visible groups. If the answer is yes, treat the conclusion as a statement about the , not the whole population.
FAQ
Can a very large sample still be biased?
Yes. Size reduces random error, but it does not fix a collection process that systematically excludes or overrepresents important groups.
What is the simplest way to spot sampling bias?
Ask who is missing and whether their outcomes would differ. If the missing cases are not random, the visible sample may describe itself more than the wider population.