Treating a non-significant result as proof of no effect without asking whether the study was precise enough to see the effect.
What power asks
Statistical power asks a design question: if an effect of a particular size is really present, how likely is this study to detect it? A powerful study has a good of spotting the effect it was built to find. A low-powered study has a weak chance, so a null result may be ambiguous. It might there is no effect, or it might mean the study was too noisy.
What the Numbers Show
20 per group · detects 24%
24%
200 per group · detects 98%
98%
The assumed real effect and noise stay the same; only the number of people changes.
Why small studies miss things
A small trial of a treatment with a modest real effect can still measure that effect on the wrong side of zero, purely from sampling noise. That wide random variation makes it hard for the result to pass a significance threshold, even when the underlying effect is real. Increasing the narrows the , which gives the signal a better to stand out from the noise.
The other side of power
More power is not automatically more meaning. A very large study can detect tiny effects that are too small to matter in practice. That is why power should be tied to the smallest worth detecting. Significance alone is a shallow target; the study should be precise enough to answer a real decision question.
How to use it
When a study finds no significant effect, ask whether it was powered for the that would matter. Look at the : if it still includes both meaningful benefit and meaningful harm, the study has not settled much. When planning a study, choose the smallest practical effect first, then calculate the needed to detect it with acceptable power.
Worked example
Suppose a training program really improves test scores by a small but useful amount. A study with 20 people per group might estimate the effect so noisily that the includes clear benefit, no effect, and even slight harm. A non-significant result from that design does not prove the program useless. It says the study did not separate the signal from the noise well enough.
When it applies
Power is a planning tool, so it works best before the study begins. It depends on the worth detecting, the , measurement noise, design, and significance threshold. After a study, avoid saying it had low power only because the result was not significant. Instead, inspect the planned design and the interval around the estimate. If the interval is still wide enough to include meaningful effects, the study has not answered the decision question.
What people get wrong
People often read "not " as "nothing is happening." Low power makes that reading unsafe. A study can miss an effect because the signal is small relative to the noise. The opposite mistake also matters: a huge study can detect an effect that is real but too small to matter.
Source note
Cohen's power primer is the trusted source behind the design framing on this page. It supports the idea that power is the of detecting an effect of a specified size under a specified design. The page applies that idea to interpretation: a null result from a low-powered design is often weak evidence, not a decisive absence of effect.
FAQ
What does low statistical power mean?
It means the study has a low chance of detecting a real effect of the size being considered. A non-significant result from such a study is weak evidence against the effect.
Can high power make tiny effects look important?
High power can make tiny effects statistically detectable, but importance still depends on the effect size, cost, risk, and decision context.
Quick Check
A tiny study finds no statistically significant effect. What should you check before concluding there is no effect?
A
Whether the p-value was exactly zero.
B
Whether the effect was already proven elsewhere.
C
Whether it had enough power to detect an effect that matters.