Sample Size Formula
These formulas give planning approximations for simple random samples when the goal is a confidence interval for one mean or proportion. They do not by themselves guarantee statistical significance or generalizability; clustered designs, finite populations, expected nonresponse, and hypothesis tests require different or additional adjustments.
1. Estimating a Population Mean
To determine the sample size needed to estimate a population mean with a desired level of confidence and precision, the formula is:
- : Required sample size
- : positive upper-tail critical value, defined by for ; equivalently the quantile (about 1.96 for 95% confidence)
- : Estimated standard deviation of the population
- : target confidence-interval half-width, in the same units as the mean; not a guaranteed upper bound on realized error
Example: Population Mean Calculation
To calculate a sample size for a study with a 95% confidence level, an estimated population standard deviation of 10, and a margin of error of 2:
import math
from statistics import NormalDist
confidence_level = 0.95
sigma = 10
E = 2
Z = NormalDist().inv_cdf((1 + confidence_level) / 2)
n = math.ceil((Z * sigma / E)**2)
print(n) # 97
2. Estimating a Population Proportion
For calculating the sample size required to estimate a population proportion within a given margin of error, the formula is:
- : Estimated proportion of the attribute present in the population
- : Margin of error
Example: Population Proportion Calculation
To estimate the sample size for a survey where you expect about 50% of the population to respond positively, with a 95% confidence level and a margin of error of 5%:
import math
from statistics import NormalDist
confidence_level = 0.95
p = 0.5
E = 0.05
Z = NormalDist().inv_cdf((1 + confidence_level) / 2)
n = math.ceil(Z**2 * p * (1 - p) / E**2)
print(n) # 385
Always round up: rounding down can miss the requested margin of error. The mean formula also assumes a reasonable planning value for ; the proportion formula uses when no prior estimate is available because that is conservative. Inflate the result for expected nonresponse, and account for design effects when observations are weighted or clustered.
Derivation and interpretation
For confidence level , a normal-based interval for a mean has half-width . Requiring this to be at most and solving for gives the first formula. A Bernoulli observation has variance , so substituting that variance gives the proportion formula. Since , maximizes the planning requirement when is unknown.
The examples require 97 completed observations for the mean and 385 for the proportion. The latter uses an absolute margin of 0.05, or five percentage points, not 5% of the estimated proportion. Both code blocks run independently with Python's standard library. Halving approximately quadruples the required sample size.
These are normal-approximation planning formulas. A mean interval with known is exact for independent normal data; for other populations it relies on a suitable central limit approximation. With unknown and small samples, planning may need an iterative t-based calculation and an uncertainty allowance for the pilot standard deviation. For proportions close to zero or one, the normal approximation may be poor; do not insert or and conclude that no data are needed. Use a binomial interval method and plan for its width instead.
Finite populations and response rates
For simple random sampling without replacement from a finite population of size , the variance of a sample proportion gains the factor . If is the unrounded proportion requirement above, solving for gives
For , 95% confidence, , and , and the corrected target is 278 completed responses. At an anticipated response rate of 0.8, inviting people gives an expected count, not a guarantee. This arithmetic does not preserve simple random sampling if nonresponse depends on the outcome. A larger invitation pool cannot by itself remove nonresponse bias.
Power analysis for detecting a difference also needs the effect size, desired power, test, and group allocation. It is a different planning question from estimating one mean or proportion to a specified interval width.
For the separate question of detecting a change, NIST’s sample-size examples include the effect to detect and the desired power. Compare those inputs with the margin-of-error formulas here before choosing a planning method.