Skip to main content

Sample Size Formula

These formulas give planning approximations for simple random samples when the goal is a confidence interval for one mean or proportion. They do not by themselves guarantee statistical significance or generalizability; clustered designs, finite populations, expected nonresponse, and hypothesis tests require different or additional adjustments.

1. Estimating a Population Mean​

To determine the sample size needed to estimate a population mean with a desired level of confidence and precision, the formula is:

n=(Zα/2×σE)2n = \left(\frac{Z_{\alpha/2} \times \sigma}{E}\right)^2
  • nn: Required sample size
  • Zα/2Z_{\alpha/2}: positive upper-tail critical value, defined by P(Z>Zα/2)=α/2P(Z>Z_{\alpha/2})=\alpha/2 for Z∼N(0,1)Z\sim N(0,1); equivalently the 1−α/21-\alpha/2 quantile (about 1.96 for 95% confidence)
  • σ\sigma: Estimated standard deviation of the population
  • EE: target confidence-interval half-width, in the same units as the mean; not a guaranteed upper bound on realized error

Example: Population Mean Calculation​

To calculate a sample size for a study with a 95% confidence level, an estimated population standard deviation of 10, and a margin of error of 2:

import math
from statistics import NormalDist

confidence_level = 0.95
sigma = 10
E = 2
Z = NormalDist().inv_cdf((1 + confidence_level) / 2)
n = math.ceil((Z * sigma / E)**2)
print(n) # 97

2. Estimating a Population Proportion​

For calculating the sample size required to estimate a population proportion within a given margin of error, the formula is:

n=Zα/22×p×(1−p)E2n = \frac{Z_{\alpha/2}^2 \times p \times (1 - p)}{E^2}
  • pp: Estimated proportion of the attribute present in the population
  • EE: Margin of error

Example: Population Proportion Calculation​

To estimate the sample size for a survey where you expect about 50% of the population to respond positively, with a 95% confidence level and a margin of error of 5%:

import math
from statistics import NormalDist

confidence_level = 0.95
p = 0.5
E = 0.05
Z = NormalDist().inv_cdf((1 + confidence_level) / 2)
n = math.ceil(Z**2 * p * (1 - p) / E**2)
print(n) # 385

Always round up: rounding down can miss the requested margin of error. The mean formula also assumes a reasonable planning value for σ\sigma; the proportion formula uses p=0.5p=0.5 when no prior estimate is available because that is conservative. Inflate the result for expected nonresponse, and account for design effects when observations are weighted or clustered.

Derivation and interpretation​

For confidence level 1−α1-\alpha, a normal-based interval for a mean has half-width Zα/2σ/nZ_{\alpha/2}\sigma/\sqrt n. Requiring this to be at most E>0E>0 and solving for nn gives the first formula. A Bernoulli observation has variance p(1−p)p(1-p), so substituting that variance gives the proportion formula. Since p(1−p)=1/4−(p−1/2)2p(1-p)=1/4-(p-1/2)^2, p=0.5p=0.5 maximizes the planning requirement when pp is unknown.

The examples require 97 completed observations for the mean and 385 for the proportion. The latter uses an absolute margin of 0.05, or five percentage points, not 5% of the estimated proportion. Both code blocks run independently with Python's standard library. Halving EE approximately quadruples the required sample size.

These are normal-approximation planning formulas. A mean interval with known σ\sigma is exact for independent normal data; for other populations it relies on a suitable central limit approximation. With unknown σ\sigma and small samples, planning may need an iterative t-based calculation and an uncertainty allowance for the pilot standard deviation. For proportions close to zero or one, the normal approximation may be poor; do not insert p=0p=0 or p=1p=1 and conclude that no data are needed. Use a binomial interval method and plan for its width instead.

Finite populations and response rates​

For simple random sampling without replacement from a finite population of size NN, the variance of a sample proportion gains the factor (N−n)/(N−1)(N-n)/(N-1). If n0n_0 is the unrounded proportion requirement above, solving for nn gives

nfinite=⌈Nn0N+n0−1⌉.n_{\mathrm{finite}}=\left\lceil\frac{Nn_0}{N+n_0-1}\right\rceil.

For N=1000N=1000, 95% confidence, p=0.5p=0.5, and E=0.05E=0.05, n0≈384.146n_0\approx384.146 and the corrected target is 278 completed responses. At an anticipated response rate of 0.8, inviting ⌈278/0.8⌉=348\lceil278/0.8\rceil=348 people gives an expected count, not a guarantee. This arithmetic does not preserve simple random sampling if nonresponse depends on the outcome. A larger invitation pool cannot by itself remove nonresponse bias.

Power analysis for detecting a difference also needs the effect size, desired power, test, and group allocation. It is a different planning question from estimating one mean or proportion to a specified interval width.

For the separate question of detecting a change, NIST’s sample-size examples include the effect to detect and the desired power. Compare those inputs with the margin-of-error formulas here before choosing a planning method.

Explore connectionsOpen network