Skip to main content

Hypothesis Tests, Effect Sizes, and Multiple Comparisons

A sample mean is 4 above its target. Is that discrepancy compatible with sampling variation, or enough to reject the claim that the population mean equals the target? Even if the claim is rejected, is a shift of 4 worth acting on? A hypothesis test addresses the first question; effect size and a practical decision threshold help with the second. See Statistics and Probability for sampling, standard errors, and interval coverage.

Turn the question into a test​

NIST's explanation of statistical tests starts with two hypotheses: the null H0H_0 states the claim to test, and the alternative H1H_1 identifies the departures to detect. A test also needs a statistic summarizing the data and its reference distribution under H0H_0 and the sampling assumptions.

For a population mean μ\mu:

Question chosen in advanceNullAlternativeMore extreme results
Has the mean changed?H0:μ=500H_0:\mu=500H1:μ≠500H_1:\mu\ne500Departures in either direction
Has the mean increased?H0:μ≤500H_0:\mu\le500H1:μ>500H_1:\mu>500Larger positive departures
Has the mean decreased?H0:μ≥500H_0:\mu\ge500H1:μ<500H_1:\mu<500Larger negative departures

Choose the question, significance level α\alpha, and sample size or stopping rule before looking at the data. For these one-sided normal mean tests, computing the tail probability at the null boundary μ=500\mu=500 controls error throughout the one-sided null.

NIST defines a p-value as the probability, under the null, of a statistic at least as extreme as the observed one. “Extreme” depends on the alternative chosen in advance. The usual rule rejects H0H_0 when p≤αp\le\alpha and otherwise does not reject it.

As the American Statistical Association explains, a p-value is conditional on the null and statistical model. It is neither the probability that the null is true nor the probability that this particular rejection is wrong. Assigning probabilities to hypotheses requires another model, such as the priors and likelihoods in Bayesian Updating.

Worked example: a mean shifts from 500 to 504​

Suppose 25 independent, identically distributed normal observations have sample mean xˉ=504\bar x=504. In this hypothetical example, the population standard deviation σ=10\sigma=10 is known. Test H0:μ=500H_0:\mu=500 against H1:μ≠500H_1:\mu\ne500 at α=0.05\alpha=0.05. NIST's test for one population mean uses the following statistic when the standard deviation is known:

SE=σn=1025=2,Z=Xˉ−μ0SE.SE=\frac{\sigma}{\sqrt n}=\frac{10}{\sqrt{25}}=2, \qquad Z=\frac{\bar X-\mu_0}{SE}.

Under H0H_0, Z∼N(0,1)Z\sim N(0,1), and the observed value is z=2z=2. With Φ\Phi denoting the standard normal cumulative distribution function, the two-sided p-value is

p=PH0(∣Z∣≥2)=2[1−Φ(2)]≈0.045500.p=P_{H_0}(|Z|\ge2)=2[1-\Phi(2)]\approx0.045500.

Under that model with H0:μ=500H_0:\mu=500 true, about 4.55% of repeated samples of 25 would produce an absolute standardized departure at least this large. Reject μ=500\mu=500 at the 5% level. This calculation uses supplied sample summaries; 504 is an example input, not an actual measurement result.

If the question chosen in advance was an increase, the upper-tail p-value is 1−Φ(2)≈0.0227501-\Phi(2)\approx0.022750. It is half the two-sided value here because the observed direction agrees with the alternative and the reference distribution is symmetric. A preselected decrease question would instead give a lower-tail p-value of about 0.977250. Choosing the upper tail after seeing a positive departure changes the test rule. Under a symmetric continuous null distribution, always choosing the smaller one-sided p-value and comparing it with 0.05 gives an actual two-sided error rate of 0.10.

Read the same result as a confidence interval​

For the same model, use NIST's mean interval for a known population standard deviation:

xˉ±z0.975SE=504±1.959964×2≈[500.080,507.920].\bar x\pm z_{0.975}SE =504\pm1.959964\times2 \approx[500.080,507.920].

The interval excludes 500, agreeing with rejection by the two-sided 5% test. NIST's explanation of the test–interval correspondence also uses the known-standard-deviation model. This correspondence requires matching the model, standard error, two-sided direction, and confidence level; a one-sided test corresponds to a one-sided bound. The 95% refers to the interval procedure's repeated-sampling coverage.

When σ\sigma is unknown, substituting the sample standard deviation ss introduces estimation uncertainty. For independent, identically distributed normal observations under H0H_0, (Xˉ−μ0)/(s/n)(\bar X-\mu_0)/(s/\sqrt n) follows tn−1t_{n-1}, so use t-based p-values and intervals together. See the Central Limit Theorem for approximations with nonnormal populations; increasing sample size does not repair dependence or sampling bias.

Error rates, power, and practical effects​

NIST's discussion of quantitative techniques distinguishes statistical significance from practical importance and defines these errors and power:

QuantityMeaning across repeated samples
Type I errorRejecting a true H0H_0; a valid level-α\alpha test bounds its probability by α\alpha
Type II error probability β\betaProbability of failing to reject H0H_0 at a specified true alternative
Power 1−β1-\betaProbability of rejecting H0H_0 at that alternative

The raw effect estimate is 504−500=4504-500=4, or 4/10=0.44/10=0.4 when standardized by the population standard deviation. If action requires an increase of at least 5, the point estimate falls short, and the shift interval [0.080,7.920][0.080,7.920] does not establish a shift above 5. Detecting a departure from zero and meeting a practical threshold are separate judgments.

Compute power for a specified true effect. In the same known-standard-deviation model, a true increase δ=4\delta=4 makes the statistic follow N(λ,1)N(\lambda,1), where λ=δn/σ\lambda=\delta\sqrt n/\sigma. For c=z1−α/2c=z_{1-\alpha/2}, two-sided power is

Pδ(∣Z∣>c)=Φ(−c−λ)+1−Φ(c−λ).P_\delta(|Z|>c)=\Phi(-c-\lambda)+1-\Phi(c-\lambda).

It is about 0.516005 at n=25n=25 and 0.979327 if n=100n=100 is chosen in advance, with the true effect still 4. Power describes a design's detection probability at a specified effect, not the probability that an observed result is true. Failure to reject does not establish equivalence either; an equivalence claim needs a prespecified acceptable difference range and a corresponding procedure, as explained in NIST’s guide to equivalence testing, Section 5. Sample Size Formula covers planning for interval precision; detecting a specified effect additionally requires power, direction, and an error rate.

This standard-library code recomputes the test, interval, and power:

import math
from statistics import NormalDist

def normal_sf(z):
# Compute the upper tail directly to avoid cancellation near 1.
return 0.5 * math.erfc(z / math.sqrt(2))


normal = NormalDist()
n, xbar, mu0, sigma, alpha = 25, 504.0, 500.0, 10.0, 0.05
se = sigma / math.sqrt(n)
z = (xbar - mu0) / se
p_upper = normal_sf(z)
p_two = 2 * normal_sf(abs(z))
critical = normal.inv_cdf(1 - alpha / 2)
lo, hi = xbar - critical * se, xbar + critical * se
print(f"SE={se:.3f}, z={z:.3f}")
print(f"p_upper={p_upper:.6f}, p_two={p_two:.6f}")
print(f"95% CI=[{lo:.3f}, {hi:.3f}]")
print(f"shift={xbar - mu0:.3f}, standardized={(xbar - mu0) / sigma:.3f}")
for size in (25, 100):
shift = 4.0 / (sigma / math.sqrt(size))
power = normal_sf(critical + shift) + normal_sf(critical - shift)
print(f"n={size}, power_at_shift_4={power:.6f}")

Output:

SE=2.000, z=2.000
p_upper=0.022750, p_two=0.045500
95% CI=[500.080, 507.920]
shift=4.000, standardized=0.400
n=25, power_at_shift_4=0.516005
n=100, power_at_shift_4=0.979327

What to control across multiple comparisons​

If 20 null hypotheses are all true and each independent test has a Type I error probability of exactly 0.05, the probability of at least one false rejection is

1−(1−0.05)20≈0.641514.1-(1-0.05)^{20}\approx0.641514.

This product requires independence. Sharing test tasks can make model comparisons dependent, so independence needs justification. Define the family nonetheless: which models, metrics, and subgroups support the conclusion, rather than retaining only significant comparisons afterward.

Benjamini and Hochberg's original paper introduced FDR and noted that it equals FWER when all nulls are true.

Control targetDefinitionQuestion answered
Familywise error rate, FWERP(V≥1)P(V\ge1)How likely is at least one false rejection in this family?
False discovery rate, FDRE[V/max⁡(R,1)]E[V/\max(R,1)]Across repetitions, what is the average proportion of false rejections among rejected hypotheses?

Here VV counts false rejections and RR counts all rejections; the proportion is zero when R=0R=0. With real effects, FDR permits some false discoveries to increase detection. Control at 5% does not guarantee that at most 5% of this particular set of discoveries is wrong.

A small Bonferroni and BH example​

For a family of mm tests defined in advance, the Bonferroni inequality gives a direct approach: test each at α/m\alpha/m, or compare adjusted values min⁡(1,mpi)\min(1,mp_i) with α\alpha. If each original p-value is valid, FWER is at most α\alpha without requiring independence. This follows by bounding the probability of the union of error events by the sum of their probabilities.

SciPy's false-discovery-control documentation retains the dependence conditions: Benjamini–Hochberg (BH) controls FDR for independent tests or those satisfying positive regression dependence (PRDS). Benjamini–Yekutieli (BY) also guarantees control under general dependence, but is more conservative. Correlation alone does not establish PRDS.

BH sorts the p-values as p(1),…,p(m)p_{(1)},\ldots,p_{(m)}, finds the largest kk with p(k)≤kq/mp_{(k)}\le kq/m, and rejects the first kk. Reject none if no such kk exists. This fixed-sample rule also appears in Section 7.2 of Johari and colleagues' paper. The sorted adjusted values are

p(i)BH=min⁡(1,min⁡j≥imjp(j)).p^{BH}_{(i)}=\min\left(1,\min_{j\ge i}\frac{m}{j}p_{(j)}\right).

Suppose five comparisons defined in advance produce these example p-values, with a target of 0.05 for either FWER or FDR:

ComparisonRaw p-valueBonferroni adjustedBH adjusted
A0.0010.0050.005
B0.0120.0600.030
C0.0180.0900.030
D0.0410.2050.05125
E0.2001.0000.200

Bonferroni rejects A; BH rejects A, B, and C. The sorted BH thresholds are 0.01, 0.02, 0.03, 0.04, and 0.05; the largest qualifying position is the third. The methods control different targets. Choose according to whether false discoveries within a set are acceptable, rather than which method produces more significant results.

The following standard-library calculation adjusts this finite list and restores BH values to input order. It requires a nonempty list of valid p-values.

ps = [0.001, 0.012, 0.018, 0.041, 0.200]
m = len(ps)
order = sorted(range(m), key=ps.__getitem__)
bh = [0.0] * m
running = 1.0
for rank in range(m, 0, -1):
index = order[rank - 1]
running = min(running, m * ps[index] / rank)
bh[index] = running
bonferroni = [min(1.0, m * p) for p in ps]
print("Bonferroni:", [round(p, 6) for p in bonferroni])
print("BH:", [round(p, 6) for p in bh])
print("reject at 0.05:", sum(p <= 0.05 for p in bonferroni),
sum(p <= 0.05 for p in bh))
print(f"20 independent true nulls, FWER={1 - 0.95**20:.6f}")

Output:

Bonferroni: [0.005, 0.06, 0.09, 0.205, 1.0]
BH: [0.005, 0.03, 0.03, 0.05125, 0.2]
reject at 0.05: 1 3
20 independent true nulls, FWER=0.641514

Adjusting p-values does not automatically adjust intervals. To obtain simultaneous coverage for the family, construct each parameter's interval at confidence level 1−α/m1-\alpha/m; Bonferroni guarantees simultaneous coverage of at least 1−α1-\alpha. BH controls false discoveries, so ordinary 95% intervals for the discoveries do not provide a simultaneous-coverage guarantee.

Model selection and repeated peeking​

Cawley and Talbot show that a model-selection criterion evaluated on a finite sample can itself be overfit, creating selection bias when performance is then reported on the same data. After comparing quantizations, prompts, context settings, and multiple metrics, selecting the best configuration and computing its p-value as if that pair alone had been chosen in advance omits the selection process.

When models share test tasks, uncertainty estimates need to preserve task pairs and any project-level grouping. Statistics and Probability explains the sampling units, task-level loss differences, and paired bootstrap.

Selection and stopping also need separate treatment. Repeating an ordinary test after every new batch of tasks and stopping when p<0.05p<0.05 removes the final p-value's original fixed-sample error guarantee. Johari and colleagues' study of continuous monitoring explains this effect and constructs always-valid p-values that permit data-dependent stopping under their statistical assumptions. Recomputing ordinary p-values does not give them that property.

For local model evaluation, organize a comparison as follows:

  1. Fix configurations, the primary metric, the practical improvement threshold, the comparison family, and the error-control target.
  2. Select configurations on development tasks; evaluate frozen candidates on fresh test tasks unused in selection. With nested cross-validation, put the entire selection process inside the inner loop.
  3. Fix test sample size in advance. If interim decisions are needed, use an explicit sequential method and retain its conditions.
  4. Report the sampling unit, effect estimate, interval, raw and adjusted p-values, and comparison and stopping rules together.

Multiple-comparison adjustment starts with valid original p-values. Applying BH to the winners at the end cannot supply the missing test design for selected data or a stopping rule driven by results.

Explore connectionsOpen network