AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Hypothesis Testing Basics

Compare an A/B gap with what the null model predicts.

One toy A/B comparison: independent normal groups, known population SD 4 in each. Illustrative observed sample means A = 20 and B = 20.00, with 25 observations per group (50 total). Editing these hypothetical summaries does not collect data or reveal true means.

H0: μB − μA = 0. H1: μB − μA ≠ 0. The alternative includes either direction.

Observed sample mean B minus A, in score units. This is not the true population difference. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Hypothetical independent observation count in each group; both group mean uncertainties contribute. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Decision threshold α, separate from the p-value. Choose it before collecting real test data. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Observed gap 0.00; size 25 per group; α = 0.05. Two-sided decision: fail to reject H0.

Test statistic if H0 is true

The standard-normal null curve stays centered at zero. Orange areas are both p-value tails beyond ±|observed Z|; indigo dashed lines are the decision’s critical boundaries, not the p-value. The dark baseline dot is the signed observed Z.

Density00.20.40.45-12-60612Z · standardized gap under H0

Observed Z = 0.0000; |Z| = 0.0000. Critical boundaries: ±1.9600. Shaded tails total p ≈ 1.0000.

The plot shows only −12 to 12; each computed tail extends to infinity. Density height is not point probability. Decisions compare full-precision computed p with α; displayed probabilities are rounded.

SE difference = √(4² / 25 + 4² / 25) = 1.1314. Z = (0.00 − 0) / 1.1314 = 0.0000.

Independent mean variances add, so both groups contribute uncertainty. This standard-normal rule is exact under the stated normal known-SD model, rather than a universal small-sample approximation.

Observed Z

Signed gap measured in difference SEs.

0.0000

Observed gap / SE = 0.00 / 1.1314

Two-sided p-value

Extreme-data probability under H0.

≈ 1.0000

P(|Z| ≥ |z observed| given H0)

Decision

A predeclared rule; neither outcome is proof.

Fail to reject

p > α = 0.05

p does not give the probability H0 is true. Failing to reject does not prove equal means. Statistical significance alone does not establish a large, practically important or causal effect. Choose α and the two-sided procedure before collecting real data.

Exact tail probabilities

Each null tail ≈ 0.5000000000; total two-sided p ≈ 1.000000000. α = 0.05. Rule: reject if computed p ≤ α, including equality. Current decision: fail to reject H0.

These numerical probabilities are shown to 10 significant digits. The decision uses full computed precision. A&S 26.2.17 supplies moderate tails; above |Z| = 4, a 100-term NIST DLMF 7.9.2 continued fraction preserves small-tail relative precision. Tiny positive p-values use scientific notation, never a displayed zero.

Model assumptions

Independent normal observations, known population SD 4 in both groups and equal group sizes give difference variance 16 / n + 16 / n. Under H0: μB − μA = 0, the standardized statistic has standard-normal law. The observed means 20 and 20.00 are hypothetical summaries; the true means are not known. Editing n holds the gap fixed for a controlled comparison and does not generate new observations. A real new sample may have a different gap.

The p-value conditions on H0 and all model assumptions. It does not invert that conditional probability, measure practical importance or validate causal design. α concerns the test’s repeated false-rejection rate when H0 is true under its model, rather than the probability this particular conclusion is wrong. Unknown SD, paired or unequal groups, non-normal small samples, dependence, biased collection, power, multiple testing and sequential peeking are outside scope. A one-sided hypothesis must be specified in advance and is a different test.

Berkeley SticiGui · Null hypotheses, p-values and decision rules
Berkeley SticiGui · Two-sided Z tails and independent mean differences
NIST · Known-SD tests and preselected significance
NIST DLMF · Small-tail numerical continued fractions