AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Type I & Type II Errors Lab

Move a cutoff; compare false alarms and missed effects.

Fixed toy A/B model: independent normal observations, known population SD 4 in each group, size 25 per group. True gap is 0 under H0 or 2.5 under one specified real alternative. Difference SE = 1.1314. Standardized Z has SD 1 in either world; its mean is 0 or 2.2097.

The toolbar’s supplied truths and observed Z values are prepared hypothetical possibilities, not random samples, frequencies or known truth in a real study. Selecting an example preserves both probability laws and their error rates.

Reject H0 if |observed Z| is at least this cutoff, including equality. Choose a real rule before collecting data. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Cutoff 1.96; prepared null truth, observed Z 2.4. Decision: reject H0. Classification: Type I false alarm.

The same cutoff applies to both worlds. Indigo dashed lines mark ±cutoff. The dark baseline dot appears only in the prepared example’s supplied-truth panel. Shaded areas describe errors over repetitions, not the chance this particular conclusion is wrong.

When H0 is true · Type I false alarm

α = P(reject H0 given H0 true) ≈ 5.00%. Orange outer tails are false alarms.

Density00.20.40.45-6-3036Z · standardized gap

When real gap 2.5 is true · Type II miss

β = P(fail to reject H0 given true gap 2.5) ≈ 40.14%. Lavender central area is misses.

Density00.20.40.45-6-3036Z · standardized gap

α ≈ 5.00%; β ≈ 40.14%. α + β ≈ 45.14%; these different conditional rates need not add to 100%.

The finite plots show −7 to 7 with the same density scale 0 to 0.45; numerical rates include the full real line. A density height is not point probability. Moving the cutoff changes the rule and shaded regions, not the curves or true gap. These are model probabilities, not empirical counts.

False alarm α

Type I rate when H0 is true.

≈ 5.00%

P(reject H0 | true gap = 0)

Miss β

Type II rate for true gap 2.5.

≈ 40.14%

P(fail to reject H0 | true gap = 2.5)

Prepared example

Type I false alarm. Truth is supplied only in this illustration.

Reject H0

|Z| = 2.4 ≥ cutoff 1.96

Type I and Type II labels combine truth with a decision. A real test does not reveal whether one result is an error. α and β are conditional repeated-test rates, not posterior truth probabilities, p-values or individual guarantees. Fail to reject does not prove no effect. No cutoff is universally best without context and consequences.

Conditional decision table
Each truth row totals 100% before rounding; the rows describe separate worlds.
Supplied population truthReject H0Fail to reject H0
H0 true · gap 0Type I false alarm · 5.00%Correct non-alarm · 95.00%
Specified real gap 2.5Correct detection · 59.86%Type II miss · 40.14%
Model assumptions and exact rates

SE = √(16 / 25 + 16 / 25) = 1.131371; real-world mean of Z = 2.5 / SE = 2.209709. α = 2Q(1.96) ≈ 0.04999565031. β = Φ(1.96 − 2.209709) − Φ(−1.96 − 2.209709) ≈ 0.4013911200.

Q is the standard-normal upper tail; Φ is its CDF. Numerical rates use A&S 26.2.17 for moderate tails and NIST DLMF 7.9.2 above standardized distance 4. The decision uses the displayed cutoff with equality rejecting; endpoints have zero probability under these continuous laws. Unknown SD, dependence, paired groups, non-normal samples, collection bias, different alternatives, error costs, repeated peeking and multiple testing require other analysis. Cutoff comparisons here are explanatory: real tests predeclare their procedure before data. Correct detection is 1 − β, also called power; effect-size and sample-size tradeoffs belong to the next lesson. Statistical significance does not establish a causal or practically important effect.

Berkeley SticiGui · Conditional Type I and Type II errors
NIST · Known-SD tests and two-sided rules
NIST DLMF · Numerical tail continued fractions