AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Rotate centered axes; decide what to keep.
The plot subtracts the mean from every original point: centered coordinates show spread around zero. Original coordinates remain in the table. Rotate the first axis; the second stays perpendicular. Keeping only one score drops the other direction’s information.
● Original centered□ Reconstructed centered┄ Residual→ First axis / dashed second axis
Rotate the manual orthonormal basis in five-degree steps. Align with PC1 uses the computed covariance direction when it is unique. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Computed PC1 line: 45°. Its variance maximum is 90.000%.
One retains only t1; two retain t1 and t2. Both scores remain visible for inspection. Two retained dimensions in this 2D toy give no dimensional compression.
Correlated cloud · Angle 0° · Keep 1 · Mean (0.000000, 0.000000) · Manual orthogonal axes; first axis is not PC1 · Retained 50.000% · Mean Error² 10.000000
u = (1.000000, 0.000000); v = (0.000000, 1.000000). c = p−mean; t1 = c·u; t2 = c·v. Reconstruct original q = mean+t1u.
| Case | Original (X,Y) | Centered (X,Y) | t1 (kept) | t2 (discarded) | Reconstructed (X,Y) | Error² |
|---|---|---|---|---|---|---|
| P1 | (-4, -2) | (-4, -2) | -4.000000 | -2.000000 | (-4.000000, 0.000000) | 4.000000 |
| P2 | (-2, -4) | (-2, -4) | -2.000000 | -4.000000 | (-2.000000, 0.000000) | 16.000000 |
| P3 | (2, 4) | (2, 4) | 2.000000 | 4.000000 | (2.000000, 0.000000) | 16.000000 |
| P4 | (4, 2) | (4, 2) | 4.000000 | 2.000000 | (4.000000, 0.000000) | 4.000000 |
Sample covariance = [[13.333333, 10.666667], [10.666667, 13.333333]]. Principal variances = 24.000000, 2.666667.
Current score variances: t1 13.333333, t2 13.333333; total 26.666667. Sample divisor: 3. Mean Error² divisor: 4 points. Discarded sample variance = (4/3) × Mean Error².
Variance fraction retained by the current axes and count.
50.000%
Σ retained score variances / total variance
Undefined at zero total; not task accuracy.
Squared Euclidean reconstruction loss, averaged per point.
10.000
Σ ||p−q||² / 4
Add the mean back; two orthogonal scores reconstruct fully.
Largest possible one-axis centered variance fraction.
90.000%
largest covariance eigenvalue / total variance
Equal eigenvalues tie; zero total makes this undefined.
Each named scenario fixes four original points. Centering subtracts each coordinate’s mean and does not scale feature units. The plot uses centered originals and centered reconstructions; the exact table reconstructs raw coordinates by adding the mean. The sample covariance divides centered product sums by 3. Its principal eigenvalues are the variances along principal directions, sorted largest first.
For symmetric covariance [[a,b],[b,d]], principal variances are (a+d ± sqrt((a−d)²+4b²))/2. When they differ, the leading line angle is half atan2(2b,a−d), modulo 180°. Reversing its sign changes score signs while preserving reconstruction. Equal positive eigenvalues make every first direction a maximum, so there is no unique PC1 line. A zero covariance cloud also has no unique direction and makes variance ratios undefined.
Manual axes u and v are orthonormal; their score variances sum to total variance. Keeping one score drops the other variance; keeping both spans the full 2D space. Mean Error² divides squared Euclidean residuals by 4 points; discarded sample variance uses divisor 3 and equals (4/3)Mean Error². Exact equivalent special-angle matrices and a complete-basis identity avoid artificial floating residuals. Other angles use trigonometry; rounded values do not determine zero denominators or identity.
Retained variance is geometry, not a probability, accuracy or assurance of downstream value. Scope excludes general-dimensional solvers, SVD derivation, standardization, whitening, probabilistic PCA, nonlinear embeddings, train/test pipelines, labels and automatic component-count selection. Actual numeric/scenario/count edits clear stale answers; prediction changes and Reset restore the current step.
scikit-learn · Centered principal component analysis