AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Tune neighborhoods; inspect the map’s limits.
Eight fixed numeric 3D points become a 2D probability map. Map axes are arbitrary and carry no original feature units. This plot auto-fits the current map with equal scales. At iteration 0 these are initial coordinates, before optimization.
Choose Inspect point to locate any ID and inspect its neighborhood. Overlapping dots keep separate rows. No point is moved to avoid an overlap.
A smooth effective neighborhood size based on probability entropy; changing it restarts at iteration 0. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Inspect fifty-step checkpoints of this actual fixed-input run; a finite checkpoint is not a certified global optimum. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
A and B are fixed, reproducible starting coordinate arrays. Changing initialization restarts at 0; source data and probabilities stay fixed.
Two clouds · Perplexity 2 · Initialization A · Iterations 0 · Inspect P1 · Achieved 2.000000 · KL 1.452285
Inspect P1: Gaussian bandwidth σ = 1.023803. Conditional p(j|P1) sums to 1 excluding self. Pij = [p(j|i)+p(i|j)]/16. Qij ∝ 1/(1+map distance²).
Entropy measures how spread probability mass is: H = −Σ p ln p; achieved perplexity = exp(H). It is an effective count, not a hard neighbor cutoff. Joint P and Q each sum to 1 over all ordered nonself pairs; their individual rows are not conditional distributions.
| Case | Original (F1,F2,F3) | Map (X,Y) | Source distance² | p(j|P1) | P1j | Q1j |
|---|---|---|---|---|---|---|
| P1 (self) | (-3, -2, 0) | (-0.800000, -0.200000) | 0.000000 | 0.000000 | 0.000000 | 0.000000 |
| P2 | (-2, -3, 1) | (0.100000, 0.700000) | 3.000000 | 0.703904 | 0.059724 | 0.013770 |
| P3 | (-3, -3, 2) | (0.600000, -0.500000) | 5.000000 | 0.271131 | 0.018882 | 0.011829 |
| P4 | (-2, -2, 3) | (-0.400000, 0.600000) | 10.000000 | 0.024966 | 0.003121 | 0.020043 |
| P5 | (2, 2, 0) | (0.800000, 0.200000) | 41.000000 | 0.000000 | 0.000000 | 0.009698 |
| P6 | (3, 2, 1) | (-0.100000, -0.700000) | 53.000000 | 0.000000 | 0.000000 | 0.020735 |
| P7 | (2, 3, 2) | (-0.600000, 0.500000) | 54.000000 | 0.000000 | 0.000000 | 0.023580 |
| P8 | (3, 3, 3) | (0.400000, -0.600000) | 70.000000 | 0.000000 | 0.000000 | 0.013876 |
Conditional row sum 1.000000; global ordered-pair P sum 1.000000; global Q sum 1.000000. KL(P||Q) = Σ Pij ln(Pij/Qij), in nats.
Effective size of the inspected source neighborhood.
2.000
exp(−Σ p(j|i) ln p(j|i))
Soft weights; not an exact k-neighbor selection.
Current global affinity-fit loss, in nats.
1.452285
Σ Pij ln(Pij / Qij)
Compare fit for fixed P; not task accuracy or original distance loss.
Finite checkpoint of this deterministic run.
0
0..400 in fifty-step checkpoints
Initialization matters; no global-optimum certificate.
The source has eight fixed three-feature points with equal Euclidean weights. Stable Gaussian entropy search chooses each row’s bandwidth to match perplexity. Self is excluded; mathematical nonself Gaussian weights are soft. Symmetrizing conditionals and dividing by 16 gives globally normalized ordered-pair P. Heavy-tailed Student-t map affinities give globally normalized Q. This is exact full-pair t-SNE on a small toy.
The gradient is 4Σj(Pij−Qij)(yi−yj)/(1+||yi−yj||²). Each step tries learning rate 20 and halves it up to 20 times until KL does not increase, then recenters the map. If no trial is accepted the map stays fixed for that step. Two fixed A/B initializations and actual 400-step trajectories make checkpoints reproducible. No early exaggeration, momentum or Barnes-Hut approximation is included; no production-speed or global-optimum claim is made.
Changing perplexity changes source P and restarts optimization, so loss comparisons across perplexities involve different objectives. Changing initialization changes the trajectory for the same source P. Map axes, orientation, global gaps and apparent cluster sizes do not reproduce original feature units, true classes or task accuracy. The plot auto-fits with equal scales; exact current map coordinates remain in the table.
Duplicate source coordinates retain separate IDs and positive mutual affinity; excluded self is a different case. A finite map does not enforce every source distance exactly, including a duplicate pair’s zero distance. Scope excludes learned semantic labels, arbitrary input data, scalable solvers, automatic best settings, new-point transforms, inverse reconstruction and downstream evaluation. Numerical/source/initialization/checkpoint edits clear stale answers; inspection alone preserves them.
van der Maaten & Hinton · Original t-SNE paper