AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Embedding Retrieval Lab

Move a query; track the ranked results.

Your dataset

An embedding is a numeric representation of an item. These six handcrafted 2D vectors illustrate exact brute-force retrieval: every item is scored. Move only the query; stored item vectors stay fixed. No learned semantic relevance is claimed. Axes show the actual toy coordinates at equal scales.

-5-50055X coordinateVertical: Y coordinateABCDEFQ

● Filled source glyph: at least one represented ID is returned. ○ Hollow: none returned. ◆ Query Q: movable. Exactly coincident IDs share a labeled source glyph; each ID retains its own table row and inclusion. Leader lines move labels only.

Drag Q, or focus it and use arrows (0.5 units; Shift: 1). Right increases X; Up increases Y. Home sets (0,0); End sets (4,4). Enter, Space or click focuses Query X. The exact editors and sliders below provide the same query changes.

Signed horizontal component, −4..4 in half units. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Signed vertical component, −4..4 in half units. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Requested count; return the first k eligible IDs without changing scores. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Similarity rule

Directions · Query (2, 0) · Euclidean · Top k 3 · Order A,B,F,C,E,D · Returned A,B,F · Query length 2.000000

Retrieval results

All IDs remain, including exact duplicates and excluded zero vectors. Lower Euclidean distance ranks first; fixed ID breaks comparison ties. Six-decimal display is not the comparison key.
RankItemCoordinatesEuclidean distanceCosineReturned
1A(3, 0)1.0000001.000000Included
2B(1, 1)1.4142140.707107Included
3F(-1, -1)3.162278-0.707107Included
4C(0, 3)3.6055510.000000Outside
5E(0, -3)3.6055510.000000Outside
6D(-3, 0)5.000000-1.000000Outside

Top result

First eligible item under the selected rule.

A

Score, then fixed ID for ties

Proximity in these vectors is not proved relevance.

Best distance

Lowest coordinate distance.

1.000000

sqrt((qx−ix)²+(qy−iy)²)

Lower first; raw vector lengths matter.

Returned

Actual count versus requested Top k.

3 of 3

First min(k, eligible count) ranked IDs

More results do not guarantee quality.

Construction and limits

Six handcrafted 2D vectors use equal coordinate weights. Query components are bounded to −4..4 and snapped to half units; Top k is 1..6. The fixed −5..5 square plot shows actual toy coordinates with equal unit scales, not a projection. Exact coincidence groups glyphs and labels only; table identities remain separate. Scenarios preserve query, scoring rule and k while changing stored vectors.

Exact brute-force scoring evaluates all six items. Euclidean sorts by squared distance (exact on this half-unit grid), then fixed ID A..F; the table displays its square root. Cosine is dot divided by both nonzero lengths, sorted highest first. Its comparison key is rounded to 12 decimals for deterministic numerical ties; fixed ID resolves equal keys. Display rounding to six decimals is separate. Tie order is reproducible, not semantic superiority.

A zero query or zero item has no direction, so geometric cosine is Undefined and this demo excludes it. A zero query yields no eligible cosine results. Some libraries assign numerical zero to these cases; that convention does not create a direction. Euclidean works at zero. Nonzero cosine 0 means perpendicular, 1 means same direction rather than identical length, and −1 means opposite. No score is a relevance probability.

The demo does not normalize coordinates when you choose Cosine. If both vectors were normalized, squared Euclidean distance would equal 2−2×cosine and rankings would agree; raw vectors need not agree. Query/scenario/metric/k edits clear stale explanations. Prediction or Reset restores the current baseline; free-exploration Reset starts Experiment 1. There is no training, semantic query text, approximate index, model inference, task ground truth or retrieval-quality metric here.

scikit-learn · Exact nearest-neighbor retrieval
scikit-learn · Cosine scoring
Faiss · Raw and normalized metric conventions