AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

LLM Loss & Perplexity Lab

Change a token probability. Compare local cost with sequence cost.

Your dataset

Four illustrative actual tokens are The, cat, sat and a period. Each next-token probability is conditioned on the previous actual tokens, called teacher forcing. <BOS> means beginning of sequence: a context marker, not a scored target. The context excludes the current and future target tokens. These are authored probabilities; no real model, tokenizer, generated text or training runs here.

Context

<BOS>

Actual next token

The

Target probability

50%

Token loss

0.693147

nats

Assigned probability of this actual next token, given its fixed previous context. Other tokens share the remaining mass. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

−ln(0.5) = 0.693147 nats · 1 bits for this token · All other tokens share 50%.

Position selection changes inspection only. All four targets always count. Editing one probability leaves the provided actual tokens, prefixes and other authored probabilities fixed. Actual probability or scenario changes clear stale explanations and transfer answers; no-ops and position selection preserve them.

Mean cross entropy

Average negative log likelihood (NLL), in nats per target.

0.693147

sum of four NLLs / 4

Total NLL=2.772589 nats. Count=4, including zero-cost targets.

Perplexity

Exponentiated mean NLL; an equivalent likelihood scale.

2

exp(mean NLL) = 2^(bits per token)

Not exp(total NLL) or the arithmetic average of 1/p.

Bits per token

Average cost in base 2, per scored token.

1

mean NLL / ln(2) = log2(perplexity)

Bits per token is not bits per character or measured file size.

Count all four targets

NLL means negative log likelihood: −ln(p), where ln is the natural logarithm and nats are its cost units. exp reverses ln: exp(x)=e raised to x. Bits use log base 2; divide nats by ln(2) to convert. Mean cross entropy here is the average NLL of the actual targets. Perplexity is also the inverse geometric mean of their probabilities. A very low target probability has a high cost. Probability 1 gives zero cost; probability 0 gives infinite cost, with no epsilon substitution or finite cap.

All four actual targets are scored. <BOS> supplies context only. Selecting a position never changes the denominator 4; calculations precede six-decimal display rounding.
PositionPrior-only contextActual tokenTarget pNLL · natsBits
1<BOS>The0.50.6931471
2<BOS> Thecat0.50.6931471
3<BOS> The catsat0.50.6931471
4<BOS> The cat sat.0.50.6931471

Every assigned target probability is positive, so all costs are finite.

Construction and limits

The four authored conditional probabilities use 0..100% in steps of one percentage point. Each context distributes p to its actual next token and 1−p to all other tokens collectively, without specifying a vocabulary or individual alternative probabilities. Prefixes use the previous supplied actual tokens even after probability edits. No model is run to recompute later conditionals.

Every target has equal weight and all four always count. There is no masking, context-window truncation or denominator change. <BOS> is a defined context marker; the first actual token The is scored given that marker. Balanced targets uses four 50% probabilities; One surprise uses 50%,50%,50%,1%; Certain targets uses four 100% probabilities. Reset and predictions restore Balanced targets for the current experiment; free Reset starts Experiment 1.

For positive probabilities, perplexity≥1. All probabilities 1 give mean 0, perplexity 1 and bits per token 0. The smallest positive probability supported here is 0.01, so the largest finite mean is ln(100), largest finite perplexity is 100 and largest finite bits per token is log2(100). Exact probability 0 gives infinite costs and does not claim a finite-precision softmax produced zero. Lower perplexity on fixed targets and protocol reflects higher assigned likelihood, not proof of better general answers, truth, safety, accuracy or calibration. Real comparisons require consistent tokenization, data, scored positions and context policy.

Hugging Face · Perplexity, tokenization and scored-token context