AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Trace first content, completion and work over time.
Four requests use one worker, fixed toy processing times and fictional credits. A token is a piece of input or output text. Each request has 120 input tokens: a 100-token prefix and a 20-token suffix. Prefill means processing that input before generation; a miss takes 400 ms and a prefix-cache hit takes 100 ms. Each output token takes 50 ms, including the first.
The worker finishes a whole request before taking the next, in arrival order; simultaneous arrivals use R1–R4 order. Every recomputation starts with an empty cache. There is no actual inference, network delay or provider billing; these are explanatory assumptions.
4 requests · 16 output tokens · Observation span 3600 ms · Prefix hits 0/4 · Mean queue 0 ms · Mean prefill 400 ms · Mean client first delay 450 ms.
Actual control/preset changes clear stale explanation/transfer answers; identical states preserve them. Cache On cannot reuse distinct prefixes.
Absolute clock in ms. Gray = queue, indigo = prefill, teal = all output-token work. Orange dot = model first token; navy diamond = client first content; end stroke = completion. The exact table below supplies every timestamp and duration.
| Request | Arrival | Start | Queue | Prefill | Model first | Completion | Model TTFT | Client first | Client delay | Latency |
|---|---|---|---|---|---|---|---|---|---|---|
| R1 | 0 | 0 | 0 | 400 | 450 | 600 | 450 | 450 | 450 | 600 |
| R2 | 1000 | 1000 | 0 | 400 | 1450 | 1600 | 450 | 1450 | 450 | 600 |
| R3 | 2000 | 2000 | 0 | 400 | 2450 | 2600 | 450 | 2450 | 450 | 600 |
| R4 | 3000 | 3000 | 0 | 400 | 3450 | 3600 | 450 | 3450 | 450 | 600 |
Model TTFT = queue + prefill + 50 ms. Full latency = queue + prefill + N×50 ms. Delivery changes only the client first-content marker: stream at model first token, buffer at full completion. Mean client first delay = 450 ms.
Generation inter-token gap = 50 ms between successive generated tokens. For one output token, TTFT and completion coincide. This is a generation measure; buffered content arrives together and does not provide a client inter-token event series.
Arrival to complete response, averaged over four requests.
600 ms
Σ(completion − arrival) / 4
Includes queueing. Streaming does not shorten full generation in this toy.
Arrival to model first token, averaged over four requests.
450 ms
Σ(queue + prefill + 50 ms) / 4
Buffering can delay client first content without changing this model clock.
All generated output tokens over the observation window.
4.444444 tokens/s
total output tokens / elapsed seconds
16 / 3.6 s. Includes idle gaps; finite window, not steady-state capacity.
Request throughput = 4 completed requests / 3.6 s = 1.111111 requests/s. The window begins at the first arrival and ends at the final completion; it includes idle time. It is not reciprocal mean request latency or a claim about real serving capacity.
Prefix-cache request-hit rate = 0/4 = 0. A hit reuses the identical 100-token input prefix of a previously completed request. The 20-token suffix and all output tokens are still processed. The cache starts cold, never evicts in this toy and stores no complete answer. This request fraction is not a cache-block or token-hit rate.
Fictional charge rule: uncached input 0.001 credits/token, cached prefix 0.00025 credits/token, output 0.002 credits/token. Timing and charges are deliberately specified independently; cache savings in this model do not promise any provider discount.
| Request | Prefix ID | Hit? | Uncached input | Cached prefix | Output tokens | Input credits | Output credits | Total credits |
|---|---|---|---|---|---|---|---|---|
| R1 | A | No | 120 | 0 | 4 | 0.12 | 0.008 | 0.128 |
| R2 | B | No | 120 | 0 | 4 | 0.12 | 0.008 | 0.128 |
| R3 | C | No | 120 | 0 | 4 | 0.12 | 0.008 | 0.128 |
| R4 | D | No | 120 | 0 | 4 | 0.12 | 0.008 | 0.128 |
Total cost = 0.512 fictional credits. Mean cost per request = total / 4 = 0.128 fictional credits/request.
The finite model has three request patterns, three output lengths, two delivery modes and two cache settings: 36 states. Spaced unique prefixes use arrivals 0,1000,2000,3000 ms with IDs A/B/C/D. Burst unique prefixes all arrive at 0 ms, in R1–R4 order. Spaced repeated prefix uses the spaced arrivals with A/A/A/A. All four requests complete successfully. One worker runs each entire request before the next; start=max(arrival,previous completion). No batching, concurrency, preemption, network overhead, client rendering, failures, retries or percentile estimation is modeled.
Every state recalculates from a cold cache. A hit requires the same prefix to have appeared in a previously completed request with cache On. Miss prefill is 400 ms, hit prefill 100 ms; the remaining input/suffix work and fixed overhead are included in those authored constants. Decode is 50 ms per token including the first. Model TTFT=queue+prefill+50; latency=queue+prefill+50N. With one token, no inter-token pair exists, so its gap is Undefined rather than a fabricated observation. The toy reports generation timing separately from client delivery.
Output throughput includes all 4N tokens and request throughput all four requests over first arrival to final completion, including idle time. Request-hit rate is hits/4. Fictional input charges are uncached tokens×0.001 + cached prefix tokens×0.00025; output charges are N×0.002. These constants are not hardware measurements or current provider prices. Reset/prediction restore Spaced requests for the current exercise; free Reset starts Experiment 1.
vLLM · Explicit latency metric conventions