Runtime lines

How AI works, one token at a time

A large language model does one thing over and over: it turns the text so far into a probability for every possible next token, picks one, and appends it. Step through that loop with real token ids, watch it invent a confident answer, fix it with retrieved context, then play with temperature and embeddings.

From text to the next token

next_token.py tiktoken, o200k_base
    Output
      Inside the model logits illustrative, softmax exact
      1. Tokenise
      2. Embed
      3. Transformer
      4. Logits
      5. Softmax
      6. Pick

      Token ids are real: tiktoken with the o200k_base encoding (the GPT-4o family’s tokenizer). The logits are made up for teaching, because a real model scores all 200,019 tokens and its numbers depend on its weights. Only the top candidates are drawn, and their probabilities are the exact softmax over those candidates. model() and generate() live in verify/ai/toy_model.py.

      The pipeline at a glance

      Every chat answer, code completion and agent step runs this loop. Only the last step involves any choice; everything before it is a fixed calculation.

      1. TokenizerSplits text into tokens: whole common words, pieces of rarer words, single characters and bytes. Each token has an integer id. The tokenizer is fixed before training and runs outside the neural network.
      2. EmbeddingsEach id selects one row of a learned matrix: a vector of a few thousand numbers. Position information is mixed in too; many current models rotate the vectors by position (RoPE).
      3. Transformer layersA stack of identical blocks. In each, attention lets every position gather information from earlier positions, and a feed-forward network then transforms each position on its own.
      4. LogitsThe final vector at the last position is multiplied by an output matrix, giving one raw score per token in the vocabulary.
      5. SamplingSoftmax turns scores into probabilities, temperature, top-k and top-p reshape them, and one token is chosen. It’s appended to the input and the loop runs again until a stop token or the token limit.

      Why models hallucinate, and what helps

      A hallucination is fluent output that isn’t supported by facts or by the sources you gave the model: an invented citation, a wrong date, a function that doesn’t exist. Nothing in the loop checks truth. It only ever asks which token is likely next.

      ?Why it happens

      • The objective is plausibility. Pretraining rewards predicting the next token of real text, not being right. A plausible name scores well whether or not it’s the true one.
      • Guessing gets rewarded. Training text is full of confident answers and rarely says “I don’t know”, and evaluations that give no credit for abstaining score a lucky guess above an honest blank.
      • Gaps and the cut-off. Rare facts, private data and anything after the training cut-off can’t be recalled. The model fills the gap with something typical.
      • Wrong or blended data. Training text contains errors and outdated facts, and the model can merge two similar facts into one that never existed.
      • Sampling and snowballing. With temperature above 0, an unlikely wrong token can be drawn. Every later token is conditioned on it, so the answer stays consistent with the mistake.
      • Long contexts. Details buried in the middle of a long prompt are used less reliably, so answers can drift from the sources you supplied.

      +What helps

      • Ground it with RAG. Retrieve relevant passages, put them in the prompt, and require citations that you can check against the sources.
      • Let it abstain. Say that “I don’t know” is allowed and preferred to a guess, and give abstaining credit in your evaluations.
      • Lower the temperature for facts. It removes sampling noise. It doesn’t fix what the model believes, as the greedy scenario above shows.
      • Structured output, then validation. Ask for JSON that matches a schema, then check ids, quotes and numbers in code before you use them.
      • Use tools. Let the model call search, a database or a calculator instead of recalling facts and doing arithmetic from memory.
      • Measure and review. Run evals on questions with known answers, check that claims are supported by the sources (with code or a second model), watch low token probabilities, and keep a human in the loop where mistakes are costly.

      Temperature, top-k and top-p

      The model hands you a probability for every token; how you pick from them is a separate step called decoding. Temperature divides the logits before the softmax, softmax(logits / T): below 1 it sharpens the distribution, above 1 it flattens it. Top-k and top-p then cut off the unlikely tail. Move the sliders and sample.

      1.00
      8
      1.00

      Prompt: My favourite programming language is. Eight candidate tokens with illustrative logits; the maths applied to them is exact.

      Cumulative probability after temperature

      Each draw uses a seeded random generator, so draw 1 is always the same sequence for the same settings.

      Why isn’t temperature 0 perfectly deterministic? Greedy decoding has no randomness, yet two runs can differ. Floating-point addition on GPUs isn’t associative, and the order of additions changes with batch size and with which other requests share the batch, so two nearly tied logits can swap. One different token changes everything after it. Providers that accept a seed describe it as best effort.

      The widget applies temperature first, then top-k, then top-p (renormalising after each cut), the order used by verify/ai/temperature.py. Its numbers, cut-offs and seeded draws match that program’s output.

      Embeddings: meaning as a direction

      An embedding model turns a piece of text into a fixed-length list of numbers, a vector. It is trained so that texts with similar meaning get vectors pointing in similar directions. The map uses toy 2-D vectors so you can see the angles; real models output hundreds to thousands of dimensions (384, 768, 1,536 and 3,072 are common sizes). Pick a query and a metric.

      Query
      Rank by
      Toy embedding space hand-placed 2-D vectors
      pets money code outdoors query
      Ranked by cosine similarity top 3 are the nearest neighbours

        Worked through for the selected pair click a row to change it

          What an embedding is

          A point in a space where direction stands for meaning. The individual numbers aren’t readable features; only distances and angles between vectors mean anything, and only between vectors from the same model.

          Why similar meanings land together

          Embedding models are trained, often contrastively, to pull related pairs (a question and its answer, a caption and its image) together and push unrelated ones apart. So feeding my kitten lands next to cat food with no shared words, and river bank lands far from bank loan despite one.

          Cosine, dot product, Euclidean

          Cosine compares directions only, from −1 to 1. The dot product also grows with length, so long vectors win (try it with the kitten query). Euclidean distance measures the gap between the tips. On vectors of length 1 all three agree: the dot product is the cosine, and the distance is √(2 − 2 cos).

          Search, RAG and vector databases

          Embed your documents once, in chunks, and store the vectors. At query time, embed the question and fetch the nearest chunks. Exact search compares against every vector; vector databases use approximate nearest neighbour (ANN) indexes such as HNSW, a layered graph walked greedily towards the query, trading a little recall for a large speed-up.

          Nearest isn’t the same as relevant. A vector search always returns k results, even when nothing in the store answers the question: pick sitting on the river bank and the third result is cat food. Useful score ranges differ between models, so tune a threshold on your own data, add a reranker, and let the model say it doesn’t know.

          In two dimensions there is only room for a few directions, so unrelated topics end up at odd angles, some even opposite with negative cosines. In hundreds of dimensions there is room for many nearly unrelated directions, and real models rarely produce strongly negative scores for ordinary text. Numbers match verify/ai/embeddings.py.

          Tokens and the context window

          Models read and write tokens, not characters or words, and every limit, price and speed is counted in them. These are real splits from o200k_base; a visible space is drawn as ␣.

          The context window

          The most tokens a model can attend to in one call: the system prompt, the conversation, retrieved documents, tool results and the answer it writes. Anything beyond it has to be cut, summarised or retrieved on demand.

          Cost

          APIs bill per token, with separate prices for input and output; output tokens usually cost more. A long system prompt is paid for on every call, which is why providers offer prompt caching for a repeated prefix.

          Latency

          The prompt is processed in parallel (prefill), which sets the time to first token. The answer is produced one token at a time (decode), so total time grows with its length. Streaming shows tokens as they arrive.

          The KV cache

          Attention needs a key and a value vector for every earlier token, in every layer. The server keeps them, so each new token only computes its own position. That makes decoding fast and makes long contexts expensive in GPU memory.

          Long isn’t perfect recall

          Models tend to use information at the start and end of a long prompt better than in the middle (“lost in the middle”). Retrieve the few chunks that matter rather than pasting everything, and put key instructions where they count.

          How a chat model is trained

          The same network goes through stages. Each changes the weights, and therefore the logits you saw in the stepper.

          1. PretrainingNext-token prediction over a huge corpus of text and code. The result, a base model, continues text and knows a lot, but doesn’t reliably follow instructions. Its knowledge stops at a training cut-off.
          2. Supervised fine-tuningAlso called instruction tuning: training on example conversations with good answers, in a chat format with roles. The model learns to follow instructions and to answer as an assistant.
          3. Preference tuningPeople (or a model) compare two answers. RLHF trains a reward model on those choices, then optimises the LLM against it with reinforcement learning (classically PPO). DPO skips the reward model and learns from preferred and rejected pairs directly.
          4. Reinforcement learning on checkable tasksReasoning models are further trained with reinforcement learning on problems whose answers can be verified, such as maths and code, which teaches them to work through longer chains of steps.

          Prompting, RAG or fine-tuning?

          Start with the cheapest tool that can work, and measure with evals before you move on to the next.

          Prompting

          Clear instructions, a few examples (few-shot) and an output format. Instant to change and free to try. Limited by what the model already knows and by the context window.

          Always the first step

          RAG

          Retrieval-augmented generation: fetch relevant passages and add them to the prompt. Knowledge stays up to date without retraining, answers can cite sources, and access control happens at retrieval time. Quality depends on chunking, embeddings and reranking.

          Private or changing knowledge

          Fine-tuning

          Further training on your examples changes behaviour: a format, a tone, a narrow task done well by a smaller, cheaper model. It is a poor way to add facts, which go stale. LoRA trains small adapter matrices instead of all the weights.

          Consistent style or a narrow task

          Tools

          Function calling: the model emits a structured call that matches a JSON schema, your code runs it, and the result goes back into the context. Use it for live data, exact arithmetic and actions.

          Fresh data and side effects

          What to remember

          One token at a timeAn LLM repeatedly predicts a probability for every next token, picks one and appends it. Chat, code and agents are all built on that loop.
          Tokens are the unitLimits, prices and latency are all counted in tokens. Numbers, rare words and many non-English languages need more of them.
          Temperature reshapes, it doesn’t informDividing logits by T sharpens or flattens the distribution, top-k and top-p cut its tail. T = 0 is greedy, and nearly but not perfectly deterministic.
          Embeddings are meaning as directionCosine similarity finds related text without shared words. ANN indexes such as HNSW make the search fast, and nearest isn’t always relevant.
          Hallucinations are plausible guessesThe model optimises likely text, not truth. Ground it, let it abstain, validate its output and evaluate it.

          Interview questions: How AI works

          Short answers you can say out loud. Try answering each question yourself before you open it.

          How does an LLM generate text?

          It runs a loop. The text so far is tokenised, passed through the network, and the network outputs a score (logit) for every token in its vocabulary. Softmax turns the scores into probabilities, a decoding rule picks one token, and that token is appended to the input for the next pass. It stops at a stop token or at the token limit.

          What is a token, and why not use words or characters?

          A token is a unit from a fixed vocabulary learned by the tokenizer, usually with byte-pair encoding: common words become one token, rarer words split into pieces, and anything can fall back to bytes. Characters would make sequences very long; whole words would need a huge vocabulary and still miss new words. Subwords balance the two. In English a token is roughly three quarters of a word, but numbers, code and many other languages need more tokens.

          What is temperature?

          A number that divides the logits before the softmax: softmax(logits / T). Below 1 it sharpens the distribution so likely tokens dominate; above 1 it flattens it so unlikely tokens get drawn more often. T = 0 is treated as greedy decoding. It changes how the model picks, not what it knows.

          What is a vector embedding?

          A fixed-length list of numbers that represents a piece of text (or an image, or code) so that similar meanings get vectors that point in similar directions. You get it from an embedding model, and compare embeddings with cosine similarity. It powers semantic search, clustering, deduplication and the retrieval step of RAG.

          What is a hallucination, and why does it happen?

          Fluent output that isn’t supported by facts or by the provided sources. The model is trained to produce likely text, not true text, so when it lacks a fact it produces a plausible one. Training data rarely says “I don’t know”, many evaluations reward guessing, knowledge has gaps and a cut-off, and sampling can pick a wrong token that later tokens then build on.

          How do you reduce hallucinations?

          Ground the model with retrieved sources and require citations, tell it that “I don’t know” is acceptable, and lower the temperature for factual tasks. Use tools for facts and arithmetic, ask for structured output and validate it in code, and measure with evals on known answers. Where errors are expensive, add a checking step or a human review.

          Temperature, top-k and top-p: what is the difference?

          Temperature reshapes the whole distribution. Top-k keeps only the k most likely tokens. Top-p (nucleus sampling) keeps the smallest set of most likely tokens whose probabilities add up to at least p, so the number kept adapts: few when the model is confident, many when it isn’t. After a cut, the survivors are renormalised and one is sampled.

          Greedy decoding or sampling: when do you use which?

          Greedy (or a low temperature) for extraction, classification, code and anything with one right answer, because you want the most likely output and repeatability. Sampling with a moderate temperature for creative writing, brainstorming or several varied candidates. Greedy decoding can get stuck in repetitive loops on long open-ended text.

          Why can temperature 0 still give different outputs?

          Greedy decoding has no randomness, but the arithmetic isn’t perfectly repeatable. GPU floating-point addition isn’t associative, and batching with other requests changes the order of operations, so two nearly tied logits can swap. One different token changes everything after it. Providers treat a seed as best effort, so design for small variations.

          What is the context window?

          The maximum number of tokens the model can process in one call, counting the system prompt, the conversation, retrieved documents, tool results and the generated answer. Anything beyond it must be truncated, summarised or retrieved when needed. Bigger windows cost more and models don’t use every part of a long context equally well.

          Why do tokens matter for cost and latency?

          APIs bill per input and output token, output usually at a higher price. Input is processed in parallel (prefill), which sets the time to first token; output is generated one token at a time, so total latency grows with answer length. Shorter prompts, prompt caching for repeated prefixes, smaller models and capped output lengths all help.

          What does attention do, at a high level?

          For each position, attention computes how relevant every earlier position is (by comparing a query vector with key vectors), then takes a weighted mix of their value vectors. That is how the word “is” can pull in “France” and “capital”. Many attention heads run in parallel in every layer, and a causal mask stops positions from seeing the future.

          What is the KV cache?

          During generation, each layer’s keys and values for all earlier tokens are stored instead of recomputed, so each new token only computes its own position. It makes decoding much faster but uses GPU memory that grows with context length and batch size, which is a main limit on long contexts and on how many requests a server can handle at once.

          What is cosine similarity, and why use it for embeddings?

          The dot product of two vectors divided by the product of their lengths: the cosine of the angle between them, from −1 to 1. It compares direction and ignores length, which matches how embedding models encode meaning. If vectors are normalised to length 1, cosine equals the dot product, which is cheaper to compute.

          Cosine, dot product or Euclidean distance?

          Use the metric the embedding model was trained for. Dot product is affected by vector length, Euclidean distance by both length and angle, cosine by angle only. On normalised vectors they all give the same ranking, since the squared distance is 2 − 2 cos.

          What is a vector database, and how does ANN search work?

          A store that indexes vectors with metadata and answers “which stored vectors are nearest to this one?”. Exact search compares against every vector, which is too slow at scale, so it uses approximate nearest neighbour indexes. HNSW builds a layered graph of neighbours and walks greedily towards the query from the sparse top layer down; IVF clusters vectors and searches only the closest clusters. You trade a little recall for a large speed-up.

          Walk me through a RAG pipeline.

          Offline: split documents into chunks, embed each chunk and store the vectors with metadata. Online: embed the question, retrieve the top-k chunks (often combined with keyword search), optionally rerank them, then put the best ones in the prompt with instructions to answer only from them and cite them. Evaluate retrieval (did we fetch the right chunk?) separately from generation (did we use it faithfully?).

          How do you choose chunk size and overlap?

          Chunks must be small enough to be specific and to fit several in the prompt, and big enough to carry complete ideas. Split on structure (headings, paragraphs) rather than fixed character counts where you can, and add some overlap so a sentence cut at a boundary still appears whole somewhere. Then tune by measuring retrieval quality on real questions.

          What is reranking, and why add it?

          A second stage that rescores the retrieved candidates with a stronger, slower model, usually a cross-encoder that reads the query and each passage together. Embedding search is fast but coarse; reranking the top 50 or so and keeping the best few improves precision, and its scores make a better threshold for dropping irrelevant passages.

          Why can RAG still hallucinate?

          Retrieval can return the wrong passage, because nearest isn’t necessarily relevant, or miss the right one. The model can ignore the context in favour of what it learned, mix sources, or answer beyond them. Mitigate with a relevance threshold or reranker, explicit permission to abstain, required citations, and a check that each claim is supported by a cited passage.

          How do grounding and citations help?

          Grounding means the answer is derived from sources you supplied. Asking for citations makes each claim traceable to a passage id, which lets users verify the answer and lets your code check that cited passages exist and actually contain the claim. It also discourages answers that have no support in the context.

          Prompting, RAG or fine-tuning: how do you choose?

          Start with prompting. Use RAG when answers depend on private or frequently changing knowledge, or need citations. Use fine-tuning to change behaviour: a consistent format or tone, or a narrow task handled by a smaller, cheaper model. Fine-tuning is a poor way to add facts. They combine well, for example a fine-tuned model inside a RAG pipeline.

          How is a chat model trained?

          Pretraining on a huge text corpus with next-token prediction gives a base model. Supervised fine-tuning on example conversations teaches it to follow instructions. Preference tuning, such as RLHF or DPO, aligns it with which answers people prefer. Reasoning models add reinforcement learning on tasks whose answers can be checked, such as maths and code.

          What is the difference between RLHF and DPO?

          Both learn from pairs of answers where one was preferred. RLHF first trains a separate reward model on those preferences and then optimises the LLM against it with reinforcement learning, classically PPO, with a penalty for drifting too far from the original model. DPO skips the reward model and the RL loop, and optimises the LLM directly on the preferred and rejected pairs with a simple classification-style loss. DPO is simpler and more stable to run.

          What is LoRA?

          Low-rank adaptation: instead of updating all the weights, you freeze the model and train small low-rank matrices added to some layers. It needs far less memory and storage, and you can keep several adapters for one base model. QLoRA does the same on a quantised base model to save even more memory.

          What is prompt injection, and how do you defend against it?

          Instructions hidden in content the model processes, such as a web page, an email or a retrieved document, that try to override your instructions, for example to leak data or call a tool. The model can’t reliably tell data from instructions, so you can’t fully prevent it with prompting. Limit what the model can do: least-privilege tools, confirmation for side effects, no secrets in the context, and output validation.

          How do structured output and function calling work?

          You give the model a JSON schema, either for its answer or for each tool it may call. The model emits JSON that matches it, and many providers can enforce the schema with constrained decoding, which only allows tokens that keep the output valid. For a tool call your code runs the function and sends the result back as a new message. Validate the arguments anyway; matching a schema doesn’t make the values correct.

          How do you evaluate an LLM application?

          Build a test set of real inputs with expected answers or grading criteria, and run it on every change of prompt, model or retrieval. Use exact checks where possible (valid JSON, correct field, cited passage contains the claim), and an LLM judge with a clear rubric where you can’t, checked against human labels. Track retrieval metrics, faithfulness, latency and cost, and monitor production traces.

          Why are LLMs bad at counting letters and doing arithmetic?

          They see tokens, not characters: strawberry is three tokens in o200k_base, so the letters aren’t directly visible. Numbers are split into chunks of up to three digits, which don’t line up with place value. Give the model a tool such as a code interpreter or a calculator instead.

          How do you cut cost and latency in production?

          Use the smallest model that passes your evals, and route only hard requests to a larger one. Shorten prompts, cache repeated prefixes, retrieve fewer but better chunks, cap output length, stream responses, and cache whole answers for repeated questions. Batch offline work, and measure time to first token and tokens per second separately.

          Token ids and counts come from tiktoken 0.14 with o200k_base; other model families use other tokenizers, so their ids and counts differ. Logits are illustrative. Every probability, similarity and sample on this page is computed exactly and matches the programs in verify/ai/.