When you press Enter on a prompt, a chat model does not "look up" an answer or "understand" you the way a colleague would. Your text is cut into small pieces called tokens, the model reads all of them together with the rest of the conversation, and it then produces its reply one piece at a time. Knowing that sequence explains most of the behavior that confuses beginners: why wording matters, why long chats drift, and why a confident answer can still be wrong.
This article walks through that sequence in five stages, using the documentation that OpenAI and Anthropic publish about their own models. It also gives you a small token-budget calculation you can reuse, and a checklist that turns each stage into a habit. If you are brand new to these tools, our 30-day AI skills plan for beginners shows how to practice; this piece explains why the practice works.
The Short Answer
Based on OpenAI's and Anthropic's own documentation, the process has four moving parts. First, your text is divided into tokens. Second, the model processes those tokens together with everything else in its context window. Third, it generates output tokens, one after another. Fourth, that output becomes part of the input for the next turn. Everything else in this article is a consequence of those four facts.
Stage 1: Your Text Becomes Tokens
OpenAI's Help Center describes a token as something that "can represent a character, part of a word, a whole word, or punctuation," and states the sequence plainly: the text is divided into tokens, the model processes those tokens, and the model generates output tokens. As a rule of thumb for English, OpenAI says 1 token is approximately 4 characters, or about three-quarters of a word. Anthropic's glossary describes tokens as the smallest individual units of a language model and says that for Claude a token represents approximately 3.5 English characters, with the exact number varying by language.
Those two figures are not identical, and neither company presents its figure as exact. Treat them as estimates. What matters for a beginner is the practical consequence: the model never sees your sentence the way you wrote it. It sees a sequence of pieces, which is why both companies say the exact count varies by language.
Stage 2: Everything Goes Into the Context Window
Anthropic's documentation defines the context window as "all the text a language model can reference when generating a response, including the response itself," and calls it a "working memory" for the model, distinct from the large body of data the model was trained on. OpenAI's Help Center puts the same idea in terms of limits: the context window limits the tokens a model can work with in a request, and models also have an output limit.
Two things follow. First, the model can only use what is in that window, so anything you want it to take into account (your audience, your constraints, the document you are working from) has to be in the conversation. Second, the window is finite. Anthropic's documentation notes that if the input alone exceeds the model's context window, the API returns a "prompt is too long" error; in consumer chat apps the details differ by product, so check the app's own help pages rather than assuming.
Stage 3: The Model Generates the Reply One Piece at a Time
Anthropic's glossary explains that autoregressive language models, which includes the model family behind Claude, are pretrained to predict the next word given the previous context of text in the document. OpenAI's tokens article describes the output side in the same terms: the model generates output tokens. In plain English, the reply is built incrementally, each new piece chosen based on everything that came before it, including the pieces the model has just written.
Anthropic's glossary also defines temperature as a parameter that controls the randomness of a model's predictions during text generation, with higher temperatures leading to more creative and diverse outputs. This is why the same prompt can return different wording on different runs. Whether you can adjust this setting depends on the product you use, so the practical takeaway is simple: if you need a different variation, ask for one instead of assuming the tool failed.
Stage 4: The Reply Joins the Conversation
Anthropic's context-window documentation says that as a conversation advances, each user message and assistant response accumulates within the context window, and previous turns are preserved completely. Each turn's input contains all previous history plus the new message, and the generated response becomes part of the input for the next turn.
This is why a long chat can feel like it drifts. Early instructions sit far behind a lot of newer text, and the window is shared between everything you have said and everything the model has replied. The documentation does not claim the model weighs all of that text equally, and this article does not either. What it does support is a habit: when a conversation has wandered, start a new one and paste in only what matters.
Stage 5: Why a Fluent Answer Can Still Be Wrong
OpenAI defines hallucinations as "plausible but false statements generated by language models." In its explanation of why they happen, OpenAI points to two causes. During pretraining there are no true/false labels attached to each statement, so arbitrary low-frequency facts cannot be predicted from patterns alone. And standard training and evaluation procedures reward guessing over acknowledging uncertainty. OpenAI also states that hallucinations are not inevitable, because language models can abstain when uncertain, but that is a statement about what better training could achieve, not a guarantee about the tool in front of you today.
For you as a user, this means fluency is not evidence. A rarely documented fact, a specific statistic, a citation, or a niche detail is exactly where a confident answer needs checking against a primary source.
The 5-Stage Prompt Journey (and What to Do About Each Stage)
Here is the sequence as a single reference table. The right-hand column is our own practical reading of the documented behavior, not a claim made by OpenAI or Anthropic.
| Stage | What happens (documented) | Habit it suggests (our reading) |
|---|---|---|
| 1. Tokenize | Text is divided into tokens (about 4 characters each in OpenAI's rule of thumb, about 3.5 in Anthropic's for Claude) | Write plain, complete sentences and define names and acronyms the first time |
| 2. Load context | Model works within a limited context window that also holds its own response | Put the audience, goal, constraints and source material in the chat instead of assuming the model knows them |
| 3. Generate | Reply is produced token by token, predicting what comes next given prior text | If the first sentence goes the wrong way, correct it early or restart rather than patching a long draft |
| 4. Accumulate | Every turn, and every reply, stays in the window and becomes input for the next turn | Start a fresh chat when a task changes; carry over only the essential facts |
| 5. Verify | Fluent output can still be a plausible but false statement | Check names, numbers, dates and quotes against a primary source before you use them |
Worked Example: A Simple Token Budget
The numbers below are an illustration built on the published rules of thumb, not measurements of any specific model. Real counts vary by model and language, so treat the results as rough estimates.
Suppose you write a 1,200-character prompt that pastes in a client brief. At about 4 characters per token, that is roughly 300 tokens (1,200 ÷ 4); at about 3.5 characters per token, roughly 343. Suppose the reply is about 2,000 characters, or roughly 500 tokens at the 4-character rule. On your second message, you add a 400-character follow-up (about 100 tokens). The model's input for that second turn now contains the first prompt, the first reply, and the follow-up: roughly 300 + 500 + 100 = 900 tokens, before it writes a single word of its answer.
The formula generalizes: tokens in the window at turn N ≈ (all earlier prompts + all earlier replies + the new message) + the reply being written. Three short exchanges are nowhere near a limit. A long document pasted in, followed by many long replies, is how a chat gets close to one. That is the practical reason to keep source material tight and to open a new chat for a new task.
A 6-Point Checklist Before You Press Enter
- Have I stated the goal, the audience and the format in words, rather than expecting the model to infer them?
- Is every fact the model needs actually in this conversation, not just in my head?
- Have I pasted only the source material that matters, not the whole file?
- Is this still the same task as the rest of this chat, or should I start a new one?
- If the answer includes a name, number, date or quote, do I know where I will verify it?
- If I want a different angle, have I asked for one explicitly?
Frequently Asked Questions
Does the AI understand what I write? The documentation describes a process: text is tokenized, processed, and used to generate output tokens. The companies' own materials describe the mechanism, not human-style comprehension, and this article stays with what they document.
Does the model remember me between chats? The context-window documentation describes memory within a conversation. Whether a given app stores anything across conversations is a product feature that differs by app, so check that product's help pages rather than assuming.
Does a better prompt fix hallucinations? A clearer prompt helps with ambiguity, but according to OpenAI, hallucinations also come from how models are trained and evaluated. Better wording does not replace verification. For the wording side, see why ChatGPT prompts fail and the CLEAR framework and system prompts versus regular prompts.
What should I do with a task too big for one prompt? Split it into steps and pass only the necessary output from one step to the next; prompt chaining covers how. If you are still choosing a tool, our beginner comparison of ChatGPT, Claude and Gemini is the place to start.
Want the Foundations in One Structured Path?
The AI Blueprint walks you through these basics step by step, with the prompts and templates already built, so you can apply them without designing your own learning plan.
Get The AI Blueprint — $27Sources: OpenAI Help Center, "What are tokens and how to count them" (token definition, ~4 characters and ~three-quarters of a word per token, processing sequence, context and output limits); Anthropic, "Context windows" (working-memory definition, accumulation across turns, "prompt is too long" error); Anthropic, glossary (tokens ~3.5 characters for Claude, next-word pretraining, temperature); OpenAI, "Why language models hallucinate" (hallucination definition, causes, not inevitable). All sources checked live on September 30, 2026. The token-budget example uses illustrative numbers derived from the published rules of thumb. The habits in the table and checklist are our own interpretation, not statements from OpenAI or Anthropic. Last reviewed: September 2026.