Summary
Key takeaways
- Tokens are the units AI models process, and they can represent whole words, parts of words, punctuation, or other characters.
- One token is not the same as one word, and the same text can produce different token counts depending on the model and tokenizer.
- For rough English planning, 1,000 tokens is often estimated at around 750 words, but actual counts vary with language, formatting, code, names, and model encoding.
- AI requests usually contain more than the visible user message because instructions, conversation history, retrieved documents, tool definitions, and other context may also consume tokens.
- Input tokens and output tokens are often priced separately, while some providers also distinguish cached input, reasoning usage, or other billable categories.
- A context window defines how much token-based information a model can work with in one request, but a larger context window does not guarantee better use of every supplied detail.
- Context is not the same as permanent memory because an application may store information outside the model and include only a selected portion in each request.
- Token budgeting should account separately for instructions, main documents, conversation history, retrieved evidence, tools, and reserved output capacity.
- The cheapest price per million tokens does not automatically produce the cheapest workflow because retries, excessive context, failed tasks, tools, and human review can raise the real cost.
- Teams should measure cost per accepted task rather than optimizing token counts in isolation, because reducing tokens is useful only when output quality remains acceptable.
When this applies
This applies when a team needs to estimate AI API costs, understand context limits, plan RAG or agent workloads, or reduce the operating cost of an LLM-powered product. It is especially useful when applications include long instructions, uploaded documents, conversation history, retrieval results, tool definitions, or repeated agent steps. The framework is also useful when comparing models or providers because it encourages measuring complete workflow usage rather than relying on word counts or headline token prices.
When this does not apply
This does not apply directly to cryptocurrency tokens, authentication tokens, or other technologies that happen to use the same term. It is also less relevant when an AI product is sold through a fixed subscription and the user does not need to manage API consumption directly, although token usage may still matter behind the scenes. Rough word-to-token conversions should also not be used for precise budgeting, hard context limits, multilingual workloads, code, images, audio, or video.
Checklist
- Record the exact model and endpoint used by the workload.
- Use a tokenizer or counting method that matches the selected model.
- Measure the full input rather than only the visible user prompt.
- Separate instructions, documents, history, retrieved evidence, and tool material in the token budget.
- Reserve realistic capacity for generated output.
- Check both the context-window limit and any separate output limit.
- Avoid double-counting message structure or tool definitions already included in provider usage totals.
- Measure actual input and output tokens across several representative tasks.
- Record retries, failed attempts, latency, and whether each task was accepted.
- Apply the provider’s current input, output, cached, and other relevant billing rates.
- Calculate total cost per completed and accepted task rather than cost per API call alone.
- Retrieve only relevant evidence instead of automatically sending entire document collections.
- Use caching where it provides a measurable cost or latency benefit.
- Set sensible limits on agent steps and retries to prevent inexpensive calls from becoming expensive loops.
- Re-test answer quality after every token-reduction optimization before deploying it broadly.
Common pitfalls
- Assuming one token always equals one word.
- Using an English word-to-token estimate for every language and content type.
- Counting only the visible prompt while ignoring instructions, history, retrieval, and tools.
- Confusing context-window capacity with permanent memory or monthly usage limits.
- Assuming a larger context window automatically improves answer quality.
- Comparing providers only by their lowest advertised price per million tokens.
- Ignoring output tokens, reasoning usage, caching categories, retries, or other billable components.
- Reducing context aggressively without checking whether task quality or success rate declines.
- Allowing AI agents to retry or loop without clear limits.
- Forecasting monthly costs from one short test prompt instead of measuring representative production workloads.
Quick answer. Tokens are the units an AI model processes. For text, a token can represent a word, part of a word, punctuation, or other characters. Tokens affect request size and often API cost. The Uvik Software token budget worksheet helps you account for instructions, documents, conversation history, and generated output before estimating a workload.
This guide explains tokens used by language models. It does not cover cryptocurrency tokens or authentication credentials, which share the name but serve different purposes. You do not need to write code to understand the worksheet.
The useful distinction is between three questions: how much information goes into a request, how much the model generates, and how the provider bills that work. They are related, but a word count alone cannot answer all three.
Is one token the same as one word?
Short answer. No. A tokenizer divides text according to the model’s encoding. Common words may fit into one token, while other words need several. The same text can have different counts with different models or encodings.
OpenAI’s token explanation notes that spaces, capitalization, and language can affect the result. Use the target model’s counting method when the number matters. A visual split shown in an article is only an illustration unless it identifies the actual tokenizer and input.
A token ID is an identifier in a vocabulary. It is not the same thing as an embedding vector, which is a numerical representation used in a model or search system. Microsoft’s token guide describes those distinct steps. Keeping the terms separate helps when an engineer explains how text moves through an application.
How many words are in 1,000 tokens?
Quick answer. For rough planning, 1,000 tokens is often around 750 English words, but this is not an exact conversion. Language, formatting, code, and the selected model can change the count. Count the actual request before relying on a hard limit or a detailed budget.
Google’s token guide gives an approximate range of 60 to 80 English words per 100 tokens for Gemini. That would make 1,000 tokens roughly 600 to 800 English words. The range is more useful than treating every page or paragraph as a fixed number of tokens.
For example, a 3,000-word English document might be estimated at about 4,000 tokens using the 750 words per 1,000 tokens rule. That is a planning estimate. A document with tables, identifiers, unusual names, or several languages needs an actual count. Do not apply the English estimate to all languages.
What are input and output tokens?
Input is the material supplied to the model. Output is the material it generates. A request may include more input than the user’s visible question: instructions, prior messages, retrieved passages, tool descriptions, and other structure can also be present.
| Category | Plain English meaning | Budget question |
|---|---|---|
| Input | Material the model receives | How much is sent on each call |
| Output | Material the model generates | How much is generated per completed task |
| Cached input | Reused input handled by a provider’s caching feature | Which part qualifies and at what rate |
| Reasoning | Internal generation used by some models | How is it reported and billed for this model |
A short visible answer can still involve additional internal work. OpenAI documents reasoning tokens as part of output usage. Other providers expose their own categories and accounting rules. Do not estimate every model’s bill from the number of words shown in the chat window.
For workplace products, a subscription may bundle usage or charge through another unit. Tokens remain useful for understanding the system, but API token rates do not automatically describe an enterprise seat plan.
What is a context window?
Short answer. A context window is the amount of token-based information a model can work with in a request. Check the model’s separate output limit as well. A large context window does not guarantee that the model will use every supplied detail correctly.
Google’s documentation describes the context window as a combined input and output limit. The exact handling of reasoning, media, history, and other material depends on the model and endpoint. Applications may shorten, summarize, or reject an oversized request. Do not assume that every app always drops the oldest message in the same way.
Context is also different from permanent memory. A product may store files or conversation history outside the model and select part of it for a later request. That storage does not mean every stored item is present in the current context.
Figure 1. AI token budget example | © 2026 Uvik Software
The Uvik Software token budget example allocates a hypothetical 128,000 token request capacity. It is not the specification of a named model. The reserved output allowance includes any internal output the model’s rules count against that limit.
Copy the Uvik Software token budget worksheet
Record the model, endpoint, counting method, and date at the top of the worksheet. The example below is hypothetical. It shows why a document upload can be only part of the total request.
| Budget component | Example token allowance | Your measured value |
|---|---|---|
| Instructions and request structure | 8,000 | Enter measured count |
| Main documents | 40,000 | Enter measured count |
| Included conversation history | 20,000 | Enter measured count |
| Retrieved evidence and tool material | 20,000 | Enter measured count |
| Reserved output | 8,000 | Enter chosen allowance |
| Total planned use | 96,000 | Sum the components |
| Remaining capacity within 128,000 | 32,000 | Capacity minus planned use |
The total is a capacity plan, not a bill. Reserving room for 8,000 output tokens does not mean the model will generate or charge for exactly that amount. Use actual usage records for the cost calculation.
Also avoid double counting. If a provider’s complete input counter already includes message structure and tool definitions, do not add them again. Use category totals that reconcile to the full request count.
Preparing and selecting data affects how much relevant evidence fits into a request. Uvik Software’s data engineering consulting service covers that upstream work. Test input reductions against the same answer quality criteria.
How do token prices turn into a bill?
Quick answer. For a simple text API request, multiply input tokens by the input rate and output tokens by the output rate, using the provider’s billing unit. Add other categories and charges where they apply. Compare cost per accepted task as well as cost per call.
Here is an invented rate card for teaching the calculation: $2 per million input tokens and $8 per million output tokens. These are not current prices for a named provider. Assume a request uses 6,000 billable input tokens and 1,000 billable output tokens, with no other charges.
| Cost component | Calculation | Result |
|---|---|---|
| Input | 6,000 ÷ 1,000,000 × $2 | $0.012 |
| Output | 1,000 ÷ 1,000,000 × $8 | $0.008 |
| Total per call | $0.012 + $0.008 | $0.020 |
| 10,000 identical calls | 10,000 × $0.020 | $200 |
If those 10,000 calls produce only 8,000 accepted completed tasks, the model charge per accepted task is $0.025. The $200 includes the unsuccessful calls. Divide that full charge by the 8,000 accepted tasks. Storage, search, tools, support, and human review would increase the full application cost if they apply.
This is why the lowest price per million tokens is not enough to choose a system. One workflow may need more calls, more context, or more corrections to finish the same task. Run a representative sample and compare the complete result.
How do you count tokens accurately?
Use a tokenizer or request counting feature that matches the selected model. For OpenAI, the Tokenizer helps inspect text; the official guide also describes complete input counting. For Gemini, the token guide describes a count before the request and usage information after it. Input counting does not predict the exact generated output.
Keep a small measurement log: task category, input count, output count, other billable usage, retries, latency, and whether the task was accepted. Measure several real examples from each important workflow. One short prompt is a weak basis for a monthly forecast.
For images, audio, and video, use the provider’s counting and billing rules for the selected model. Record relevant settings, such as image resolution or audio duration, with the measured usage. Text word counts do not capture the full media workload.
How can a team reduce token cost without damaging quality?
Remove repeated material that does not help the task. Retrieve relevant passages instead of attaching every document by default. Ask for the output format and level of detail the user actually needs. Test whether the shorter request still passes the same evaluation.
Use caching where the provider supports it and the workload benefits. Check eligibility, storage duration, and the actual rate categories. Caching is a billing and performance feature, not a reason to send confidential data outside an approved path.
Set a sensible limit on agent steps and retries, and provide a useful fallback when the limit is reached. A cheap individual call can become an expensive loop. Uvik Software’s guide to using AI in software development provides context for testing the whole workflow, while its generative AI consulting service covers the system design decisions behind it.
A measured example of reducing tool context
In a November 2025 engineering example, Anthropic reported reducing the tokens used for tool definitions from 150,000 to 2,000 by loading tools as needed through a code-based approach. That is a 98.7% reduction in that part of the example. It is not a general reduction in the complete bill, and the approach adds secure execution requirements.
The useful budgeting lesson is to separate the parts of the request. Tool descriptions and intermediate results can consume a large share even when the user’s question is short. Record them separately before deciding which text to reduce.
Use this Uvik Software comparison record for a small optimization experiment. Keep the same tasks and acceptance criteria in both runs; replace the recording prompts in each column with your measured results.
| Measure | Before the change | After the change |
|---|---|---|
| Input tokens by source | Instructions, documents, history, tools | Same categories and task set |
| Complete usage cost | Include retries and other billable categories | Same cost boundary |
| Accepted tasks | Accepted count divided by tasks tested | Same review rules |
| Response time | Median and 95th percentile | Same start and finish events |
Reducing tokens is useful only if the workflow still meets its purpose. A change that saves input tokens but adds failed attempts may raise cost per accepted task. Inspect the failures before expanding the change, and keep any extra execution or maintenance costs in the budget.
Use and cite the worksheet
Suggested citation: Uvik Software, What are tokens in AI and how do they affect cost, 2026. Credit the Uvik Software token budget worksheet and original visual when adapting them. All capacity allocations and prices in the worked examples are hypothetical.