How do I estimate token usage for my AI agent?
Calculate tokens per interaction
Tokens are chunks of text; roughly 1 token = 0.75 words in English. Use the tokenizer provided by your model (e.g., OpenAI's tiktoken) to count tokens in your system prompt, user input, and expected output. Sum these for a single interaction.
If your agent uses memory or context, include the tokens for retrieved documents or conversation history. For multi-step agents, each step (e.g., tool call, reasoning) adds tokens. Estimate the average number of steps per user request.
- Count tokens in system prompt
- Count tokens in user input
- Estimate output length
- Add tokens for context/memory
- Multiply by steps per request
Scale to your expected volume
Multiply tokens per interaction by the number of interactions per day, then by 30 for monthly usage. Add 20-30% buffer for retries, errors, and variability. Monitor actual usage after launch and adjust.
Remember that input and output tokens may be priced differently. Also, some models have different tokenization (e.g., Claude vs. GPT). Test with real examples from your users to get accurate averages.
Common mistakes
- Using word count as a proxy for tokens; it's inaccurate, especially for code or non-English text.
- Forgetting to include system prompts and context, which can be large.
- Underestimating output length; agents often generate more text than expected.
