Build an AI usage budget from measured requests
Use observed usage rather than prompt character counts.
Define one workload
A short classification task and a long document answer have different token patterns. Group requests by use case, model and relevant configuration before estimating cost. Record the source of the price and whether the billing unit is per token, per million tokens or another meter.
Measure the whole request
Include repeated instructions, retrieved context and conversation history in input usage. Include all billable output categories identified by the provider. A visible answer length can be different from the billable output total. Use the documented usage fields instead of assuming a fixed characters-to-tokens ratio.
Keep cache meters separate
Provider field names do not always mean the same thing. In the Claude API, input_tokens excludes input counted in cache_creation_input_tokens and cache_read_input_tokens. Add those three categories to describe total input, but price each category at its applicable rate. Do not treat the ordinary input field as the full request or charge cached tokens a second time. Other providers may report a total with cached usage as a subset; check their own definitions.
Work a small category example
Suppose one request reports 200 ordinary input tokens, 800 cache-read tokens, no cache writes and 100 output tokens. With invented per-million rates of $2, $0.20 and $8 respectively, cost is (200 × 2 + 800 × 0.20 + 100 × 8) ÷ 1,000,000 = $0.00136. Ten thousand such requests cost $13.60 before other charges. These are illustrative rates, not a provider quote. Our two-rate calculator excludes separate cache categories; use a category worksheet for this example.
Model the tail as well as the average
Record a representative average and a high-usage case. Include failed attempts and retries where they are billed. A budget based only on the shortest successful request can understate demand. Keep a separate estimate for tools, retrieval, storage and other services.
Reconcile after release
Compare the estimate with billing exports over the same time interval and account scope. Investigate unit mismatches, pricing tiers, caching and unexpected request volume. A spending alert is useful, but an alert alone does not stop new work; use the platform’s supported controls where available.
Keep working through the question
MyDev Codes · Published . Updated . Prepared with AI assistance; editorial approach and corrections are described on our method page. Examples are illustrative, not case-study results.
References and further reading
- Anthropic: tracking prompt-cache usage — Defines the three Claude input-usage fields. The example rates and calculation above are our own.