Token & Cost Planner
Estimate context usage, request cost, and monthly spend with assumptions you can edit and inspect.
- Sample workflow
- Exportable
- Runs in this browser
Runs in your browser. Nothing you enter is sent anywhere.
Built in pricing effective 2026-07-26. Every price below is a starting point you can overwrite.
These are planning estimates, not a bill
Every number on this page comes from the assumptions you enter, run through the visible formulas below. This tool has no connection to any billing account and cannot see what a provider actually charged you.
Awaiting input
Enter a workload, or start from a realistic example
Fill in a system prompt size, average tokens, and a request volume below, or load a sample comparing a flagship model against a cheaper one with prompt caching turned on.
Plan settings
Both scenarios read their request volume in this unit.
Compare mode runs the same formulas twice, side by side.
Scenario A
Picking a profile fills in the three prices below. Edit any of them afterward.
USD per 1,000,000 uncached input tokens.
USD per 1,000,000 output tokens.
USD per 1,000,000 tokens served from a prompt cache.
Tokens spent on the fixed system prompt for every call.
Visitor facing input per request, not counting the system prompt.
Requests handled per the time period selected above.
Percent of requests that need one extra full retry.
Percent of system prompt plus input tokens served from cache.
Scenario A: derived assumptions
Every value below is stated, not hidden.
- Effective input tokens per call
- 0 from system prompt plus input tokens fixed
- Cached tokens per call
- 0 from from your cache hit rate fixed
- Uncached tokens per call
- 0 from from your cache hit rate fixed
- Model calls per request
- 1.000 from from your retry rate fixed
- Requests per day
- 0 from normalized from your period fixed
Scenario B
Picking a profile fills in the three prices below. Edit any of them afterward.
USD per 1,000,000 uncached input tokens.
USD per 1,000,000 output tokens.
USD per 1,000,000 tokens served from a prompt cache.
Tokens spent on the fixed system prompt for every call.
Visitor facing input per request, not counting the system prompt.
Requests handled per the time period selected above.
Percent of requests that need one extra full retry.
Percent of system prompt plus input tokens served from cache.
Scenario B: derived assumptions
Every value below is stated, not hidden.
- Effective input tokens per call
- 0 from system prompt plus input tokens fixed
- Cached tokens per call
- 0 from from your cache hit rate fixed
- Uncached tokens per call
- 0 from from your cache hit rate fixed
- Model calls per request
- 1.000 from from your retry rate fixed
- Requests per day
- 0 from normalized from your period fixed
Cost breakdown
Per call, per request, and projected over the selected period.
Idle
Enter a workload above to see a cost estimate.
Token allocation
Where tokens go on a single model call, and over the period.
Idle
No prompt caching yet
High retry rate
Large projected monthly spend
Note How this planner works
Every model call is billed for the tokens it reads and the tokens it writes. This planner splits the tokens a call reads into a system prompt plus your average input, then splits that total again by your cache hit rate, because a cached token and a fresh token are priced differently. A retry rate above zero adds a fraction of an extra call to every request, since a retried request pays for both attempts.
Token counts here are whatever you enter. A rough rule of thumb for English prose is about four characters per token, but the exact count depends on the model's tokenizer, so use a real count from your own prompts when you have one.
Token & Cost Planner stopped responding.
Something in this tool threw an error. Your input was not sent anywhere, and nothing was saved. Reloading the page usually clears it.