Wapgee Logowapgee
JavaScript

Count Tokens for a Claude Prompt and Estimate API Cost

Count tokens for a Claude prompt in your browser, then estimate the API cost per call and per month. Free token counter, with a worked example.

Qamar Abbas
· 6 min read
Count Tokens for a Claude Prompt and Estimate API Cost
Count Tokens for a Claude Prompt and Estimate API Cost

If you build on the Anthropic API, you pay by the token, not by the month. That makes one question worth answering before you ship: how many tokens is this prompt, and what will it cost me at scale?

You can get a number in seconds. To count tokens for a Claude prompt, paste it into the free Token Counter, pick a Claude model, and it shows the token count and the cost per call and per month. This post explains what the number means, walks through a real example, and shows how to bring the bill down.

What a token is, and why word count misleads

Claude does not read words. It reads tokens, which are chunks of text the model has learned as units. Common English words are often one token. Longer or rarer words split into several. Punctuation, whitespace, code and JSON all cost tokens too.

A rough rule for English prose is that one token is about three quarters of a word. Do not rely on it for code, JSON or non-English text, which usually need more tokens per word. Counting beats guessing.

Why the count matters: cost and context

Two things depend on the token count. The first is the context window, the total the model can handle in one request. Your system prompt, the conversation history, the user message, any tool definitions and the reply all share it.

The second is price. Input and output tokens are billed at different rates, and output costs several times more. On the Claude models in the counter, output is five times the input price.

What counts toward the input

People often count only the user message and get a number that is far too low. Everything you send in a request is billed as input:

  • the system prompt, which is repeated on every call

  • every earlier message in the conversation that you send back

  • the new user message and any pasted documents

  • tool definitions and the results your tools return

This is why a chat feature that looks cheap on the first message gets steadily more expensive. Paste the whole payload you actually send, not just the last message.

How to count tokens for a Claude prompt

You do not need a library or a script.

  1. Open the Token Counter.

  2. Paste your full prompt, including the system prompt.

  3. Choose how many output tokens you expect per call, and how many calls you make per month.

  4. Read the token count and compare the cost across Claude Haiku 4.5, Sonnet 5 and Opus 5.

One caveat, stated plainly. The counter runs an open tokenizer (the one used by GPT-4o) in your browser. Claude uses its own tokenizer, so treat the result as a close estimate, good for budgeting and for comparing models. For an exact number, see the section on Anthropic's counting endpoint below.

A worked example: a code review prompt

Suppose you are building a code review bot. To count tokens for a Claude prompt like this one, paste it into the tool:

You are a senior Python developer. Review the function below and suggest improvements for performance and readability. Return the refactored code and a bulleted list of changes.

def calculate_total(items):
    total = 0
    for item in items:
        total = total + item['price'] * item['quantity']
    return total

That is 45 words. The counter estimates it at 104 tokens. Now assume each review returns about 500 tokens of refactored code and notes. Using the list prices in the tool (August 2026, in line with Anthropic's pricing page), one call costs:

Model             Input price   Output price   Cost per call
Claude Haiku 4.5  $1 / M        $5 / M         $0.0026
Claude Sonnet 5   $2 / M        $10 / M        $0.0052
Claude Opus 5     $5 / M        $25 / M        $0.0130

A fraction of a cent per call looks harmless. Scale shows the difference.

Estimate the monthly API cost

Say the bot runs 10,000 reviews a month. On Claude Sonnet 5, the arithmetic is:

Input:   104 tokens x 10,000 calls = 1,040,000 tokens  -> 1.04 M x $2  = $2.08
Output:  500 tokens x 10,000 calls = 5,000,000 tokens  -> 5.00 M x $10 = $50.00
Total:   $52.08 per month

Two things stand out. Output is 96 percent of the bill even though the prompt is a fifth of the size of the reply. And switching the same workload to Haiku 4.5 gives $26.04 a month, while Opus 5 gives $130.20. Choosing the smallest model that does the job is the biggest lever you have.

When you need an exact count

For billing-grade numbers, Anthropic provides a token counting endpoint that returns the exact input tokens for a given model without running it:

import anthropic

client = anthropic.Anthropic()

count = client.messages.count_tokens(
    model="claude-sonnet-5-5",
    messages=[{"role": "user", "content": "Your prompt here"}],
)
print(count.input_tokens)

A practical split: use the browser counter to count tokens for a Claude prompt while drafting and comparing models, and the endpoint when you need to audit real traffic.

Four ways to cut the token count

  1. Trim the system prompt. It is sent on every request. Removing 50 redundant tokens from a prompt that runs 10,000 times a month saves 500,000 input tokens.

  2. Stop resending the whole conversation. Summarise old turns or retrieve only the relevant ones, because every message you resend is billed again.

  3. Use prompt caching. If a long block repeats across requests, such as a document or a fixed instruction set, Anthropic's prompt caching bills the repeated part at a lower rate.

  4. Cap and shape the output. Because output is the expensive side, ask for concise answers and set a sensible max_tokens.

Tool definitions count too. Every tool you pass is part of the input, so a large set adds tokens to every call. If you are building agents, see how to write a Claude tool definition that stays focused, then paste it into the counter to see what it adds.

Why in the browser matters

Prompts often contain customer data, internal code or unreleased copy. The Token Counter runs in your browser, so the text you paste is not uploaded, logged or stored. That is the reason to prefer a client-side tool for anything sensitive.

FAQ

Do Claude and GPT count tokens the same way?

No. Each model family has its own tokenizer, so the same text gives different counts. The tool uses an open GPT-4o style tokenizer, which is why its Claude numbers are estimates.

Are tokens per word the same in every language?

No. English is relatively efficient. Other languages, and code, often need more tokens for the same meaning, so count the real text instead of converting from words.

How accurate is the estimate for Claude?

Close enough to budget and to compare models, not exact. The tokenizer differs from Claude's, so expect a small gap. Use Anthropic's counting endpoint when you need the precise figure.

Can I count tokens for a Claude prompt without an API key?

Yes. The browser counter needs no key and no account, so you can count tokens for a Claude prompt before you write any code.

Does the Token Counter send my prompt anywhere?

No. The count and the cost maths run locally in your browser.

Can I count a whole document?

Yes. Paste it in and compare the total with your model's context window before you send it.

Ready to count tokens for a Claude prompt of your own? Open the Claude Token Counter. If your prompt includes a tool definition, build it with the LLM Tool Schema Builder first. For typed validation of the JSON your model returns, try the JSON Schema to Zod converter or read how to turn a JSON Schema into Zod and TypeScript types.

#claude#token counter#count tokens#claude api cost#llm pricing
Written by
Qamar Abbas

Comments

Sign in to join the conversation.

No comments yet. Be the first.

More in JavaScript

All posts →