Observed Signal · Jun 28, 2026 · Technical Guidance · Source: DEV Community · Impact: 2/5 · Sentiment: Positive
Stop Using tiktoken for Claude; Use countTokens
The author discovered a 15–20% undercount when using OpenAI's tiktoken to estimate Claude token usage and costs. Claude uses a different, model-specific tokenizer, so token counts (and costs) diverge — especially on code or non-English text — and can change between Claude model versions. Anthropic provides a messages.countTokens endpoint (and SDK wrappers) that returns the exact input token count for a specified Claude model; developers should call it for accurate cost/context estimates, avoid caching counts across model versions, and remember output tokens typically drive generation costs. The post includes SDK examples and published per-million pricing for Haiku 4.5, Opus 4.8, and Fable 5. Publication date: 2026-06-28.
Accurate token counting materially affects cost and context budgeting for LLM usage; the article documents a common developer mistake (using tiktoken for Claude) and points to a model-specific API (Anthropic's countTokens) that corrects it. This is practical, developer-facing guidance rather than a major platform policy or industry-shifting event.
Track OpenAI Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- tiktoken tokenizes for OpenAI models and undercounts Claude tokens by roughly 15–20% on typical English prose.
- Anthropic provides a messages.countTokens endpoint (wrapped in its SDK) that returns model-specific token counts for Claude.
- Token counts can change between Claude model versions (e.g., Opus 4.7 produces a higher count than Opus 4.6 for the same input).
- Published per-million token rates (2026): Haiku 4.5 input $1 / output $5; Opus 4.8 input $5 / output $25; Fable 5 input $10 / output $50.
- Recommendation: always call countTokens against the exact model used for inference and do not reuse cached counts across model versions.
Connected Companies & Entities
2 Entities mapped“I was counting Claude tokens with `tiktoken`, which is OpenAI's tokenizer....”
“import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic(); // The SDK wraps a dedicated count_tokens endpoint....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
5 Tips to Reduce Claude Code Token Costs by 30%
A DEV Community post by Alaric (published 2026-05-18) shares five practical habits to cut token consumption when using Anthropic’s Claude Code. Recommendations include adding a concise CLAUDE.md at the project root so Claude Code can load durable context, scoping each session to a single task, using prompt caching aggressively, preferring the Read tool over pasting large files, and using smaller model variants (Sonnet or Haiku) for routine work. The author reports typical token savings of 25–35% and gives concrete examples (a ~70% cache hit rate and session input cost dropping from $0.60 to $0.18). The post also lists relative model-output costs and warns against ultra-cheap third-party relays and manual prompt compression.
Reduce AI Agent Token Costs via CLI (2026 Guide)
A 2026 technical guide (published 2026-05-20) explains how CLI-based coding agents (examples: Claude Code and Codex) waste tokens and offers practical tactics to cut costs without changing models or lowering output quality. Recommended measures include narrowing file/directory scope, keeping project memory files (e.g., CLAUDE.md) short, compressing or clearing long sessions, enabling prompt (system-prefix) caching, routing simple subtasks to cheaper models, filtering and silencing noisy tool outputs, limiting RAG retrieval sizes, and measuring tokens/costs per run. The article provides command examples, estimated token-savings ranges for each tactic, a checklist for implementation, and sample cost-calculation formulas. It also links to tooling (Apidog) and provider-specific notes (OpenAI/Codex/Claude) where relevant.
Benchmark: Claude Code Sends 33k Tokens vs OpenCode 7k
A Systima benchmark measured token overhead from two agent harnesses: Claude Code 2.1.207 and OpenCode 1.17.18 (both targeting claude-sonnet-4-5). Claude Code injected roughly 33,000 tokens of system prompt, tool schemas and scaffolding before the user prompt, compared with about 7,000 tokens for OpenCode. Claude Code also rewrote many more prompt-cache tokens per session (up to 54x), increasing billable cache-write costs. The study shows real production setups (instruction files, MCP servers, gateways) can escalate bootstrap payloads to 75,000–85,000 tokens before user input, and session shape (batching vs repeated small turns) determines total cost.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
