Observed Signal · Jul 29, 2026 · Technical Release · Source: Nates Substack · Impact: 2/5 · Sentiment: Positive

Author Releases Token Saver Skill to Cut LLM Token Reuse

Executive Signal Summary

The author measured extreme token reuse in their local LLM usage (3.77 billion total tokens; 95.73% of input reported as reused) and built a "Token Saver" skill to aggressively reduce reported reused input by up to 90% without increasing errors or rework. The piece describes the measurement approach (one job run two ways), presents fifteen changes with conditions and measurements, lists nine immediate habits readers can adopt, explains what the Token Saver enforces in Codex and Claude Code, and discusses caching behavior and remaining unproven limits. The article is published 2026-07-29.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Provides practical techniques and a tool to significantly reduce LLM token reuse and costs for heavy users; relevant to teams managing LLM billing but not a major platform policy or industry-wide technical standard change.

SIGNAL RADAR

Track OpenAI Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Author recorded 3.77 billion tokens in a local Token Burn tracker.
  • Of 3.75 billion input tokens in the log, 3.59 billion were reported as reused input (95.73%).
  • Local usage data spanned 143 Codex threads and 28,877 local records.
  • Author's goal: reduce reported reused input by 90% without increasing mistakes, retries, review time, or repeated work.
  • The Token Saver skill implements changes for Codex and Claude Code and documents fifteen possible changes and nine immediately actionable habits.

Connected Companies & Entities

2 Entities mapped

“The log includes cumulative usage updates across 143 Codex threads and 28,877 local records....”

“The Token Saver skill. What it carries for you in Codex and Claude Code, the four changes it enforces, and the one boundary it can’t cross....”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: Nates Substack•Published: Jul 29, 2026
Original Coverage Title: “I Built The Token Saver Skill To Cut My Token Use By 90%. Here Is What It Can And Cannot Do For You.”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIMay 8, 2026

3 Hidden Token Sinks in Claude Code

A Dev.to technical follow-up by Prakash Ponali describes three additional sources of token waste in Claude Code after applying a Skill Vault pattern. Starting from a ~51K token per-session baseline, the author identified and fixed a bloated root CLAUDE.md (saving ~1.5K tokens), disabled claude-mem's SessionStart timeline auto-injection (~2K tokens), and re-vaulted 27 unused skills (~1.4K tokens). Combined, these changes shave roughly 4.9K tokens per session without losing capability. The post details file locations, audit scripts, and configuration edits (including a caution about plugin upgrades), and advocates auditing auto-loaded context and moving episodic content to on-demand access to improve LLM attention and efficiency.

Read assessment
Large Language Models (LLM) & AIJun 20, 2026

Developer Cuts AI Token Use by 82% with Tools

A developer published a hands-on guide showing how careful context management and tooling can dramatically reduce LLM token usage. Using a command-proxy and context-compression plugins across 6,000+ commands, the author recorded 7.4 million tokens saved—an 82% reduction. The post details three levers: trimming a resident rules file (CLAUDE.md), installing automatic context-compression plugins (RTK, claude-mem, codegraph), and model tiering to run grunt tasks on cheaper models. The author also explains prompt caching for billing discounts and warns of trade-offs (index build time, memory recall errors, over-compression). Publication date: 2026-06-20.

Read assessment
Large Language Models & AIMay 18, 2026

5 Tips to Reduce Claude Code Token Costs by 30%

A DEV Community post by Alaric (published 2026-05-18) shares five practical habits to cut token consumption when using Anthropic’s Claude Code. Recommendations include adding a concise CLAUDE.md at the project root so Claude Code can load durable context, scoping each session to a single task, using prompt caching aggressively, preferring the Read tool over pasting large files, and using smaller model variants (Sonnet or Haiku) for routine work. The author reports typical token savings of 25–35% and gives concrete examples (a ~70% cache hit rate and session input cost dropping from $0.60 to $0.18). The post also lists relative model-output costs and warns against ultra-cheap third-party relays and manual prompt compression.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.