Save 30-90% on Claude Code costs with an installable skill, CLI tools, and 12 deep-dive guides.
30-60% is the typical result for a mixed real workload. Up to 90% is the ceiling when you stack every lever -- prompt caching, Batch API, model routing, and context discipline -- against an unoptimized all-Opus baseline. Every number is sourced: see How far can you actually go?.
Claude Code (official plugin system):
/plugin marketplace add Sagargupta16/claude-cost-optimizer
/plugin install cost-mode@claude-cost-optimizerMulti-agent (Cursor, Cline, Codex, 40+ agents):
npx skills add Sagargupta16/claude-cost-optimizerThen activate in any session:
/cost-mode # Standard (40-60% output token reduction)
/cost-mode lite # Professional brevity (20-40% output reduction)
/cost-mode strict # Telegraphic, max savings (60-70% output reduction)
/cost-mode off # Resume normal behavior
| Feature | How It Saves Tokens |
|---|---|
| Strips filler | Drops pleasantries, hedging, restating your question, trailing summaries |
| Suggests cheaper models | "Haiku handles this -- /model haiku" for simple tasks |
| Suggests CLI tools | "Use prettier/eslint --fix directly" instead of burning LLM tokens |
| Session awareness | Reminds to /compact after 20+ turns, fresh sessions for new tasks |
| Minimal code gen | Diffs over rewrites, no obvious comments, no speculative error handling |
| Auto-deactivates | Full clarity for security warnings, destructive ops, and when you're confused |
Technical accuracy is never sacrificed. Code in commits and PRs is written normally.
claude-rate runs on your filesystem -- no signup, no GitHub upload, no network round-trip. Pick whichever runner fits your shell:
# curl one-shot (no Node, no install)
curl -sSL https://raw.githubusercontent.com/Sagargupta16/claude-cost-optimizer/main/tools/claude-rate/install.sh | sh -s -- .
# curl, persistent install
curl -sSL https://raw.githubusercontent.com/Sagargupta16/claude-cost-optimizer/main/tools/claude-rate/install.sh | sh -s -- --installAdd --fix to print copy-pasteable fix commands, --strict to fail CI when the grade drops below B, or --json for machine-readable output. See tools/claude-rate/README.md for the full breakdown.
The local rater works on private and uncommitted repos and inspects things the web analyzer can't reach: settings.local.json, whether your Read(...) deny rules cover the lock files actually on disk, and secrets in your working tree. The web analyzer applies the same 7-category rubric (CLAUDE.md, file-read exclusions, settings, MCP, hooks, security, skills/agents/commands) to any public repo.
| Tool | What It Does |
|---|---|
| Repo Analyzer | Paste a GitHub URL to get a cost audit, grade (A+ to F), and recommendations |
| Cost Calculator | Estimate monthly spend with interactive charts: per-turn cost curve, savings breakdown, model comparison |
| Badge Checker | Score your setup and get a shields.io badge for your repo |
Worked example, 30-turn session, priced at legacy Opus 5 rates (identical math on Opus 4.8 -- same $5/$25). On Opus 5.5, the current Opus flagship ($4/$20, cache hits at 0.05x), every line item is cheaper, so both columns come out lower. The dollar figures are arithmetic from the posted rates, not a metered bill. The MCP share of the drop assumes MCP tool search is off (full tool schemas loaded up front); under the default tool search only tool names and server instructions enter context, so that part of the saving is smaller:
BEFORE (no optimization): AFTER (5 minutes of setup):
CLAUDE.md: 6,200 chars CLAUDE.md: 2,800 chars
Read deny rules: none Read deny rules: 12 patterns
MCP servers: 6 active MCP servers: 2 active
System prompt: ~15,000 tokens/turn System prompt: ~5,500 tokens/turn
Session cost: $2.85 Session cost: $1.12
Monthly (3x/day): $188.10 Monthly (3x/day): $73.92
Savings: $114.18/month (61%)
Short answer: 30-60% is what a typical mixed workload saves. Up to ~90% is the ceiling when every lever stacks against a naive all-Opus, no-cache, verbose baseline. The headline numbers below are each real and sourced -- but each one applies only to the favorable slice of your spend, so the compound rarely holds across an entire real workload. Treat 90% as a ceiling, not a promise.
| Lever | Published savings | Applies to | Source |
|---|---|---|---|
| Prompt caching | up to 90% cost (95% on Opus 5.5, 97.5% on Fable 5.1), 85% latency | cached input only (0.1x input price; 0.05x on Opus 5.5; 0.025x on Fable 5.1 / Mythos 5.1) | Anthropic: Prompt caching, pricing docs |
| Batch API | flat 50% off input and output, stacks with caching | any async (24h) work | Anthropic: Message Batches API, batch docs |
| Model routing | ~75% (Opus 5.5->Haiku is a flat 4x ratio; 5x from legacy Opus 5); RouteLLM up to 85% at 95% quality | tasks a cheaper model handles well | RouteLLM (arXiv 2406.18665), LMSYS |
| Context management | 84% fewer tokens in a 100-turn eval | long agentic sessions | Anthropic: Context management |
| Subscription vs API | ~90%+ for heavy users (Pro $20 / Max $100-200 flat) | power users vs pay-as-you-go | ksred cost tracker (n=1) |
- "up to 90%" is a ceiling, not a typical result. Anthropic itself publishes 90% for caching alone, and the levers genuinely multiply against an unoptimized baseline. But each 90% is a best case on its favorable slice (cached input for caching; conversational traffic for routing), so a whole mixed workload lands well below the sum.
- The stacked total has no source. Each lever above is individually sourced; the combined figure is not, and this repo holds no controlled multi-lever measurement. For a number that applies to your workload, run claude-rate on your own repo and price only the levers you can apply.
- The cost-mode skill on its own delivers 30-60% -- it does output-token reduction and model-routing hints, not batch/caching/subscription. The 90% ceiling needs the full playbook in the guides, not just the skill.
- Caching and batch are the firmest floors (first-party Anthropic pricing; batch is a documented flat 50% that provably stacks with caching). Routing and subscription savings are the most workload-sensitive.
Prices verified against the pricing reference below (2026-09-29). Cache hit = 0.1x base input on every model except Fable 5.1 and Mythos 5.1 (0.025x) and Opus 5.5 (0.05x); Batch = 50% off both input and output.
Even without installing the skill, these 5 changes cut costs immediately:
| # | Strategy | Savings | Guide |
|---|---|---|---|
| 1 | Keep CLAUDE.md under 200 lines -- it loads in full every session; Anthropic's guidance is that longer files "consume more context and reduce adherence". Move workflow-specific instructions into skills or path-scoped .claude/rules/ |
No published figure | Context Optimization |
| 2 | Use Haiku for simple tasks (--model haiku) -- 4x cheaper than Opus 5.5 |
20-40% | Model Selection |
| 3 | Use Plan Mode before coding (Shift+Tab) -- prevents wasted iterative cycles | 15-25% | Workflow Patterns |
| 4 | Add Read(...) deny rules to permissions.deny in .claude/settings.json -- keep Claude's file tools out of node_modules, dist, lock files (.claudeignore is not a Claude Code feature) |
No published figure | Context Optimization |
| 5 | Delegate to subagents -- isolate expensive searches from main context | 20-40% | Workflow Patterns |
Full walkthrough: Getting Started in 5 Minutes
cost-mode is the first skill. More are planned:
| Skill | Status | What It Does | Cost Impact |
|---|---|---|---|
| cost-mode | Live | Concise responses, model routing suggestions, session awareness | 30-60% (skill alone) |
| deny-rules-gen | Planned | Generates permissions.deny Read(...) rules for your project's tech stack |
No published figure |
| context-compress | Planned | Rewrites your CLAUDE.md to be shorter while keeping all essential info | 10-20% input reduction |
| cache-optimizer | Planned | Detects cache-busting patterns and suggests fixes to maximize prompt cache hits | 10-25% input reduction |
| budget-guard | Planned | Per-session and per-day spending limits with warnings before you blow past them | Prevents overspend |
Want to build one? Skills are just SKILL.md files -- see CONTRIBUTING.md and the skills/cost-mode/ directory for the pattern.
12 deep-dive guides covering every optimization area:
| Guide | What You'll Learn |
|---|---|
| 00 - Getting Started | Zero to optimized in 5 minutes -- the essential setup |
| 01 - Understanding Costs | How billing works, what costs the most, where money goes |
| 02 - Context Optimization | Reduce input tokens: CLAUDE.md, Read deny rules, file reads |
| 03 - Model Selection | When to use Opus vs Sonnet vs Haiku (with decision tree) |
| 04 - Workflow Patterns | Plan mode, subagents, commands, batch operations |
| 05 - Team Budgeting | Per-developer budgets, cost tracking, ROI calculation |
| 06 - Access Methods & Pricing | Compare API vs Bedrock vs Vertex AI vs Claude Code pricing |
| 07 - MCP & Agent Cost Impact | MCP server overhead, subagent costs, Agent SDK patterns |
| 08 - Prompt Caching Deep Dive | Cache mechanics, TTL economics, maximizing hit rates, ROI math |
| 09 - Subscription Plan Value | Choose the right plan, maximize allowance, upgrade/downgrade signals |
| 10 - Three-Tier Task Routing | Skip the LLM for Tier 0 tasks, route cheap tasks to Haiku, save Opus for complex work |
| 11 - Speed vs Cost | Make Claude faster without burning money -- free latency levers first, Fast Mode economics last |
Also: Visual Diagrams (Mermaid flowcharts) | One-Page Cheatsheet
Copy-paste configs that are already optimized:
CLAUDE.md: Minimal | Standard | Comprehensive | Monorepo
By Stack: React+Vite | Next.js | FastAPI | MERN | Terraform | Go | Rust | Django | Rails | Spring Boot
Settings: Cost-Conscious | Balanced | Performance-First
Commands: /cost-check | /budget-mode | /quick-fix | /optimize
The 8 tools in tools/, plus the budget hooks in hooks/. Full tools documentation
| Tool | What It Does |
|---|---|
| claude-rate | Grade your local setup on the 7-category rubric (recommended entry point) |
| Token Estimator | Estimate token count and cost for any file |
| Usage Analyzer | Find cost hotspots across your sessions |
| Badge Generator | Grade your project config (A+ to F) from the CLI |
| MCP Cost Server | In-session cost estimation via MCP |
| VS Code Extension | Token count and cost in the status bar |
| GitHub Action | Automated cost audit on PRs |
| /optimize Command | Claude Code custom command that audits the current project |
| Budget Hooks | Track tool calls, log costs, warn at thresholds |
Verified 2026-09-29 against Anthropic's pricing, models overview, Opus 5.5 overview, Sonnet 5.5 overview, and model deprecations pages.
| Model | Input / 1M | Output / 1M | Cache Hit / 1M | 5m Cache Write / 1M | 1h Cache Write / 1M | Context | Max Output | Min cacheable prompt |
|---|---|---|---|---|---|---|---|---|
| Fable 5.1 (highest capability) | $10.00 | $50.00 | $0.25 | $12.50 | $20.00 | 1M | 128K | 512 |
| Mythos 5.1 (limited, Glasswing) | $10.00 | $50.00 | $0.25 | $12.50 | $20.00 | 1M | 128K | 512 |
| Fable 5 (legacy) | $10.00 | $50.00 | $1.00 | $12.50 | $20.00 | 1M | 128K | 512 |
| Mythos 5 (limited, Glasswing) | $10.00 | $50.00 | $1.00 | $12.50 | $20.00 | 1M | 128K | 512 |
| Opus 5.5 (Opus flagship, recommended default) | $4.00 | $20.00 | $0.20 | $5.00 | $8.00 | 1M | 128K | 512 |
| Opus 5 (legacy) | $5.00 | $25.00 | $0.50 | $6.25 | $10.00 | 1M | 128K | 512 |
| Opus 4.8 (legacy) | $5.00 | $25.00 | $0.50 | $6.25 | $10.00 | 1M | 128K | 1,024 |
| Opus 4.7 | $5.00 | $25.00 | $0.50 | $6.25 | $10.00 | 1M | 128K | 2,048 |
| Opus 4.6 | $5.00 | $25.00 | $0.50 | $6.25 | $10.00 | 1M | 128K | 4,096 |
| Opus 4.5 | $5.00 | $25.00 | $0.50 | $6.25 | $10.00 | 200K | 64K | 4,096 |
| Opus 4.1 (retired 2026-08-05, still on Bedrock + Google Cloud) | $15.00 | $75.00 | $1.50 | $18.75 | $30.00 | 200K | 32K | 1,024 |
| Sonnet 5.5 (Sonnet flagship) | $2.00 | $10.00 | $0.20 | $2.50 | $4.00 | 1M | 128K | 512 |
| Sonnet 5 (legacy) | $2.00 | $10.00 | $0.20 | $2.50 | $4.00 | 1M | 128K | 1,024 |
| Sonnet 4.6 | $3.00 | $15.00 | $0.30 | $3.75 | $6.00 | 1M | 64K | 1,024 |
| Sonnet 4.5 | $3.00 | $15.00 | $0.30 | $3.75 | $6.00 | 200K | 64K | 1,024 |
| Haiku 4.5 | $1.00 | $5.00 | $0.10 | $1.25 | $2.00 | 200K | 64K | 4,096 |
| Mythos Preview (deprecated, no retirement date) | $25.00 | $125.00 | $2.50 | $31.25 | $50.00 | 1M | -- | 2,048 |
Mythos 5.1's rates come from the pricing page; its context window, max output and cache floor are not listed there. The Fable 5.1 model page states Mythos 5.1 "shares Claude Fable 5.1's specifications and pricing", and the prompt-caching page lists it at the 512-token floor.
Opus 5.5 (claude-opus-5-5, released 2026-09-22) is the current Opus flagship, and Anthropic's models overview now says to "start with Claude Opus 5.5 for most workloads". It costs $4/$20 -- 20% below the $5/$25 of Opus 5 and Opus 4.8, the first Opus release to lower the rate, and cache hits read at $0.20, 0.05x base input (95% off). It is not a drop-in swap of the model string -- see what actually changes first. Opus 5 (claude-opus-5, GA 2026-07-24) is now legacy at $5/$25. 1M context on Fable 5.1, Mythos 5.1, Fable 5, Mythos 5, Opus 5.5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5.5, Sonnet 5, and Sonnet 4.6 bills at standard rates across the full window (no long-context premium). Sonnet 5.5 (claude-sonnet-5-5, released 2026-09-28) is the current Sonnet flagship at $2/$10 -- identical to Sonnet 5, including caching and batch ($1/$5), with the same tokenizer, so moving to it is free at the posted rate; it still rejects some Sonnet 5 request shapes -- see what changes. Sonnet 5 (claude-sonnet-5, GA 2026-06-30), now legacy, is $2/$10 per MTok permanently -- the launch rate was labelled introductory through 2026-08-31, but Anthropic made it standard and cancelled the increase to $3/$15, so Sonnet 5 costs half as much as Opus 5.5 (and 60% less than legacy Opus 5). Batch API: 50% off both input and output, and it stacks with caching -- Opus 5.5 batch is $2/$10 (Opus 5: $2.50/$12.50), with up to 300K output via the output-300k-2026-03-24 beta. Fast Mode (research preview, Opus 5.5, Opus 5 and Opus 4.8 only): 2x each model's own base rate -- $8/$40 on Opus 5.5, $10/$50 on Opus 5 and 4.8 -- up to 2.5x output tokens/sec. Regional endpoints (Bedrock / Vertex AI / Claude API inference_geo: "us" for 4.6+ models): +10%. Subscriptions: Pro $20/mo (or $200/yr ≈ $16.67/mo, ~17% off), Max 5x $100/mo, Max 20x $200/mo. Web search: $10 per 1,000 searches plus token costs. Web fetch: free beyond token costs. Code execution: free with web search/fetch; otherwise 1,550 free hours/month then $0.05/hour per container. Bash tool: +325 input tokens on Opus 5 / 4.8 / 4.7 (+244 on Opus 4.6 and earlier). Text editor tool: +700 input tokens.
Minimum cacheable prompt is a real cost lever. A
cache_controlblock below the model's threshold is silently ignored -- you pay full input price every turn and see no error. Opus 5.5, Opus 5 and Sonnet 5.5 all sit at 512 tokens, half the 1,024 threshold of Opus 4.8 and Sonnet 5, so system prompts and CLAUDE.md files that never cached on those start caching on Opus 5.x and Sonnet 5.5. Haiku 4.5, Opus 4.6, and Opus 4.5 sit at 4,096, the worst of the current lineup.Opus 5.5 / 5 / 4.8 / 4.7 tokenizer caveat: The tokenizer introduced with Opus 4.7 may use up to 35% more tokens for the same fixed text. Effective per-task cost is higher than posted pricing suggests -- factor this into budgets, especially when comparing against Opus 4.6 / Sonnet 4.6.
Fast Mode (research preview): Now Opus 5.5, Opus 5 and Opus 4.8 only, via the
fast-mode-2026-02-01beta header (speed: "fast"). All three are 2x their own base rate: $8/$40 per MTok on Opus 5.5, $10/$50 on Opus 5 and Opus 4.8. Up to 2.5x more output tokens/second; the speed gain is on output tokens/sec, not time-to-first-token. Opus 4.7 now returns an error onspeed: "fast"with no fallback, and Opus 4.6 silently runs at standard speed and standard rates (usage.speedcomes back"standard") -- if you were paying 6x for Fast Mode on either, that option is gone. Claude API + Managed Agents only: not on Claude Platform on AWS, Bedrock, Vertex AI, Microsoft Foundry, Batch API, or Priority Tier. Switching speeds invalidates prompt cache. Join the waitlist.Claude Code's
sonnetalias depends on your provider: it is Sonnet 5.5 only on the Anthropic API; on Bedrock, Google Cloud and Microsoft Foundry it resolves to Sonnet 4.5 ($3/$15, 200K context -- 1.5x the price of Sonnet 5.5 for a smaller window), and on Claude Platform on AWS to Sonnet 4.6 ($3/$15, 1M). Pin it with--model claude-sonnet-5-5(Bedrock:anthropic.claude-sonnet-5-5) orANTHROPIC_DEFAULT_SONNET_MODEL; Sonnet 5.5 needs Claude Code v2.1.284 or later.Fable 5 / Mythos 5 (GA 2026-06-09): Anthropic's Mythos-class tier above Opus, at $10/$50 -- 2x legacy Opus 5's price and 2.5x Opus 5.5's. Same specs for both: 1M context at standard rates, 128K max output, always-on adaptive thinking (control depth with
effort;thinking: disablednot supported), 4.7-generation tokenizer. Fable 5 is GA everywhere (Claude API, Claude Platform on AWS, Bedrock, Vertex AI, Microsoft Foundry) and includes safety classifiers that can decline requests -- a refusal returns HTTP 200 withstop_reason: "refusal", pre-output refusals are not billed, and the betafallbacksparameter plus fallback credit make retrying on another model cheap. Mythos 5 is the same model without the classifiers, limited to approved Project Glasswing customers. No Fast Mode on either; Batch API supported ($5/$25). Requires 30-day data retention (no zero-data-retention option).Mythos Preview: superseded by Mythos 5 and deprecated -- still functional, no longer recommended, and no retirement date is published. The invite-only defensive-cybersecurity research preview under Project Glasswing.
Looking for older model IDs and pricing? See the Legacy & Retired Models section below for migration context.
Opus 5.5 is 20% cheaper than Opus 5, but the bill does not simply fall 20%. These changes move it in both directions or break the request:
| Change | Why it matters for cost |
|---|---|
| $4/$20, cache hit $0.20 | Down from $5/$25 and $0.50. Cache hits are 0.05x base input (95% off) instead of 0.1x. Batch $2/$10, Fast Mode $8/$40. |
| Thinking is always on | thinking: {type: "disabled"} and thinking: {type: "enabled", budget_tokens} both return 400. Code that disabled thinking on Opus 5 now pays for thinking tokens, billed as output at $20/1M, that it did not pay for before. |
Default effort drops to medium |
Opus 5 defaulted to high. A request that omits effort now thinks less than it did on Opus 5. Effort is the only depth control; all five levels are supported. |
| Forced tool use returns 400 | tool_choice any/tool is rejected -- use auto plus strict tool use or structured outputs. The tool-use system prompt is 286 tokens. |
| Sampling params and prefill return 400; no Priority Tier | Non-default temperature/top_p/top_k and assistant prefill fail. Opus 4.8 keeps Priority Tier; Opus 5.5 has none. |
| Thinking blocks are tied to the model and conversation | Text between tool calls now comes back in thinking blocks (empty at the default display setting). On the Claude API and Google Cloud, computer_20251124 is rejected -- use computer_toolset_20260801. |
Re-baseline cost after migrating rather than assuming the 20%: the lower rate pulls the bill down, and thinking you used to disable pushes it up. Audit your prompts at the same time -- in Anthropic's published runs, prompts written for Opus 4.8 cost 36% more per ticket on Opus 5 for no accuracy gain, and once audited they were 14% cheaper and more accurate (cost and intelligence guide). Full details: Anthropic's Opus 5.5 migration guide.
Sonnet 5.5 costs exactly what Sonnet 5 does and uses the same tokenizer, so the posted rate does not move. Each change below either returns a 400 when Sonnet 5 code is carried over or shifts what you pay:
| Change | Why it matters for cost |
|---|---|
| Same $2/$10, same tokenizer | Cache hit $0.20 (0.1x base input), 5m write $2.50, 1h write $4.00, Batch $1/$5. The same text gives the same token count, so migrating is free at the posted rate. No Fast Mode. |
| Min cacheable prompt drops to 512 tokens | Sonnet 5 needed 1,024, so short system prompts that never cached on Sonnet 5 now do. The tool-use system prompt is 286 tokens (Sonnet 5: 354), a small per-request input saving. |
| Thinking is on by default and cannot be disabled | Adaptive thinking is on at default effort high, and reasoning tokens bill as output at $10/1M. thinking: {type: "disabled"} and manual budget_tokens return 400; the lowest setting is thinking: {type: "between_tools"}, which turns off up-front thinking and is accepted only at low, medium and high effort. |
| Effort levels are recalibrated | The same level does not produce the same amount of thinking as on Sonnet 5, so re-run your effort sweep instead of carrying a setting over. Anthropic's starting points: high in general, medium for well-specified agentic coding and multistep tool use. |
| Forced tool use and sampling params return 400 | tool_choice any/tool is rejected -- use auto plus strict tool use or structured outputs. Non-default temperature/top_p/top_k fail. |
| Thinking blocks are tied to the model and conversation | Keep conversations append-only: replaying a Sonnet 5.5 thinking block after editing earlier history can return 400. On the Claude API and Google Cloud, computer_20251124 is rejected -- use computer_toolset_20260801. The advisor tool rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors for a Sonnet 5.5 executor. |
| Tools and instructions can change mid-conversation | Mid-conversation tool changes (beta), mid-conversation system messages and per-message effort (beta) work on Sonnet 5.5 and not on Sonnet 5, so changing them no longer costs you the prompt cache. |
Full details: What's new in Sonnet 5.5.
Opus 5 is now legacy; if you are moving off Opus 4.8 today, go straight to Opus 5.5. The notes below still apply to anyone pinned on Opus 5. Opus 5 is the same price as Opus 4.8, but it is not a drop-in swap of the model string. Four things change your bill or break your request:
| Change | Why it matters for cost |
|---|---|
| Thinking is ON by default | Omit thinking and Opus 5 thinks adaptively. Reasoning tokens bill as output at $25/1M, and max_tokens is a hard cap on thinking plus text -- an unchanged max_tokens: 4096 now gets eaten by thinking before your answer is written. Raise it to 64K+ if you run xhigh/max effort. |
thinking: {type: "disabled"} is effort-gated |
Legal only at effort high or below. Pairing it with xhigh or max returns a 400, so a config that worked on 4.8 can hard-fail. |
| Min cacheable prompt drops to 512 tokens | Prompts too short to cache on 4.8 (1,024) now cache on 5. Free savings if you re-check your cache_control placement. |
| Cybersecurity classifiers ship on Opus 5 | Security-adjacent work can hit stop_reason: "refusal". Set the server-side fallbacks param (beta server-side-fallback-2026-07-01) to auto-retry on Opus 4.8 inside the same call rather than paying for a failed round-trip. |
Two prompt-level cleanups worth doing at the same time: Opus 5 writes longer output than 4.8 by default, so re-tune your verbosity instructions; and it self-verifies, so any "double-check your work before answering" instruction you carried over from an older model is now paying twice for the same behavior. New beta mid-conversation-tool-changes-2026-07-01 also lets you change tool definitions between turns without invalidating the prompt cache -- previously a cache-busting move.
Full details: Anthropic's Opus 5 migration guide.
Reference only -- don't use these for new work. Kept for migration context if you're unwinding code that still pins old model IDs.
Recently retired (requests now fail):
| Model | Retired on | Migrate to |
|---|---|---|
Claude Opus 3 (claude-3-opus-20240229) |
2026-01-05 | Opus 5.5 |
Claude Sonnet 3.7 (claude-3-7-sonnet-20250219) |
2026-02-19 | Sonnet 5.5 |
Claude Haiku 3.5 (claude-3-5-haiku-20241022) |
2026-02-19 (still on Bedrock + Vertex AI) | Haiku 4.5 |
Claude Haiku 3 (claude-3-haiku-20240307) |
2026-04-20 | Haiku 4.5 |
Claude Sonnet 4 (claude-sonnet-4-20250514) |
2026-06-15 | Sonnet 5.5 |
Claude Opus 4 (claude-opus-4-20250514) |
2026-06-15 | Opus 5.5 |
Claude Opus 4.1 (claude-opus-4-1-20250805) |
2026-08-05 (still on Bedrock + Google Cloud) | Opus 5.5 |
| Claude Sonnet 3.5 v1 / v2, Sonnet 3, Claude 2.x, Claude 1.x, Instant 1.x | 2024-2025 | See deprecations page |
Deprecated: Claude Mythos Preview (claude-mythos-preview) is the one model in Deprecated state -- still functional, no longer recommended, and Anthropic publishes no retirement date for it. Migrate to Mythos 5 (Glasswing). Every other model that has not retired reads Active -- Sonnet 4.5, Haiku 4.5 and Opus 4.5 included -- so no dated forced migration is outstanding. Published retirement dates are "not sooner than" dates; the nearest is Sonnet 4.5, not sooner than 2026-09-29.
Older snapshots still callable (not retired, but not the headline tier):
| Snapshot | Pricing | Context | Earliest retirement | Why use |
|---|---|---|---|---|
| Opus 5 | $5/$25 | 1M | 2027-07-24 | Previous flagship -- pin if you need thinking disabled or forced tool_choice, both of which Opus 5.5 rejects |
| Opus 4.8 | $5/$25 | 1M | 2027-05-28 | Older flagship -- pin if prompts are tuned to it, if you need thinking off at xhigh/max (Opus 5 rejects that combination), or if you need Priority Tier (Opus 5.5 has none). Also the fallback target for Opus 5 cyber refusals |
| Opus 4.7 | $5/$25 | 1M | 2027-04-16 | Pinned workloads. No Fast Mode -- speed: "fast" now errors |
| Opus 4.6 | $5/$25 | 1M | 2027-02-05 | Pinned workloads / older tokenizer. speed: "fast" silently runs standard |
| Opus 4.5 | $5/$25 | 200K | 2026-11-24 | Pinned workloads only |
| Sonnet 5 | $2/$10 | 1M | 2027-06-30 | Previous Sonnet flagship, same price as Sonnet 5.5 -- pin if you need thinking disabled or forced tool_choice, both of which Sonnet 5.5 rejects |
| Sonnet 4.6 | $3/$15 | 1M | 2027-02-17 | Pinned workloads -- migrate to Sonnet 5.5 |
| Sonnet 4.5 | $3/$15 | 200K | 2026-09-29 | Pinned workloads only |
Authoritative source: platform.claude.com/docs/en/about-claude/model-deprecations.
The cheatsheet has a more detailed legacy table including last-known pricing for every retired tier.
- Task Comparison -- same task, optimized vs not
- Model Comparison -- Opus vs Sonnet vs Haiku
- Context Size Impact -- how CLAUDE.md size affects cost
- Community Leaderboard -- crowdsourced cost-per-task data
- Case Studies -- submission template and format; no stories published yet, contributions welcome
- GitHub Discussions -- ask questions, share strategies
- Contributing -- from starring the repo to building new skills
- Issue Templates -- tips, benchmarks, case studies
Stars read on GitHub 2026-09-28. The savings column is each project's own claim, not something this repo has measured; -- means the project publishes no savings figure.
| Project | Stars | What It Does | Savings (project's own claim) |
|---|---|---|---|
| caveman | 108k | Brevity skill plus proxy -- strips filler from responses | Claims "cuts 65% of tokens" |
| claude-mem | 95k | Persistent compressed context across sessions | -- |
| rtk | 82k | CLI proxy that compresses common dev-command output before Claude reads it | Claims "60-90%" on common dev commands |
| claude-code-router | 37k | Routes Claude Code requests across models | -- |
| ccusage | 18.7k | npx ccusage usage and cost reports from your local logs |
-- |
| Claude-Code-Usage-Monitor | 8.7k | Real-time usage monitor with predictions | -- |
| claudetop | 212 | htop-style cost and cache monitor | -- |
How much does Claude Code actually cost?
With Pro ($20/mo or $200/yr ≈ $16.67/mo with annual billing -- 17% off), Max 5x ($100/mo), or Max 20x ($200/mo), you get included usage. For enterprise deployments, Anthropic's own average is about $13 per developer per active day and $150-250 per developer per month, with 90% of users under $30 per active day (Claude Code costs docs); run /usage in Claude Code to see your own session cost. Opus 5, 4.8, 4.7, and 4.6 at $5/$25 are 3x cheaper per token than Opus 4.1 ($15/$75), and Opus 5.5 at $4/$20 is cheaper still -- but the tokenizer introduced with Opus 4.7 (shared by Opus 5.5) can use up to 35% more tokens, so the effective gap for the $5/$25 models is closer to ~2x.
Does this apply to the Claude API too?
Context optimization, model selection, and prompt engineering apply to both. The skill and commands are Claude Code-specific.
Will optimization reduce output quality?
No. These strategies eliminate waste (duplicate context, unnecessary file reads, expensive models for simple tasks). Quality stays the same or improves -- less noise means better reasoning.
What's the biggest single change I can make?
Install cost-mode (npx skills add Sagargupta16/claude-cost-optimizer) and switch to Haiku for routine tasks. Combined: 50-70% savings.
If this repo helped you save money, consider giving it a star!
If you found this useful, check out my other AI/Claude tools:
| Project | Description |
|---|---|
| claude-code-recipes | 47 copy-paste recipes for Claude Code - commands, subagents, hooks, skills |
| claude-skills | Custom Claude Code plugin marketplace with dev-workflow, FARM stack, and more |
| agent-recipes | AI agent workflows for real-world dev tasks - code review, testing, security |
| ai-git-hooks | AI-powered git hooks - auto-review diffs, generate commit messages, security scanning |
| mcp-toolkit | Production-ready middleware for MCP servers - auth, caching, rate limiting |
MIT - use these strategies, templates, and tools however you want.