%PDF-1.4 %âãÏÓ 1 0 obj << /Type /Catalog /Pages 2 0 R >> endobj 2 0 obj << /Type /Pages /Count 6 /Kids [5 0 R 7 0 R 9 0 R 11 0 R 13 0 R 15 0 R] >> endobj 3 0 obj << /Type /Font /Subtype /Type1 /BaseFont /Helvetica >> endobj 4 0 obj << /Type /Font /Subtype /Type1 /BaseFont /Helvetica-Bold >> endobj 5 0 obj << /Type /Page /Parent 2 0 R /MediaBox [0 0 595.28 841.89] /Resources << /Font << /F1 3 0 R /F2 4 0 R >> >> /Contents 6 0 R >> endobj 6 0 obj << /Length 5450 >> stream BT /F2 22 Tf 0.06 0.08 0.12 rg 1 0 0 1 46 789.89 Tm (How Claude Code Usage Limits Work, and How to) Tj ET BT /F2 22 Tf 0.06 0.08 0.12 rg 1 0 0 1 46 762.89 Tm (Cut Your Token Usage) Tj ET BT /F2 11 Tf 0.72 0.14 0.18 rg 1 0 0 1 46 725.89 Tm (TechRounder PDF Edition) Tj ET BT /F1 9.5 Tf 0.36 0.39 0.46 rg 1 0 0 1 46 709.89 Tm (Live article: https://www.techrounder.com/ai/claude-code-usage-limits-token-optimization/) Tj ET q 0.82 0.85 0.9 RG 1 w 46 691.39 m 549.28 691.39 l S Q BT /F1 10 Tf 0.24 0.27 0.32 rg 1 0 0 1 46 679.39 Tm (By Vipin PG | Published August 20, 2026 | Updated August 20, 2026 | Format: Deep Dive | 11 min read) Tj ET BT /F2 13 Tf 0.72 0.14 0.18 rg 1 0 0 1 46 656.39 Tm (In brief) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 636.39 Tm (Claude Code's usage limits, a rolling five-hour session window and a weekly cap, draw from the same) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 621.39 Tm (shared allowance as regular Claude.ai chats and Cowork sessions, with heavy coding tasks consuming) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 606.39 Tm (far more tokens than ordinary conversations due to file contents, tool calls, and multi-step reasoning.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 591.39 Tm (Understanding how prompt caching works, what breaks it, and which habits genuinely reduce token) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 576.39 Tm (consumption \(like clearing between tasks, using Plan Mode, keeping CLAUDE.md under 200 lines, and) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 561.39 Tm (matching models to task complexity\) can help you stay within limits that would otherwise drain in a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 546.39 Tm (single afternoon of pasting whole files or running large parallel workflows.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 521.39 Tm (If you have hit a Claude Code usage limit after what felt like a handful of prompts, you are not) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 506.39 Tm (imagining it. The five-hour session window and the weekly cap draw from the same pool as your) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 491.39 Tm (regular Claude.ai chats and Cowork sessions, and Anthropic's own cost guidance points out that a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 476.39 Tm (single heavy debugging session can outweigh an entire day of ordinary chat use, since a coding turn) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 461.39 Tm (carries file contents, tool calls, and multi-step reasoning that a plain chat message does not. One long) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 446.39 Tm (conversation, one large parallel workflow, or an afternoon spent pasting whole files instead of the) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 431.39 Tm (relevant lines can drain a budget that would otherwise last for days.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 409.39 Tm (This guide breaks down how the limits actually work, where the tokens go during a normal session,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 394.39 Tm (and which habits and settings genuinely change the outcome. Most of it comes down to paying attention) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 379.39 Tm (to numbers Claude Code already shows you, not finding a clever way around them.) Tj ET BT /F2 15 Tf 0.08 0.1 0.14 rg 1 0 0 1 46 351.39 Tm (The two clocks that control your access) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 327.39 Tm (Claude Code enforces two allowances at the same time. The session limit is a rolling five-hour) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 312.39 Tm (window. It starts counting from your first prompt and clears five hours later, not at a fixed clock) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 297.39 Tm (time. The weekly limit sits above it and resets on a fixed day assigned to your account. Both draw on) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 282.39 Tm (the same underlying allowance, and usage counts against both at once, so a single burst of heavy) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 267.39 Tm (activity, such as a large parallel workflow, can exhaust your weekly allowance before the session) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 252.39 Tm (window has even reset.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 230.39 Tm (When you hit one, Claude Code tells you exactly when it clears:) Tj ET BT /F1 9.5 Tf 0.18 0.2 0.24 rg 1 0 0 1 54 206.39 Tm (You've hit your session limit · resets 3:45pm) Tj ET BT /F1 9.5 Tf 0.18 0.2 0.24 rg 1 0 0 1 54 189.515 Tm (You've hit your weekly limit · resets Mon 12:00am) Tj ET BT /F1 9.5 Tf 0.18 0.2 0.24 rg 1 0 0 1 54 172.64 Tm (You've hit your Opus limit · resets 3:45pm) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 154.14 Tm (The session and weekly limits are shared across every model, so switching with `/model` will not) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 139.14 Tm (restore access to either one. The Opus limit works differently: it only covers Opus requests, so) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 124.14 Tm (dropping to Sonnet with `/model` gets you working again immediately.) Tj ET q 0.86 0.88 0.92 RG 1 w 46 42 m 549.28 42 l S Q BT /F1 8.4 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 30 Tm (TechRounder | Page 1 of 6) Tj ET BT /F1 7.2 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 19 Tm (https://www.techrounder.com/pdf/blog/claude-code-usage-limits-token-optimization.pdf) Tj ET endstream endobj 7 0 obj << /Type /Page /Parent 2 0 R /MediaBox [0 0 595.28 841.89] /Resources << /Font << /F1 3 0 R /F2 4 0 R >> >> /Contents 8 0 R >> endobj 8 0 obj << /Length 6232 >> stream BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 789.89 Tm (The allowance itself is also shared more broadly than most people expect. It covers Claude Code, the) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 774.89 Tm (Claude.ai chat interface, and Cowork together, and its size depends on your seat tier rather than one) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 759.89 Tm (number that applies to everyone. A long research session in the regular Claude.ai chat window already) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 744.89 Tm (eats into what is left for your afternoon in the terminal.) Tj ET BT /F2 15 Tf 0.08 0.1 0.14 rg 1 0 0 1 46 716.89 Tm (Why nobody can give you an exact token number anymore) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 692.89 Tm (Search around and you will still find posts estimating a Pro plan at somewhere around 44,000 tokens) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 677.89 Tm (per five-hour window, with the Max tiers scaled up from there. Treat numbers like that as dated,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 662.89 Tm (third-party estimates rather than official figures. Anthropic's current documentation deliberately) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 647.89 Tm (avoids advertising a fixed count, because how much a developer actually spends depends heavily on) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 632.89 Tm (which model they run, how large the codebase is, and habits like keeping several sessions or) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 617.89 Tm (automated jobs going at once. The one hard number Anthropic does publish: across enterprise) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 602.89 Tm (deployments, average cost lands around $13 per developer per active day, and 90% of users stay) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 587.89 Tm (under $30 a day.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 565.89 Tm (The reliable way to check your own number is to run `/usage` inside a session. It shows your current) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 550.89 Tm (plan bars, when each one resets, and a breakdown of what recently consumed the allowance,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 535.89 Tm (attributed by skill, subagent, plugin, and MCP server, so you can see which specific piece of your) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 520.89 Tm (setup is the expensive one.) Tj ET BT /F2 15 Tf 0.08 0.1 0.14 rg 1 0 0 1 46 492.89 Tm (Where the tokens actually go) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 468.89 Tm (Every message you send re-sends the entire conversation so far: the system prompt, your project's) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 453.89 Tm (`CLAUDE.md`, every prior message, and every tool result. Prompt caching is what keeps this) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 438.89 Tm (affordable. The API matches the start of each new request against what it cached from the one) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 423.89 Tm (before, and only the new material at the end gets processed at full price. A cache hit costs a fraction,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 408.89 Tm (roughly a tenth, of what a full-price input token costs, which is why a long-running session is far) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 393.89 Tm (cheaper per turn than its raw size would suggest, right up until something breaks the match.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 371.89 Tm (A specific set of actions invalidates that cache and forces a full, expensive re-read: switching models) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 356.89 Tm (with `/model`, changing the effort level, turning on fast mode partway through a session, an MCP) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 341.89 Tm (server connecting or disconnecting when its tools load directly into the prompt, denying an entire tool,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 326.89 Tm (running `/compact`, or upgrading Claude Code itself. Editing a file, changing permission mode,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 311.89 Tm (invoking a skill, and rewinding the conversation all leave the cache intact, so none of those cost you a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 296.89 Tm (rebuild.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 274.89 Tm (Cache lifetime matters just as much as what breaks it. On a subscription, Claude Code requests a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 259.89 Tm (one-hour cache window by default, so a coffee break will not force a full re-read. It only drops to) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 244.89 Tm (five minutes once you have gone past your plan's usage limit and are drawing on paid usage credits,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 229.89 Tm (and even then you can hold the one-hour window by setting `ENABLE_PROMPT_CACHING_1H=1` as an) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 214.89 Tm (environment variable.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 192.89 Tm (Beyond caching, Anthropic's own guidance names the specific reasons a long session climbs faster) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 177.89 Tm (than your activity would suggest: the full history gets resent with every tool call, a break longer than) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 162.89 Tm (the cache window means the next message reprocesses everything from scratch, a scheduled task) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 147.89 Tm (fires on its own timer and sends full context each time it runs, and a live agent-team member keeps) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 132.89 Tm (drawing on the budget for as long as it stays active, even when it is doing very little. None of that is a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 117.89 Tm (bug. It is the direct cost of keeping a long, stateful conversation alive.) Tj ET q 0.86 0.88 0.92 RG 1 w 46 42 m 549.28 42 l S Q BT /F1 8.4 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 30 Tm (TechRounder | Page 2 of 6) Tj ET BT /F1 7.2 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 19 Tm (https://www.techrounder.com/pdf/blog/claude-code-usage-limits-token-optimization.pdf) Tj ET endstream endobj 9 0 obj << /Type /Page /Parent 2 0 R /MediaBox [0 0 595.28 841.89] /Resources << /Font << /F1 3 0 R /F2 4 0 R >> >> /Contents 10 0 R >> endobj 10 0 obj << /Length 5792 >> stream BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 789.89 Tm (Codebase exploration adds its own tax on top of all this. Developers on the Claude Code subreddit have) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 774.89 Tm (described the first few minutes of a new task, when the agent reads a dozen or more files just to) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 759.89 Tm (understand what it is looking at, as the single most expensive stretch of an entire session, sometimes) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 744.89 Tm (costlier than the fix that follows it. That is an anecdotal pattern rather than an official figure, but it) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 729.89 Tm (lines up with a simple mechanical fact: Claude has to read a file before it can reason about it, and a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 714.89 Tm (ten-thousand-line file costs the same to read whether you plan to touch one line of it or fifty.) Tj ET BT /F2 15 Tf 0.08 0.1 0.14 rg 1 0 0 1 46 686.89 Tm (Small habits that make the real difference) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 662.89 Tm (None of these require new tooling. They are mostly a matter of interrupting a default.) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 640.89 Tm (- Clear between unrelated tasks. Run `/rename` to label a session before you leave it, then `/clear` to start) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 627.09 Tm (the next task with nothing to re-read. Stale history from a finished task gets billed on every message of the) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 613.29 Tm (next one.) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 596.49 Tm (- Write prompts that name the file and the function. A request like "improve this codebase" forces Claude to) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 582.69 Tm (scan broadly before it can act; "add input validation to the login handler in auth.ts" lets it work from a) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 568.89 Tm (narrow, known starting point.) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 552.09 Tm (- Use Plan Mode before a large change. Press Shift+Tab to switch into it. Claude explores and proposes an) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 538.29 Tm (approach for a few thousand tokens instead of writing, testing, and rewriting hundreds of lines because) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 524.49 Tm (the first direction was wrong.) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 507.69 Tm (- When Claude gets something wrong, resist the instinct to type a conversational correction. A follow-up like) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 493.89 Tm ("no, only fix section 3" stacks on top of the mistake and forces a second full generation. Double-tap Escape) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 480.09 Tm (or run `/rewind` to restore the checkpoint before the error and try again from there.) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 463.29 Tm (- Match the model to the task. Sonnet handles most day-to-day coding well and costs less than Opus, which is) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 449.49 Tm (worth reserving for genuinely hard architectural calls. For a subagent doing simple, well-defined work,) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 435.69 Tm (set `model: haiku` in its configuration rather than inheriting the parent's model by default.) Tj ET BT /F2 15 Tf 0.08 0.1 0.14 rg 1 0 0 1 46 412.89 Tm (Set these once and stop thinking about them) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 388.89 Tm (A handful of configuration choices pay for themselves across every future session, not just the) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 373.89 Tm (current one.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 351.89 Tm (Keep `CLAUDE.md` short. Anthropic's own guidance suggests staying under roughly 200 lines and) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 336.89 Tm (moving anything workflow-specific, like a database migration procedure or a release checklist, into a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 321.89 Tm (Skill instead. A skill only loads into context when it is actually invoked; a bloated `CLAUDE.md` gets) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 306.89 Tm (billed at the start of every single session whether that day's task needs it or not. On a genuinely large) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 291.89 Tm (codebase, a single root-level file also stops being useful. Layering smaller, directory-specific) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 276.89 Tm (`CLAUDE.md` files closer to the code they describe keeps each one relevant to what Claude is actually) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 261.89 Tm (touching that session.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 239.89 Tm (MCP servers are convenient, but every connected tool has to be described in the request. Where a CLI) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 224.89 Tm (tool already covers the job, such as `gh` for GitHub or `aws` for AWS work, it is more) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 209.89 Tm (context-efficient than an equivalent MCP server, since Claude can run it directly without a tool listing) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 194.89 Tm (attached. Run `/mcp` periodically and disconnect anything you configured once and stopped actually) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 179.89 Tm (using.) Tj ET q 0.86 0.88 0.92 RG 1 w 46 42 m 549.28 42 l S Q BT /F1 8.4 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 30 Tm (TechRounder | Page 3 of 6) Tj ET BT /F1 7.2 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 19 Tm (https://www.techrounder.com/pdf/blog/claude-code-usage-limits-token-optimization.pdf) Tj ET endstream endobj 11 0 obj << /Type /Page /Parent 2 0 R /MediaBox [0 0 595.28 841.89] /Resources << /Font << /F1 3 0 R /F2 4 0 R >> >> /Contents 12 0 R >> endobj 12 0 obj << /Length 6099 >> stream BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 789.89 Tm (A PreToolUse hook can filter noisy command output before it ever reaches the model. A test suite that) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 774.89 Tm (produces ten thousand lines of passing-test noise for one real failure is a common offender. A short) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 759.89 Tm (script that greps for `FAIL` or `ERROR` and returns only the relevant lines can cut that particular cost) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 744.89 Tm (from tens of thousands of tokens down to a few hundred, and Anthropic's own documentation ships a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 729.89 Tm (working example of exactly this pattern for filtering test output. Installing a code intelligence plugin for) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 714.89 Tm (a typed language has a similar effect at a different layer: it gives Claude precise "go to definition") Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 699.89 Tm (navigation instead of a blind grep across the repository followed by reading several candidate files to) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 684.89 Tm (find the right one.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 662.89 Tm (If you use agent teams, keep them small on purpose. Anthropic's own cost documentation notes that a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 647.89 Tm (team running in plan mode uses roughly seven times the tokens of a standard session, since each) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 632.89 Tm (teammate maintains its own separate context window. That multiplier is fine for a task that genuinely) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 617.89 Tm (benefits from parallel review; it is expensive overhead for something one focused session could have) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 602.89 Tm (handled alone.) Tj ET BT /F2 15 Tf 0.08 0.1 0.14 rg 1 0 0 1 46 574.89 Tm (How to check your own numbers) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 550.89 Tm (Guessing at optimization is a waste of the very budget you are trying to protect. Claude Code has native) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 535.89 Tm (tools for this that most people never open.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 513.89 Tm (`/usage` is the fastest check: a live breakdown of the current session's spend, plus, on a paid plan,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 498.89 Tm (what fraction of your recent usage came from a specific skill, subagent, or MCP server, and a flag) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 483.89 Tm (for any single behavior, such as long context or repeated cache misses, that accounts for 10% or) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 468.89 Tm (more of recent activity. `/insights` goes further, analyzing up to your last 200 sessions on that) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 453.89 Tm (machine and writing an HTML report to `~/.claude/usage-data/report.html` that surfaces recurring) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 438.89 Tm (friction, like requests that needed clarifying or code that kept failing review.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 416.89 Tm (For a team, native local commands only cover one machine at a time. OpenTelemetry export streams) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 401.89 Tm (session counts, token totals, and cost per user into whatever observability stack your organization) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 386.89 Tm (already runs, and it works the same way regardless of whether developers authenticate through a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 371.89 Tm (subscription, the Console, or a cloud provider. If you just want a fast terminal view without setting up) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 356.89 Tm (an observability pipeline, ccusage is a popular open-source CLI that parses your local session logs) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 341.89 Tm (offline and prints daily and per-session cost breakdowns. It is not an Anthropic product and has not) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 326.89 Tm (been reviewed by Anthropic, so treat it, like any third-party tool that reads your local Claude Code data,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 311.89 Tm (with the same care you would give any script from outside your organization.) Tj ET BT /F2 15 Tf 0.08 0.1 0.14 rg 1 0 0 1 46 283.89 Tm (What to actually do when you hit the wall) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 259.89 Tm (Once the limit lands, your options are narrower, but they are still worth knowing before it happens) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 244.89 Tm (rather than while you are staring at the error.) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 222.89 Tm (1. Read the reset time in the message itself. It is exact, and Claude Code will not accept further requests) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 209.09 Tm (until that moment regardless of what you try.) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 192.29 Tm (2. If it is specifically the Opus limit, switch models with `/model` . Session and weekly limits will not budge) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 178.49 Tm (this way, but the Opus-only cap will.) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 161.69 Tm (3. Run `/usage-credits` to keep working past the allowance. On a Pro or Max plan this opens your) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 147.89 Tm (usage-credit settings directly; on Team or Enterprise it either opens your organization's usage settings or) Tj ET BT /F1 10.5 Tf 0.2 0.23 0.28 rg 1 0 0 1 46 134.09 Tm (sends a request to an admin, depending on your billing access.) Tj ET q 0.86 0.88 0.92 RG 1 w 46 42 m 549.28 42 l S Q BT /F1 8.4 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 30 Tm (TechRounder | Page 4 of 6) Tj ET BT /F1 7.2 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 19 Tm (https://www.techrounder.com/pdf/blog/claude-code-usage-limits-token-optimization.pdf) Tj ET endstream endobj 13 0 obj << /Type /Page /Parent 2 0 R /MediaBox [0 0 595.28 841.89] /Resources << /Font << /F1 3 0 R /F2 4 0 R >> >> /Contents 14 0 R >> endobj 14 0 obj << /Length 6446 >> stream BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 789.89 Tm (You will also find community-built scripts, such as unsnooze and similar auto-resume tools, that watch) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 774.89 Tm (for the reset timestamp in the limit error and automatically send a "continue" once it passes. These) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 759.89 Tm (are not part of Claude Code and Anthropic does not maintain them, so running one unattended,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 744.89 Tm (especially overnight with your account actively signed in, deserves the same caution you would give) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 729.89 Tm (any third-party script with standing access to your credentials. A separate tip that circulates in the) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 714.89 Tm (same threads, using a scheduled routine with a minimal prompt to somehow reset or influence usage) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 699.89 Tm (counters, has no confirmation behind it beyond anecdote. I would not build a workflow around it.) Tj ET BT /F2 15 Tf 0.08 0.1 0.14 rg 1 0 0 1 46 671.89 Tm (For teams and API integrations, a different set of limits applies) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 647.89 Tm (Everything above covers the subscription-plan session and weekly windows. If you are calling the) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 632.89 Tm (Claude API directly, whether from Claude Code on the Console, a custom integration, or a cloud) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 617.89 Tm (provider, you are on a completely different system: per-minute rate limits measured in requests per) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 602.89 Tm (minute, input tokens per minute, and output tokens per minute, enforced per model class at the) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 587.89 Tm (organization level. Exceeding any one of the three returns a 429 error naming which limit you hit,) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 572.89 Tm (along with a `retry-after` header telling you exactly how long to wait.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 550.89 Tm (A few details are worth building into any production integration from the start. Only uncached input) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 535.89 Tm (tokens count against your ITPM limit, so a high cache-hit rate effectively raises your usable throughput) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 520.89 Tm (well above the raw number on your account. Usage tiers advance automatically as your organization's) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 505.89 Tm (cumulative spend crosses each threshold, so a workload that gets rate-limited today may simply) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 490.89 Tm (outgrow the problem in a few weeks without any configuration change. And when a request does get) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 475.89 Tm (rejected, exponential backoff with a randomized jitter, rather than an immediate retry, is what keeps a) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 460.89 Tm (wave of failed requests from retrying in lockstep and tripping the same limit again a few seconds) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 445.89 Tm (later.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 423.89 Tm (If you are rolling Claude Code out to a team on the Console rather than through seat-based) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 408.89 Tm (subscriptions, Anthropic publishes starting-point recommendations for per-user rate limits, scaled) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 393.89 Tm (down as the team grows since fewer people tend to be active at exactly the same moment in a larger) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 378.89 Tm (organization:) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 356.89 Tm (Team size: 1 to 5 users | TPM per user: 200k to 300k | RPM per user: 5 to 7) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 339.89 Tm (Team size: 5 to 20 users | TPM per user: 100k to 150k | RPM per user: 2.5 to 3.5) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 322.89 Tm (Team size: 20 to 50 users | TPM per user: 50k to 75k | RPM per user: 1.25 to 1.75) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 305.89 Tm (Team size: 50 to 100 users | TPM per user: 25k to 35k | RPM per user: 0.62 to 0.87) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 288.89 Tm (Team size: 100 to 500 users | TPM per user: 15k to 20k | RPM per user: 0.37 to 0.47) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 271.89 Tm (Team size: 500+ users | TPM per user: 10k to 15k | RPM per user: 0.25 to 0.35) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 254.89 Tm (Full details on how the tiers, the rate-limit headers, and workspace-level caps fit together are in) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 239.89 Tm (Anthropic's API rate limits reference and the shorter support explainer on the same system.) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 217.89 Tm (I would start with `/usage`, not a third-party tool or a workaround script. It takes about ten seconds to) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 202.89 Tm (tell you whether you are fighting a genuinely tight plan, a bloated `CLAUDE.md`, or a session that has) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 187.89 Tm (simply been left open since yesterday, and that answer decides which part of this guide is actually) Tj ET BT /F1 11 Tf 0.14 0.16 0.2 rg 1 0 0 1 46 172.89 Tm (worth acting on next.) Tj ET BT /F2 13 Tf 0.08 0.1 0.14 rg 1 0 0 1 46 144.89 Tm (References) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 124.89 Tm (1. reddit.com - r / ClaudeCode -) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 111.39 Tm (https://www.reddit.com/r/ClaudeCode/comments/1rlimtx/claude_code_beginner_best_practice_token_usage/) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 93.89 Tm (2. code.claude.com - docs / en - https://code.claude.com/docs/en/hooks-guide) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 76.39 Tm (3. code.claude.com - docs / en - https://code.claude.com/docs/en/monitoring-usage) Tj ET q 0.86 0.88 0.92 RG 1 w 46 42 m 549.28 42 l S Q BT /F1 8.4 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 30 Tm (TechRounder | Page 5 of 6) Tj ET BT /F1 7.2 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 19 Tm (https://www.techrounder.com/pdf/blog/claude-code-usage-limits-token-optimization.pdf) Tj ET endstream endobj 15 0 obj << /Type /Page /Parent 2 0 R /MediaBox [0 0 595.28 841.89] /Resources << /Font << /F1 3 0 R /F2 4 0 R >> >> /Contents 16 0 R >> endobj 16 0 obj << /Length 667 >> stream BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 789.89 Tm (4. reddit.com - r / ClaudeAI -) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 776.39 Tm (https://www.reddit.com/r/ClaudeAI/comments/1uv1qjj/unsnooze_autoresumes_claude_codex_and_other_ai/) Tj ET BT /F1 10 Tf 0.18 0.2 0.24 rg 1 0 0 1 46 758.89 Tm (5. platform.claude.com - docs / en - https://platform.claude.com/docs/en/api/rate-limits) Tj ET q 0.86 0.88 0.92 RG 1 w 46 42 m 549.28 42 l S Q BT /F1 8.4 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 30 Tm (TechRounder | Page 6 of 6) Tj ET BT /F1 7.2 Tf 0.42 0.45 0.5 rg 1 0 0 1 46 19 Tm (https://www.techrounder.com/pdf/blog/claude-code-usage-limits-token-optimization.pdf) Tj ET endstream endobj xref 0 17 0000000000 65535 f 0000000015 00000 n 0000000064 00000 n 0000000154 00000 n 0000000224 00000 n 0000000299 00000 n 0000000441 00000 n 0000005942 00000 n 0000006084 00000 n 0000012367 00000 n 0000012510 00000 n 0000018354 00000 n 0000018498 00000 n 0000024649 00000 n 0000024793 00000 n 0000031291 00000 n 0000031435 00000 n trailer << /Size 17 /Root 1 0 R >> startxref 32153 %%EOF