Claude Code Keeps Running Out of Context — How to Fix It (2026)

Claude Code Keeps Running Out of Context — How to Fix It (2026) — illustration

Last updated: September 5, 2026 — every command, config key and context-window figure below re-checked against Anthropic’s own documentation on that date. The only change affecting anything below was to the 1M-context model list, which gained Fable 5.1.

Correction, August 29, 2026. This article previously said the 1M-token context window “costs more per token.” That was true when it was written and is not true now: Anthropic’s pricing page states that Claude 4.6 and later models include the full 1M token context window at standard pricing, and that a 900K-token request is billed at the same per-token rate as a 9K-token request. The long-context surcharge is gone. The advice below is unchanged — sprawling sessions still cost you in tokens and in answer quality — but the price penalty specifically no longer exists.

You’re halfway through a refactor, three files deep, and Claude Code pops up a “context low” warning. Or worse: it quietly summarizes the conversation and forgets a decision you made twenty minutes ago, so now it’s “fixing” the thing you already fixed. It always happens at the wrong moment. Below is why it happens and how to make it stop, ordered by what gives you the most room back the fastest.

Why this keeps happening

Claude Code keeps everything from your session in one place: every message, every file it read, every command’s output. How much room that is depends entirely on which model you’re on, and Anthropic documents it per model rather than as one number.

Claude Code’s docs put the split like this: Fable 5.1, Fable 5, Sonnet 5, Opus 4.6 and later, and Sonnet 4.6 support a 1 million token context window for long sessions with large codebases. Of the models that don’t, Haiku 4.5 is the one still in the current lineup at 200K; Sonnet 4.5 and Opus 4.5 are 200K too. (Fable 5.1 was added to that sentence at some point before 5 September 2026, which is when I re-checked it.)

Whether you can use a 1M window is a separate question from whether the model has one. That depends on your plan and on a couple of configuration traps, and it’s covered further down.

Either number sounds enormous until you notice what fills it. And the 1M models are not a way out of this: the same habits that exhaust 200K exhaust 1M, they just take longer to do it.

Long conversations are the obvious one. Every exchange stays in there. A two-hour session with a lot of back-and-forth is a lot of tokens.

Big file reads are the sneaky one. Ask Claude to “take a look at the auth module” and it might pull in a 2,000-line file in full. That’s gone now, sitting in context for the rest of the session.

Then there’s tool output. A complete npm run build log, a giant git diff, an accidental ls of node_modules — all of it lands in context verbatim. MCP servers can be the worst offenders here: a database server that returns 500 rows, or a docs server that hands back an entire API reference page, can blow a hole in your budget in a single call. (If you haven’t set any of those up yet, our walkthrough of how MCP servers work covers what they actually do.)

And finally, the logs you paste “just in case.” That 400-line stack trace you dropped in? Those are 400 lines Claude is now carrying around.

When the window fills, Claude Code auto-compacts: it summarizes what’s happened so far and keeps going with the summary instead of the raw history. By default it does this when the conversation reaches the model’s context limit, and you can now move that threshold yourself — /autocompact takes a window size anywhere from 100K to 1M tokens, as in /autocompact 500k. There are several documented exceptions to that default, so check your own case rather than assuming: Sonnet 5 compacts early, at about 967K tokens; Sonnet 4.6 and Opus 4.6 without extended context compact at the 200K boundary; and cloud sessions and unrecognised gateway model IDs each behave differently again. Your session survives, which is the point. But summaries lose things. The exact line numbers. The reason you rejected the first approach. The “we agreed to use middleware, not decorators” decision. So the real goal isn’t just avoiding a crash — it’s staying in control of what Claude actually remembers.

The fixes, fastest first

Here’s the short version. The details are below.

WhatWhen you reach for itEffort
/clearSwitching to an unrelated taskInstant
/compact with an instructionMid-task, need room, want to keep specificsInstant
One task per sessionAlways — it’s preventativeA habit
Point Claude at exact pathsEvery requestA habit
Keep a CLAUDE.mdOnce per project10 minutes
Hand big searches to subagentsHeavy explorationPer task
Quiet down noisy MCP serversOne-time setupOne-time
Don’t paste big logs — save to a fileAnytime you’d paste 50+ linesA habit
/context to see what’s eating tokensWhen the warning surprises you5 seconds
The 1M-token windowGenuinely huge codebases, last resortPlan-dependent

/clear between unrelated things

You just finished the login bug. Now you’re starting on the CSV export. Type /clear. It dumps the conversation and gives you a full window again. There is no reason to drag login-debugging history into export work. Honestly, most “I keep running out of context” complaints are really “I never cleared between three different jobs.” This is the habit that fixes it.

/compact, and tell it what to keep

When you’re still in the middle of something and just need breathing room, /compact it. Left alone, it summarizes the whole conversation and continues. Better to give it a hint:

/compact keep the auth refactor details — which files changed, the new token flow, and the decision to use middleware not decorators

That instruction steers what survives. Without it, Claude decides what’s important. With it, you do. Reach for /compact when you need continuity, /clear when you don’t.

One task per session

A session that starts as “fix this failing test” and grows into “also refactor the API client, oh and update the docs, and let’s talk about the deploy pipeline” is going to run out of room — and the work gets worse on the way down, too. One session, one task. Done? /clear. If you want to carry something forward, ask for an end-of-session summary (there’s a pattern for this in our Claude Code tips post) and paste it as the first message of the next session.

Be specific so it reads less

“Fix the bug in the checkout flow” sends Claude exploring: it reads the cart component, the checkout component, the payment service, the order model, probably three of them in full. “In src/checkout/PaymentForm.tsx, validateCard rejects valid Amex numbers — the regex is too strict” gets it to read one function. Same fix, a fraction of the tokens. Whenever you know the path and the symbol, say so.

A CLAUDE.md so it doesn’t re-learn your project every time

Drop a CLAUDE.md in the repo root with your conventions, a quick architecture note, and your “always do X” rules. Claude reads it automatically at the start of every session, so it doesn’t have to go spelunking to discover that you use Swift 6 concurrency or that tests live under Tests/. You pay for that context once, cheaply, instead of re-burning it every session. The 17 Claude Code tips post goes into what’s worth putting in there.

Push big searches to a subagent

“Find every place we still call the legacy auth endpoint” is exactly the kind of noisy, file-heavy job that doesn’t belong in your main context. Hand it to a subagent via the Agent tool — this was called the Task tool until Claude Code 2.1.63, and existing Task(...) references still work as aliases, so older write-ups (including an earlier version of this one) will say Task. The subagent does the grepping and reading in its own throwaway window and comes back with just the answer. Your session stays lean. Works the same for “audit all our error handling” or “list every component importing this deprecated module.”

Turn down the chatty MCP servers

MCP servers are great, but a verbose one is a tax you pay on every call. If your database server returns 200 columns when you wanted three, fix the query. If you’ve got a docs server, a Jira server, a Slack server, and a Postgres server all loaded but you’re only touching one today, disable the rest for this session. Some servers let you cap result size — do it. Our best MCP servers roundup flags which ones behave themselves about output.

Don’t paste the whole log

Pasting a 600-line build log puts all 600 lines in context, permanently. Instead, send it to a file (npm run build &> build.log) and tell Claude “the build failed — read the last 40 lines of build.log.” Or filter it before it ever reaches the chat: npm test 2>&1 | grep -A5 -i fail. Claude Code reads your terminal output natively anyway, so even just running the failing command and letting it pick up the tail beats copy-pasting the lot.

/context when the warning blindsides you

If the warning shows up earlier than you expected, run /context (and keep an eye on the context indicator in the status line). It renders your window as a coloured grid — system prompt, CLAUDE.md, MCP tool definitions, conversation, file reads — and flags context-heavy tools and memory bloat; pass all to expand it. Usually the culprit is obvious, and it’s usually one enormous file read.

Worth knowing if you last looked at this a while ago: loaded MCP servers no longer cost you a fortune before you’ve said a word. Anthropic’s docs state that MCP tool definitions are deferred by default, so only tool names and server instructions enter context until Claude actually uses a specific tool. The old advice to prune servers purely because their schemas were eating your window is out of date — prune them because their output is verbose, which is the next section.

/usage is worth a glance too, since token spend and context pressure tend to track each other. If you learned this as /cost, that still works — the docs now list it as an alias for /usage, which does considerably more: session token and cost totals, prompt-cache statistics, and a breakdown of recent usage attributed to skills, subagents, plugins and individual MCP servers. That last breakdown is the useful one here, because it names which MCP server is costing you.

The 1M-token window — it exists, don’t lean on it

Fable 5.1, Fable 5, Sonnet 5, Sonnet 4.6, and Opus 4.6 and later all support a 1M-token context window. It’s real and it genuinely helps with large codebases.

It no longer costs extra per token. Anthropic’s pricing page says Claude 4.6 and later models include the full 1M token context window at standard pricing, spelling out that a 900K-token request is billed at the same per-token rate as a 9K-token request. If you have been rationing the long window on the assumption that it carries a surcharge — as an earlier version of this article told you to — that assumption is out of date.

What it costs instead is access, and the rules differ by model rather than by plan alone. On Max, Team and Enterprise, Opus is automatically upgraded to 1M context with no configuration. On Pro it requires usage credits. Sonnet 4.6 at 1M requires usage credits on every subscription plan, Max included. On the API and pay-as-you-go there is full access to both, and Sonnet 5 in particular always runs at 1M there — no 200K variant, nothing to select, no credits required.

There is a switch to turn the long window off, and it does more than it sounds like it does, so read the whole of this before you use it. CLAUDE_CODE_DISABLE_1M_CONTEXT=1 removes 1M model variants from the model picker — but on a model with a native 1M window, such as Sonnet 5 or the Fable models, it also treats that model as having a 200K context window.

Both branches then bite. With auto-compaction on, which is the default, sessions compact at the 200K boundary — and raising /autocompact above 200K does not lift that, because Claude Code caps the auto-compact window at the model’s context window. With auto-compaction off, sessions stop dead at the 200K boundary with a context-limit error instead of compacting. In an article about running out of context, that is worth spelling out: this setting can be the thing that causes the problem you came here to fix.

One more configuration to know about: if ANTHROPIC_BASE_URL points at an LLM gateway, Claude Code can’t verify 1M support and budgets the window at 200K. To get the full window there you select “Sonnet 5 (1M context)” in the model picker, which maps to sonnet[1m].

None of which changes the advice above. A bigger window still gets lost in the middle of a giant context, and it does nothing about the habit of letting sessions sprawl — it just raises the ceiling you eventually hit. Treat it as headroom for a hard problem, not as permission to skip everything above.

/clear vs /compact, in one table

/clear/compact
What it doesWipes the conversation; full fresh windowSummarizes the conversation; continues from the summary
What you keepNothing. Clean slate.The gist. Specifics can vanish.
Use it whenStarting something unrelatedMid-task, low on room, want to keep going
Can you steer it?NoYes — /compact <what to keep>
The riskClearing too eagerly and losing useful contextThe summary drops a detail you needed

If you remember one thing: different task, /clear; same task, no room, /compact with an instruction. And if you’re about to do something where you’ll definitely want the earlier detail, write it down somewhere first.

When auto-compaction bites you

Auto-compaction is fine when what you need going forward is the gist. “Keep building the feature we’ve been working on” survives it without trouble. Where it hurts is when you need precise earlier detail: the exact diff from step two, the reason approach A was a dead end, the line numbers you were about to touch. A summary papers over those.

So before it triggers — you’ll usually get the warning, or you can watch the indicator creep up — checkpoint the stuff that matters. Ask Claude: “Summarize what we’ve changed so far, every file modified with a one-line note, plus the key decisions, and write it to NOTES.md.” Or drop the decisions straight into CLAUDE.md (“auth is middleware-based, not decorators”). Then clear or compact freely, because the record that matters is on disk now, not at the mercy of a summarizer. Treat the context window like RAM, not storage. Anything you’d hate to lose, write it down.

TL;DR

The context window is finite — 200K on Haiku 4.5, Sonnet 4.5 and Opus 4.5, 1M on Fable 5.1, Fable 5, Sonnet 5, Sonnet 4.6 and Opus 4.6 or later (on eligible plans for the 4.6s), and no longer priced at a premium on any of them — and Claude Code auto-compacts when it fills, which is lossy. Different task: /clear. Same task, out of room: /compact <what to keep>. To stop hitting it at all: be specific about file paths, keep one task per session, maintain a CLAUDE.md, push big searches to subagents, quiet down noisy MCP servers, and never paste a log you could’ve grepped first. Use /context to see what’s eating tokens. And checkpoint important state to a file before auto-compaction hits.

One caveat: slash commands and limits move between Claude Code versions, so if /compact or /context behaves differently than what’s here, run /help for your version’s current list.

If you’re also wondering whether a different tool handles long sessions better, our Claude Code vs Cursor vs Copilot comparison gets into how each one deals with context.


Want more out of Claude Code than just “stop running out of context”? 17 Claude Code Tips That 10x Your Productivity covers the CLAUDE.md patterns, subagent tricks, and custom slash commands — the same habits that keep your context lean in the first place. More in the AI Tools section.

Frequently Asked Questions

Is Claude Code's context window unlimited?

No. It runs on a fixed-size window whose size depends on the model. Claude Code's docs state that Fable 5.1, Fable 5, Sonnet 5, Opus 4.6 and later, and Sonnet 4.6 support a 1 million token context window; Haiku 4.5, Sonnet 4.5 and Opus 4.5 run 200K. On Sonnet 4.6 and Opus 4.6 the 1M window is extended context rather than the default, so it depends on your plan. When the window fills, Claude Code summarizes the conversation so far and keeps going, and that summary can lose detail.

What does /compact do in Claude Code?

It compresses the current conversation into a summary and continues from there, which frees up tokens. You can steer what it keeps by adding an instruction, like /compact keep the auth refactor details.

What's the difference between /clear and /compact in Claude Code?

/clear throws the conversation away and starts fresh. /compact keeps a summarized version and continues. Use /clear when you're switching to something unrelated; use /compact when you're mid-task and just need more room.

How big is Claude Code's context window?

It depends on the model. Haiku 4.5, Sonnet 4.5 and Opus 4.5 have a 200K-token window. Fable 5.1, Fable 5, Sonnet 5, Sonnet 4.6, and Opus 4.6 and later support 1M — natively on Fable 5.1, Fable 5 and Sonnet 5, and on eligible plans for Sonnet 4.6 and Opus 4.6. Anthropic no longer charges a premium for the long window — its pricing page states that Claude 4.6 and later include the full 1M token context window at standard pricing, so a 900K-token request is billed at the same per-token rate as a 9K-token one. It still doesn't replace good habits. Run /context to see what's actually using up your window right now.

Written by Hirak Banerjee

Indie dev and maker. I build AI-powered apps and write about the tools I actually use. Follow on X · GitHub

Get told when these numbers change

Every figure here is re-checked against the vendor's own docs. Leave your email and we'll tell you when one moves — only when something actually changes.

Join builders who ship faster. No spam.

Comments