How to Reduce Claude Code Token Waste Without Lowering Your Engineering Standards
Practical lessons from building a personal operating system with AI agents, local models, and automated workflows.
There is an awkward moment when you are halfway through building something, the work is progressing, and Claude tells you that you have reached your usage limit.
I used to reach that point within three or four days of the week.

I have a Claude Max subscription: $200 a month, or $242 after tax. With several projects running in parallel, I kept exhausting the available allowance. I wanted to keep the quality of the work and the pace of delivery, so I began measuring where the capacity was going.
At first, I found the expected problems: repeated status checks, oversized conversations, unnecessary coordination, and investigations without a clear stopping point. I changed the workflow and published what I had learned.
Then, over the next forty-eight hours, I watched the system use another ten to twelve percent of a weekly token allowance in a single night.
The first changes had helped. They had also missed a large source of waste: the activity between useful tasks. This is the fuller account of what I found and what I changed.
The context: building LifeOS
I am building LifeOS around a personal philosophy: the key to happiness is health, wealth, and wisdom.

LifeOS is my personal operating system, not a commercial product. It helps me organise work and decisions around those priorities.
I currently have 62 agents that I treat as employees across Authority, Revenue, Wisdom, Content Studio, Delivery, Health, Governance, and System. Every employee has a contract defining its responsibilities and boundaries. Its work is measured and governed against that contract.
Those employees are a mixture of deterministic n8n workflows, shell-based automation, and model-powered agents, including Qwen running locally on my Mac mini. Together, they support an operation that runs around the clock.
Local execution keeps much of that operation cost effective. For tasks that need different capabilities, I also maintain subscriptions or API balances for models from providers such as Gemini, DeepSeek, Kimi, and GLM.
Three offices use Claude Code to build and maintain the system:
Architect handles architecture, shared interfaces, and decisions across projects.
Builder implements changes and coordinates delivery.
Auditor independently checks the work and its evidence.
My Delivery pod runs multiple projects in parallel through their own agentic workflows.
You do not need anything like this setup to benefit from the lessons. The same waste can appear in one Claude Code session working on one project. Running more agents simply made the patterns easier for me to see.
The problem: useful work was surrounded by overhead
When I first reviewed the workflow, some scheduled checks found nothing actionable nine times out of ten. One change went through eleven review rounds. Offices passed lengthy explanations to one another. Senior sessions absorbed raw logs and broad code searches when they needed a specific finding.
Each activity looked reasonable in isolation. Together, they consumed capacity without producing enough progress.
Long conversations added overhead too. Claude Code uses caching and compaction, so it would be inaccurate to say every earlier token is processed at full price on every turn. Nevertheless, repeated processing and unnecessary context still consume capacity. Anthropic’s Claude Code guidance recommends managing context, choosing models deliberately, and moving suitable work into tools or separate agents.
I began evaluating the workflow around one question:
How much work does it take to produce a verified result?
That includes implementation, investigation, coordination, review, and correction. A low-token implementation followed by several rounds of rework is not an efficient result.
What I changed inside the work
1. Replace repeated model checks with deterministic tools
I had agents repeatedly checking whether something had changed, whether a process was healthy, or whether work was ready.
Many of those questions had answers a small tool could produce. I built tools that gather status, detect changes, check prerequisites, and return concise results. Formatting, configuration checks, known tests, and many repository checks now run through deterministic tools as well.
The tool can perform the check without model inference. Asking Claude to invoke it and interpret its output still uses tokens, so I also keep that output focused.
For your own project, you could ask:
Review the checks we repeat during this workflow. Identify which can be performed reliably by a script, test, hook, or existing tool. Propose the smallest useful automation. Return a concise result with evidence, and distinguish a failed check from one that could not run.
Start with the check you repeat most often. There is no need to build an elaborate automation system first.
2. Keep the main conversation focused on decisions
A large log can be useful evidence. It rarely needs to become the centre of a long-running conversation.
I moved substantial, noisy investigations into narrowly scoped sub-agents. They inspect the relevant material and return a short finding with evidence locations and unresolved questions. Claude Code supports separate sub-agent contexts for this kind of work. Claude Code sub-agent documentation
There is an important qualification, which I learned later: a sub-agent has startup and context overhead. Launching one to read two files or answer a one-line question can cost more than using a direct search or query in the current session.
Now I ask whether the investigation is large enough to justify that separation.
A useful instruction is:
Investigate this specific failure. Start with the named files and logs, and widen the search only if necessary. Return the finding, evidence locations, uncertainties, and next action in approximately 150 words. Keep detailed output available for follow-up.
The short report keeps the main session focused. It must still disclose any material risk or unresolved failure.
3. Define the assignment before starting the work
Treating agents as employees changed how I brief them.
My standing contracts define each employee’s role and authority. An individual assignment then defines the outcome, scope, constraints, acceptance criteria, and evidence required to finish.
A vague instruction gives an agent room to make assumptions. Those assumptions can lead to unnecessary exploration or, worse, a confidently implemented answer to the wrong question.
For a coding task, I use the equivalent of:
Implement [specific outcome] within [scope]. Preserve [important constraints]. Completion requires [observable acceptance criteria]. Run the relevant checks and report their results. If a missing decision prevents a correct implementation, identify that exact decision. Record unrelated findings separately.
This makes the work easier to execute and much easier to review.
4. Put a checkpoint on investigations
An agent can keep investigating long after it has stopped making useful progress.
In my workflow, a short-lived work unit has a tool-turn checkpoint. The original limit was fifteen; narrower read-only investigations now have a shorter ceiling. At the checkpoint, the agent records what it completed, what it measured, what remains uncertain, and what should happen next.
Those numbers are rules for my environment, not universal settings. The important part is exposing a stalled approach before it consumes an afternoon.
You can give Claude Code a similar instruction:
At the agreed checkpoint, pause and report what has been established, what remains unknown, and whether continuing the current approach is justified. If the work needs a handoff, preserve the evidence and name the next concrete step.
If you require a hard ceiling in an automated workflow, enforce it in the runner. A prompt alone is a request, not a mechanical limit.
5. Match capability to the task
I reserve strong reasoning for architecture, ambiguity, and difficult trade-offs. Suitable implementation can use Sonnet. Narrow mechanical edits can use a lighter model, while repeatable operations belong in tools when possible.
LifeOS follows the same principle beyond Claude Code. Local models handle appropriate routine work; other providers are available when the task justifies them.
I now use a routing ladder:
A deterministic script, test, direct search, or database query.
A local model for suitable batch reading or summaries.
A cost-effective subscribed or API model for bounded, high-volume work.
A Claude Code agent for substantial exploration, implementation, or review.
Stronger reasoning for difficult rulings and unresolved decisions across projects.
This is an order of consideration, not a rule that the cheapest model must always be used. I compare the effort needed to reach an accepted result, including corrections.
For a project, you can ask:
Classify this work into deterministic operations, bounded implementation, and decisions requiring substantial reasoning. Recommend an approach for each. Keep the acceptance criteria and verification requirements consistent across models.
6. Store decisions where the next session can find them
Long conversations can become informal project databases. That makes decisions expensive to recover and easy to misremember.
I keep durable records of decisions, current project state, and handoffs. Agents read the authoritative record rather than relying on a message relayed through another office. A fresh session begins from a compact account of where the work stands.
For a smaller project, a short decision log and a current handoff file may be enough.
Try this instruction:
Before ending this task, update the handoff with the current state, decisions and their rationale, changed files, verification results, blockers, and next action. Link to detailed evidence. Remove superseded guidance.
The record should be concise. Copying an entire conversation into a file merely moves the clutter.
7. Make review answer a specific question
Eleven review rounds on one change made me examine what each round was accomplishing.
I moved toward a focused review followed by targeted verification of the fixes. Further review needs a reason: an unresolved finding, a new risk introduced by the fix, or an area the first review did not cover.
Automated gates handle repeatable checks. Model and human judgement can focus on behaviour, design, security, and gaps in the evidence.
During this period, our main verification gate grew from 86 to 104 checks. That number alone does not prove quality, but it reflects the direction of the process: more enforceable verification alongside less repetitive discussion.
You can ask:
Review this change against its acceptance criteria and relevant risks. Report actionable findings with evidence. After fixes, verify those findings and any newly affected behaviour. Expand the review if the change reveals a broader defect.
Saving tokens must never mean leaving a known correctness issue unresolved.
8. Require evidence before calling work complete
Some of my most expensive failures began with a reassuring status message.
An agent announced an action that had not completed. A failed query returned an empty result that looked healthy. A fix existed in the repository while the running system still used the old version.
I now distinguish an intention, an attempted action, and a verified outcome.
The corresponding instruction is:
Report completed actions with their evidence. Distinguish “passed,” “failed,” and “could not verify.” For deployed changes, check the running version or behaviour. State any unverified outcome explicitly.
This is where governance becomes practical. I have published research on AI governance and continue collecting operational data for that research. Contracts, measurements, and verifiable outcomes give me something concrete to study. They also prevent costly rework.
What the next forty-eight hours taught me
After making those eight changes, I still saw heavy overnight usage. Roughly 150 turns produced perhaps a dozen decisions.
Some of that work was justified. The memory system completed important consolidation, a race in a shared lock was found, and a scraper that had stopped working recovered.
But the audit showed substantial activity between those useful outcomes.
A long-running Claude session woke every thirty minutes to re-arm a watch. Offices woke to receive peer messages. Sessions wrote notes to themselves, then sometimes used another turn to correct a timestamp. Several sub-agents were launched to answer questions that a targeted query could have handled.
My earlier advice to “wake on events rather than poll” was incomplete. I had reduced broad polling but left a model session waking to maintain the event watcher.
I changed that. Local shell scripts now perform routine status checks, timestamp checks, and scheduled sweeps. A Claude session wakes when a script detects something that needs interpretation or a decision. The scripts still use local computing resources, but routine monitoring no longer requires model inference.
I also changed how I dispatch agents. Before launching parallel work, I check whether tasks depend on one another or touch the same files. Sequential dependencies go through one lane in order. Parallel lanes work in separate areas. A brief states who dispatches the work, the model, the checkpoint, and the expected output.
Coordination received a limit too. Peer messages must be short and point to the record containing the decision. If two exchanges do not resolve a disagreement, the discussion stops and moves to a defined decision process. I am measuring that process before treating every part of it as permanent.
These changes address a kind of waste that is easy to miss because no individual task appears to be failing.
The failures that changed my definition of “efficient”
One morning, a gate had been refused but its process chain still held a shared lock. Other work queued behind it for two hours while recovery waited for me to wake up. I clarified the authority boundary: the Architect can inspect and recover a dead or runaway internal engineering process, with the action recorded. Decisions involving money, production changes, or publication remain with me.
Another failure exposed the danger of treating an empty response as a healthy one. A monitoring tool may return nothing because there is nothing to report—or because it could not ask the question. Important probes now report three states: success, failure, and inability to check.
I also learned to distrust manually maintained lists of things the system can discover itself. A hand-written deployment list omitted modules for weeks. Where possible, the workflow now derives the set from its source and checks for omissions.
These are engineering controls, but they are also token controls. A false “done,” a hidden blockage, or an incomplete list sends agents into unnecessary investigation and repair.
How to start on your own project
You can apply this without multiple offices or dozens of agents.
For your next few tasks, record the outcome you wanted, the time and usage consumed, repeated checks, investigation and review rounds, and the evidence that established completion. Include activity between tasks: scheduled wake-ups, coordination messages, and agent launches that produced no decision.
Then change one recurring source of waste. Automate a repeated check, narrow an assignment, replace a tiny sub-agent job with a direct tool call, or move routine monitoring out of the model session.
Compare similar tasks using the same quality criteria. Watch usage, elapsed time, and defects together. Token consumption alone cannot tell you whether the workflow improved.
You can begin with this request:
Audit how we use Claude Code on this project. Look for repeated deterministic work, unnecessary session wake-ups, oversized context, small questions delegated to separate agents, vague assignments, duplicate investigation, and review loops. Identify the three most useful improvements using evidence from our available history. Preserve required testing and acceptance criteria. Explain how to measure each change without assuming a savings percentage.
What optimisation means to me
My subscription still costs $242 a month. Improving the workflow does not lower that fixed price. It helps me use the included capacity more effectively. Separately metered API spending is a different measure and should be tracked separately. Anthropic’s usage and cost guidance makes that distinction as well.
I do not have a controlled benchmark that supports a precise savings percentage or a universal speed claim.
What I have is a practical process shaped by decades in software development: clear assignments, appropriate tools, bounded investigations, durable decisions, and evidence-based completion. It helps me deliver production-grade software while keeping the work measurable.
If you keep running out of Claude capacity, examine what happens both inside the useful coding tasks and between them. Count the checks that find nothing, the messages that change no decision, and the agent jobs that could have been one tool call. Then protect the tests and evidence that tell you the software is actually ready.
That is where I found much of my waste—and where I would start again.
The article is also available in Linkedin https://www.linkedin.com/pulse/how-reduce-claude-code-token-waste-without-lowering-your-delwar-6drxc
