πŸ“ Tooling Guide

OpenClaw Context Budget Guide: What Fills the Window Before You Type

β€’12 min read

My main agent once spent 41,000 tokens on its first turn of the day. I had typed six words. Everything else was furniture: workspace files, tool definitions from MCP servers I had forgotten were connected, a skill list, and a pile of recalled memories, all loaded before the model read a single character I wrote.

That overhead costs money on every turn. It also costs attention, which is harder to see on a bill. A model working through a context that is two-thirds boilerplate follows instructions less reliably, and the instruction it drops is usually the one buried in the middle of a long file you wrote months ago and never reread.

This guide covers how to measure what your agent loads and which cuts are safe to make.

Measure Before You Cut

Guessing is how people end up deleting the one paragraph that mattered. Get numbers first. Recent OpenClaw builds ship a /context command you can run inside a session; it prints a breakdown of what the runtime injected, grouped by source. If your version has it, start there and save the output somewhere so you can compare after every change.

For the workspace files themselves, a dumb script is enough. Characters divided by four is a rough token estimate for English prose, and rough is fine when the question is β€œwhich file is the problem.”

#!/usr/bin/env bash
# context-audit.sh [workspace-dir] [per-file-cap]
WS="${1:-$HOME/.openclaw/workspace}"
CAP="${2:-20000}"
total=0

for f in "$WS"/*.md; do
  chars=$(wc -c < "$f" | tr -d ' ')
  total=$((total + chars))
  note=""
  [ "$chars" -gt "$CAP" ] && note="   <-- over cap, tail gets cut"
  printf "%-16s %7s chars  ~%6s tokens%s\n" \
    "$(basename "$f")" "$chars" "$((chars / 4))" "$note"
done

echo "workspace total: ~$((total / 4)) tokens"

The cap argument matters more than it looks. OpenClaw limits how many characters of each bootstrap file it injects (the setting is bootstrapMaxChars on the versions I run, with a default around 20,000), and anything past the limit is silently dropped. No warning in the chat. No error in the logs unless you go looking. If your AGENTS.md has grown to 31,000 characters, the model has never seen its last third, and any rule you added there recently is a rule that does not exist.

I found two of mine that way.

Where the Tokens Actually Go

Here is the breakdown from that 41,000-token morning, rounded, because it is a fairly typical shape for an agent that has been running for a while without anyone pruning it.

MCP tool schemas: ~19,000 tokens

Four connected servers exposing 63 tools. The agent used seven of them regularly.

Workspace bootstrap files: ~11,500 tokens

Mostly AGENTS.md and a MEMORY.md that had become a diary.

Recalled memories: ~5,000 tokens

Retrieval returning its maximum on every turn, relevant or not.

Skill list: ~2,800 tokens

Names and descriptions for every installed skill.

Runtime system prompt and the actual message: the rest

The part you cannot control, and the part you care about.

Tool schemas winning by a wide margin surprises almost everyone the first time. It should not. Each tool carries its description plus a full JSON schema for its arguments, and servers written for general use tend to expose everything their API can do.

Trim Tool Schemas First

This is the cheapest win and the least risky one, because a tool the agent never calls cannot be missed. Look through a week of logs, list the tools that actually got called, and restrict each agent to those. OpenClaw supports per-agent tool allow and deny lists; the shape below matches the config I run, but key names have moved between releases, so check against your version before pasting.

{
  "agents": {
    "list": [
      {
        "id": "briefing",
        "tools": {
          "allow": ["calendar_list_events", "gmail_search", "weather_*"]
        }
      },
      {
        "id": "repo-janitor",
        "tools": {
          "allow": ["github_*", "exec", "read", "write"],
          "deny": ["github_delete_*"]
        }
      }
    ]
  }
}

Wildcards help, but be careful with them on big servers. github_* on a full GitHub MCP server can still pull in forty-odd tools, and you might want six. If a server lets you choose toolsets at launch (the GitHub one does, and so do several of the popular MCP servers), filter there, before the schemas ever reach OpenClaw.

Disconnect servers you are not using. Obvious advice, rarely followed. I had a Notion server attached for a project that ended in March, costing about 3,000 tokens a turn for six months.

Rewrite Bootstrap Files Every Quarter

Workspace files grow by accretion. Something goes wrong, you add a rule, and the rule stays forever, even after the thing it guarded against stopped being possible. After a year, AGENTS.md reads like a legal code with no index.

The fix I use is a quarterly pass with one question per paragraph: would the agent behave differently this week if this paragraph vanished? If the honest answer is no, it goes. Rules that only apply to one kind of task move into a skill, where they load only when that task comes up. Long reference material (API quirks, a list of every client and their time zone) moves into files the agent can search with memory tools instead of reading on every turn.

Put the rules that matter at the top

Truncation eats from the bottom. Models also tend to weight the start of a long document more heavily than its middle. Both facts point the same way, so the three or four rules you would be furious to see broken belong in the first screen of the file, and the history of why you added them belongs somewhere else entirely.

Keep MEMORY.md short

If you let the agent append to MEMORY.md, it will, cheerfully and forever. Keep that file to current facts and move dated entries into daily notes that sit behind search. The memory system deep dive explains how the two layers are meant to work together.

Make Skill Descriptions Earn Their Tokens

The skill list is small in tokens and large in consequences. Every skill contributes its description to every turn, and the model picks skills based on that text alone, so a vague description is expensive twice: it takes space and still fails to get the skill chosen. Put the trigger in the first sentence (β€œUse when the user asks about today's calendar or upcoming meetings”) and cut the marketing copy that tends to follow. Uninstall skills you tried once. The skills guide has more on writing descriptions that trigger reliably.

Cap Memory Recall

Retrieval that always returns its maximum result count is padding with a search engine attached. Lower the result limit (and raise the similarity threshold, if your memory backend exposes one). Then read what actually comes back for a handful of real prompts. If three of five recalled snippets are irrelevant on ordinary questions, the threshold is too loose.

Understand What Compaction Keeps

Long sessions eventually hit the model's limit, and OpenClaw compacts the history: older turns get summarized into a shorter block so the session can continue. You can trigger it yourself with /compact before a big task, which is often smarter than letting it fire in the middle of one.

The trap is where you put instructions. Bootstrap files are reinjected on every turn, so they survive compaction intact. Something you told the agent in chat forty turns ago only survives if the summarizer decided it was important, and summarizers are bad at knowing which offhand sentence was actually a standing order. If you find yourself repeating an instruction after long sessions, that instruction belongs in a workspace file.

A Budget That Holds

After pruning, the same agent starts its day at about 14,000 tokens, and the behavior improved noticeably: fewer ignored rules, fewer wrong tool picks. The per-turn savings showed up on the bill within a week, which the cost optimization guide covers from the pricing side.

Pruning once is easy. Staying pruned takes a tripwire, so I run the audit script from a weekly cron and have it message me when any workspace file crosses 80% of the cap or the total crosses a number I picked. The cron automation guide shows how to wire that up.

# weekly check: warn at 80% of the per-file cap
CAP=20000
WARN=$((CAP * 8 / 10))
for f in ~/.openclaw/workspace/*.md; do
  c=$(wc -c < "$f" | tr -d ' ')
  [ "$c" -gt "$WARN" ] && echo "$(basename "$f") at $c / $CAP chars"
done

Rerun the measurement after every upgrade, too. New releases sometimes change what the runtime injects by default, and the upgrade guide has a story about a truncation change that cost me four days.

Common Failure Modes

A rule that lives past the cap

The model never saw it. Check file sizes before assuming the agent is disobeying.

Every MCP tool on every agent

The biggest single source of overhead in most setups, and the easiest to fix with an allowlist.

Standing orders given in chat

They fade after compaction. Anything permanent goes in a workspace file.

Trimming without a baseline

Without before and after numbers you cannot tell a real saving from a lucky day. Save the /context output each time.

Final Verdict

Your agent's context window fills up long before your first message arrives, and most of that fill is stuff you added once and forgot. Measure it with /context and a character count. Cut unused tool schemas first, then rewrite bootstrap files so the important rules sit at the top and under the cap. Keep permanent instructions out of chat.

Then check again next month. It grows back.

⚑

Ready to build?

Get the OpenClaw Starter Kit β€” config templates, 5 production-ready skills, deployment checklist. Go from zero to running in under an hour.

$14 $6.99

Get the Starter Kit β†’

Also in the OpenClaw store

πŸ—‚οΈ
Executive Assistant Config
Buy
Calendar, email, daily briefings on autopilot.
$6.99
πŸ”
Business Research Pack
Buy
Competitor tracking and market intelligence.
$5.99
⚑
Content Factory Workflow
Buy
Turn 1 post into 30 pieces of content.
$6.99
πŸ“¬
Sales Outreach Skills
Buy
Automated lead research and personalized outreach.
$5.99

Get the free OpenClaw quickstart guide

Step-by-step setup. Plain English. No jargon.