OpenClaw Upgrade Guide: Version Pinning, Canaries, and Rollback
Upgrades are where healthy agents go to die. Nobody writes a postmortem titled βwe ran the installer on a Tuesday,β yet that is the root cause of more broken OpenClaw deployments than any bug I have personally chased. This guide is the routine I use now: pin everything, snapshot the state, let one agent take the new version first, and keep a rollback that works without thinking.
None of it takes long. The whole ritual adds maybe fifteen minutes to an upgrade, and it has turned three separate would-be outages into shrugs.
Why Agent Upgrades Hurt More Than App Upgrades
When you upgrade a web framework, the breakage is usually loud. A type error or a failed build, usually within seconds. An agent runtime has a quieter failure mode, because a lot of its behavior lives in tool schemas and default settings that a changelog describes in a single line if it mentions them at all. A renamed hook event, or a slightly different way the runtime formats tool results before the model reads them, can change what your agent decides to do while every health check stays green.
The worst one I hit was a minor release that changed how skill descriptions were truncated when the skill list got long. Nothing errored. My morning-briefing agent simply stopped choosing the calendar skill, because the part of its description that said βuse this for anything involving todayβ had been cut off. It took four days and one missed dentist appointment to notice.
Four days.
Step One: Pin the Exact Version
If your install command says βlatest,β your agent upgrades whenever a container rebuilds or a teammate runs setup on a fresh laptop. That is an upgrade policy, just an accidental one. Pin to an exact version wherever OpenClaw gets installed, and commit that pin next to your config so the version and the config that was tested against it travel together.
# openclaw-version.txt (committed alongside openclaw.json)
2026.9.2
# install script reads the pin instead of guessing
OPENCLAW_VERSION="$(cat openclaw-version.txt)"
npm install -g "openclaw@${OPENCLAW_VERSION}"
openclaw --versionAdapt the install line to however you actually run OpenClaw (a Docker tag, or a vendored binary checked into a tools repo). The principle does not change. The same applies to MCP servers you pull from npm or a registry, and it applies double to anything launched with a floating npx call, since that fetches whatever is newest every single time the process starts.
Step Two: Snapshot Before You Touch Anything
A rollback is only as good as the thing you roll back to. Before an upgrade I copy the whole OpenClaw state directory. openclaw.json alone is nowhere near enough. Skills, hook scripts, memory files, and any local databases the runtime migrates on startup all belong in the snapshot, because a new version that rewrites a memory index on first boot leaves the old version unable to read it.
#!/usr/bin/env bash
# pre-upgrade-snapshot.sh
set -euo pipefail
STAMP="$(date +%Y%m%d-%H%M%S)"
SRC="$HOME/.openclaw"
DEST="$HOME/openclaw-snapshots/$STAMP"
mkdir -p "$DEST"
openclaw --version > "$DEST/VERSION"
rsync -a --exclude 'logs/' --exclude 'cache/' "$SRC/" "$DEST/state/"
echo "snapshot: $DEST"Excluding logs and cache keeps the snapshot small. Stop the agent before you run it if the runtime writes memory continuously; a snapshot taken mid-write can be subtly corrupt in ways you only discover during the rollback you needed it for. If you keep your config in git (you should, and the config walkthrough shows one way to lay it out), tag the commit too.
Step Three: Read the Changelog Like a Suspect
Skimming release notes for new features is the wrong read. You are looking for anything that changes a default, renames something, or touches how the model sees tools. I search every changelog between my pinned version and the target for a short list of words: default, renamed, deprecated, removed, schema, timeout, truncat. Each hit gets checked against my config.
Skipping versions makes this much harder. Jumping five minor releases means five changelogs, and deprecation warnings that were printed for two releases before the removal never reached your logs at all. Upgrading one minor version at a time is slower on paper and faster in practice.
Deprecations are free warnings
Run your current version with warnings visible for a day before upgrading and grep the logs for deprecation notices. Whatever shows up there is the list of things the next release is most likely to break.
Step Four: Run Your Evals Against Both Versions
This is where the upgrade actually gets tested. If you followed the skill testing guide, you already have a set of recorded prompts with expected tool calls. Run the suite on the pinned version, then on the candidate, and diff the tool-call traces instead of just comparing pass counts. A suite can pass at the same rate while the agent picks different tools to get there, and that shift is exactly the early signal you want.
# run the same eval set under each version, keep the traces
OPENCLAW_BIN=~/openclaw-versions/2026.9.2/bin/openclaw \
./run-evals.sh --out traces/current.jsonl
OPENCLAW_BIN=~/openclaw-versions/2026.9.3/bin/openclaw \
./run-evals.sh --out traces/candidate.jsonl
# compare which tools were called, per case
jq -c '{case: .case_id, tools: [.tool_calls[].name]}' traces/current.jsonl > a.jsonl
jq -c '{case: .case_id, tools: [.tool_calls[].name]}' traces/candidate.jsonl > b.jsonl
diff a.jsonl b.jsonlAn empty diff is the goal. A non-empty one is not automatically a blocker, since some releases genuinely improve tool selection, but every changed line needs a human to look at it and decide.
Step Five: Canary on One Agent
If you run more than one agent, never upgrade them together. Pick the least critical one (for me it is the agent that summarizes RSS feeds, which nobody would miss for a day) and move only that agent to the new version. Let it run through at least one full cycle of its schedule, including overnight jobs, because a surprising number of regressions only appear in the long tail, like the weekly job that only fires on Sundays.
Running two versions side by side is easiest when each agent has its own install prefix and its own launch script pointing at it. A shared global install makes a canary impossible, which is another argument for the pin file above.
Watch the boring metrics while the canary runs. Token usage per run, tool calls per run, error rate, wall-clock duration. The observability guide covers how to collect them. A version that suddenly uses 30% more tokens on the same work is telling you something changed in how context gets assembled, even if every output looks fine.
Step Six: Make Rollback Boring
The test of a rollback plan is whether you can execute it at 2 a.m., half awake, from a phone. Mine is one script. It stops the agent, restores the snapshot, reinstalls the previous pinned version, and starts the agent again.
#!/usr/bin/env bash
# rollback.sh <snapshot-dir>
set -euo pipefail
SNAP="${1:?usage: rollback.sh <snapshot-dir>}"
PREV="$(cat "$SNAP/VERSION" | awk '{print $NF}')"
openclaw-agent stop || true
rsync -a --delete "$SNAP/state/" "$HOME/.openclaw/"
npm install -g "openclaw@${PREV}"
echo "$PREV" > openclaw-version.txt
openclaw-agent start
openclaw --versionSwap openclaw-agent start and stop for whatever supervises your process (launchd on a Mac mini, systemd on a VPS). Then actually run the script once, on purpose, on a day when nothing is wrong. An untested rollback is a hope.
One wrinkle deserves its own warning. Any work the agent did on the new version, memory it wrote, notes it saved, is lost when you restore the snapshot. For most setups that is an acceptable trade. If yours writes something you cannot afford to lose, export it before the restore.
Upgrading MCP Servers and Skills
The runtime is only one moving part. MCP servers ship on their own schedules, and a server that renames a tool or changes an argument from a string to an object breaks every skill that calls it. Treat each server like a dependency with its own pin, and upgrade them one at a time, never in the same change as the runtime. When two things change at once and the agent misbehaves, you have doubled the search space for no benefit.
Third-party skills are the same story with a security angle. A skill update can add new tool permissions or a new outbound network call, so read its diff before accepting it. The security best practices post goes further into reviewing what a skill is allowed to touch.
Common Failure Modes
Floating versions in production
An unpinned install means the next reboot is an unscheduled upgrade. Pin the runtime and every MCP server.
Snapshotting only openclaw.json
State migrations touch memory and local databases. Restoring the config alone can leave the old version unable to start.
Judging the upgrade by pass rate
Same score, different tools. Diff the traces.
Upgrading the runtime and servers together
When behavior shifts after a combined change, you cannot tell which piece caused it without undoing both and trying again one at a time, which is the slow path you were trying to skip.
A rollback script nobody has run
It will fail on some hardcoded path the first time. Better that first time happens on a calm afternoon.
Final Verdict
Treat an OpenClaw upgrade as a deploy, because that is what it is. Pin the version so upgrades only happen when you choose them. Snapshot the full state directory. Read the changelog for changed defaults instead of new features, run your evals on both versions and diff the traces, then let your least important agent carry the new release for a full cycle before anyone else gets it.
And write the rollback first. Once it exists and you have run it, upgrades stop feeling risky, and you will take them more often, in smaller steps, which is the actual trick. If something slips through anyway, the troubleshooting guide is where I start, and the production deployment patterns post covers the supervisor setup that makes a clean restart possible.
Ready to build?
Get the OpenClaw Starter Kit β config templates, 5 production-ready skills, deployment checklist. Go from zero to running in under an hour.
$14 $6.99
Get the Starter Kit βAlso in the OpenClaw store
Get the free OpenClaw quickstart guide
Step-by-step setup. Plain English. No jargon.