🧰 Tooling Guide

OpenClaw Model Failover: Pick Fallbacks That Survive the Outage, Then Prove It

•7 min read

A fallback model you have never watched take over is a line of JSON you are hoping about. The common version sits on the same provider as the primary, so the outage that takes out one takes out both.

OpenClaw handles model failure in two stages. It first rotates auth profiles inside the current provider, and only then moves to the next model in your fallback list. The guides that rank for this topic cover the keys and the CLI well enough. None of them tell you how to choose the fallback, and none of them make it fail on purpose. That is most of this page.

The Shape

In the annotated config the model is a single string. To get failover, make it an object with a primary and an ordered fallbacks array:

{
  "agents": {
    "defaults": {
      "model": {
        "primary": "anthropic/claude-opus-4-6",
        "fallbacks": [
          "openai/gpt-5",
          "ollama/gemma3:12b"
        ]
      }
    }
  }
}

Or from the CLI, which is harder to typo:

openclaw models fallbacks list
openclaw models fallbacks add openai/gpt-5
openclaw models fallbacks remove openai/gpt-5
openclaw models fallbacks clear

The ollama/gemma3:12b alias is the one registered in the cost optimization guide. Restart the gateway after any edit, then confirm with openclaw config get agents.defaults.model that the tooling reads what you meant to write. Key names have moved between releases. Check yours.

Order Fallbacks by What Breaks

The usual list is ordered by taste: second-favorite model, then third. Each entry should instead earn its place by surviving a failure the entries above it cannot.

The first fallback should be on a different provider. Auth rotation already covers the case where one of your keys for Anthropic is rate limited and another is not. What it cannot cover is Anthropic being down, or your billing being disabled there, and a second Anthropic model will hit the same wall in the same second. Putting Sonnet behind Opus protects you from almost nothing.

The last fallback should not need the internet. A local Ollama model is slow and limited, and it answers when every cloud provider on your list is having a bad afternoon. For a chat agent that mostly triages messages, "slow and limited" beats silence by a wide margin.

Then check every fallback against the two things that actually break agents on a model switch. Tool calling is the first. A model that cannot call tools reliably will take over a session and then fumble every skill in it, and you will spend an evening debugging skills that were fine. Context size is the second, and it is the one I would bet nobody checked. According to the OpenClaw failover docs, a context overflow is not a reason to try the next candidate. So if your longest chat session is bigger than the fallback's window, the fallback errors on overflow and the chain stops right there, one step after the outage, with every model behind it untried. Look at your biggest sessions with /context (the context budget guide covers reading it), and put fallbacks that cannot hold them at the end of the list or off it.

Where Fallbacks Quietly Do Not Apply

A model you picked by hand. When you switch a session with /model, that is treated as an exact selection and it is strict. The session will not fail over, and the override is easy to forget: type /model once to test something in the Telegram chat and that chat stays pinned until /new or /reset clears it.

Per-agent models. An agent in agents.list that names its own primary is strict unless it also names its own fallbacks. You can give an agent only a fallback chain and let it keep inheriting the shared primary, or set "fallbacks": [] to say, on purpose, that this agent should fail loudly. That second option is right for anything doing money or legal work, where a weaker model guessing is worse than no answer.

Cron jobs use the configured fallbacks unless the job payload supplies its own. That is usually what you want, with one exception covered in the gateway doctor guide: a cron that fires every few minutes and fails over to a paid fallback will now eat that provider's quota too.

The Fire Drill

Five minutes. Do it on a scratch agent so the agent people actually talk to never notices.

Give the scratch agent a primary that does not exist and a real fallback chain:

{
  "id": "failover-drill",
  "model": {
    "primary": "anthropic/claude-this-model-does-not-exist",
    "fallbacks": ["openai/gpt-5", "ollama/gemma3:12b"]
  }
}

A missing model comes back as model_not_found, which is on the documented list of errors that advance to the next candidate. Restart the gateway, open a fresh session with the drill agent, and send it a request that needs a tool. Ask it to read a file in its workspace and quote the first line.

Then run /status. Pass looks like this: the selected model and the active model are shown as different, a fallback notice appeared in the session (it starts with "Model Fallback" and names the reason), and the file line came back correct. A correct reply with no notice means the primary somehow resolved, so check the ID. No reply at all is a fail, and the gateway log will say which candidate it gave up on.

Now break the first fallback too. Change openai/gpt-5 to another nonexistent ID and repeat. This tests the local model under the same tool-using prompt, which is the run that tells you whether your last line of defense can do the job or can only say hello.

Delete the drill agent when you are done. Keep the JSON in a file next to openclaw.json and run it again after every upgrade, alongside the canaries from the allow and deny list guide. The upgrade guide has the checklist they both belong on.

After a Real Outage

Run git diff on openclaw.json. The primary in the file should still be the primary you chose, and if anything rewrote it during the failover you want to know today, before the provider recovers and your agent keeps running on the slow model for a month. If the file is not in git yet, that is the fix.

Cooldowns

A failed provider sits out for a while before OpenClaw tries it again, and the wait grows with repeated failures, so a recovered primary can stay benched for several minutes after its status page goes green. Leave it alone.

Common Failure Modes

The provider went down and the agent went silent anyway

Check for a /model override on that session first. Then check whether every fallback sat on the same provider as the primary.

Failover happened, then every tool call broke

The fallback model is weak at tool calling. Move it down the list and rerun the drill with a tool-using prompt.

The chain stopped at the first fallback with a context error

The fallback's window is smaller than the session. Overflow does not advance the chain, so reorder by window size or start a fresh session.

One agent fails over and another does not

The second agent names its own primary in agents.list without its own fallbacks, which makes it strict.

Final Verdict

Put a different provider first in the chain and a local model last, and drop anything that cannot call tools or hold your longest session. Then break the primary on a scratch agent and watch /status show the switch. The config takes two minutes. The drill is the only part that tells you the next outage will be a slower reply and not a quiet afternoon where nobody notices the agent stopped answering.

⚡

Ready to build?

Get the OpenClaw Starter Kit — config templates, 5 production-ready skills, deployment checklist. Go from zero to running in under an hour.

$14 $6.99

Get the Starter Kit →

Also in the OpenClaw store

🗂️
Executive Assistant Config
Buy
Calendar, email, daily briefings on autopilot.
$6.99
🔍
Business Research Pack
Buy
Competitor tracking and market intelligence.
$5.99
⚡
Content Factory Workflow
Buy
Turn 1 post into 30 pieces of content.
$6.99
📬
Sales Outreach Skills
Buy
Automated lead research and personalized outreach.
$5.99

Get the free OpenClaw quickstart guide

Step-by-step setup. Plain English. No jargon.