🩺 Tooling Guide

OpenClaw Gateway Doctor Guide: What to Run When Your Agent Goes Quiet

β€’11 min read

The message showed as delivered. Telegram gave me the two grey ticks, the typing indicator never appeared, and twenty minutes later I was still staring at a chat with an agent that had apparently decided to stop existing.

Nothing had crashed in any way I could see. The Mac mini was on, the process list showed something called openclaw, and my first instinct was to restart the gateway and move on with my evening. That works often enough to become a habit, and it is a bad habit, because a restart wipes the evidence of whatever actually went wrong and guarantees you meet the same failure again next Tuesday at a worse time.

This guide is the order I check things in now. It takes about five minutes.

Start With One Question: Is the Gateway Alive?

Almost every β€œmy agent is ignoring me” report falls into one of two buckets. Either the gateway process is down or wedged, or it is running fine and something between it and you (a channel connection, a model provider, a config value) is broken. The fixes for those buckets have nothing in common, so sort first.

openclaw gateway status
openclaw status

gateway status tells you whether the service is installed, whether it is running, and whether it answers on its local port. A process that exists but does not answer the health probe is the wedged case, and it deserves different treatment from a process that is simply gone. openclaw status is the wider view, including which model each agent is set to use. Output format shifts between releases, so read it rather than grepping for exact strings in scripts.

If the gateway runs under launchd or systemd, check the supervisor too. On macOS, launchctl print gui/$(id -u)/<your-label> shows the last exit status, and an exit code of -15 or 143 means something sent it SIGTERM. That is a completely different story from a crash, usually involving memory pressure or another tool cleaning up processes it thought were orphans.

Run Doctor Before You Touch Anything

openclaw doctor is the most underused command in the CLI. It walks the install and reports problems it knows how to recognize: config keys that no longer exist, services pointing at an old binary path, workspace files the runtime cannot read, channel credentials that are missing, stale lock files left behind by an unclean shutdown.

# read-only pass first
openclaw doctor

# only after you have read the report and backed up config
cp ~/.openclaw/openclaw.json ~/.openclaw/openclaw.json.bak-$(date +%Y%m%d-%H%M)
openclaw doctor --fix

Read the plain report before reaching for --fix. The fix mode rewrites config to match what the current version expects, and most of the time that is exactly what you want. Sometimes it is not. I once watched it migrate a renamed key correctly and also drop a custom field my own scripts read, because doctor had no way of knowing that field mattered to anything. The backup took four seconds to make and saved me an hour.

Keep the timestamped copies. Disk is cheap.

Tail the Logs While You Reproduce

Reading old logs is archaeology. Watching live logs while you send a test message is diagnosis, and it is far faster.

# terminal one
openclaw logs --follow

# terminal two, or just your phone
# send the agent a short message and watch what arrives

What you see (or fail to see) narrows things quickly. No inbound event at all means the channel never delivered your message to the gateway. An inbound event followed by a model request that hangs or returns a 401, 429, or 529 means the provider is the problem. An inbound event, a model response, and then an error on the outbound send means the agent did answer and the reply died on its way back to you, which is the most maddening version because the work was actually done.

If your build does not have logs --follow, the log files live under the OpenClaw state directory (~/.openclaw/logs/ on my machines) and tail -F does the same job. The observability guide covers what to capture permanently once the fire is out.

Probe the Channels Individually

Channel connections fail quietly. A Telegram bot token gets revoked, a WhatsApp session gets logged out from the phone, a Slack app loses a scope after someone in the workspace edits its permissions, a Discord bot gets kicked during a server cleanup, and the gateway keeps running happily with one fewer ear than it had yesterday.

openclaw channels status --probe

The probe flag makes the gateway actually test each connection instead of reporting whatever state it cached at startup. That distinction matters. A cached β€œconnected” can be hours stale. If one channel fails the probe and the others pass, you have your answer and can stop suspecting the model. The chat integration guide walks through re-pairing each channel type.

Check for Config Drift

The gateway reads openclaw.json at startup. Edit that file while it is running and, depending on the key and your version, the change may apply on the next reload or may do nothing until a full restart. People lose afternoons to this. They edit a value, test, see no difference, conclude the value is wrong, edit it again, and by the time they restart they have changed four things and cannot tell which one mattered.

My rule now: every manual edit gets a comment in a changelog file next to the config, and every edit is followed by a reload and a test before the next edit. Tedious. It works. openclaw config get <path> shows you the value the tooling sees, which is a useful sanity check when you suspect a typo in a nested key.

Also check whether something else writes to the file. Doctor in fix mode modifies config, and so do setup wizards. If you keep the file in git you can see exactly what changed with a plain git diff. If you do not keep it in git, this is a good week to start.

The Model Provider Is Down More Often Than You Think

When the logs show a model request going out and nothing coming back, check the provider's status page before debugging your own setup. Rate limits are the other common culprit, especially after you add a cron job that fires every few minutes and quietly eats the quota your chat agent needed. A fallback model in the agent config turns a provider outage into a slower reply instead of silence, and the multi-model routing guide shows a setup that has survived two major outages for me.

Restart Last, and Restart Properly

Once you know what broke, restarting is fine. Use the service manager so the supervisor and the gateway agree about what is running:

openclaw gateway restart

# then confirm, do not assume
openclaw gateway status
openclaw channels status --probe

Killing the process by PID and starting a fresh one by hand is how you end up with two gateways fighting over the same port, or one gateway running an older binary than the one you just upgraded to. If that sounds familiar, the upgrade guide has the cleanup steps.

A Five-Minute Triage Script

I keep this in ~/bin/claw-triage and run it before anything else. It changes nothing, which is the point: you can run it half asleep without making the problem worse.

#!/usr/bin/env bash
# claw-triage: read-only snapshot of gateway health
set -u
OUT="$HOME/.openclaw/triage-$(date +%Y%m%d-%H%M%S).txt"

{
  echo "== $(date)"
  echo "== gateway status";  openclaw gateway status 2>&1
  echo "== status";          openclaw status 2>&1
  echo "== channels";        openclaw channels status --probe 2>&1
  echo "== doctor";          openclaw doctor 2>&1
  echo "== process count";   ps aux | wc -l
  echo "== last 80 log lines"
  tail -n 80 "$HOME/.openclaw/logs/"*.log 2>/dev/null
} > "$OUT"

echo "saved $OUT"
grep -iE "error|fail|refused|timeout|sigterm" "$OUT" | head -20

The saved file is the useful part. When the same failure comes back a month later, you can compare two snapshots instead of trying to remember what the logs said last time. Adjust the log path if yours lives somewhere else, and wire it into a cron job if you want a snapshot taken automatically whenever a health check fails.

Common Failure Modes

Process running, health probe failing

A wedged gateway. Capture logs first, then restart through the service manager.

Exit status -15 in launchd

Something sent SIGTERM. Look at process counts and memory before blaming OpenClaw itself.

One channel silent, others fine

Expired or revoked credentials. channels status --probe finds it in seconds.

Edits to openclaw.json that change nothing

The running gateway has not reloaded. One edit, one reload, one test.

A custom config field vanished after doctor --fix

Restore it from the timestamped backup you made first. You did make one.

Final Verdict

A silent agent is almost never mysterious. It feels mysterious because the first thing most of us do is restart, and the restart erases the clues. Check gateway status, read a plain doctor report, and watch the logs while you send a test message. Probe the channels. Back up config before any automated fix. Then restart, and confirm the restart actually worked.

Save the snapshot. You will want it again.

⚑

Ready to build?

Get the OpenClaw Starter Kit β€” config templates, 5 production-ready skills, deployment checklist. Go from zero to running in under an hour.

$14 $6.99

Get the Starter Kit β†’

Also in the OpenClaw store

πŸ—‚οΈ
Executive Assistant Config
Buy
Calendar, email, daily briefings on autopilot.
$6.99
πŸ”
Business Research Pack
Buy
Competitor tracking and market intelligence.
$5.99
⚑
Content Factory Workflow
Buy
Turn 1 post into 30 pieces of content.
$6.99
πŸ“¬
Sales Outreach Skills
Buy
Automated lead research and personalized outreach.
$5.99

Get the free OpenClaw quickstart guide

Step-by-step setup. Plain English. No jargon.