Production image upgrade flow
Overview
scripts/upgrade-production-images updates the OpenClaw Control Plane (OCC),
Agent runtimes, or both. The controller-only command ends after the OCC API and
worker recover; it does not request Agent deployments. A runtime release deploys
a new revision for every Agent that was running when the command began and ends
after the selected Pods are ready and each replacement gateway passes read-only
Doctor lint. Model and external integration checks remain operator tasks.
Entry Points
- Trigger: an operator runs
scripts/upgrade-production-imageswith one or both image options and explicit cluster, release, source, and protected-file inputs. - Required state: a healthy production Helm release and matching OCC and Kubernetes Installation identity. Runtime releases additionally require a complete authorized fleet inventory with no deployment in progress.
- Source:
scripts/upgrade-production-images,packages/occ/src/index.ts:OpenClawController.getInstallationDeploymentInventory, andpackages/occ/src/index.ts:OpenClawController.deployAgent.
Flow
graph TD
A["Validate inputs and freeze fleet"] --> Q{"Repository broker enabled?"}
Q -->|Yes| P["Check image protocol and node architectures"]
Q -->|No| B["Render candidate and save recovery record"]
P --> B
B --> C["Read live Secret and Helm state"]
C --> D{"Candidate Helm release deployed?"}
D -->|No| E["Stop API and worker Pods"]
E --> F["Update Secret and run Helm migration"]
F --> G{"Helm completes?"}
G -->|No| H["Inspect migration and release before retry"]
H --> C
G -->|Yes| I["Verify OCC rollout"]
D -->|Yes| I
I --> J{"Runtime release?"}
J -->|No| K["Hand off application checks"]
J -->|Yes| L["Deploy recorded Agents"]
L --> M{"Dispatch response known?"}
M -->|No| N["Read Agent and stop for reconciliation"]
N --> L
M -->|Yes| O["Check revisions, Pods, and Doctor"]
O --> KExecution Trace
1. Prepare and freeze the target
scripts/upgrade-production-images:245
The script verifies the protected files, cluster, deployed Helm release, and matching OCC and Secret Installation IDs. It compares protected and live configuration in full, including selected image fields. It captures separate reviewed candidate files and accepts changes only to the Slack directory proxy and the curated Codex PluginDriver selection. All other Helm and Installation settings remain protected, including new fields. The guard compares broker resource references, not the contents of referenced ConfigMaps or Secrets. Image flags select images after this comparison. A runtime release also reads complete authorized inventory and records every running Agent's baseline revision. Nonterminal deployment work, a missing active revision, or an unready Namespace stops preparation. Stopped and deleting Agents are excluded.
For repository-enabled releases, the helper requires explicit immutable
controller and broker images. It inventories every node matching the control
plane selector, including unready and cordoned nodes, and requires a single
native architecture. scripts/upgrade-repository-image-probe.mjs runs both
selected images with synthetic inputs and a private receipt listener. It checks
recovery and refused reservation through the actual Driver and broker, and
records the selected digest, platform manifest, configuration, requests, and responses.
The check proves wire compatibility, not Kubernetes image availability, receipt
durability, or disposal. The
helper also preserves the live broker hostname; broker restart recovery remains
an operator task in the broker procedure.
The script renders the chart and performs a server-side Helm dry run. It saves candidate inputs, inventory, target identity, and parameter hashes in the private evidence directory before marking preparation complete. A per-directory lock prevents two helpers from using that record at once. The operator must freeze other writers, autoscalers, Helm changes, and runtime draft edits as specified in the production guide.
2. Reconcile a prior attempt
scripts/upgrade-production-images:506
On every attempt the script rereads Helm status and values and the Installation Secret. Only the recorded baseline or candidate values are accepted; the Secret's UID, Installation annotation, and other data must match the baseline. Resume also binds the original kubeconfig contents, inputs, OCC URL, scripts, flow, chart, and pair evidence. It requalifies the pair and rejects a changed eligible-node set. Candidate files must still match the saved reviewed inputs; the helper applies the saved candidate. Unexpected drift or an in-progress Helm release stops the command.
If a previously started Helm release is deployed at a newer revision with the candidate values, the script continues without repeating Helm. Otherwise it requires the operator's migration-history check and refuses a retry while an initialization Job or Pod remains active. A failed or disconnected migration may have committed; the check and Job inspection are operator-owned and are not a rollback. The guide describes the required attestation and recovery.
3. Quiesce writers and run the candidate release
scripts/upgrade-production-images:624
Before changing the Secret or invoking Helm, the script scales the selected API and worker Deployments to zero, waits until their Pods disappear, and confirms both desired replica counts remain zero. Kubernetes requests in this phase share a bounded quiescence deadline. This covers the selected Helm release; it does not detect independent database writers, autoscalers, or partitioned nodes. The operator must stop those writers and keep nodes reachable.
The protected files are atomically replaced with the saved candidate. When the Installation changes, including a controller-only settings release, the script reads the Secret before updating its Installation key, preserving other data and metadata with a resource-version precondition. If the candidate is already present after a lost response, it does not write it again. The candidate Installation checksum is included in both OCC Pod templates.
A durable marker precedes helm upgrade. The candidate migrator and bootstrap
run in the Helm initialization hook before API and worker rollout. If Helm
fails, resumption reads its current status and migration state before a retry;
it does not start the old image to undo a committed schema change.
4. Verify control-plane recovery
scripts/upgrade-production-images:694
The script waits for both OCC Deployments, checks their controller image and replica count, and checks the Installation checksum when its configuration changed. It requires exactly one named container for each component. For a broker-enabled worker it also accepts a restartable init container, provided no worker exists in the ordinary container list. It rejects a non-restartable init worker or ambiguous placement. It retries authenticated OCC access and verifies the same Installation ID. A controller-only release then ends without requesting Agent deployments.
For a repository-enabled release, it also verifies the ready API and worker Pods, their owning ReplicaSets, node architecture, and runtime controller and broker image IDs against the qualified pair. The deployed Driver checks the real broker's admission capability without opening a session. Pod identities must remain stable across that check and before dispatch and completion.
5. Deploy and verify the recorded fleet
scripts/upgrade-production-images:797
Before sending each ordinary exact-Agent deployment request, the script records an intent. Successful responses are saved atomically. On resume, existing responses are reused; an intent without a response triggers an Agent readback and stops for operator reconciliation. The helper does not infer rejection from an unchanged active revision or replay an unknown request. An operator can record a verified accepted response and resume.
The script polls each returned deployment through its authorized status
operation, confirms active revision selection, and waits for all revision Pods
to be Running and Ready on the candidate runtime digest. Embedded execution
requires one runtime container; dedicated execution requires both gateway and
Agent containers. Each replacement gateway then runs read-only
openclaw doctor --lint --json --severity-min error. Failures retain dispatch,
status, Pod, and Doctor evidence for inspection.
6. Hand off application verification
docs/guides/deploy/production-upgrade.md:Verify the release
Successful script completion proves the selected Helm rollout and OCC access. For a runtime release it also proves the recorded deployments and Pods reached the checked states and Doctor reported no error. The operator next verifies model responses, providers, channels, credentials, workspace continuity, native access, and required restore behavior.
Debugging and Verification
- Inspect
server-dry-run.txtfor chart or admission failures before mutation. - For OCC rollout failures, inspect
helm-upgrade.txt, initialization Job logs, and API and worker rollout status. - For runtime failures, inspect
dispatch/*.error, revision history, andstatus/*.jsonbefore retrying anything. Doctor failures are recorded instatus/*.doctor.jsonandstatus/*.doctor.error. - Compare
before-workloads.jsonandafter-workloads.jsonfor unexpected workload changes. The controller-only helper requests no Agent deployments; separately check repository-bound revisions affected by broker restart. Runtime proof should show the intended replacements. - Use the credentialed production Kubernetes integration with distinct baseline and candidate images for end-to-end proof. Mocked commands prove only script control flow.
