OpenClaw EnterpriseDOCSGitHub

Controller reconciliation

The controller worker advances Namespaces from provisioning to ready, finishes deleting empty Namespaces, stops or deletes Agents, and prepares and activates admitted Agent revisions. It runs separately from the OpenClaw Control Plane (OCC) HTTP API, polls durable PostgreSQL work, and calls its selected Compute Driver. The default development Docker Compute Driver creates one Docker network per Namespace. It rejects the Harness authentication required by the public deployment API and cannot deploy Agents. Select Kubernetes to deploy Agents locally; it supports embedded OpenClaw, dedicated Codex, and dedicated native OpenClaw when a full-facet provisioning SandboxDriver is selected. Reviewed bundled or installed Drivers can reconcile their supported operations in both development and production. See the deployment guide.

This reference owns durable reconciliation, queue states, and recovery guarantees. The worker flow explains their execution through the source; the quickstart and deployment guide own process startup procedures.

Requirements

The worker does not support process-local state. Never start it with migration or PostgreSQL administrator credentials.

Configuration

The worker requires its own NODE_ENV and OCC_DATABASE_URL. Production also requires the same absolute OCC_CONFIG_PATH startup YAML as the API. Development omits OCC_CONFIG_PATH to select the bundled Docker Compute Driver; the default filesystem Configuration Driver is API-only and uses the controller's OCC_DEVELOPMENT_CONFIGURATION_ROOT. Set OCC_CONFIG_PATH only to choose another trusted Driver set explicitly. The worker resolves the singleton Installation internally. Driver IDs, implementations, and closed-schema settings come from the YAML; worker environment variables can tune the poll interval, claim lease, maximum attempts, and optional readiness-marker path. Defaults and validation are defined in the worker configuration reference.

The API's listener and Better Auth settings are not worker inputs. Both processes must use the same migrated PostgreSQL database.

Startup and readiness contract

Both processes resolve the persisted singleton Installation internally. Their YAML selects approved IAM, Compute, and Configuration Drivers and the bundled Kubernetes Secret Driver. Reviewed bundled and installed implementations are available according to the Driver selection reference. Driver-owned closed schemas are validated before construction; missing files, unknown fields, unavailable implementations, or plaintext credentials fail closed. See the Configuration guide.

The API additionally requires OCC_AUTH_SECRET and OCC_AUTH_BASE_URL. User sessions authenticate controller API callers; ordinary exact-resource IAM permissions and Restrictions still authorize every operation. Operators must expose the API only through an internal ClusterIP Service and enforce default-deny ingress with explicitly approved namespace and Pod selectors.

Before serving requests or claiming work, both processes verify the existing Installation and persisted IAM state. When the bundled Kubernetes Compute Driver is selected, they also verify explicit Kubernetes credentials, TLS trust, and exact Kubernetes Namespace access. The bundled Kubernetes Configuration Driver validates its authentication settings at startup but checks tenant ConfigMap access only when its first CRUD request runs; a driver-managed provisioning Namespace can return 503 until its tenant namespace and API RoleBinding exist. Explicitly selected external Namespaces instead reject Configuration creation with 409 until ready. AgentRevisions retain their selected Compute identity and immutable Configuration snapshot and explicit Harness execution mode. Production worker claim, stale-claim recovery, and backlog queries include Namespace and both approved AgentRevision pairs: dedicated Codex and embedded OpenClaw. Each Agent-owned gateway serves only its own active revision; unsupported Harness/mode combinations and external ingress remain unavailable.

Reconciliation lifecycle

Namespace creation and deletion queue infrastructure work. Agent stop and deletion queue exact-Agent lifecycle work. Agent deployment queues an immutable revision; the worker prepares it, activates its route, and retires its predecessor. Kubernetes uses a single-replica gateway with Recreate, so replacement can interrupt serving. For embedded OpenClaw, the predecessor can stop before the replacement passes authentication and readiness. Read Namespace and Agent reconciliation for lifecycle, queue states, authorization, and recovery.

Observability

Optional OCC metrics expose request, reconciliation, Agent inventory, and process measurements through separate private API/worker listeners. Metrics remain active independently of log level.

Use the observability guide to set log levels, configure export, and verify delivery. The API and worker share the OCC Pino logger. The API disables Fastify's default request logging and emits one sanitized http.completed record per response with the generated request ID, method, route template, status, and duration. Unexpected internal failures add http.unexpected_error with a bounded error code.

The worker emits fixed operational event classes through the same logger:

Bootstrap and migration scripts use the same level and write machine-protocol success records to stdout. Their structured failure diagnostics go to stderr. Startup failures write startup-error, worker.startup-error, installation.bootstrap-failed, or migration.failed and exit before serving or processing work. Log sanitization keeps only reviewed scalar fields and drops credentials, provider payloads, request objects, and unbounded error values. The worker does not expose an HTTP health endpoint.

Failures and diagnostics

Search documentation