Controller reconciliation
The controller worker advances Namespaces from provisioning to ready,
finishes deleting empty Namespaces, stops or deletes Agents, and prepares and
activates admitted Agent revisions. It runs separately from the OpenClaw Control Plane
(OCC) HTTP API, polls durable PostgreSQL work, and calls its selected Compute
Driver. The default development Docker Compute Driver creates one Docker
network per Namespace. It rejects the Harness authentication required by the
public deployment API and cannot deploy Agents. Select Kubernetes to deploy
Agents locally; it supports embedded OpenClaw, dedicated Codex, and dedicated
native OpenClaw when a full-facet provisioning SandboxDriver is selected. Reviewed
bundled or installed Drivers can reconcile their supported operations in both
development and production. See the deployment guide.
This reference owns durable reconciliation, queue states, and recovery guarantees. The worker flow explains their execution through the source; the quickstart and deployment guide own process startup procedures.
Requirements
- Node.js 24 or newer and the repository's existing workspace dependencies.
- For default Compose development, Docker Engine and
docker compose. To deploy Agents locally, follow Local Setup. - A migrated local PostgreSQL database and an API using the same application connection; the full Compose stack starts both. See the configuration reference.
- For default PostgreSQL-backed development, the API must also use the
filesystem Configuration Driver root. Compose mounts
occ_configuration_dataonly into the controller at/app/.development/configurations. - A bootstrapped Installation. Compose's
bootstrapservice and the Helm initialization Job runscripts/bootstrap-installation.mjsafter migration and before the API or worker. Direct-process setups run that initializer first. NODE_ENV=developmentorNODE_ENV=productionand the application-roleOCC_DATABASE_URL.- In production, the shared absolute
OCC_CONFIG_PATHto trusted Installation startup YAML selecting approved IAM, Compute, and Configuration Drivers and the bundled Kubernetes Secret Driver.
The worker does not support process-local state. Never start it with migration or PostgreSQL administrator credentials.
Configuration
The worker requires its own NODE_ENV and OCC_DATABASE_URL. Production also
requires the same absolute OCC_CONFIG_PATH startup YAML as the API.
Development omits OCC_CONFIG_PATH to select the bundled Docker Compute
Driver; the default filesystem Configuration Driver is API-only and uses the
controller's OCC_DEVELOPMENT_CONFIGURATION_ROOT. Set OCC_CONFIG_PATH only to
choose another trusted Driver set explicitly. The worker resolves the singleton
Installation internally. Driver IDs,
implementations, and closed-schema settings come from the YAML; worker
environment variables can tune the poll interval, claim lease, maximum
attempts, and optional readiness-marker path. Defaults and validation are
defined in the worker configuration reference.
The API's listener and Better Auth settings are not worker inputs. Both processes must use the same migrated PostgreSQL database.
Startup and readiness contract
Both processes resolve the persisted singleton Installation internally. Their YAML selects approved IAM, Compute, and Configuration Drivers and the bundled Kubernetes Secret Driver. Reviewed bundled and installed implementations are available according to the Driver selection reference. Driver-owned closed schemas are validated before construction; missing files, unknown fields, unavailable implementations, or plaintext credentials fail closed. See the Configuration guide.
The API additionally requires OCC_AUTH_SECRET and OCC_AUTH_BASE_URL.
User sessions authenticate controller API callers; ordinary
exact-resource IAM permissions and Restrictions still authorize every
operation. Operators must expose the API only through an internal ClusterIP
Service and enforce default-deny ingress with explicitly approved namespace and
Pod selectors.
Before serving requests or claiming work, both processes verify the existing
Installation and persisted IAM state. When the bundled Kubernetes Compute
Driver is selected, they also verify explicit Kubernetes credentials, TLS
trust, and exact Kubernetes Namespace access. The bundled Kubernetes
Configuration Driver validates its authentication settings at startup but
checks tenant ConfigMap access only when its first CRUD request runs; a
driver-managed provisioning Namespace can return 503 until its tenant
namespace and API RoleBinding exist. Explicitly selected external Namespaces
instead reject Configuration creation with 409 until ready. AgentRevisions
retain their selected Compute identity and immutable
Configuration snapshot and explicit Harness execution mode. Production worker
claim, stale-claim recovery, and backlog queries include Namespace and both
approved AgentRevision pairs: dedicated Codex and embedded OpenClaw. Each
Agent-owned gateway serves only its own active revision; unsupported
Harness/mode combinations and external ingress remain unavailable.
Reconciliation lifecycle
Namespace creation and deletion queue infrastructure work. Agent stop and
deletion queue exact-Agent lifecycle work. Agent deployment queues an immutable
revision; the worker prepares it, activates its route, and retires its
predecessor. Kubernetes uses a single-replica gateway with Recreate, so
replacement can interrupt serving. For embedded OpenClaw, the predecessor can
stop before the replacement passes authentication and readiness. Read
Namespace and Agent reconciliation for lifecycle,
queue states, authorization, and recovery.
Observability
Optional OCC metrics expose request, reconciliation, Agent inventory, and process measurements through separate private API/worker listeners. Metrics remain active independently of log level.
Use the observability guide to set log levels,
configure export, and verify delivery. The API and worker share the OCC Pino
logger. The API disables Fastify's default request logging and emits one
sanitized http.completed record per response with the generated
request ID, method, route template, status, and duration. Unexpected internal
failures add http.unexpected_error with a bounded error code.
The worker emits fixed operational event classes through the same logger:
worker.started: confirms the selected Compute Driver and optional SandboxDriver.worker.health: reports readiness and pending work count at debug level.worker.completed: includesnamespaceId, work identity, attempt, outcome, and a stable result code; AgentRevision operations also includeagentIdandrevisionId. Each deployment pass adds worker wall-clock milliseconds:durationMsfor the pass,deployPassesand summed ComputeprepareMsso far,readinessWaitMsfrom the first unready observation to the first ready one (or to now while pending),activationMsfrom the ready observation to completion, andelapsedMssince admission. Totals cover the passes this worker process ran; a restart starts them again. Maintenance, cleanup, and stop work carry no deployment timing.worker.error: reportsCLAIM_LOSTorWORKER_UNAVAILABLEwithout exposing credentials.worker.stopped: confirms graceful shutdown.
Bootstrap and migration scripts use the same level and write machine-protocol
success records to stdout. Their structured failure diagnostics go to stderr.
Startup failures write startup-error, worker.startup-error,
installation.bootstrap-failed, or migration.failed and exit before serving
or processing work. Log sanitization keeps only reviewed scalar fields and drops
credentials, provider payloads, request objects, and unbounded error values. The
worker does not expose an HTTP health endpoint.
Failures and diagnostics
- Namespace stays
provisioning: Start the separate worker, verify both processes use the same database, and inspectworker.healthandworker.completedoutput. - Docker Namespace does not become ready: Confirm Docker Engine access from the worker and inspect the per-Namespace Docker network labels. Docker Agent container tooling is for contributor verification; it cannot deploy through the public API. Use Docker Compute Driver troubleshooting.
- Configuration creation returns
503: Confirm the API, not the worker, hasOCC_DEVELOPMENT_CONFIGURATION_ROOTset and can write the/app/.development/configurationsmount backed byocc_configuration_data. - AgentRevision does not activate: Confirm its Namespace is
ready, the original actor retains exact-Agentdeploypermission andreadon any account in its immutable revision, and its pinned Harness descriptor and Compute implementation match the worker. - Kubernetes Namespace or Agent workload does not become ready: Confirm the API and worker selected the same configured driver and the worker uses the bootstrapped singleton Installation; check explicit cluster authentication, externally provisioned tenant-local RBAC, enforced NetworkPolicies, image availability, and gateway EndpointSlices. See the Kubernetes Compute Driver guide.
- Installation is not bootstrapped: Confirm the Compose
bootstrapservice or Helm initialization Job succeeded against the API/worker database. For direct-process setup, runscripts/bootstrap-installation.mjswith the selected environment's protected output settings before starting either process. Resolve failed initialization manually before another attempt; bootstrap does not clean up or retry. - Startup YAML is missing or rejected: Set production
OCC_CONFIG_PATHto the same absolute, readable file for API and worker. Remove unknown Driver fields and plaintext credentials; verify all selected Driver implementations and exact Kubernetes access. - Configuration operations fail: Verify exact Namespace or Configuration authorization, tenant-local ConfigMap CRUD, and a native JSON configuration document; see Configuration troubleshooting.
A valid PostgreSQL connection URL must be explicitly configured.: SetOCC_DATABASE_URLto the migrated application'spostgresql:connection.- Worker mode is rejected: Set
NODE_ENV=developmentorNODE_ENV=productionexplicitly. Production additionally requires valid trusted startup YAML selecting approved Drivers, including the bundled Kubernetes Secret Driver. AUTHORIZATION_DENIEDorACTOR_REVOKED: Inspect the initiating actor's current role, binding, exact-Namespace Restrictions, andreadpermission on any service account captured in the immutable revision; see IAM.CLAIM_LOST: Another valid claim recovered the operation. The stale attempt cannot publish lifecycle state; inspect subsequent worker events.
Related
- API operations, authentication, and permission reference
- Namespace lifecycle and deletion
- Agent Configuration, revisions, and deployment
- Native service accounts and account authorization
- Controller and PostgreSQL configuration
- Docker Compute Driver and Compose development
- Namespace Configuration and Kubernetes ConfigMaps
- Kubernetes Compute Driver and local-cluster verification
- Identity and access management
- Platform architecture
