OpenClaw EnterpriseDOCSGitHub

Namespace and Agent reconciliation

The controller worker reconciles durable Namespace and AgentRevision operations. This reference defines lifecycle transitions, queue ownership, retries, and recovery.

Namespace lifecycle

Creating a Namespace saves provisioning and queues its provisioning operation in the same transaction. An Installation administrator can also specify existingNamespace to persist the exact existing Kubernetes namespace before provisioning starts. Before acting, the worker reloads current IAM policy, reauthorizes the original actor, and confirms the operation still belongs to its exact Namespace. External selection additionally requires Installation administer authorization at admission and immediately before adoption. The worker asks the Compute Driver to ensure backing infrastructure and sets ready once it is ready; selecting an existing namespace never requires stopping the shared worker.

Deleting an empty Namespace saves deleting and queues a distinct teardown operation. The worker rechecks the original actor's permission, asks the same Compute Driver to delete the Namespace and its owned Agent gateways, and records a tombstone after Namespace deletion. Tombstoned Namespaces disappear from public reads.

Selected non-Compute Drivers may run hooks after Namespace infrastructure readiness, before workload start, before workload retirement, and before Namespace removal. Compute owns every transition; revocation failures block teardown, launch values are restricted to opaque- placeholders, and production workers process Namespace operations plus embedded OpenClaw and dedicated Codex Agent revisions. See ComputeDriver lifecycle hooks.

Backing infrastructure behavior is defined by the selected Docker or Kubernetes Compute implementation. A successful queue transition does not itself establish enforcement of cluster admission, NetworkPolicy, or a SandboxDriver facet; those guarantees require the selected implementation and its documented infrastructure.

Agent lifecycle

Stopping an Agent sets its desired runtime state to stopped and queues an exact-Agent stopped target. The worker reauthorizes the original actor, stops the current revision, and clears the active pointer only if it still identifies that revision. Revision history, credentials, and persistent state remain; a later deployment starts a new revision.

Deleting an Agent sets its lifecycle status to deleting, sets desired runtime state to stopped, and queues an exact-Agent deleted target. Synchronous Agent mutations reject this state. The worker reauthorizes delete, retires every revision, and removes runtime credentials before a claim-protected database finalizer removes the Agent, revisions, service principal, API keys, exact IAM references, and Agent work rows. The function records durable success evidence; an expired claim or failed external cleanup leaves the rows intact for safe retry. Namespace-owned Configurations and Secrets are not Agent teardown state.

AgentRevision lifecycle

An authorized bodyless Agent deployment reads its exact Namespace-owned native Configuration and permits the selected SandboxDriver to transform a copy before validation. It snapshots the admitted document, including unresolved inline SecretRefs, alongside the native-selected Harness identity, server-approved version, explicit Agent execution mode, and Compute implementation. The required Agent harnessAuth selects either an OCC Secret API-key source or an issued managed ChatGPT account. The revision freezes the Secret reference and selected Driver, or the account's exact access-token reference and verified private Backend binding. OCC separately authorizes the Configuration and harness source before queueing one revision operation in the same transaction. The source Configuration identity and generation remain pinned even when the admitted copy differs. Later Configuration, account, or Agent placement changes never mutate an admitted revision; see Agent references and deployment. PostgreSQL enforces the exact admitted snapshot shape, so the worker trusts persisted structure instead of revalidating it.

Before processing that operation, the worker reloads current IAM policy, reauthorizes the original actor for the exact Agent and Configuration, and checks the frozen harness source. API-key delivery requires exact Secret operate for both actor and Agent service principal; managed ChatGPT delivery requires actor read on the exact account and matching credential/Backend ownership. It also checks the owning ready Namespace, stable Agent service principal, and approved Harness, version, mode, and Compute implementation. These checks use the immutable revision, not a later Agent draft. Revoked source access fails permanently with AUTHORIZATION_DENIED, records attributable deployment-denial audit evidence, and prevents Compute calls and activation. The harness credential flow owns the complete source-resolution sequence. Production accepts dedicated Codex and embedded OpenClaw. Dedicated Codex prepares its revision-specific workload before the Agent Service selects it. Embedded OpenClaw replaces and checks the shared Agent gateway during activation. The worker records the exact active revision, activates the Kubernetes route when applicable, and retires the prior revision. Its live claim remains unfinished until it atomically records one attributable activation audit and completes the durable operation. Already-active recovery repeats route activation and predecessor retirement before that audit and finalization. Kubernetes dedicated replacement stops all earlier runtimes before preparing the candidate and reuses the Harness-only RWO claim. This interrupts serving, including the gateway; a failed candidate needs retry or a new revision, not automatic rollback. See the exclusive replacement contract. Pod termination does not fence independent processes during node partitions or manual replacement. See the Harness execution topology flow for the full placement, runtime, and recovery sequence.

A recovered older operation never replaces a newer active revision: the worker marks it superseded without calling Compute. Sibling Agents have independent queue lanes, while revisions for the same Agent serialize. The default PostgreSQL-backed development Compute Driver provisions Docker Namespace resources but rejects harness bindings for Agent deployment; manual host-process debugging also needs PostgreSQL for a durable worker path. The explicitly selected Kubernetes driver creates a hardened Deployment and dedicated Kubernetes ServiceAccount for the revision's existing Agent ServicePrincipal. Its audience-scoped projected token is required in production but does not implement ServicePrincipal token verification or exchange. The dedicated Codex Agent Service does not select a replacement workload until it is ready and the revision is active; embedded OpenClaw reuses and replaces its existing gateway. Selected SandboxDriver facets are pinned at admission and enforced by the selected Driver. See the SandboxDriver contract for provider-specific preparation and failure boundaries.

Controller queue states

Every controller operation persists in PostgreSQL and moves through the following states:

stateDiagram-v2
    [*] --> queued: Namespace or Agent operation committed
    queued --> claimed: Worker acquires claim and lease
    claimed --> claimed: Heartbeat renews lease
    claimed --> succeeded: Effect and lifecycle update commit
    claimed --> queued: Pending convergence, retryable failure, or expired lease
    claimed --> failed_permanent: Access denied or attempts exhausted
    queued --> failed_permanent: Recovery finds attempts exhausted
    succeeded --> [*]
    failed_permanent --> [*]

Terminal results

Terminal work stores its overall outcome in reasonCode and optional structured success or failure details in resultData (the PostgreSQL result_data column). Successful activation keeps REVISION_ACTIVATED or REVISION_ALREADY_ACTIVE even when resultData.warnings contains different plugin failure codes. Convergence deadline failures store their allowed timeoutMs in the same field. Queued and claimed work have no result data.

Warnings contain only an allowed code and an admitted plugin ID. The deployment status API derives error and warnings from this saved outcome; a successful deployment with plugin warnings still returns error: null. Only the current live claim can publish the result.

Deferred Namespace and Agent convergence

The worker defers a Namespace or AgentRevision operation when its Compute Driver successfully observes infrastructure that is not ready yet and reports no operational failure. Examples include waiting for an operator-provisioned tenant RoleBinding, Agent image startup, a ready gateway Pod or EndpointSlice, or completion of Kubernetes Namespace deletion.

defer() is a transition, not an additional queue state. It returns the work to queued, releases its claim, schedules bounded backoff, records audit evidence, and restores the attempt consumed when the work was claimed. The worker can observe ordinary convergence repeatedly without exhausting its failure budget.

Actual dependency failures instead use retry(), which also returns work to queued but retains the consumed attempt. Once OCC_WORKER_MAX_ATTEMPTS is exhausted, the operation becomes failed_permanent. Pending convergence has its own limit: OCC_WORKER_CONVERGENCE_TIMEOUT_MS, measured from the original operation creation time. Exceeding it fails the operation with CONVERGENCE_DEADLINE_EXCEEDED. A runtime that reports a deterministic credential rejection fails the deployment earlier with RUNTIME_AUTHENTICATION_FAILED. See the worker configuration reference for defaults and supported overrides.

Replacement behavior depends on the Harness. Dedicated Codex prepares its revision-specific workload before the worker switches the Agent Service, but the shared single-replica gateway can still interrupt serving during its rollout. Embedded OpenClaw replaces that shared gateway using Kubernetes Recreate: the predecessor can stop before the replacement passes startup authentication and readiness. A failed embedded rollout can interrupt serving; the worker retries according to its queue policy but does not guarantee that the predecessor stays available or restore it automatically.

If a worker exits or stops renewing its lease, stale-claim recovery either requeues the operation or marks it failed_permanent after its final attempt. Recovery can also terminalize an already queued operation whose attempts are exhausted. When that work targets Namespace creation, recovery changes a still provisioning Namespace to failed in the same atomic statement as the terminal work state and audit evidence. Retryable recovery leaves it provisioning; Agent work, Namespace deletion work, and Namespaces already past provisioning do not change Namespace status through this recovery path.

Authorization, retries, and scope

Search documentation