OpenClaw EnterpriseDOCSGitHub

Harness Execution Topology Flow

Overview

An authorized deployment resolves its harness from native selected-model/provider policy, freezes the Agent's explicit embedded or dedicated placement and harness authentication binding in its AgentRevision, and asks Compute to start that topology. The flow ends after guarded route publication, predecessor retirement, and exactly-once activation audit.

Entry Points

Flow

graph TD
  A["Authorize Agent and Configuration"] --> B["Resolve explicit native runtime and placement"]
  B --> C["Freeze configuration, harness identity, and harness authentication binding"]
  C --> D["Claim and reauthorize revision work"]
  D --> E{"Approved topology"}
  E -->|embedded OpenClaw| F["Create gateway or stage replacement"]
  E -->|dedicated Codex| G["Start Gateway and Codex in separate namespaces"]
  E -->|dedicated OpenClaw| Q{"Full-containment provisioning Sandbox?"}
  Q -->|no| H
  Q -->|yes| R["Start Gateway; SandboxDriver provisions native Harness"]
  E -->|unsupported or mismatched| H["Reject before workload creation"]
  F --> I["Activate shared gateway; Recreate on replacement"]
  I --> K{"Gateway ready after startup authentication?"}
  K -->|no| L["Stay unready; Agent may be unavailable until repair"]
  K -->|yes| J["Complete activation, retire predecessor, and commit audit"]
  G --> N{"Predecessor Gateway can enroll node?"}
  N -->|yes| M["Activate authenticated dedicated revision"]
  N -->|no| O["Start candidate Gateway as bootstrap endpoint"]
  O --> P["Enroll and observe workspace node"]
  P --> M
  R --> S{"Gateway and enrolled Harness ready?"}
  S -->|no| L
  S -->|yes| M
  M --> J

Execution Trace

1. Resolve and freeze the native harness

packages/occ/src/index.ts:OpenClawController.deployAgent

OCC authorizes and locks the exact Agent and Configuration. Selected-model/provider agentRuntime.id explicitly selects codex or openclaw; only an unambiguous built-in configuration defaults to embedded OpenClaw. Missing ambiguous/plugin runtime policy, conflicting routes, unsupported IDs, and harness/mode mismatches fail closed. OCC validates each primary and fallback model through the same resolver; fallbacks must keep the primary provider and Harness. It preserves their order in the native configuration. The admitted revision immutably captures its native configuration, approved harness identity/version, explicit mode, Compute selection, and Agent ServicePrincipal. Production admits approved openclaw/embedded and codex/dedicated, and openclaw/dedicated only when the selected SandboxDriver provisions Harnesses with networking, filesystem, and process containment. An associated access_token additionally requires dedicated Codex; the frozen account contains only its OCC identity, credential kind, and opaque Secret reference.

2. Claim work and realize the approved topology

apps/controller/src/worker.ts:ControllerWorker

The worker claims exact revision work, reauthorizes its actor and ownership, revalidates its frozen approved harness, and calls ComputeDriver.prepareRevision.

apps/controller/src/drivers/compute/docker/index.ts:DockerComputeDriver.prepareRevision

Docker starts an embedded gateway or dedicated Codex container but does not support the harness-auth binding contract; unsupported bindings fail before deployment. In the underlying container path, dockerGatewayConfigurationDocument admits only supported authentication fields and modes. Omitted mode renders password mode. An omitted password or explicit managed reference selects OPENCLAW_GATEWAY_PASSWORD; other password settings are preserved. Explicit trusted proxy retains its native configuration and can also request the managed password. reconcileGateway generates the managed credential only for a new container, leaving a reused container's credential intact. Dedicated Codex app-server authentication remains independent. These implementation checks do not establish a currently deployable Docker Agent path.

Kubernetes supports managed bindings. SSH supports { "method": "runtime" } only for embedded OpenClaw: operator credentials remain on the host and OCC checks gateway readiness without model validation. See the SSH flow.

apps/controller/src/drivers/compute/kubernetes/index.ts:KubernetesComputeDriver.prepareRevision

Kubernetes workload rendering calls prepareHarnessAuth once for the resolved source. It projects the OCC Secret key only into embedded OpenClaw or a dedicated Harness. Canonical sources live in CP; Compute delivers selected fields into an exact revision-owned DP Secret, including the account token/workspace for ChatGPT. Dedicated gateways receive neither model source. This namespace-local delivery also applies to fixture images without native runtime configuration; only the native dedicated transport token depends on that configuration. See the harness authentication flow for admission, immutable source snapshots, and worker reauthorization.

Kubernetes ensureNamespace prepares the data-plane namespace and a distinct managed Gateway runtime namespace. requireGatewayNamespace verifies the latter's exact logical owner. prepareRevision and activateRevision place dedicated Gateway Deployments, private PVCs, Services, native configuration and routes there; Harness resources stay in the data-plane namespace. deliverGatewaySecrets validates direct references to canonical CP sources for dedicated Gateways; deliverHarnessAuth creates the selected DP runtime projection. Dedicated app-server DNS includes the Harness namespace, and NetworkPolicy peers combine namespace and exact Agent/revision selectors. The active dedicated Harness Service selector carries the same Namespace, Agent, revision, and workload-role labels before adding a Compute-owned workload-name selector, so Service-IP traffic remains compatible with NetworkPolicy implementations that check Service selectors before destination translation. Active Gateway Services carry the Namespace, Agent, and gateway workload-role labels, satisfying gateway policy selectors without tying the stable Gateway route to a revision. runtime.gatewayNodeSelector independently places the Gateway Pod and private-state initializer on trusted nodes. During a dedicated replacement, preparation keeps a healthy predecessor Gateway in place while the candidate Harness enrolls its workspace node. If the predecessor Gateway is the same Agent but cannot become ready, preparation starts the candidate Gateway after the candidate Harness is otherwise ready. That candidate Gateway provides the bootstrap endpoint; the revision remains not ready until the workspace node is enrolled and observed.

Dedicated Codex and dedicated OpenClaw keep separate Agent-owned Gateway and Harness ServiceAccounts. Compute owns the Gateway Pod; the selected SandboxDriver owns the native Harness Pod. The OpenClaw Harness enrolls as a paired node, owns its identity and workspace, and alone receives the model key. It reads the one-use enrollment target from a private file; later starts reuse the persisted device token. Compute pins the enrolled device in a generated dedicated-native profile with inference: "worker", so a missing or disconnected Harness fails the turn rather than using Gateway inference. An exact callback route and session-bound worker admission scope the transport to the owning Agent. Embedded OpenClaw uses one combined workload with its exact Agent identity and model key. The worker has scoped Secret permissions for admitted delivery and node enrollment. Its trusted workload-writing authority also projects tenant Secrets. Gateway Pods receive no controller or Harness Kubernetes credentials.

The selected Sandbox consumes the same rendered projections and explicit login mode in HarnessWorkloadRequirements. Unsupported upstream projection fails without a test-only credential bridge.

Every Pod template Kubernetes Compute renders carries the ordinary network profile. Ordinary allow policies and Gateway/Harness peers require it, and readiness rejects a template without it.

When a selected SandboxDriver provisions the dedicated Harness, providerHarnessReady lists Pods using the same Agent/revision/role labels as the active Service. It validates the complete observation and requires exactly one nonterminating candidate with the supplied Harness labels and Ready=True. An unready second live candidate blocks readiness even when the first is Ready. Malformed or incomplete observations throw through the existing preparation cleanup path. activateRevision repeats this check before changing routing. See the Kubernetes readiness contract for candidate rules and the limits of this observation.

3. Publish safely and complete activation once

apps/controller/src/worker.ts:ControllerWorker

For dedicated Kubernetes execution, Compute declares requiresStoppedPredecessors. ControllerWorker.prepareRevision stops every earlier runtime and waits for Pod termination before preparing the replacement. The worker records each predecessor it stopped and skips it on later pending passes and maintenance, which avoids repeating every stop on each readiness poll. Compute reports a predecessor that came back (for example, a lost claim's late write) as an unready successor, not an error, so the worker stops each recorded predecessor again after one claim lease, then after two, four and so on. A failed preparation pass, or preparing or activating that predecessor, drops the record. Old reconciliation and maintenance cannot restart a predecessor after a newer exclusive revision is admitted. Both PVCs survive this downtime window; a failed candidate is recovered by retry or a new revision, not automatic rollback. Dedicated Codex and dedicated OpenClaw must complete a bounded native authentication/model probe before their Harness becomes ready. The candidate Gateway repair in step 2 changes no unrelated Gateway and never activates a revision without its exact workspace node. Embedded preparation does not validate the replacement's credentials. See the authentication flow.

The worker commits the database activeRevisionId with an exact compare-and-set before Kubernetes default after-commit activation. KubernetesComputeDriver.activateRevision updates the shared gateway's Recreate Deployment and Service. Embedded cutover can stop the serving gateway before the replacement validates credentials in its own startup. The same bounded check runs for initial and replacement gateways. A failed check, including a provider timeout or rate limit, holds the gateway unready until repair and restart or a new deployment. Readiness polling does not repeat model requests; worker retries do not restart an unchanged Pod. No automatic rollback restores the predecessor.

If activation, readiness, predecessor retirement, or audit completion fails, the worker requeues the revision with REVISION_FINALIZATION_INCOMPLETE; recovery retries activation and retirement for the already-active revision. Lost claims and foreign/stale workloads fail closed.

When stopping a revision, the Driver stops its Gateway while leaving the Harness available for active work. Gateway supervision and Pod termination allow the pinned runtime's 330-second service stop budget; the controller waits for Pod disappearance before stopping the Harness. Idle shutdown, or one before OpenClaw starts (wrappers run under tini), completes promptly. Forced termination can delay the successor until the persistent owner lease expires.

Kubernetes gateways in both modes mount their own persistent SQLite and media directories. Embedded gateways also retain their attested default workspace on the same private claim so continued turns survive Pod replacement. Dedicated Harnesses receive only the Harness workspace claim, where an Agent-scoped subdirectory keeps the node identity across Pod and revision replacement. The gateway's nested Codex home remains ephemeral. The driver creates separate Harness and gateway claims before their consuming Pods and relies on workload readiness instead of waiting for Bound, which would deadlock WaitForFirstConsumer storage classes. A nonroot gateway-image init container prepares private SQLite and media directories without credentials or elevated privileges, plus a node-owned mode-0700 /tmp so the fsGroup-writable emptyDir root never becomes a worker workspace ancestor.

Each image initializes its own bundled and plugin assets. Workspace-file access uses the enrolled Harness node; generated-image bytes return through the remote media reader. Gateway and Harness share no workspace, session, skill, or image mounts. See the storage contract. The OpenClaw node host keeps Gateway-issued worker bundles in its own state and workspaces below /home/node/workspace, away from gateway state, CODEX_HOME, and credentials. A restart republishes image-owned runtime assets and reconnects with the paired identity; readiness waits for the bounded identity check. The compile cache and model-probe state stay in node state and TMPDIR, which a Sandbox Driver can grant.

For a selected Sandbox Driver, stopping or retiring a revision always runs its required cleanup after stopping a Compute-owned ordinary Harness, or delegates provider-owned Harness removal to that cleanup. An absent ordinary Deployment does not skip cleanup, so a cleanup failure remains retryable. Revision retirement retains both owned claims even after stop removed the gateway. When another revision's Gateway or route survives in the other physical namespace, retirement removes only the old Gateway's resources and preserves the shared data-plane Agent identity, Service and policies. apps/controller/src/worker.ts:ControllerWorker.processAgentDeletion retires every revision before calling apps/controller/src/drivers/compute/kubernetes/index.ts:KubernetesComputeDriver.deleteAgentRuntimeCredentials to delete exact-owned private and shared claims by UID. Final deletion checks both physical targets, independently of the Agent draft's current execution mode. Cleanup failures retry before the worker removes the Agent's database identity. The storage contract owns claim sizes, mount paths, StorageClass requirements, and final teardown.

Debugging and Verification

Search documentation