OpenClaw EnterpriseDOCSGitHub

Production Startup Flow

Overview

Production startup begins after an operator supplies approved images, PostgreSQL credentials, authentication material, trusted Installation startup YAML, network policy inputs, and protected bootstrap storage. The supported path prepares a fresh bootstrap PVC, installs the Helm chart, waits for the private API and worker, then proves authenticated /installation access with the retrieved bootstrap service key. This flow ends at control-plane access; tenant Agent deployment and model-backed TUI proof are later flows.

For the operator commands, use the deployment guide. The chart owns migration/bootstrap ordering and controller readiness. It does not provision cloud infrastructure, publish images, create TLS, retrieve keys, or prepare application Secrets automatically.

Entry Points

Flow

graph TD
    subgraph Operator["Operator-owned preparation"]
        A["Edit native values, Installation YAML, and bootstrap PVC manifest"]
        B["Create system namespace and file-backed Secrets"]
        C["Create fresh bootstrap PVC"]
        D["prepare-bootstrap-volume verifies empty output and root permissions"]
    end
    subgraph Helm["Helm-owned startup"]
        D --> E["Render chart with native values"]
        E --> F["Run initialization Job with migrator and application roles"]
        F --> O{"Canonical migration history?"}
        O -->|No| P["Refuse initialization before migration DDL"]
        O -->|Yes| G["Apply migrations, bootstrap administrators and write protected key output"]
        G --> H["Start private API Deployment"]
        G --> I["Start independent worker Deployment"]
        I --> O{"Repository credentials enabled?"}
        O -->|Yes| P["Sidecar copies protected inputs and starts private control"]
        P --> Q["Sidecar probe gates worker Pod readiness"]
        H --> L["Run Kubernetes Compute preflight"]
        I --> L
        L --> M{"Kubernetes older than 1.35?"}
        M -->|Yes| N["Emit advisory warning and continue"]
    end
    subgraph Proof["Operator-owned authenticated proof"]
        M -->|No| J["Retrieve service-key response from protected storage"]
        N --> J
        J --> K["occ installation get from approved client"]
    end

Execution Trace

The migration, shared bootstrap, API, and worker entrypoints use createPostgresPool. See connection authentication settings for password and Azure workload-identity configuration. When PostgreSQL ends an idle pooled connection (failover, maintenance restart, idle_session_timeout or a proxy reset), the pool discards that client and writes one database.idle-client-error warning with only the error code to stderr; the process keeps running and the next query opens a new connection.

1. Prepare native production inputs

deploy/helm/openclaw-enterprise/values.yaml:1

The operator copies and edits the production example values, Installation YAML, and bootstrap PVC manifest outside the checkout. Helm values select the controller image, API endpoint, Secret names, bootstrap claim, API-client selectors, control-plane node selector, and egress destinations. The Installation YAML selects IAM, Configuration, Compute, optional Backend, gateway/Agent images, projected workload identity, and runtime networking/storage.

The operator creates file-backed Kubernetes Secrets for Installation startup, database URLs, optional database CA bundles, Better Auth signing material, and optional ChatGPT Backend administrator credentials. These are prepared inputs, not recurring synchronization targets. The chart does not infer gateway/Agent images from Helm values or rewrite Driver configuration.

2. Prepare the fresh bootstrap volume

scripts/prepare-bootstrap-volume:124

Before the first install, the operator creates the bootstrap PVC named by bootstrap.password.claimName and runs the helper with explicit kubeconfig, context, namespace, claim, approved Node-capable image, and optional repeated --node-selector KEY=VALUE labels. The helper launches a bounded preparation Pod, applies the selectors before WaitForFirstConsumer storage binds, verifies the mounted root is fresh except for filesystem-owned lost+found, sets UID/GID 1000 with mode 0700, and refuses to continue on any other entry.

If cluster policy forbids the helper Pod, storage administration owns the same state transition through an approved storage workflow. A preprepared claim goes directly to Helm. The helper does not create the PVC, repair a used claim, retrieve generated credentials, or change controller configuration.

3. Run Helm initialization

deploy/helm/openclaw-enterprise/templates/jobs.yaml:8

helm upgrade --install --wait --timeout 5m renders the chart with native values. If database.caSecretName is set, the Pod mounts that CA Secret read-only into both containers before they connect. The initialization hook first runs migrations with the dedicated migrator credential, then runs bootstrap with the lower-privilege application credential, Better Auth settings, first administrator email, Installation name, and protected output paths.

scripts/migrate-production.mjs:1, scripts/migration-history.mjs:migrateWithHistory

The migration command verifies the complete SQL source manifest, checks the dedicated role and canonical receipt/catalog state, and holds one advisory lock on the connection used by Drizzle's normal transaction. It accepts a fresh database, canonical history through migration 0023, or the completed history through 0025. Unsupported or mixed development histories fail before migration DDL, preventing bootstrap from running. The same preflight serves development and production; see migration history and recovery for the read-only check and developer-selected recreation procedure.

scripts/bootstrap-installation.mjs creates or verifies the singleton Installation, human administrator, service administrator, IAM seed, audit evidence, and initial service key. On fresh bootstrap, it creates the initial default Namespace through OpenClawController.createNamespace, authorized as the bootstrap Principal. The Namespace and its queued reconciliation commit with Installation/IAM state and bootstrap audit; existing Installations receive no new Namespace. The worker later provisions normal Driver-owned infrastructure; operators still provide the tenant RoleBindings described in the deployment guide. The platform name does not select Kubernetes' default namespace. It writes password and service-key files only from the bootstrap container to the protected PVC. Existing output, unsafe storage permissions, inconsistent accounts, or mismatched IAM identity fail the Job; Helm failure does not imply the database hook was rolled back.

4. Start private API and worker Deployments

apps/controller/src/server.mjs:138, apps/controller/src/worker.ts:312

apps/controller/src/drivers/compute/kubernetes/index.ts:KubernetesComputeDriver.preflight

After successful initialization, Kubernetes starts separate API and worker Deployments. The API validates production listener settings, Better Auth, database access, trusted Installation YAML, selected Drivers, Backend membership, and Kubernetes Compute preflight before readiness. It serves private controller routes, /healthz, and database-backed /readyz behind the operator-managed endpoint.

When controlPlane.nodeSelector is non-empty, the chart places the API and worker Pods with that selector. The same selector applies to the initialization Job that runs the migration init container and bootstrap container, so production operators can keep migration, bootstrap, API, and worker Pods on a reviewed control-plane node pool. deploy/helm/openclaw-enterprise/templates/gateway-routing.yaml also projects that selector into EnvoyProxy.spec.provider.kubernetes.envoyDeployment.pod, so the credential-checking private proxy stays on the trusted pool. Empty chart defaults omit the field for clusters that do not label a dedicated control-plane pool. When database.caSecretName is set, API and worker also mount the CA Secret read-only at database.caMountPath. Tenant gateway and Agent placement remain in the selected Compute Driver configuration.

The shared egress policy selects only api, worker, and initialization Pods with the release identity. It allows DNS, database egress to every database.cidrs host, and Kubernetes API egress to every cluster.cidrs host. Collectors use their separate DNS, API, and exporter policy; unknown or missing component labels retain default-deny. Pre-install initialization has only its hook DNS/database grants. Each configured database or API destination must be an explicit IPv4 /32; operators must refresh the values when a managed database or API endpoint resolves to a different address set.

deploy/helm/openclaw-enterprise/templates/networkpolicies.yaml also renders an API-only TCP 443 egress policy when api.modelDiscoveryCidrs contains provider IPv4 /32 hosts. Empty defaults grant no provider egress. Operators maintain those addresses for the optional model-discovery API; Console model selection and Harness egress do not depend on this policy.

When api.channelDirectoryProxyUrl names an approved HTTP(S) proxy at a literal IPv4 address and port, the chart passes it to the API and grants only that Pod egress to the proxy's exact /32 and TCP port. The proxy must permit CONNECT to slack.com:443. The empty default renders no rule and leaves production Slack directory lookup unavailable with manual exact-ID entry.

The Kubernetes Compute Driver queries the API server version and verifies authenticated Namespace access. Kubernetes 1.35 or later is the supported baseline. An older server returns a structured preflight warning instead of blocking startup; the API logs compute.preflight-warning with the observed and minimum versions in its message and continues. An invalid version response, unreachable API, or failed Namespace access still fails preflight.

The worker independently validates production settings, opens the same application-role database, loads the selected Driver bundle, validates IAM, runs Compute preflight, emits the same advisory warning for an older Kubernetes server, emits worker.started, and polls durable Namespace and AgentRevision work. Worker readiness depends on fresh queue-health observations. Neither process mounts the bootstrap PVC.

apps/controller/src/composition/repository-credentials/platform.ts:composeRepoDriver

When selected, both processes load the same canonical registry and public CA, construct the GitHub Backend with a lazy Unix client, and register GitHubRepoDriver under the optional repo capability. The configured Backend member and drivers.repo must select the same Driver ID. API startup performs no control operation and loads no App key or session engine. The worker owns subsequent session lifecycle calls through the selected Driver; API readiness remains database-backed.

apps/controller/src/composition/repository-credentials/projected-inputs.ts:prepareProjectedInputs

The optional service runs beside the single worker. Its own process pins a projection generation, copies the known config/key/TLS/registry files into owned private regular files, validates the selected Backend and exact Service origin, then uses the protected loader and starts both listeners. App/TLS private material stays in that container. Only the worker receives an explicitly projected Kubernetes API token. The service's private health probe checks the Unix listener; it performs no provider operation. The service and worker share only the private control volume, and the Pod uses Recreate with a 75-second termination grace period. Both containers must pass readiness for the Pod to be ready; the worker keeps its queue-health probe and retries unavailable session operations through its ordinary Driver contract. This is one service owner, without independent worker/service high availability. Repository runtime reconciliation continues in the Agent repository credential flow.

5. Retrieve the key and prove authenticated access

internal/occclient/client.go:Client.GetInstallation

After Helm readiness, the operator retrieves initial-admin-service-key.json from protected bootstrap storage through an approved reader path and stores it in an owner-readable file. A completed Job is not an exec endpoint, and the API and worker cannot retrieve this file for the operator.

From an approved client environment, occ installation get uses the protected key file through the OCC client and displays the Installation. The production startup proof succeeds only when its ID matches the key response's meta.installationId. The operator records that ID in the openclaw.dev/installation-id annotation on the Installation startup Secret; coordinated upgrades use the marker to bind their OCC endpoint to the selected Kubernetes Installation. Agent runtime, gateway WebSocket authentication, and model calls remain unproven until the tenant deployment and TUI procedures run.

Debugging and Verification

Search documentation