Compute Driver Lifecycle Hooks Flow
Overview
Selected non-Compute Drivers can participate in Namespace and workload lifecycles without owning infrastructure or receiving deployment authority. This flow starts when trusted startup composition selects hook owners, follows Compute readiness, workload preparation, revocation, and teardown, and stops when Compute returns its existing lifecycle observation to OCC or the independent worker.
Entry Points
- Trigger: OCC reconciliation or an independently claimed Namespace/AgentRevision lifecycle job.
- Source:
apps/controller/src/composition/installation-config.ts:loadInstallationConfiguration,packages/occ/src/index.ts:OpenClawController.selectDriver, andapps/controller/src/worker.ts:ControllerWorker. - Assumptions: one server-owned Installation, exact selected Driver identities, established IAM authorization, an immutable Namespace or AgentRevision, and an unchanged concrete ComputeDriver.
Flow
graph TD
subgraph Startup["Trusted Installation startup"]
A["Select exact non-Compute Driver owners"] --> B["Freeze ordered hook callbacks"]
B --> C["Inject hooks into concrete ComputeDriver"]
end
subgraph Lifecycle["Compute-owned lifecycle"]
C --> D{"Lifecycle operation"}
D -->|ensure Namespace| E["Prepare tenant infrastructure and run namespace hooks"]
D -->|prepare revision| F["Prepare or stage revision resources"]
D -->|retire revision| G["Revoke workload access before stopping"]
D -->|delete Namespace| H["Revoke workloads and namespace before deletion"]
E -->|success| I["Report ready"]
F -->|ready| J["Return readiness for worker activation"]
E -->|failure| K["Compensate completed owners in reverse"]
F -->|failure| K
G -->|revocation fails| L["Preserve owned runtime for retry"]
H -->|revocation fails| L
endExecution Trace
1. Attach selected hook owners before Compute starts
apps/controller/src/composition/installation-config.ts:loadInstallationConfiguration
The Driver package loading flow resolves exact IAM, Configuration, and Compute identities. OCC and worker startup attach selected non-Compute lifecycle owners once before the first operation. The selected IAM Driver remains stable and loads current policy for every authorization decision.
Trusted startup enforces the production revision-stage contract. The worker invokes each stage only when the Harness lifecycle requires it and fails closed if a required stage or owner is unavailable.
2. Snapshot owners and bind cancellation
apps/controller/src/drivers/compute/lifecycle-hooks.ts:ComputeLifecycleDispatcher
The hook dispatcher snapshots exact Driver capability, identity, and callbacks once during registration. It preserves controller selection order for preparation and reverses that order for teardown. Registered callbacks remain stable even if their trusted owner changes. Hooks receive the worker's existing claim-owned abort signal; direct Compute calls receive a nonaborted fallback.
3. Gate Namespace readiness on completed hooks
apps/controller/src/drivers/compute/kubernetes/index.ts:KubernetesComputeDriver.ensureNamespace
Both Kubernetes and
Docker Compute implementations
prepare tenant infrastructure first, then run afterNamespacePrepared before reporting ready.
No gateway exists until Agent revision preparation. Failed preparation compensates completed
owners with beforeNamespaceDelete in reverse.
4. Validate launch contributions before starting a workload
apps/controller/src/drivers/compute/lifecycle-hooks.ts:ComputeLifecycleDispatcher.beforeWorkloadStart
Kubernetes prepareRevision invokes selected workload hooks for initial embedded gateway creation
and for dedicated Codex workload preparation. Before a dedicated Codex workload starts, preparation
stages the gateway-to-Agent runtime NetworkPolicies, the temporary authentication egress policy,
and workspace-node enrollment material; the gateway cannot enroll its node or reach the Agent
app-server until those policies exist. Embedded replacement revisions are different: after staging
the immutable configuration, Service, and private claim, prepareRevision can return ready without
starting the replacement gateway. After the worker commits the new active revision,
KubernetesComputeDriver.activateRevision invokes beforeWorkloadStart, updates the Recreate
gateway Deployment and Service, then checks gateway readiness.
runtime.nodeSelector schedules Harness and embedded Pods. Dedicated real Gateways require
runtime.gatewayNodeSelector and run in the logical Namespace's managed control-plane runtime
namespace. The Pod-level selector also schedules the Gateway's private-state initializer there.
Compute owns both targets through the same revision lifecycle; teardown selects each resource's
physical namespace and preserves newer revisions and durable Agent claims.
Dedicated replacement stops earlier Harnesses before preparing the successor.
The Agent-owned authentication NetworkPolicy selects only the successor revision;
superseded reconciliation cannot move that grant back to a predecessor.
SSH stages embedded snapshots without starting the candidate gateway. After the
worker commits the active revision, SshComputeDriver.activateRevision invokes
beforeWorkloadStart, then projects accepted launch placeholders into the
Agent's systemd unit before restart. Failed activation compensates prepared
bindings through beforeWorkloadStop.
The dispatcher rejects reserved environment keys and values outside the explicit opaque-
placeholder format, then freezes a detached launch snapshot using the shared
immutability helpers. Kubernetes and Docker Compute project
accepted placeholders only into the combined embedded gateway or separate dedicated Codex container,
never a separate dedicated gateway. A failed launch or activation revokes successfully prepared
workload bindings.
5. Revoke access before stopping owned resources
apps/controller/src/drivers/compute/kubernetes/index.ts:KubernetesComputeDriver.retireRevision
Workload retirement runs beforeWorkloadStop in reverse owner order before stopping the exact
revision. Namespace deletion is limited to empty tenants and runs
beforeNamespaceDelete before requesting Kubernetes deletion. A
failed revocation preserves the owned resource for retry; aborted preparation receives a fresh,
bounded cleanup signal so cancellation cannot suppress compensation.
Debugging and Verification
- Run
node --test tests/conformance/utils.test.mjs tests/conformance/compute-lifecycle-hooks.test.mjsfor shared immutable copies, callback capture, preparation/teardown ordering, opaque placeholders, cancellation, and rollback. - Run
node --test --test-name-pattern='Kubernetes lifecycle owners cannot be replaced|Kubernetes lifecycle hooks never run' tests/conformance/kubernetes-compute.test.mjsfor concrete Kubernetes owner and lifecycle-boundary behavior. - Run
node --test tests/conformance/plugin-compute.test.mjs --test-name-pattern "embedded plugin preparation applies runtime egress before gateway readiness"for Kubernetes preparation ordering of embedded and dedicated runtime NetworkPolicies before Gateway readiness. - Errors include only the failing hook phase and owner. Investigate exact selected identities, operation cancellation, unsafe placeholder values, and pending revocation without printing credentials or sensitive endpoints.
- PostgreSQL, Docker, live Kubernetes, and OpenShell proof require their real dependencies and infrastructure; unavailable integrations must be skipped explicitly.
