Kubernetes networking and isolation
Configure tenant network boundaries, private Agent routes, and existing namespace ownership for the Kubernetes Compute Driver.
Networking
Configure the cluster DNS namespace and Pod labels and the gateway port.
Set network.gatewayTrustedProxyCidrs to a
nonempty list of valid CIDRs for the actual proxy socket sources. This is trusted
Installation configuration; the Driver has no production CIDR default and rejects
all-source ranges, including IPv4-mapped equivalents. Without
private routing, also configure the namespace and Pod selectors in
network.gatewayClients for your authenticated proxy.
Each tenant starts with default-deny ingress and egress. Explicit policies allow DNS, approved gateway clients, and required communication between an Agent's gateway and dedicated Harness. Cross-tenant traffic, traffic between different Agents, Kubernetes API access, and cloud metadata access remain denied.
For Compute-owned startup failure evidence, plugin reporting, and on-demand
deployment diagnostics, set
network.pluginStatusProxySourceCidrs to the precise source addresses used by the
Kubernetes API server when proxying requests to workload Pods. The policy allows
those sources only to the private status port, TCP/18791. Both worker and API
ServiceAccounts need namespace-local get on pods/proxy for their respective
reads. The ingress rule also applies when an Agent has no enabled plugins.
Prefer individual /32 or /128 addresses. On an
overlay network, the observed source may be the control-plane node's overlay
address rather than its node IP. Verify it across nodes with enforced policies.
An omitted list adds no API-proxy ingress rule and leaves status unavailable
where the cluster blocks that traffic. This setting does not expose the native
gateway or grant workloads Kubernetes API access.
When private Agent routing is enabled, Compute derives the only allowed peer
from gatewayRouting: the Envoy namespace and the Gateway's exact owning name
and namespace labels. Omit network.gatewayClients; startup rejects explicit
clients in routed mode. The native gateway trusts
the proxy's source range; NetworkPolicy distinguishes the authenticated proxy
from other Pods in that range. Do not retain direct API or tenant-workload
access to the native gateway port for this mode.
Gateway authentication
Kubernetes Compute supports trusted-proxy gateway authentication only, for
embedded and dedicated Agents, with or without private routing. At deployment,
it renders gateway.trustedProxies from network.gatewayTrustedProxyCidrs,
gateway.auth.mode: trusted-proxy, userHeader: x-occ-identity, the allowed
identity occ-workspace-files with operator.admin, and
gateway.allowRealIpFallback: true. Agent Configuration and Console starters
can omit those fields. Unsupported gateway authentication fields or conflicting
tenant trust fields fail deployment; matching explicit CIDR lists are accepted
regardless of order. trustedProxy.allowLoopback must be omitted or false:
loopback access uses the separate password, not proxy identity headers. Native
required-header and device auto-approval settings retain their separate purposes.
An optional loopback password
supports operator verification; it does not change the gateway's authentication mode.
Readiness uses a Pod-local HTTP request to
127.0.0.1:$OPENCLAW_GATEWAY_PORT/readyz; TLS terminates at Envoy, so native
readiness probes remain unchanged. Docker and SSH default to managed password
authentication and also support explicit trusted proxy.
Operators must verify that the configured CIDRs contain the proxy's actual source addresses and exclude untrusted sources. CIDRs do not authenticate a proxy: retain the exact Envoy NetworkPolicy peer, TLS verification, service-key authentication, and identity/header sanitization. Direct embedded access still requires a trusted proxy or the optional operator loopback password.
For repository-bearing revisions, Compute grants credential-service egress to
the embedded gateway/Harness or dedicated Codex Pod. The separate dedicated
gateway receives no repository egress rule. With repositoryCredentials.enabled,
Helm admits TCP/8443 ingress to the worker's credential sidecar from managed
gateway Pods carrying an Agent label, and managed dedicated Agent Pods carrying
both Agent and revision labels. Each peer also requires the tenant namespace
label. These selectors permit transport; the credential service still validates
the session and repository grant. Verify the effective policies in the installed
cluster; rendered rules alone do not prove traffic enforcement.
Compute projects repository broker policy to the actual Codex consumer: the
Agent Pod for dedicated Codex, or the gateway for embedded OpenClaw with
plugins.driver.implementation: occ/codex-plugin. Embedded OpenClaw using
occ/openclaw-plugin or no PluginDriver selection receives no Codex projection.
Selected Codex plugins receive a filesystem-only profile that grants read-only
access to the stock runtime package at /app/node_modules/openclaw and
published plugin skills at /home/node/.openclaw/plugin-skills and
/home/node/openclaw-runtime-assets/plugin-skills. This lets sandboxed skill
reads use the installed runtime and packaged skills without enabling proxy
networking, granting repository credential paths, granting whole-filesystem
reads, or changing project write permissions.
For Codex consumers with repository bindings, Compute also adds the exact broker
hostname from admitted session material to the tool proxy's domain allowlist and
sets stock Codex allow_local_binding = true and mode = "full". An explicit
deny matching the broker hostname fails closed. The repository-bound filesystem
profile additionally grants read-only access to the repository client at
/opt/oce/repository-credentials and admitted session material at
/run/oce/repository-credentials. Unbound Agents receive none of those network
or repository-material changes; their existing policy remains in effect.
These settings apply to the Agent's whole tool proxy: local binding is allowed, Codex's additional private-address guard is disabled, and every HTTP method is allowed at otherwise allowed destinations. Domain rules match hostnames, not ports: an allowed host is reachable on any port permitted by the lower network layers. This is not a broker-only port or method exception. Managed requirements that forbid local binding or require limited mode reject the conflicting configuration. Domain allowlisting, explicit denies, Kubernetes NetworkPolicy, TLS verification, and broker session/repository authorization remain separate boundaries. The workspace sandbox remains enabled.
Compute supplies the broker's public CA to that consumer before Codex starts. Stock full mode
normally tunnels HTTPS, so Git verifies the broker certificate directly. If
Codex separately requires HTTPS interception, it retains platform and startup
roots upstream and supplies child tools with its managed CA bundle. Preserve
inherited GIT_SSL_CAINFO; TLS verification remains enabled in both paths.
Production currently permits public TCP/443 egress for model access; a
restricted model proxy is not yet available. Before readiness, each dedicated
revision receives its own authentication-only egress policy. Concurrent pending
candidates cannot replace each other's grant; stop and retirement remove the
exact revision's policy after its Harness terminates. Channels require an
approved HTTP(S) proxy in runtime.channels: a literal IP endpoint, or the exact
Helm-managed proxy Service URL paired with runtime.channels.managedProxy.
Direct public channel-provider access is denied.
Explicit network profiles
Ordinary DNS, model, repository-credential, authentication, channel, workspace-node,
plugin-status, sandbox-preview ingress and gateway/Harness allow policies require the reserved Pod label
openclaw.dev/network-profile=broad-egress-v1, together with their existing
role, Agent, namespace and revision selectors. Gateway/Harness peer selectors require the
same profile. Missing, empty or unknown profiles receive no ordinary grant;
the tenant and separate Gateway namespace default-deny policies still select every Pod.
Compute assigns this profile when creating ordinary embedded and dedicated workload templates. Their existing routes and ports remain unchanged. Deployment readiness requires the expected template profile.
Harness Pods provisioned by a SandboxDriver, such as OpenShell, carry
provider-fenced-v1 instead. The provider fences their egress, so Compute grants
them only Gateway transport and plugin-status ingress (allow-agent-runtime is
ingress-only for them): no DNS, workspace-node, model or authentication egress.
Provider Harness readiness and activation reject a Pod with any other profile.
Existing policy names remain stable, and the upgrade restarts no Pod. New
namespaces receive the narrowed allow-dns, allow-gateway-ingress and
allow-node-gateway. Earlier namespaces keep their previous versions, which
ignore the profile, until recreated: Compute never narrows them in place.
Running Pods keep their templates until Compute next prepares a revision of
their Agent:
- Preparing a revision re-renders that Agent's grants and templates with the profile; other Agents are untouched. Re-preparing an active revision (as repository-credential maintenance does) rolls its Pods once.
- An unprofiled embedded or dedicated Gateway keeps serving, with its Gateway grants, during the next preparation, which exempts that predecessor from the profile check and leaves the stable Agent Service alone. Activation replaces it.
Existing OpenShell Sandboxes are not relabeled because Sandbox names are per revision: redeploy the Agent revision. For development, follow the development recovery procedure.
The separately installed OpenShell gateway needs its own scoped DNS/API and
callback policies. Its caller is the OpenShell supervisor Pod
(openshell.ai/managed-by=openshell, openshell.ai/boundary-role=supervisor),
which carries no openclaw.dev labels; supervisor-labelled peers admit it, not
the ordinary profile. Platform services retain their existing Helm policy selectors.
Profile assignment is a trusted controller decision. The label qualifies a Pod for network grants; it does not supply workload identity or authorization to request those grants. Operators must control workload creation, profile-label mutation and NetworkPolicy writes. This component does not install admission controls for those privileges.
Kubernetes combines grants from every matching policy, so stale or additional allow policies can bypass this restriction. Inspect installed policies and verify allowed and denied connections on a cluster with NetworkPolicy enforcement.
Private Agent gateway routes
See gateway routing with Envoy for shared infrastructure, service-key bootstrap, TLS, and network enforcement.
Runtime-enabled dedicated Harnesses require private routing and node enrollment before Compute can prepare or activate them. Missing wiring raises a configuration error before changing workloads; there is no Gateway-local workspace fallback. Embedded Harnesses can still use direct access.
Installation Compute settings enable stable Agent routes:
gatewayRouting: gatewayName: oce-agent-gateways gatewayNamespace: openclaw-system envoyNamespace: envoy-gateway-systemThe Gateway name and namespace must match the Helm-managed Gateway;
envoyNamespace identifies its Envoy data-plane Pods. The chart always creates
the Gateway in its release namespace. These three settings are required when
routing is enabled; hostname is optional. envoyHttpsTargetPort defaults to
10443 and must match Helm. Compute grants Harness egress only to this
installation's Envoy Pods on that port, before waiting for node enrollment.
endpointPort defaults to 443 and changes only the port in generated WSS URLs.
When set, the external load balancer must forward that port to the HTTPS listener.
When hostname is omitted or empty, Compute and Helm derive the same Service
name: occ-gateway- followed by the first 12 hexadecimal characters of the
SHA-256 of <gatewayNamespace>/<gatewayName>. The hostname is
<serviceName>.<envoyNamespace>.svc. It uses standard Linux Pod DNS search and
does not assume a cluster.local suffix. Set the same explicit hostname in
Compute and Helm for custom DNS or clients outside that cluster DNS context.
The default needs no existing Agent or Kubernetes lookup.
The operator installs Envoy Gateway and cert-manager and configures the
private gateway infrastructure.
Do not put an Agent endpoint, service key, certificate, or file contents into
native Configuration or an AgentRevision.
getGatewayEndpoint derives
wss://<hostname>/namespaces/<namespaceId>/agents/<agentId> without Kubernetes
API access. During preparation and activation, Compute reconciles an owned
HTTPRoute in the Gateway's physical namespace (control plane for dedicated,
data plane for embedded), attached to the configured Gateway's
https listener. Both rules match the configured private hostname and target
the existing same-namespace gateway Service:
- The exact Agent path rewrites to
/, preserving workspace-file WSS access. - A prefix rule below that Agent path rewrites the prefix to
/and retains the suffix for native UI assets, deep links, and WebSocket paths.
OCC bounds proxy requests to the selected Agent base. Public native UI browser
traffic enters through OCC; Envoy and gateway Services remain private. See
Agent native admin UI.
Namespaces receive the Gateway membership label used by allowedRoutes.
Runtime-enabled dedicated revisions also receive a /node route and a
route-specific SecurityPolicy for native device authentication. The
routing reference owns its
credential boundary and the remaining Harness lifecycle requirements.
The Service and route remain stable across revision cutover. Retiring an old revision preserves a newer gateway's route; final gateway cleanup removes the owned route. Reconciliation runs through the existing revision lifecycle; this Driver does not add periodic route drift repair. Missing CRDs or denied worker permissions fail reconciliation rather than disabling routing silently.
Envoy's Gateway-level SecurityPolicy authenticates the OCC service key before
forwarding. The route overwrites the native identity and real-IP headers and
removes caller forwarding and scope headers. Native allowRealIpFallback
accepts Envoy's direct downstream connection address when OCC and Envoy share a
Pod CIDR. That source address must be nonloopback; a loopback port-forward alone
is not a working native attribution path.
Namespaces and isolation
Each OpenClaw Namespace has a data-plane Kubernetes namespace and a managed
Gateway runtime namespace, oce-gateways-<hash>, where hash is the first 24
hexadecimal characters of sha256(namespaceId). The latter is discovered by
openclaw.dev/gateway-namespace=<namespaceId>; it deliberately omits the
data-plane discovery label openclaw.dev/namespace. Existing data-plane
namespace adoption does not adopt or reuse OCC's own namespace for Gateways.
Compute prepares restricted Pod security, quotas, defaults, default-deny and DNS
policies in both targets. Dedicated Gateway resources, private PVCs, configuration,
Services and HTTPRoutes live only in the Gateway target; Harness resources and
model credentials remain in the data target. Explicit namespace and Pod
selectors allow only the same Agent's selected Harness revision on app-server
and private plugin-status ports. DNS uses agent-<hash>.<harness-namespace>.svc.
The stable dedicated Harness Service keeps the same Namespace, Agent, revision,
and workload-role labels as the gateway egress and Harness ingress policies
while a revision is active. A prepared successor does not change that Service
selector until activation; deactivation moves the Service back to an inactive
selector. Active Gateway Services include Namespace, Agent, and gateway-role
labels, satisfying gateway policy selectors without tying the stable route to a
revision. These Service selectors support the
AWS VPC CNI pre-DNAT policy resolution requirement.
Current app-server transport is capability-token ws://, not mTLS; this change
does not implement cross-cluster transport or runtime attestation.
Stop and revision retirement inspect both targets and retain durable claims. Agent deletion removes its owned claims; Namespace deletion deletes only the exact managed Gateway namespace and preserves an adopted data namespace. Neither operation may remove shared OCC infrastructure. A missing or foreign Gateway target fails preparation rather than falling back to data-plane placement.
Identity labels under openclaw.dev/ contain the full platform Namespace,
Agent, revision, ServiceAccount, ServicePrincipal, or Configuration ID, not a
hash. Ownership checks, discovery, Service selectors, and NetworkPolicies use
those same raw IDs. Generated Kubernetes resource names still use bounded
hashes to satisfy their naming constraints.
An Installation administrator can select an existing, exclusively dedicated Kubernetes namespace when creating the OpenClaw Namespace:
{ "name": "customer-support", "existingNamespace": "customer-support-prod"}Prepare the namespace by annotating
openclaw.dev/namespace-lifecycle=external, applying
pod-security.kubernetes.io/enforce=restricted,
pod-security.kubernetes.io/audit=restricted, and
pod-security.kubernetes.io/warn=restricted, and granting tenant-local worker
and API RoleBindings. The running worker rechecks Installation administrator
authorization, rejects foreign NetworkPolicies and competing tenant claims, and
binds the generated tenant identity through one resource-version-guarded,
non-forced Kubernetes patch. No worker pause or restart is required. Missing
worker permissions keep provisioning pending; missing API permissions prevent
Configuration access. Docker and external Compute Drivers reject
existing-namespace selection with 409.
See tenant RoleBindings for the required worker and API grants.
Workload Pods run as nonroot, use RuntimeDefault seccomp by default, drop
Linux capabilities, disable privilege escalation, and use read-only root
filesystems. When runtime.codexSeccompProfile is configured, only the
dedicated Codex Agent container uses
seccompProfile: { type: "Localhost", localhostProfile: <profile> }; the Pod,
gateway container, embedded runtime, controller, and init containers keep their
default seccomp settings. The profile path must be relative to the kubelet's
localhost seccomp profile root and cannot be empty, absolute, traversing, or
unconfined. Agent identity is provided through an audience-scoped, short-lived
projected ServiceAccount token. Workloads never receive controller credentials.
The driver deletes Kubernetes namespaces it created when their corresponding OpenClaw Namespaces are deleted. For an operator-owned existing namespace, it removes only its exact-owned quota, limit, and three tenant NetworkPolicies; the Kubernetes namespace, ownership markers, RoleBindings, and unrelated resources remain intact.
