Configure platform observability
Send operational logs from OpenClaw Control Plane (OCC), managed gateways, and Codex workloads to your log backend, then verify that records arrive. The OpenTelemetry Collector also exposes metrics about its own delivery pipeline; it does not collect application metrics, traces, or audit records. Run commands from the repository root.
To give Installation administrators a shortcut to an observability UI, set
observability.url in the trusted startup YAML
and restart the API. The console opens that URL in a separate tab after an
Installation administer check. Configure authentication at the destination.
The link does not change Collector export.
For the demonstration Grafana stack, point the link to /d/occ-observability;
that landing page lists its metrics and operational logs views. The demo does
not provide traces.
| Signal | Available path |
|---|---|
| Operational logs | Local container output; optional OpenTelemetry Collector export over OTLP/HTTP to your log backend. |
| Collector metrics | Prometheus endpoint on port 8888 for the collection pipeline itself. |
| Audit records | Stored separately in PostgreSQL; the Collector does not export them. See Audit Log. |
| Application metrics | Private OCC Prometheus endpoints enabled by default in Helm; see production scraping and the development dashboard. |
| Distributed traces | No application tracing pipeline is installed. |
The default Helm installation needs no telemetry backend: operational logs go to
container output at info, and private metrics listeners run with no scraper
ingress grant. Connect your collectors using this guide and the metrics guide.
For disposable visualization, use the demonstration stack,
which is not recommended for production.
The Collector exports predefined operational events and approved fields. It
excludes arbitrary messages, prompts, responses, and Codex protocol output, even
at debug. See the
security boundary.
If your Compute Driver declares deployment-managed runtime logging, follow its runtime platform's collection and verification procedures. Continue using this guide for OCC logs; the Driver does not change how OCC records audits.
Requirements
- A working development stack, or the protected YAML inputs and namespace from production setup.
- An operator-owned OTLP/HTTP Logs receiver: full
/v1/logsendpoint, authentication, and trusted TLS chain. The Collector includes no storage or viewer. - For Docker: Docker Compose and an endpoint reachable from the Collector container. The Docker Engine must also reach the Fluent Forward receiver.
- For Kubernetes: Helm,
kubectl,yqv4, an explicit kubeconfig/context, enforcing NetworkPolicies, and permission to install the Collector DaemonSet, its Pod-read ClusterRole, and dedicated Secrets.
Steps
1. Choose the log level
Set the shared level in the startup YAML:
logging: level: infoUse debug, info, warn, or error. The default is info. For the
Docker logging override, edit deploy/logging/occ.yaml.
For production, edit the protected Installation YAML and update its mounted
startup Secret through your deployment process. This is separate from Helm's
logging.collector values.
Restart the API and worker after a level change. With the Docker logging
override, migration and bootstrap read the YAML on their next execution. The
current Helm initialization Job does not mount that YAML or set OCC_CONFIG_PATH,
so its migration and bootstrap processes use info.
Existing AgentRevisions keep their original level. Deploy an Agent again to apply the new level to its gateway or Codex runtime. See the startup settings for the accepted configuration.
2. Configure the exporter
Use deploy/logging/exporter.yaml as the
native Collector exporter configuration. It reads OTEL_EXPORTER_OTLP_LOGS_ENDPOINT
and limits the queue size and retry window. Use HTTPS with verified server
identity for real backends; use plain HTTP only for a local test receiver.
For authentication or a custom CA, prepare a protected exporter.yaml with native
Collector header/TLS settings. Keep credentials in Collector-only Secrets or
protected mounts. Explicitly supply referenced environment variables: Docker
forwards only the endpoint; Helm loads its dedicated exporter environment Secret.
Additional mounts require deployment changes: Helm projects only three named
configuration files and offers no extra-mount value. Never reference an unmounted CA.
Keep the exporter named otlp_http, or update the receiver pipeline's exporter
reference too. Preserve the shared filtering policy and bounded queue/retry
settings. Do not put exporter credentials in Installation YAML, Agent
Configurations, SecretBindings, lifecycle hooks, or runtime images.
3. Enable collection for your deployment
Docker Compose
Start your receiver. For a receiver reachable through Docker's host alias:
export OTEL_EXPORTER_OTLP_LOGS_ENDPOINT='http://host.docker.internal:4318/v1/logs'./scripts/dev-up -- -f compose.yaml -f compose.logging.yamlUse a same-network DNS name if the receiver is another container. If you prepared
a custom exporter file, add a private Compose override mounting it at
/etc/otel/exporter.yaml in the collector service and pass that override last.
The override routes OCC and newly created managed runtime containers through
Docker's nonblocking fluentd driver. It publishes Fluent Forward on loopback
port 24224 and Collector metrics on loopback port 8888. Redeploy existing
Agents to recreate their containers with logging enabled. When Docker runs in a
VM, verify that the Engine can reach the receiver; resolving the container's
DNS name alone does not prove this.
If you change the Fluent Forward port, set both OTEL_COLLECTOR_PORT and
OCC_DOCKER_LOGGING_ADDRESS to matching values. OTEL_COLLECTOR_METRICS_PORT
changes only the host metrics port. See Docker settings
for defaults and environment precedence.
Kubernetes and Helm
Use production setup's KUBECONFIG_FILE, CONTEXT, OCC_INPUT_DIRECTORY, and
existing openclaw-system namespace. Retain the complete values.yaml; the
logging block alone cannot install OCC.
If an existing cluster Collector already reads the OCC and tenant CRI files, use it only after applying the same native receiver and shared policy: trusted metadata, privacy, filtering, egress, and bounded state. Use one collection route per stream. Otherwise, enable the bundled Collector below.
Create its two dedicated Secrets before installing or upgrading the chart:
kubectl --kubeconfig "$KUBECONFIG_FILE" --context "$CONTEXT" -n openclaw-system \ create secret generic occ-otel-collector-config \ --from-file=collector.yaml=deploy/logging/collector.yaml \ --from-file=kubernetes.yaml=deploy/logging/kubernetes.yaml \ --from-file=exporter.yaml=deploy/logging/exporter.yamlkubectl --kubeconfig "$KUBECONFIG_FILE" --context "$CONTEXT" -n openclaw-system \ create secret generic occ-otel-collector-exporter \ --from-literal=OTEL_EXPORTER_OTLP_LOGS_ENDPOINT='https://otel.example.internal/v1/logs'Replace the endpoint. For authenticated export, select your protected
--from-file=exporter.yaml=... and add referenced credentials from protected
files to the exporter Secret. Update existing Secrets through your normal
Secret-management workflow.
Set the exact approved exporter or proxy IPv4 address and port in the protected
values copy; 203.0.113.10/32 below is a placeholder:
yq -i '.logging.collector.enabled = true | .logging.collector.exporter.cidr = "203.0.113.10/32" | .logging.collector.exporter.port = 443' \ "$OCC_INPUT_DIRECTORY/values.yaml"Keep a digest-pinned approved Collector image. For custom Secret names, set
logging.collector.configSecretName and logging.collector.envSecretName.
Do not reuse application Secrets. See Helm settings
for resource and storage limits.
Before first install, finish bootstrap PVC preparation.
For existing installations, apply reviewed values through the same Helm upgrade. The Collector reads
/var/log/pods read-only and needs get/list/watch on Pods across workload
namespaces. Namespace and node names come from Pod fields; it does not need
Namespace, Node, Secret, or pods/log API access. Its node filter limits queries,
but is not an RBAC security boundary.
Restart the Collector after changing either Secret to load its configuration:
kubectl --kubeconfig "$KUBECONFIG_FILE" --context "$CONTEXT" -n openclaw-system \ rollout restart daemonset/openclaw-enterprise-collectorTests
Check delivery to the backend
- With
logging.level: infoordebug, run the authenticated API check for your deployment: Docker development or Kubernetes production. - Find a new
service.name=occ-api,event.name=http.completedrecord in your backend. Confirm its status, timestamp, and request ID match the request. - For runtime coverage, deploy an Agent and exercise its gateway or Codex
app-server. Check the corresponding
openclaw-gatewayorcodex-app-serverrecords andopenclaw.agent.id/openclaw.revision.idresource attributes.
A healthy Collector and local container output do not prove that the backend received the records. Only approved runtime events appear. Receiving API logs does not verify model turns or other integrations.
Check Collector metrics
For Docker, inspect Collector errors and its Prometheus endpoint:
docker compose -f compose.yaml -f compose.logging.yaml logs --tail=100 collectorcurl --fail "http://127.0.0.1:${OTEL_COLLECTOR_METRICS_PORT:-8888}/metrics"For Kubernetes, check rollout and errors, then forward one Collector Pod's metrics port to your machine (this checks only that node):
kubectl --kubeconfig "$KUBECONFIG_FILE" --context "$CONTEXT" -n openclaw-system \ rollout status daemonset/openclaw-enterprise-collector --timeout=120skubectl --kubeconfig "$KUBECONFIG_FILE" --context "$CONTEXT" -n openclaw-system \ logs -l app.kubernetes.io/component=collector --tail=100COLLECTOR_POD=$(kubectl --kubeconfig "$KUBECONFIG_FILE" --context "$CONTEXT" \ -n openclaw-system get pods -l app.kubernetes.io/component=collector \ -o jsonpath='{.items[0].metadata.name}')kubectl --kubeconfig "$KUBECONFIG_FILE" --context "$CONTEXT" -n openclaw-system \ port-forward "pod/$COLLECTOR_POD" 8888:8888In another terminal, run curl --fail http://127.0.0.1:8888/metrics. Inspect
whether the receiver accepts or refuses records, whether exports succeed,
queue usage, and process memory. otelcol_exporter_queue_size should not grow indefinitely;
compare it with otelcol_exporter_queue_capacity. An increasing
otelcol_processor_filter_logs_filtered can reflect expected privacy filtering.
Privately scrape every Collector instance. The chart provides no dashboards, metrics Service, or scrape discovery; add a narrowly scoped ingress allow rule for your scraper.
Production readiness
Use production handoff to assign alert recipients and response procedures alongside these collection checks.
- Keep runtime native OTLP export disabled and preserve Collector filtering. Local container logs and remotely exported records have different privacy boundaries; restrict access to both.
- Keep exporter traffic within the approved
/32and port, with DNS and Kubernetes API access configured by the chart. Use an approved fixed proxy when your backend cannot be represented by that egress policy. NetworkPolicies are additive: the current shared dependency policy also permits Collector traffic to the configured database destination; the dedicated Collector policy does not remove that access. - Alert on failed exports, refused records, queue saturation, and Collector restarts. Verify retention and access controls in your selected backend.
- Treat delivery as best-effort. Docker keeps exporter queues in the
occ_otelcol_datavolume and bounded runtime log caches; its push-based Fluent Forward receiver has no file offsets. Kubernetes keeps file offsets and exporter queues in/var/lib/otelcolon boundedemptyDirstorage, which survives container restart but is lost on Pod or node replacement. An outage can lose operational logs without blocking OCC work. Audit records are stored separately in PostgreSQL.
Troubleshooting
| Symptom | Check and recovery |
|---|---|
| OCC will not start after changing the level | Use only the supported logging.level values; unknown logging keys fail startup. |
| Runtime level did not change | Restart OCC with the new startup YAML, then deploy a new AgentRevision. |
| No Docker records reach the Collector | Check the Engine-reachable Fluent Forward address and matching published port; recreate existing runtime containers through Agent deployment. |
| Kubernetes Collector is pending or cannot read files | Check image pull access, Secret names, Pod security admission, and node CRI file permissions. |
| Metadata is missing or Kubernetes requests are forbidden | Check the Collector ServiceAccount binding and Pod-read ClusterRole; retain the shipped Pod association and extraction rules. |
| Collector receives records but the backend does not | Check endpoint path, TLS trust, credentials, exporter errors, allowed egress address/port, and whether records pass the operational filter. |
Expected records disappear at warn or error |
The source level suppresses lower-severity events; use info for the delivery check. |
