OpenClaw EnterpriseDOCSGitHub

OCC metrics

The production Helm chart enables metrics on each OpenClaw Control Plane (OCC) API and worker by default; direct process and development defaults remain off. Each enabled process serves GET /metrics on a private listener. Use the development walkthrough for Prometheus and Grafana, or production scraping for Kubernetes.

Configuration and access

Setting Contract
OCC_METRICS_ENABLED true or false; defaults to false.
OCC_METRICS_HOST Required when enabled: 127.0.0.1 or ::1 in development; explicit Pod IP in production.
OCC_METRICS_PORT Required when enabled: integer 1–65535, distinct from the API port.

Supplying host or port while disabled fails startup. An invalid configuration or failed bind fails startup. Listener shutdown follows process shutdown. Metrics are independent of log level and the logging Collector's port 8888.

The listener serves no sessions or resource operations and does not invoke IAM. Network isolation protects the aggregate operational data. Exposition uses Prometheus text, Cache-Control: no-store, and the client content type. Other paths return 404; other methods on /metrics return 405. Connections are limited to 16, with five-second socket/request handling limits. No public ingress.

Application families

All families have service="api|worker". No tenant, Agent, request, revision, actor, credential, Driver-instance, payload, raw URL, or user-defined labels.

Family Type Additional labels Meaning
occ_http_requests_total Counter route, method, status_class Completed API-process responses, including auth, console, denied and failing requests.
occ_http_request_duration_seconds Histogram route, method Request-hook to response-completion seconds.
occ_reconciliation_attempts_total Counter work_kind, outcome Finished passes processing claimed work.
occ_reconciliation_attempt_duration_seconds Histogram work_kind Processing seconds, including Driver calls and finalization.
occ_work_pending Gauge None Shared PostgreSQL count of queued/claimed work, including delayed retries.
occ_agents Gauge lifecycle_state Persisted Agent reconciliation lifecycle; see below.
occ_agent_operation_duration_seconds Histogram operation Admission to successful deployment or stop completion, including queue wait and retries.
occ_work_oldest_pending_age_seconds Gauge None Age since admission of the oldest queued/claimed item; zero when no work is pending.

HTTP excludes health probes, metrics scrapes, and disconnected requests without a completed response. route is the registered template or unmatched; wildcard routes stay templates. Methods are GET|HEAD|POST|PUT|PATCH|DELETE|OPTIONS|OTHER; status classes are 1xx|2xx|3xx|4xx|5xx|other. Status classes cannot distinguish 401/403 from other 4xx failures. These are not model-turn or WebSocket durations.

Work kinds are namespace_ensure|namespace_delete|agent_revision|agent_stop|agent_delete. Outcomes are success|pending|retry|permanent|claim_lost|error. Pending convergence, retries, maintenance, and superseded work may generate multiple passes per deployment. Durations exclude queue wait and time between passes. Idle polling and stale recovery without a claimed processing pass are excluded. Process death can lose observations; these counters are not durable audit evidence.

HTTP histogram boundaries in seconds: 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10. Work boundaries: 0.01, 0.1, 0.5, 1, 5, 15, 30, 60, 120, 300, 900. Both add +Inf, count, and sum.

Shared snapshots and outages

Each worker scrape collects all gauges in one read-only PostgreSQL statement against the singleton Installation. Every lifecycle category exists even at zero; each Agent counts once. The projection uses desired runtime state, the latest admitted revision, and deployment/stop/delete work:

State Meaning
draft No revision, stop, or deletion request has been admitted.
deploying Running is desired and the latest deployment is pending or not selected. Includes redeployment while an older revision remains active.
running The selected latest deployment has completed. This does not prove continuous runtime health.
stopping Deletion is pending, or stopped is desired with an active pointer or unfinished deployment/stop work.
stopped Stop has converged; revision history and Agent records remain.
failed The latest operation for the desired state failed permanently. Includes retained Agents whose deletion cleanup failed.

Successful deletion removes the Agent from inventory.

Maintenance work does not change these lifecycle categories. Collection never polls Compute or updates resource state. Queue depth and oldest age include delayed retries and scheduled maintenance, even before their next eligible time. Age uses the original work admission timestamp, not the latest attempt, and is clamped at zero for clock skew.

Operation durations have operation="deploy|stop". The completing worker records one observation after successful queue finalization, including final activation and predecessor retirement for deployments. Retry and convergence delays count; maintenance, superseded operations, and permanent failures do not. Repeated stops that find the Agent already stopped still count as completed stop requests. Both bounded operation label sets start at zero so the first completion can contribute to rates. Buckets in seconds are 0.1, 0.5, 1, 2, 5, 10, 15, 20, 30, 45, 60, 90, 120, 180, 240, 300, 450, 600, 900, 1800 plus +Inf, count, and sum, resolving deployments from one second to the 900-second convergence deadline. Per-phase deployment timing is in the worker's worker.completed log. The timer uses wall-clock admission time and clamps negative elapsed time to zero. Process death between commit and observation can lose a sample; observations are operational metrics, not durable audit evidence.

Worker startup supplies a separate read-only application-role pool with maximum one connection, a 500 ms connection deadline, and 1500 ms query/statement deadlines. Concurrent scrapes share collection. Failed collections destroy the affected connection and return 503 without partial exposition or a stale-value fallback. Reconciliation uses its own pool and continues. API scrapes do not query the database.

Process families and series budget

The pinned @prometheus-io/client@0.16.1 supplies defaults. OCC retains only:

Unsupported platform measurements may be absent. No handle/resource breakdowns, heap-space labels, custom labels, configurable buckets, or exemplars. Client upgrades must explicitly review this list; new defaults are filtered out.

For R observed registered route/method pairs, HTTP has at most 20R series (six statuses plus fourteen histogram series). Unmatched methods add at most eight pairs. Worker application metrics have at most 136 series (30 outcomes, 70 pass-duration series, 28 operation-duration series, and eight gauges). Process collectors add at most 53 series per process. Do not preallocate the route/status Cartesian product.

Replica aggregation

Prometheus scrapes every Pod directly and assigns instance/deployment labels. Sum rate() of counters; sum histogram buckets before histogram_quantile(). Use max for duplicated shared gauges, filtering targets by up == 1 and retaining deployment identity. Independent snapshot times can temporarily overestimate decreasing values; missing/failed targets mean unknown, not zero. Memory can be summed for fleet use or maximized for the largest process. Inspect event-loop percentiles per process; do not average them.

See production queries and discovery. The metrics testing page owns fixtures and proof limits.

Search documentation