OpenClaw EnterpriseDOCSGitHub

GitHub Actions testing flow

Overview

GitHub Actions selects explicit test lanes, prepares disposable resources, runs the real Node test runner, and rejects missing or skipped required coverage. This flow ends at the aggregate check and resource cleanup. A PR check proves its ten selected noncredentialed lanes; it does not establish that protected model or service integrations passed.

Entry Points

Flow

graph TD
  subgraph Actions["GitHub Actions"]
    A["PR or main event"] --> B["Ten PR-safe jobs"]
    A --> N["Suite audit"]
    N --> L
    C["Manual integration dispatch"] --> D["Environment protection preflight"]
    D -->|main-only provider or approved other lane| E["Protected jobs"]
    D -->|missing protection| X["Failed check"]
  end
  subgraph Runner["Disposable job runner"]
    B --> F["Prepare lane resources"]
    E --> F
    F -->|prepared| G["Prepare file prerequisites"]
    G --> H["Node tests and structured reporter"]
    H --> I["Case and skip validation"]
    F -->|fixture cluster startup fails| P["Save bounded setup diagnostics"]
    P --> J["Owned-resource cleanup"]
    F -->|other preparation fails| J
    I --> J
  end
  subgraph Results["Check results"]
    I --> K["Sanitized lane result"]
    J --> L["Aggregate expected jobs and results"]
    K --> L
    L --> M["Pass or fail for named coverage"]
  end

Execution Trace

1. Select one source revision and coverage group

.github/workflows/ci.yml:jobs, .github/workflows/full-integration.yml:jobs, scripts/ci/full-integration-preflight.mjs:validateFullIntegrationPreflight, and scripts/ci/test-suites.mjs:loadTestSuites

The suite index, scripts/ci/test-suites.json, holds ordered lane references and coverage groups. loadTestSuites loads each referenced scripts/ci/test-suites/<lane>.json into the shared suite map. Each lane file owns its test inventory, environment, required inputs, and preparation settings. The runner and preparation tools consume the assembled map.

The PR workflow uses the event checkout and supplies no external service credentials. Suite Audit and all ten lanes start independently on ephemeral runners. Kubernetes fixture lanes use ubuntu-22.04 for bridge netfilter support; other lanes and the audit use blacksmith-8vcpu-ubuntu-2404. Its aggregate uses ubuntu-22.04 and requires a successful audit plus checks-baseline, postgres, postgres-application, images-packaging, k3d-fixture-configuration, k3d-fixture-state, k3d-fixture-plugins, logging-collector, repository-credentials-container, and repository-credentials-platform. A failed audit still fails CI Required even when the lanes pass. Full Integration checks configured environment protection and checks out the immutable event SHA. It admits refs/heads/main for every lane. Only k3d-model may use another branch: preflight requires an exact branch rule in integration-model, and GitHub still requires reviewer approval with self-review prevention. Wildcards, tags, and other non-main lanes are rejected. The administrator removes the temporary branch rule after verification. A manual dispatch selects its requested lane or all; pushes and merges do not start this workflow. Manual runs share one concurrency group and do not cancel an in-progress run. The provider environment must allow exactly the main branch and needs no per-run reviewer approval. Other credentialed environments still require reviewers with self-review prevention. No PR event enters this credentialed workflow. A targeted integration run has a narrower claim than a full inventory run.

PostgreSQL migration and application suites own separate servers. Each of the three Kubernetes fixture files owns a separate cluster and PostgreSQL server. For these Kubernetes fixture lanes, the shared action enables bridge netfilter on the ephemeral runner before creating k3d nodes, which share its kernel. Missing bridge filtering fails setup rather than running with unenforced Pod network policies. The repository credential platform lane uses Blacksmith for its full-image HTTP, PostgreSQL, Unix-control and credential-material proof; NetworkPolicy enforcement remains the fixture lanes' separate responsibility. Lane state and cleanup stay local to its runner; files within each lane remain sequential. The suite map retains one owner per file in both workflow groups.

The native IAM barrier test receives its own migrated PostgreSQL database through the application lane's per-file preparer. Only that test receives the matching application and migrator connection details, and the CLI refuses to write those details to GITHUB_ENV. The test installs its unregistered supplier only in that disposable database. This fixture does not establish that production writers participate in the barrier.

Both workflows call the shared run-ci-lane action after checkout. It owns tool and dependency setup, baseline checks when selected, lane preparation, execution, unconditional cleanup, and sanitized result upload. Callers keep the source revision, timeout, protected environment and explicit credentials.

Ordinary PR dependency caches may be restored and saved within GitHub's PR merge-ref scope. Main jobs use main-scoped caches. Test results and credential-bearing state are not dependency caches, and protected jobs do not promote PR build artifacts.

The provider job selects the shared blacksmith-8vcpu-ubuntu-2404 runner for disk headroom during runtime image build and k3d import. The standard Ubuntu runner reached DiskPressure and evicted the seccomp probe before it could start. The repository must retain access to this organization runner label. Image preparation copies the saved archive into each owned k3d node and runs node-local ctr image import; k3d tools-node can log per-node import failures while returning success. The imported manifest and CRI checks remain required before any test starts.

2. Prepare resources under the job owner

scripts/ci/prepare.mjs:main and scripts/ci/prepare.mjs:ensureK3dCluster

CI resource preparation traces tool setup, image and cluster preparation, protected credentials, and resource ownership. Continue below when preparation has produced the lane state.

For the three Kubernetes fixture lanes, cluster startup records phase timings and host snapshots. On failure, bounded diagnostic reads save <state-file>.diagnostics.json outside the cluster directory before cleanup. Creation uses --no-rollback for these lanes so the workflow owns teardown after capture; local callers still invoke cleanup with their failed run's state file. Collection preserves the original error, including when an observation fails or times out. The CI guide describes the retained evidence.

Tests delete their Agent namespaces, and those namespaces' events, before a file exits. scripts/ci/run-tests.mjs:runFile therefore watches Compute-managed Pods and Kubernetes events in each ready k3d cluster while the file runs, then scripts/ci/k3d-diagnostics.mjs:projectAgentNamespaceActivity appends the Pod status transitions and those namespaces' events to the same report under agentNamespaces, passing or failing. Each file keeps at most 200 Pod and 200 event records and the report keeps 40 files; messages are redacted and truncated, Pod specs are dropped, and raw watch streams stay in the cluster directory that cleanup removes. The artifact is uploaded for every lane that writes it.

Dedicated Codex preparation and the operator's offline profile generator share scripts/lib/codex-seccomp-profile.mjs:deriveCodexBwrapProfile. Preparation requires an actual workspace write and denied write to a container-writable outside path before publishing the selected Localhost profile to the live suite. Native runtime-image tests trust a dynamic Codex Docker seccomp profile only when OPENCLAW_ENTERPRISE_CI_STATE records the exact prepared cluster.codexDockerSeccompProfile path and SHA. A self-hashed profile without that state is not CI proof; the standalone fallback remains the pinned reviewed manual profile. Production node provisioning remains outside CI ownership; see Codex sandbox setup.

3. Execute and account for actual cases

scripts/ci/run-tests.mjs:main and scripts/ci/reporter.mjs:jsonLinesReporter

The runner discovers active test files and verifies that the map assigns each file to exactly one lane. Tests with different prerequisites live in separate files. The runner invokes whole files with invocation-scoped environment inputs. A custom Node reporter exposes case names, locations and outcomes; arbitrary test output and credential-bearing error payloads are excluded from published results. Failed provider-test HTTP assertions also retain numeric actual and expected status codes, an allowlisted OCC error code, and the upstream ChatGPT operation and status when available. Denied-traffic failures retain only an allowlisted traffic category, without target addresses or response data. Plugin-status fixture failures retain an allowlisted readiness or rollout stage. Rollout diagnostics include bounded Pod phases, readiness and scheduling flags, container restart counts and exit codes, and allowlisted reasons. Response bodies, credentials, and identities remain excluded.

Required named cases must pass. Every skip or TODO fails the selected lane; there are no counterpart-skip lists or CI name filters. A synthetic file-wrapper success, missing result output, zero executed cases or an interrupted run without final reporter output cannot establish coverage. The runner retains failure, timeout and cleanup outcomes in the lane result.

4. Clean up and publish the bounded result

scripts/ci/cleanup.mjs:main and scripts/ci/run-tests.mjs:main

.github/actions/run-ci-lane/action.yml uploads one sanitized result artifact per lane and workflow run. A job retry replaces that lane's earlier artifact; other lanes retain their results. This prevents aggregation from selecting a stale failed result after a successful retry. The earlier job logs remain the failure record; retain a result separately before retrying when needed.

For images-packaging, scripts/ci/export-image-reconciliation.mjs attempts to retain attempt-specific cleanup records for the two controller and runtime tags prepared by the lane. A planned record does not prove an image was created. The run-and-attempt component of each tag name is metadata, not authentication or permission to delete an image. Missing state is reported as unavailable; neither that result nor an empty inventory proves cleanup. The separate tag created by the runtime-images test, other resource kinds, and private environment values are excluded. Export or upload failure and runner loss can prevent retention.

Fixture bootstrap failures also upload diagnostics-<artifact-prefix>-<lane> separately from test results. Cleanup removes the cluster and its private state; the diagnostic file remains available for upload and does not satisfy the aggregate's required test results.

Per-file cleanup releases its disposable database. Job cleanup removes only the state-owned resources. A whole owned k3d-cluster resource owns Kubernetes API object deletion for its Collector Namespace and RBAC. Logging cleanup cleans the local Docker backend container and JSONL/config directory independently, so a dead Kubernetes API does not block local log backend teardown. Cleanup failure fails the check and keeps the private state file usable only while that runner host and path remain available. User databases, contexts, unrelated containers and global images remain outside that ownership.

The aggregate runs after success or failure and checks expected job outcomes plus same-revision lane results. Case validation belongs to the runner; the aggregate checks lane identity and success, required evidence, and cleanup outcomes without interpreting cases again. Missing, failed, cancelled or skipped selected jobs cannot pass. A full-suite result accounts for every lane selected by the full group. The explicitly selected ssh-host lane remains outside the automatic groups until an operator prepares its disposable host; see SSH raw-host testing. Abrupt hosted-runner loss can prevent teardown and also loses the private RUNNER_TEMP state at job end. External resource reconciliation is deferred until an approved resource ledger exists.

Debugging and Verification

Search documentation