GitHub Actions testing flow
Overview
GitHub Actions selects explicit test lanes, prepares disposable resources, runs the real Node test runner, and rejects missing or skipped required coverage. This flow ends at the aggregate check and resource cleanup. A PR check proves its ten selected noncredentialed lanes; it does not establish that protected model or service integrations passed.
Entry Points
.github/workflows/ci.yml:jobs: PR, main push, merge-group and manual checks on ephemeral runners..github/workflows/full-integration.yml:jobs: manual integration from main or an explicitly approved Kubernetes model branch, bound to the dispatched commit.scripts/ci/run-tests.mjs:main: local or workflowaudit,runandaggregatecommands; the suite map is the coverage owner.
Flow
graph TD
subgraph Actions["GitHub Actions"]
A["PR or main event"] --> B["Ten PR-safe jobs"]
A --> N["Suite audit"]
N --> L
C["Manual integration dispatch"] --> D["Environment protection preflight"]
D -->|main-only provider or approved other lane| E["Protected jobs"]
D -->|missing protection| X["Failed check"]
end
subgraph Runner["Disposable job runner"]
B --> F["Prepare lane resources"]
E --> F
F -->|prepared| G["Prepare file prerequisites"]
G --> H["Node tests and structured reporter"]
H --> I["Case and skip validation"]
F -->|fixture cluster startup fails| P["Save bounded setup diagnostics"]
P --> J["Owned-resource cleanup"]
F -->|other preparation fails| J
I --> J
end
subgraph Results["Check results"]
I --> K["Sanitized lane result"]
J --> L["Aggregate expected jobs and results"]
K --> L
L --> M["Pass or fail for named coverage"]
endExecution Trace
1. Select one source revision and coverage group
.github/workflows/ci.yml:jobs, .github/workflows/full-integration.yml:jobs,
scripts/ci/full-integration-preflight.mjs:validateFullIntegrationPreflight, and
scripts/ci/test-suites.mjs:loadTestSuites
The suite index, scripts/ci/test-suites.json, holds ordered lane references and
coverage groups. loadTestSuites loads each referenced
scripts/ci/test-suites/<lane>.json into the shared suite map. Each lane file owns
its test inventory, environment, required inputs, and preparation settings. The
runner and preparation tools consume the assembled map.
The PR workflow uses the event checkout and supplies no external service credentials. Suite Audit and all ten lanes start independently on ephemeral runners. Kubernetes fixture lanes use ubuntu-22.04 for bridge netfilter support; other lanes and the audit use blacksmith-8vcpu-ubuntu-2404. Its aggregate uses ubuntu-22.04 and requires a successful audit plus checks-baseline, postgres, postgres-application, images-packaging, k3d-fixture-configuration, k3d-fixture-state, k3d-fixture-plugins, logging-collector, repository-credentials-container, and repository-credentials-platform. A failed audit still fails CI Required even when the lanes pass. Full Integration checks configured environment protection and checks out the immutable event SHA. It admits refs/heads/main for every lane. Only k3d-model may use another branch: preflight requires an exact branch rule in integration-model, and GitHub still requires reviewer approval with self-review prevention. Wildcards, tags, and other non-main lanes are rejected. The administrator removes the temporary branch rule after verification. A manual dispatch selects its requested lane or all; pushes and merges do not start this workflow. Manual runs share one concurrency group and do not cancel an in-progress run. The provider environment must allow exactly the main branch and needs no per-run reviewer approval. Other credentialed environments still require reviewers with self-review prevention. No PR event enters this credentialed workflow. A targeted integration run has a narrower claim than a full inventory run.
PostgreSQL migration and application suites own separate servers. Each of the three Kubernetes fixture files owns a separate cluster and PostgreSQL server. For these Kubernetes fixture lanes, the shared action enables bridge netfilter on the ephemeral runner before creating k3d nodes, which share its kernel. Missing bridge filtering fails setup rather than running with unenforced Pod network policies. The repository credential platform lane uses Blacksmith for its full-image HTTP, PostgreSQL, Unix-control and credential-material proof; NetworkPolicy enforcement remains the fixture lanes' separate responsibility. Lane state and cleanup stay local to its runner; files within each lane remain sequential. The suite map retains one owner per file in both workflow groups.
The native IAM barrier test receives its own migrated PostgreSQL database through the application lane's per-file preparer. Only that test receives the matching application and migrator connection details, and the CLI refuses to write those details to GITHUB_ENV. The test installs its unregistered supplier only in that disposable database. This fixture does not establish that production writers participate in the barrier.
Both workflows call the shared run-ci-lane action after checkout. It owns tool and dependency setup, baseline checks when selected, lane preparation, execution, unconditional cleanup, and sanitized result upload. Callers keep the source revision, timeout, protected environment and explicit credentials.
Ordinary PR dependency caches may be restored and saved within GitHub's PR merge-ref scope. Main jobs use main-scoped caches. Test results and credential-bearing state are not dependency caches, and protected jobs do not promote PR build artifacts.
The provider job selects the shared blacksmith-8vcpu-ubuntu-2404 runner for disk headroom during runtime image build and k3d import. The standard Ubuntu runner reached DiskPressure and evicted the seccomp probe before it could start. The repository must retain access to this organization runner label. Image preparation copies the saved archive into each owned k3d node and runs node-local ctr image import; k3d tools-node can log per-node import failures while returning success. The imported manifest and CRI checks remain required before any test starts.
2. Prepare resources under the job owner
scripts/ci/prepare.mjs:main and scripts/ci/prepare.mjs:ensureK3dCluster
CI resource preparation traces tool setup, image and cluster preparation, protected credentials, and resource ownership. Continue below when preparation has produced the lane state.
For the three Kubernetes fixture lanes, cluster startup records phase timings
and host snapshots. On failure, bounded diagnostic reads save
<state-file>.diagnostics.json outside the cluster directory before cleanup.
Creation uses --no-rollback for these lanes so the workflow owns teardown after
capture; local callers still invoke cleanup with their failed run's state file.
Collection preserves the original error, including when an observation fails or
times out. The CI guide describes the retained evidence.
Tests delete their Agent namespaces, and those namespaces' events, before a
file exits. scripts/ci/run-tests.mjs:runFile therefore watches Compute-managed
Pods and Kubernetes events in each ready k3d cluster while the file runs, then
scripts/ci/k3d-diagnostics.mjs:projectAgentNamespaceActivity appends the Pod
status transitions and those namespaces' events to the same report under
agentNamespaces, passing or failing. Each file keeps at most 200 Pod and 200
event records and the report keeps 40 files; messages are redacted and
truncated, Pod specs are dropped, and raw watch streams stay in the cluster
directory that cleanup removes. The artifact is uploaded for every lane that
writes it.
Dedicated Codex preparation and the operator's offline profile generator share
scripts/lib/codex-seccomp-profile.mjs:deriveCodexBwrapProfile. Preparation
requires an actual workspace write and denied write to a container-writable
outside path before publishing the selected Localhost profile to the live suite.
Native runtime-image tests trust a dynamic Codex Docker seccomp profile only when
OPENCLAW_ENTERPRISE_CI_STATE records the exact prepared
cluster.codexDockerSeccompProfile path and SHA. A self-hashed profile without
that state is not CI proof; the standalone fallback remains the pinned reviewed
manual profile. Production node provisioning remains outside CI ownership; see
Codex sandbox setup.
3. Execute and account for actual cases
scripts/ci/run-tests.mjs:main and scripts/ci/reporter.mjs:jsonLinesReporter
The runner discovers active test files and verifies that the map assigns each file to exactly one lane. Tests with different prerequisites live in separate files. The runner invokes whole files with invocation-scoped environment inputs. A custom Node reporter exposes case names, locations and outcomes; arbitrary test output and credential-bearing error payloads are excluded from published results. Failed provider-test HTTP assertions also retain numeric actual and expected status codes, an allowlisted OCC error code, and the upstream ChatGPT operation and status when available. Denied-traffic failures retain only an allowlisted traffic category, without target addresses or response data. Plugin-status fixture failures retain an allowlisted readiness or rollout stage. Rollout diagnostics include bounded Pod phases, readiness and scheduling flags, container restart counts and exit codes, and allowlisted reasons. Response bodies, credentials, and identities remain excluded.
Required named cases must pass. Every skip or TODO fails the selected lane; there are no counterpart-skip lists or CI name filters. A synthetic file-wrapper success, missing result output, zero executed cases or an interrupted run without final reporter output cannot establish coverage. The runner retains failure, timeout and cleanup outcomes in the lane result.
4. Clean up and publish the bounded result
scripts/ci/cleanup.mjs:main and scripts/ci/run-tests.mjs:main
.github/actions/run-ci-lane/action.yml uploads one sanitized result artifact
per lane and workflow run. A job retry replaces that lane's earlier artifact;
other lanes retain their results. This prevents aggregation from selecting a
stale failed result after a successful retry. The earlier job logs remain the
failure record; retain a result separately before retrying when needed.
For images-packaging, scripts/ci/export-image-reconciliation.mjs attempts
to retain attempt-specific cleanup records for the two controller and runtime
tags prepared by the lane. A planned record does not prove an image was created.
The run-and-attempt component of each tag name is metadata, not authentication
or permission to delete an image. Missing state is reported as unavailable;
neither that result nor an empty inventory proves cleanup. The separate tag
created by the runtime-images test, other resource kinds, and private environment
values are excluded. Export or upload failure and runner loss can prevent retention.
Fixture bootstrap failures also upload diagnostics-<artifact-prefix>-<lane>
separately from test results. Cleanup removes the cluster and its private state;
the diagnostic file remains available for upload and does not satisfy the
aggregate's required test results.
Per-file cleanup releases its disposable database. Job cleanup removes only the state-owned resources. A whole owned k3d-cluster resource owns Kubernetes API object deletion for its Collector Namespace and RBAC. Logging cleanup cleans the local Docker backend container and JSONL/config directory independently, so a dead Kubernetes API does not block local log backend teardown. Cleanup failure fails the check and keeps the private state file usable only while that runner host and path remain available. User databases, contexts, unrelated containers and global images remain outside that ownership.
The aggregate runs after success or failure and checks expected job outcomes plus same-revision lane results. Case validation belongs to the runner; the aggregate checks lane identity and success, required evidence, and cleanup outcomes without interpreting cases again. Missing, failed, cancelled or skipped selected jobs cannot pass. A full-suite result accounts for every lane selected by the full group. The explicitly selected ssh-host lane remains outside the automatic groups until an operator prepares its disposable host; see SSH raw-host testing. Abrupt hosted-runner loss can prevent teardown and also loses the private RUNNER_TEMP state at job end. External resource reconciliation is deferred until an approved resource ledger exists.
Debugging and Verification
node scripts/ci/run-tests.mjs auditchecks the actual checkout inventory against the suite map.node --test tests/integration/ci-runner.test.mjsexercises the runner with real child Node processes and controlled pass/fail/skip cases.- Use the failing test's file, name and location in the sanitized result to reproduce its exact invocation with approved local prerequisites. Treat the named aggregate as its coverage boundary.
- On local Docker Desktop or equivalent VM-backed Docker hosts, run one Kubernetes lane at a time when disk or network pressure has caused measured instability. GitHub Actions still runs the configured matrix; this local guidance is for reproducible operator runs.
- Missing protected environments, tools, images or credentials are setup failures. Configure the approved resource; do not mark its required test skipped or replace it with a fixture.
- Retain sanitized results for seven days. Keep private cleanup state and credential files outside uploaded artifacts. On local runs, follow the run-owned state when recovering a failed teardown while that host and state path still exist.
