Skip to content

Latest commit

 

History

History
229 lines (211 loc) · 128 KB

File metadata and controls

229 lines (211 loc) · 128 KB

Self-Hosted Runner Failure Modes

Use this catalog when diagnosing AWF failures on self-hosted, ARC/DinD, GHEC, and GHES environments.

Platform fingerprint checklist

Establish these facts before matching a failure mode:

  • DOCKER_HOST scheme and socket path
  • ARC markers: ACTIONS_RUNNER_POD_NAME, ACTIONS_RUNNER_CONTAINER_HOOKS
  • GITHUB_SERVER_URL
  • runner HOME
  • daemon libc / runtime (glibc, musl, runc, runsc, kata)
  • Docker IPv6 state

Category A — ARC / DinD

ID Signal Root cause Fix / flag Probe Citations
A1 Bind-mounted files are missing in DinD containers Bind mounts point at runner paths that the DinD daemon cannot see --docker-host-path-prefix /tmp/gh-aw, container.dockerHostPathPrefix, AWF_DIND=1 Create a /tmp sentinel on the runner, then docker run -v /tmp:/tmp ... ls <sentinel> #2833, #2945, #3553, #3845, #3906, #4023, #4271, #4399, #4727, #4737, #4787, #5753
A2 DOCKER_HOST=tcp://... disappears and compose falls back to /var/run/docker.sock TCP DOCKER_HOST was stripped instead of propagated Preserve DOCKER_HOST into compose and container env docker -H "$DOCKER_HOST" info and inspect generated compose env #4830
A3 Non-standard unix DOCKER_HOST does not trigger split-fs handling Split-filesystem auto-detection only keyed on tcp:// Treat non-default unix sockets and ARC fingerprints as split-fs Check whether DOCKER_HOST uses a unix socket outside /var/run/docker.sock #3553, #3906, #4023
A4 capsh missing, /bin/bash missing, or node missing in DinD chroot Alpine/musl daemon host lacks glibc tooling expected by chroot mode Use ghcr.io/github/gh-aw-firewall/dind-ubuntu:latest; fail fast on musl ldd --version, inspect daemon /etc/os-release, verify capsh #3393, #3397, #4567, #4737, #4787, #2535
A5 one-shot-token.so ... __fprintf_chk: symbol not found The token-protection shared object was built for glibc, not musl Use the build with _FORTIFY_SOURCE=0; expect graceful fallback warning ldd --version, set AWF_ONE_SHOT_TOKEN_DEBUG=1 #2535
A6 getent passwd <UID> fails, HOME=/, USER=root inside chroot DinD daemon lacks the runner UID in /etc/passwd and /host/etc is read-only Use staged passwd/group synthesis; set chroot.identity.* when needed docker run --rm <dind-image> getent passwd $(id -u) #4829, #4831
A7 HOME, USER, or LOGNAME do not survive into the engine command capsh user switching resets identity variables Re-apply chroot.identity.{home,user,uid,gid} after capsh Inspect env inside chroot after the user switch #4567, #4787
A8 /host/usr/local/bin/copilot missing or node: command not found Runner-installed binaries are not visible from the daemon filesystem chroot.binariesSourcePath; dind.stageEngineBinary.{path,targetPath} Test visibility of runner-installed binaries in a DinD container #4271, #4399, #4727, #4737, #4787
A9 GitHub MCP tools disappear under --disable-builtin-mcps mount_mcp_as_cli.cjs hardcoded the GitHub server as internal Remove or make the internal-server list configurable Check generated mcp-* shims and INTERNAL_SERVERS in the image #4271, #4399, #4727, #4737, #4787
A10 Docker socket not found plus Invalid container ID format: arc-... MCP gateway assumed /var/run/docker.sock, group 0, and Docker-style container IDs Propagate DOCKER_HOST, detect socket GID, relax pod-name handling stat -c '%g' ${DOCKER_HOST#unix://}, cat /proc/self/cgroup #2267, #2292, #2664, #2706, #2808
A11 Threat detection passes even though the engine binary is missing GH_AW_DETECTION_CONTINUE_ON_ERROR suppressed a real setup failure Reconsider default or log the skipped check explicitly printenv GH_AW_DETECTION_CONTINUE_ON_ERROR; inspect agent logs for ENOENT #4787
A12 mkdirat ... : read-only file system during agent chroot startup on ARC/DinD chroot.binariesSourcePath set to the same root as --docker-host-path-prefix (e.g. both /tmp/gh-aw); Docker mounts /tmp/gh-aw/usr:/host/usr:ro first, then the attempt to mkdir /host/usr/local/bin as a nested overlay mount point fails because the parent is read-only Fixed in firewall v0.27.10: upgrade AWF; the overlay is now mounted at /host/tmp/awf-runner-bin:ro (writable /host/tmp parent) instead of /host/usr/local/bin:ro Check awf --version; inspect agent container logs for mkdirat; verify chroot.binariesSourcePath equals docker-host-path-prefix root #5481, #5482
A13 chroot: failed to run command '/bin/sh': No such file or directory or [entrypoint][ERROR] capsh not found on host system on a glibc/Debian daemon (not musl/Alpine) ARC/DinD split-fs: system-mount source dirs (/tmp/gh-aw/{usr,bin,lib,...}) are empty because nothing populates them. The entrypoint "musl/Alpine" warning is misleading — it fires because no dynamic loader is found, not because the daemon is musl. Fixed in AWF v0.27.15: set runner.topology: "arc-dind" in the AWF config JSON. AWF emits a sysroot-stage init container that copies the signed build-tools image filesystem (bash, capsh, gcc, dev libs, coreutils) into a named sysroot volume mounted at /host:ro before the agent starts. Use runner.sysrootImage to pin a specific image. Check awf --version ≥ v0.27.15; verify runner.topology: "arc-dind" is set; inspect compose output for sysroot-stage service and sysroot volume #5541, #5693, #5696
A14 unknown shorthand flag: 'd' in -d / Command failed with exit code 125: docker compose up -d --pull never ARC/DinD sidecar image lacks docker-compose-plugin; AWF uses docker compose (v2 plugin syntax) to orchestrate containers but the DinD sidecar only has legacy standalone Docker or no Compose support Add docker-compose-plugin to the DinD sidecar Dockerfile: RUN apt-get update && apt-get install -y docker-compose-plugin (Debian/Ubuntu) or RUN apk add docker-cli-compose (Alpine) docker compose version inside the DinD sidecar — v2 output confirms plugin is present; inspect sidecar Dockerfile for docker-compose-plugin #5729
A15 [WARN] Rootless artifact permission repair failed for .../sandbox/firewall/logs (exit 1); squid log files unreadable after ARC/DinD run; awf logs summary returns Failed to load logs: EACCES fixArtifactPermissionsForRootless() binds the log directory into a repair container but does not apply dockerHostPathPrefix translation to the bind mount source path; the DinD daemon cannot resolve the runner-local path, so chmod exits non-zero Fixed in PR #5963: fixArtifactPermissionsForRootless() now calls applyHostPathPrefixToVolumes() so the repair container bind mount is correctly translated for the DinD daemon. Upgrade AWF to the version that includes #5963. Workaround (older AWF): run chmod -R a+rX inside the squid container before docker compose down. ls -la <proxy-logs-dir> after run — files owned by uid 13 (squid) confirm the mode; check AWF logs for [WARN] Rootless artifact permission repair failed #5816, #5817, #5963
A16 ARC/DinD with runner.topology: arc-dind: custom --mount paths (e.g. ${RUNNER_TEMP}/gh-aw:...) are silently dropped; agent command fails with node: command not found or other binary-not-found errors even when the tool is correctly installed and the mount was confirmed daemon-visible AWF sysroot mount filter (filterAgentVolumesForSysroot) was too aggressive: it dropped any mount whose source or target fell under effectiveHome (/home/runner), incorrectly including daemon-visible workspace paths such as ${RUNNER_TEMP}/gh-aw (/home/runner/_work/_temp/...) Fixed in AWF (PR #5739): filter now only drops dot-directories and the home root; workspace paths under _work/ now pass through. Upgrade AWF to the version including #5739. Inspect generated agent compose YAML for the expected --mount entry; docker run --rm -v ${RUNNER_TEMP}:${RUNNER_TEMP}:ro alpine ls ${RUNNER_TEMP} to confirm daemon-visibility #5739
A17 On ARC/DinD with runner.topology: arc-dind, --image-tag build-tools=sha256:<digest> throws Error: invalid key 'build-tools'; the sysroot-stage init container image cannot be digest-pinned IMAGE_DIGEST_KEYS in src/image-tag.ts does not include 'build-tools'; buildSysrootStageService() constructs the image ref as a template string, bypassing buildRuntimeImageRef() entirely Fixed in AWF (PR #5986): 'build-tools' is now in IMAGE_DIGEST_KEYS; buildSysrootStageService() accepts ParsedImageTag via SysrootServiceParams and calls buildRuntimeImageRef() instead of a hardcoded template string. Upgrade to the AWF version that includes #5986. awf --image-tag build-tools=sha256:abc ... — on a patched version the command succeeds and the sysroot-stage image ref includes the pinned digest (visible in generated compose YAML); on older AWF without the fix the command throws Error: invalid key 'build-tools' #5985, #5986
A18 XDG-respecting tools (Flutter, etc.) fail with EACCES / permission errors; writes land directly under /home/runner (for example /home/runner/tool_state) under runner.topology: arc-dind; the actual write target is under /home/runner (root-owned) rather than the writable ${RUNNER_TEMP}/gh-aw/home The copilot engine entrypoint (gh-aw v0.79.8+) emits export XDG_CONFIG_HOME="$HOME" before reassigning HOME to the writable arc-dind path ${RUNNER_TEMP}/gh-aw/home. Any XDG-respecting tool sees the stale, unwritable value. engine.env is sourced before this export, so XDG_CONFIG_HOME set there is silently overwritten by the later shell export. Fixed in gh-aw (PR github/gh-aw#48658, merged 2026-07-28): XDG_CONFIG_HOME is now exported after HOME is reassigned to the writable arc-dind path. Upgrade gh-aw to the version including github/gh-aw#48658. Workaround (older gh-aw): Setting XDG_CONFIG_HOME in engine.env is ineffective because the later shell export overwrites it; override HOME to the writable path instead (e.g. add HOME=${RUNNER_TEMP}/gh-aw/home to engine.env). Inside the arc-dind agent container: echo "$XDG_CONFIG_HOME" — if it shows /home/runner rather than a path under $RUNNER_TEMP, the ordering bug is present; ls -la /home/runner — root ownership confirms the mode #6684, github/gh-aw#48658
A19 create_pull_request fails with No patch file found / No patch or bundle files found in: /tmp/gh-aw on ARC/DinD even though the safeoutputs MCP server inside the agent container reports it successfully wrote aw-.patch/aw-.bundle The /tmp/gh-aw:/tmp/gh-aw:rw bind mount used for the safeoutputs patch/bundle handoff was not passed through AWF's existing translateBindMountHostPath() / --docker-host-path-prefix normalization (src/services/agent-volumes.ts). In DinD split-filesystem topologies the Docker daemon resolves the bind source against its own filesystem, not the runner's staged path, so writes made inside the container land somewhere the runner-side ingestion step never sees. Fixed in AWF (PR #6959, merged 2026-08-05): the safeoutputs exchange mount source is now built through the same docker-host-path-prefix translation path as other agent bind mounts (the generated Compose target remains /host/tmp/gh-aw, which is /tmp/gh-aw inside the chroot; only the host-side bind source changes when a prefix is configured). Upgrade AWF to include #6959. After a create_pull_request failure on ARC/DinD, check whether --docker-host-path-prefix is set and inspect the generated Compose bind mount for /host/tmp/gh-aw — on unpatched AWF the host source is untranslated (e.g. /tmp/gh-aw:/host/tmp/gh-aw:rw instead of <prefix>/tmp/gh-aw:/host/tmp/gh-aw:rw) github/gh-aw#50217, #6948, #6958, #6959
A20 Under runner.topology: arc-dind, awf-agent fails to start (runc cannot create the ~30 credential-hiding /dev/null overlay mountpoints under /host$HOME), or once worked around, the entrypoint aborts with mkdir -p /host$HOME/.m2 failing under set -e filterAgentVolumesForSysroot() (src/services/optional-services.ts) dropped every mount targeting /host$HOME, including the compiler-supplied writable home (${RUNNER_TEMP}/gh-aw/home), because it could not distinguish AWF's own unshared ${workDir}-chroot-home mount (correctly dropped) from a caller-supplied, daemon-visible home mount Fixed in AWF (PR #7244, merged 2026-08-11): home mounts whose target matches an explicitly supplied --mount/config.volumeMounts spec now survive the sysroot filter (the caller vouches for daemon-visibility); AWF's own workDir-based chroot-home mount is still dropped. If no writable /host$HOME survives, /dev/null credential overlays under that path are skipped with a warning instead of failing runc startup (overlays at the un-prefixed $HOME path are unaffected). containers/agent/entrypoint.sh's JVM proxy pre-seeding now guards its mkdir -p .../.m2 call and logs+skips instead of aborting under set -e when the chroot home is read-only. buildCustomVolumeMounts() (src/services/agent-volumes/workspace-mounts.ts) also stops re-prefixing targets that already start with /host, fixing a related double-/host bug for --mount src:/host/path:ro specs. Upgrade AWF to include #7244. Not addressed: gh-aw emitting ${RUNNER_TEMP}/gh-aw read-only over the chroot home, and its DOCKER_HOST gate on the chroot config patch — both require changes in github/gh-aw. Inspect docker-compose.redacted.yml for a writable /host$HOME (or its explicit-mount target) under runner.topology: arc-dind; check agent startup logs for the "no writable home survived, skipping overlays" warning vs. a runc mountpoint-creation failure; check entrypoint logs for the "Cannot create .../.m2 (read-only home)" skip message #7239, #7244
A21 awf-agent fails to start with runc create failed: ... mkdirat /var/lib/docker/overlay2/<layer-id>/merged/tmp/awf-init: read-only file system (or equivalent for /tmp/awf-runner-bin) when a filesystem.allowWrite policy narrows /tmp to read-only; most reliably reproduced on ARC/DinD split-filesystem topologies using --docker-host-path-prefix runc creates missing bind mountpoints with mkdirat against whichever bind already covers the destination. AWF control-plane mountpoints (/tmp/awf-init, /tmp/awf-runner-bin) were nested under the user-narrowable /tmp bind, so narrowing /tmp to ro blocked nested mountpoint creation and failed startup with EROFS. /tmp/awf-lib was helper-copy staging rather than a nested mountpoint; narrowing /tmp could silently prevent those copies. On ARC/DinD with a /tmp-rooted --docker-host-path-prefix, shared-prefix detection also misclassified AWF workDir-derived binds as daemon-only and failed closed. Fixed in AWF (PR #7679, merged 2026-08-24): init-signal moved to /run/awf-init; a new planNestedMountpoints()/ensureNestedMountpoints() pass pre-creates mountpoints that would land inside read-only covers (or fails closed); isSharedDockerHostPathPrefix now treats only the literal /tmp prefix as shared for ARC/DinD detection; legacy /tmp/awf-init compatibility binds remain for older pinned agent images; and /tmp/awf-lib helper staging (one-shot token protection library, Claude API key helper, gh CLI proxy wrapper, CA bundles, runner shims) moved to /run/awf-lib, eliminating silent degradation under filesystem.allowWrite. Startup now fails closed if the one-shot token library cannot be staged or if CLI proxying is enabled but the gh wrapper cannot be installed. Upgrade AWF to include #7679. Check awf --version for #7679; inspect startup logs for mkdirat ... read-only file system with active filesystem.allowWrite; inspect entrypoint logs for [entrypoint][WARN] Could not copy one-shot-token library to /tmp/awf-lib — on older AWF this confirms the silent-degradation mode; on ARC/DinD verify whether --docker-host-path-prefix is exactly /tmp (shared) vs. daemon-only (for example /host) #7678, #7679, #7681, #7728
A22 arc-dind topology fails to start with Docker rejecting the compose cap_drop list: invalid CapDrop: capability not supported by your kernel or not available in the current environment: "CAP_SYS_MODULE" (or similar) on hosts, such as Talos Linux, that trim capabilities from the container capability bounding set src/services/squid-service.ts and src/services/agent-service.ts hardcoded cap_drop lists (NET_RAW, SYS_ADMIN, SYS_PTRACE, SYS_MODULE, MKNOD, AUDIT_WRITE, SETFCAP for Squid; a similar list for the agent) for both the Squid and agent containers with no filtering against what the host/daemon's capability bounding set actually supports, so hosts with a trimmed bounding set (for example Talos gha-runner-scale-set with docker:29-dind) can never satisfy Docker's compose validation Fixed in AWF (PR #7795, merged 2026-08-28): cap_drop is now filtered against the effective host capability bounding set (read daemon-side from /proc/self/status CapBnd via a privileged probe container) before writing docker-compose.yml; capabilities already absent from the bounding set are silently omitted from cap_drop (a safe no-op, since a capability that can't be granted can't be exploited either). Upgrade AWF to include #7795. getHostCapabilityBoundingSet() (src/capability-filter.ts:78-89) runs docker run --rm --privileged --network=none alpine:latest cat /proc/self/status against the daemon and decodes CapBnd; reproduce with docker compose up on a host missing CAP_SYS_MODULE from the bounding set — pre-fix this fails with invalid CapDrop, post-fix compose starts normally github/gh-aw#56127, #7788, #7795
A23 On runner.topology: arc-dind with --docker-host-path-prefix set, docker compose up fails with error mounting "/dev/null" to .../home/.npmrc: create mountpoint ...: read-only file system (also seen for .docker/config.json, .composer/auth.json) filterAgentVolumesForSysroot() (src/services/optional-services.ts) is meant to drop the bogus AWF-owned chroot-home mount that the DinD daemon can't resolve, but it compared already-prefixed mount sources (from buildAgentVolumes(), which applies --docker-host-path-prefix as its final step) against the raw, unprefixed config.workDir/effectiveHome. Once a host-path prefix is set — always true on real ARC/DinD — the comparison silently stopped matching, so the bogus chroot-home mount survived filtering and Docker tried to create a .npmrc credential-hiding overlay mountpoint inside a path the daemon couldn't write to (EROFS). Distinct from A20 (which was about legitimate caller-supplied home mounts being dropped); A23 is the inverse case — the bogus mount not being dropped. Fixed in AWF (PR #7998, merged 2026-09-02): extracted prefixHostPath() in src/services/host-path-prefix.ts from translateBindMountHostPath() so bare paths can be prefixed consistently; filterAgentVolumesForSysroot() now prefixes config.workDir/effectiveHome before comparing against mount sources, restoring correct detection of daemon-invisible mounts. The existing safe fallback (dropUnbackedHostHomeOverlays, warn + skip masking) now engages correctly instead of silently failing; explicit writable --mount for the home root still preserves credential masking. Upgrade AWF to include #7998. Reproduce with runnerTopology: 'arc-dind' + dockerHostPathPrefix set (e.g. /host) and inspect generated compose for a surviving prefixed ${workDir}-chroot-home:/host$HOME mount together with /dev/null:/host$HOME/<credential>:ro overlays; on unpatched AWF, docker compose up fails with the EROFS mounting error above github/gh-aw#57468, #7994, #7998
A24 On runner.topology: arc-dind, docker compose up -d --pull never fails with error mounting "/dev/null" to rootfs at ".../gh-aw/home/.npmrc": create mountpoint for .../.npmrc mount: ... openat .npmrc: read-only file system — distinct from A23 in that it persists even after #7998 (A23's fix), specifically when the credential mountpoint is missing under a declared-rw home bind whose real backing directory is genuinely read-only on the ARC/DinD-staged filesystem pruneUnmountableCredentialOverlays (src/services/agent-volumes/credential-hiding.ts) decided whether a /dev/null credential mask could be mounted purely from the declared compose bind mode (ro/rw), never checking the real filesystem when the mode was rw. AWF's own home-directory mount is declared rw, but under --docker-host-path-prefix it can resolve to a directory that is genuinely read-only on the runner's staged filesystem; Docker doesn't remount that case read-only inside the container (unlike a declared ro bind), so it touches the real host path directly when creating a missing mountpoint and hits EROFS, crashing the agent container before it starts. Fixed in AWF (PR #8086, merged 2026-09-04): an overlay whose mountpoint already exists is always kept (mounting over an existing path succeeds regardless of declared mode or real writability); a missing mountpoint under a declared-ro bind is still always dropped; a missing mountpoint under a declared-rw bind is now probed against the real filesystem (walking up to the nearest existing ancestor, mirroring how Docker creates missing intermediate directories) and the overlay is dropped if that real directory isn't writable. Upgrade AWF to include #8086. Reproduce with an absent credential path (for example, .npmrc) beneath a chmod-based real read-only home directory under runner.topology: arc-dind + --docker-host-path-prefix; on unpatched AWF (even with #7998 applied) docker compose up fails with the EROFS mounting error above; on patched AWF the missing overlay is skipped and the agent starts. Existing credential paths remain mountable. github/gh-aw#57468, #8076, #8086
A25 On runner.topology: arc-dind, workloads inside the AWF sandbox need a GitHub Actions services: container's native protocol (DB drivers, migration tools, etc.) but cannot reach it — the services: container runs on the runner's own bridge network while the AWF agent runs on the isolated awf-net, and the two bridges are unrouted; existing host-iptables service-port routing doesn't help because ARC/DinD network isolation never programs host iptables rules No AWF mechanism previously joined a services: container to awf-net; raw-protocol clients (e.g. psql) have no route from the sandbox to the service Documented in AWF (PR #8085, merged 2026-09-04): new docs/arc-dind.md section "Joining services: containers to awf-net for direct protocol access" documents a verified workaround — a pre-step waits for awf-net to exist, then docker network connect --alias <name> awf-net <service_container> attaches the service container (never the agent) with a resolvable alias. Security invariant: only the service joins awf-net; joining the agent to the runner bridge would bypass the Squid egress firewall. Longer-term direction (services.<name>.attach: true compiler sugar) is not yet implemented. Confirm the services: container, not the agent, is the one calling docker network connect --alias <name> awf-net <container>; verify the agent can resolve/reach <name> after the join; confirm the agent itself never appears attached to the runner's default bridge #8075, #8085
A26 On runner.topology: arc-dind deployed on OpenShift/ARO clusters, Squid fails all CONNECT requests with 503 HIER_NONE; the awf-net bridge subnet (172.30.0.0/24, Squid pinned at 172.30.0.10) collides with the cluster's default service CIDR (172.30.0.0/16) and specifically with the CoreDNS ClusterIP (172.30.0.10), so Squid sends its own DNS queries to itself src/docker-manager.ts hardcoded the awf-net subnet to 172.30.0.0/24 in generateDockerCompose() with no override, ignoring the Docker daemon's --default-address-pool because Compose-declared subnets take precedence; src/squid-config.ts templated dns_nameservers from the runner pod's /etc/resolv.conf, which resolves inside the claimed subnet on OpenShift/ARO, creating a self-referential route once awf-net exists Fixed in AWF (PR #8398, merged 2026-09-10): new --network-subnet <cidr> CLI flag / network.subnet config key relocates awf-net off the default 172.30.0.0/24 (accepts /16–/26; fixed host offsets — .1 gateway, .10 Squid, .20 agent, .30 api-proxy, .40 DoH, .50 cli-proxy — are rebased into the chosen block via resolveNetworkAddressing() in src/network-subnet.ts). assertNetworkSubnetUsable() now fails loudly at startup if the chosen (or default) subnet contains a detected/explicit DNS resolver or overlaps a non-Docker-managed host route, instead of silently producing broken DNS. The override is rejected with `--container-runtime sbx cloud-hypervisor` because those runtimes have fixed guest network plans. Upgrade AWF to include #8398. Check awf --version for #8398; on OpenShift/ARO inspect the startup error for The awf-net subnet 172.30.0.0/24 contains the DNS resolver(s) ...; relocate with --network-subnet 10.88.0.0/24 (or {"network":{"subnet":"10.88.0.0/24"}}) and confirm Squid's CONNECT no longer returns 503 HIER_NONE
A27 On runner.topology: arc-dind, safe-output payload files staged under /tmp/gh-aw/agent (patch/bundle handoff, PR body, etc.) silently fail to write inside the AWF chroot — the heredoc/write succeeds on the runner side against a path the Docker daemon never sees, producing empty safe-output fields (e.g. an empty PR body) instead of a hard failure ARC/DinD's "Create gh-aw temp directory" step and AWF's dind-bootstrap.ts pre-staging only created /tmp/gh-aw/agent on the runner filesystem. The Docker daemon in a split-filesystem DinD topology cannot see that path, so when RUNNER_TEMP is set (the daemon-visible staging root), the corresponding ${RUNNER_TEMP}/gh-aw/agent directory was never created or pre-staged, leaving no daemon-visible landing path for the agent's safe-output writes Fixed in AWF (PR #8933, merged 2026-09-24, fixes #8932): ensureAgentStagingDirectories() in src/dind-bootstrap.ts now creates both /tmp/gh-aw/agent (existing behavior) and ${RUNNER_TEMP}/gh-aw/agent when RUNNER_TEMP is set; agent was added to DEFAULT_PRE_STAGE_DIRS so DinD bootstrap also pre-stages it through Docker; ${RUNNER_TEMP}/gh-aw is pre-staged as a second work-directory tree when it differs from the configured dind.workDir. Non-DinD/no-RUNNER_TEMP behavior is unchanged. Upgrade AWF to include #8933. Set RUNNER_TEMP and dind.preStageDirs: true and inspect whether ${RUNNER_TEMP}/gh-aw/agent exists and was pre-staged through Docker (not just /tmp/gh-aw/agent); on unpatched AWF only the runner-side path exists and safe-output payloads written to /tmp/gh-aw/agent inside the chroot silently vanish; on patched AWF both paths exist and are pre-staged github/gh-aw#63045, github/gh-aw#62924, #8932, #8933, #8938
A28 On runner.topology: arc-dind, an engine's streaming log beneath ${RUNNER_TEMP}/gh-aw fails to write with read-only file system, turning a successful engine run into a failure and losing safe outputs The log path was directly under the read-only ${RUNNER_TEMP}/gh-aw parent mount Fixed in AWF (PR #9188, merged 2026-09-29, fixes #9183): AWF retains a writable sandbox/agent child beneath the read-only parent. Use ${RUNNER_TEMP}/gh-aw/sandbox/agent/pi-streaming.jsonl consistently for the writer, parser, and artifact upload. Moving Pi's log and consumers requires a separate gh-aw compiler change. Inside the agent, touch "${RUNNER_TEMP}/gh-aw/sandbox/agent/awf-doctor-probe" should succeed, while touch "${RUNNER_TEMP}/gh-aw/awf-doctor-probe" should fail with EROFS #9183, #9188
A29 On arc-dind, post-run consumers cannot find the api-proxy token-usage.jsonl (hardcoded /tmp/gh-aw/...); gh-aw.aic / gen_ai.usage.* telemetry is missing even though the run succeeds The compiler passes ${RUNNER_TEMP}/gh-aw/... log dirs. With --docker-host-path-prefix /tmp, a --proxy-logs-dir outside /tmp is rewritten to /tmp<dir>, so AWF reported and repaired the original, empty dir Fixed in AWF v0.28.31 (PR #9357, merged 2026-10-02, fixes #9352): AWF logs Token usage log available at: <path> and exports AWF_TOKEN_USAGE_LOG to $GITHUB_ENV when the file exists. It pre-creates the rewritten api-proxy log dir (mode 0777) with a warning. Read the exported path instead of hardcoding /tmp/gh-aw; gh-aw's parse_token_usage.cjs must still read AWF_TOKEN_USAGE_LOG (unresolved) In a later job step, printf '%s\n' "${AWF_TOKEN_USAGE_LOG:-unset}"; ls -l "${AWF_TOKEN_USAGE_LOG:-/dev/null}"; expect the variable to be set and the file to exist when token usage was written #9352, #9357

Category B — Self-hosted runners

ID Signal Root cause Fix / flag Probe Citations
B1 /home/runner/... paths are wrong on a custom runner home The runner uses a non-standard HOME Use the real HOME; when configuring stdin, set chroot.identity.home echo "$HOME" and inspect mounted home paths #2109, #2290
B2 All outbound traffic fails behind a mandatory corporate proxy AWF must chain Squid through the upstream proxy Set https_proxy / http_proxy on the host or use --upstream-proxy. Note: AWF ≤ v0.27.32 (before PR #6267) silently blocked proxy environment variables (NO_PROXY, HTTP_PROXY, HTTPS_PROXY, etc.) when passed via --env if enableApiProxy is active, because PROXY_ENV_VARS were in the credential exclusion set. Fixed in PR #6267: proxy vars are now checked against the PROXY_ENV_VARS allowlist and passed through additionalEnv even when credential isolation is active. Additional fix (PR #8887, merged 2026-09-22, fixes #8877): AWF's upstream-proxy URL parser (parseProxyUrl() in src/upstream-proxy.ts) silently discarded an explicitly-supplied default HTTP port. URL normalization strips :80 from http://proxy.corp.example.com:80 before AWF read the port, so a corporate proxy genuinely listening on port 80 was rewritten to Squid's 3128 default and all cache_peer chaining failed. The fix captures the raw authority port (getExplicitProxyPort()) before URL normalization removes it, and keeps the 3128 fallback only when no port was supplied at all. Validation was also hardened to reject an explicitly-empty port (for example http://proxy.corp.com:) instead of silently falling back. This affects both host https_proxy/http_proxy auto-detection and --upstream-proxy. Upgrade AWF to include #8887. `env grep -i proxy; inspect Squid config for cache_peer; on unpatched AWF, set https_proxy=http://proxy.corp.example.com:80` and inspect the generated squid.conf cache_peer line — a wrong 3128 port (instead of 80) confirms the bug; on patched AWF the cache_peer port matches the explicit :80
B3 Squid exits with FATAL: http_port: IPv6 is not available Docker IPv6 is disabled but Squid tries to bind an IPv6 listener Enable Docker/kernel IPv6 (required with current AWF builds), or use a custom AWF build that removes the [::] listener `docker info
B4 node: command not found after actions/setup-node on self-hosted Node was installed in $HOME/work/_tool and that toolcache is not visible Mount / expose the runner toolcache; use AWF_EXTRA_TOOLCACHE_DIRS if needed which node; inspect $HOME/work/_tool/node #3544, #3545
B5 getaddrinfo EAI_AGAIN <awmg-cli-proxy> → awf-cli-proxy could not connect to the external DIFC proxy → The agent was never invoked in --network-isolation + --topology-attach runs, and the topology peer becomes reachable once attached to awf-net Startup ordering deadlock: connectTopologyContainers() runs only after startContainers() succeeds, but startContainers() blocks on the cli-proxy health gate that requires the topology peer to be reachable on awf-net (which internal: true). The peer is never attached → transient EAI_AGAIN/ENOTFOUND → fail-fast → deadlock. Deterministic, not flaky. Resolved in AWF: attach topology peers to awf-net before the health-gated bring-up (Fix A: split up -d, network first → attach → remaining); also harden cli-proxy to treat transient EAI_AGAIN/ENOTFOUND as not-yet-ready (Fix B). If ARC/DinD nslookup awmg-cli-proxy still fails even after the ordering fix, match B12 instead. Confirm topologyAttach is non-empty; inspect whether the topology peer is attached to awf-net before the cli-proxy health check runs. On ARC/DinD, run docker run --rm alpine nslookup awmg-cli-proxy first — if it returns NXDOMAIN/SERVFAIL, prefer B12 over B5. #5543, #5542, #6326, #6328
B6 EACCES in upload-artifact step after a sudo: false (--network-isolation) AWF run; firewall log/audit dirs present but unreadable Sidecars write files as non-runner UIDs (squid → uid 13, cli-proxy → cliproxy, agent/iptables-init → root). AWF's chmod -R a+rX repair runs as the unprivileged runner and silently fails at debug level on files it doesn't own Resolved in AWF: (a) run Node sidecars as runner UID via compose user:; (b) root perm-fixer container at cleanup (daemon-run, mounts log dir, chowns to runner UID, skipped when --keep-containers); (c) promote swallowed-chmod failure from debug to warn. PR #6328 (merged 2026-07-17): Benign permission errors (Operation not permitted, EPERM, EACCES) from the artifact repair container are now logged at debug with an "expected on restricted runners" note. Only genuinely unexpected failures still emit [WARN]. Post-#6328: a [WARN] Rootless artifact permission repair failed message indicates a genuine failure, not a restricted-runner non-issue. ls -la <firewall-logs-dir> after run — look for root or uid-13 owned files; check AWF logs for the swallowed chmod warning #5545, #5542, #6328
B7 AWF < v0.27.13: unhandled EACCES stack trace shows unlink ... /tmp/awf-<ts>-chroot-home/<path> (e.g. .aws/config, cloud credentials). AWF ≥ v0.27.13: removeWorkDirectories() catches the error and emits [WARN] Failed to remove chroot home directory after permission repair instead of crashing In rootless Docker mode the agent container runs with UID namespace remapping. Files created by the agent inside the chroot-home temp directory are owned by remapped UIDs. AWF's removeWorkDirectories() runs as the unprivileged host runner and fs.rmSync fails on these files. Partially fixed in AWF v0.27.13 (repair container with CHOWN/DAC_OVERRIDE/FOWNER capabilities); further fix merged post-v0.27.15 (#5717): in rootless Docker the repair container's chown operates within the user namespace and may not change host-level ownership. The post-v0.27.15 fix adds chmod -R a+rwX so the host can delete the directory regardless of ownership. Non-fatal if unfixed — leaves an orphan /tmp/awf-*-chroot-home dir. additional hardening (#5766): changes chown && chmod to chown 2>/dev/null; chmod so chmod always runs as a fallback even when chown fails within the rootless UID namespace. ls -la /tmp/awf-*-chroot-home/ after a rootless run — files owned by non-runner UIDs confirm the mode; upgrade to AWF ≥ v0.27.13; check AWF logs for [WARN] Failed to remove chroot home directory after permission repair #5653, #5708, #5717, #5766
B8 EACCES: permission denied, mkdir '/tmp/gh-aw/sandbox/firewall/logs' (or any /tmp/gh-aw/... path) — failure occurs before any container starts, at writeConfigs time, on a persistent self-hosted runner A previous AWF run or the Docker daemon left /tmp/gh-aw/sandbox/firewall/ (or a parent) owned by root. With --network-isolation now the default, AWF runs without sudo, so mkdirSync on the root-owned parent fails with EACCES. Fixed in AWF (PR #5983): added preflight-reclaim.ts — on non-root invocation, walks upward from the target path to find the first non-writable ancestor and removes it via sudo rm -rf with fs.rmSync fallback; protected paths (/, /tmp, /home/runner) are never touched. Workaround (older AWF): sudo rm -rf /tmp/gh-aw/sandbox before re-running. ls -la /tmp/gh-aw/sandbox/firewall/ — dirs owned by root (uid 0) confirm the mode; `docker info grep -i rootless`
B9 No CA certificates were loaded from the system — Copilot CLI or other HTTPS tools fail inside AWF chroot on RHEL, Fedora, or Amazon Linux runners; all HTTPS traffic returns TLS verification errors AWF chroot mounts only Debian/Ubuntu CA paths (/etc/ssl:ro, /etc/ca-certificates:ro). On RHEL/Amazon Linux the system CA bundle lives under /etc/pki/ca-trust/ which is not mounted. Fixed in AWF (PR #5783): copy_system_ca_bundle() in agent entrypoint detects the CA bundle from 5 candidate paths (Debian, RHEL, Fedora, macOS, Alpine), copies it to /tmp/awf-lib/system-ca-certificates.crt if not directly accessible, and sets SSL_CERT_FILE, NODE_EXTRA_CA_CERTS, REQUESTS_CA_BUNDLE, CURL_CA_BUNDLE, GIT_SSL_CAINFO. Further fixed in AWF (PR #6460, merged 2026-07-21): /etc/pki/ca-trust and /etc/pki/tls are now included in the chroot mount policy, making RHEL/Amazon Linux CA paths directly accessible inside /host/etc/pki without requiring the staging fallback. Workaround (older AWF): copy /etc/pki/ca-trust/extracted/pem/tls-ca-bundle.pem to a chroot-visible path and set those env vars. ls /etc/pki/ca-trust/extracted/pem/tls-ca-bundle.pem — present on RHEL/Amazon Linux confirms the mode #5733, #5783, #6460
B10 [WARN] Rootless artifact permission repair failed for .../sandbox/firewall/logs (exit 1) × 3 directories, each taking exactly ~30 s; total ~90 s wasted; cascading EACCES: permission denied / Failed to remove chroot home directory after permission repair fixArtifactPermissionsForRootless() builds compound tag@digest image refs (e.g. agent:0.27.22@sha256:55f065...) for its docker run --pull never repair container. Despite --pull never, Docker still attempts registry manifest verification for compound refs; when GHCR credentials are unavailable (cleaned in an earlier workflow step or expired) the verification TCP-connects and times out (~30 s per directory). Fixed in AWF (PR #6025): resolvePermFixerImageRef() now strips the @sha256:... digest and uses a tag-only ref. --pull never with a tag-only ref skips all registry I/O. Upgrade to AWF version that includes #6025. Workaround (older AWF): ensure GHCR credentials are available until after the AWF cleanup step. Additional fix (PRs #6342 / #6356, merged 2026-07-18): Even with a tag-only image ref, docker run was missing --entrypoint sh, causing AWF's entrypoint.sh to execute in place of the repair command. entrypoint.sh waits ~30 s for an iptables-init container that never starts in this context, then fails. The fix adds --entrypoint sh and passes the chown/chmod command via -c; stdout is now captured alongside stderr so timeout messages surfaced on stdout are not silently swallowed. [WARN] Rootless artifact permission repair failed messages each followed by exactly ~30 s delay in the workflow log; docker pull --dry-run agent:...@sha256:... times out while docker inspect agent:... succeeds #6025, #6342, #6356
B11 AWF exits with code 1 and logs show [WARN] Rootless artifact permission repair failed ... (exit 1) plus cleanup warnings, but no actionable stderr detail from the repair container fixArtifactPermissionsForRootless() previously discarded repair-container stderr, so the warning hid the real failure context. The non-zero exit code comes from runAgentCommand() before cleanup; cleanup warnings do not override it. Improved in AWF (PR #6072, merged 2026-07-10): stderr from the repair container is now captured in the warning, and chroot-home cleanup noise is reduced by downgrading that log from warn to debug. Treat this mode as diagnostic-opacity around a pre-existing non-zero command exit, not cleanup-driven exit-code mutation. Compare agent-command logs with Command completed with exit code: <n>; on affected versions, repair warnings lack stderr context. On AWF including #6072, the warning includes stderr detail for root-cause triage. #6070, #6072
B12 getaddrinfo EAI_AGAIN <awmg-cli-proxy> or ENOTFOUND <awmg-cli-proxy> → awf-cli-proxy could not connect to the external DIFC proxy → The agent was never invoked on ARC/DinD, including when the topology peer is attached to awf-net ARC/DinD DNS configuration can cause either of two failures: Kubernetes search domains inherited by DinD, often with ndots:5, make Alpine/musl try expanded names for the single-label Docker peer before its direct lookup; alternatively, DinD cannot reach the cluster resolver at all. Diagnosed in AWF (PR #6328, merged 2026-07-17): detectDnsResolutionFailure() scans cli-proxy logs for EAI_AGAIN/ENOTFOUND, extracts the unresolved hostname, and augments the startup error. Fix in PR #9100: cli-proxy disables inherited search domains with dns_search: [], allowing direct Docker peer-name lookup. On older AWF, set dns_search: [] on the cli-proxy Compose service. If the cluster resolver itself is unreachable, address the DIFC proxy by IP or configure dockerd --dns <cluster-dns-ip>. Inspect cli-proxy's /etc/resolv.conf for Kubernetes search entries and options ndots:5. Compare docker run --rm alpine nslookup awmg-cli-proxy with docker run --rm alpine nslookup awmg-cli-proxy.: if the bare lookup fails but the trailing-dot lookup resolves and search/ndots are configured, match the search-domain cause; if both fail with timeout/SERVFAIL, check resolver reachability and peer attachment. #6326, #6328, #9100
B13 In --network-isolation mode, MCP tool calls to topology-attached peers (e.g. awmg-mcpg) or to a difcProxyHost configured independently of topologyAttach return connection errors or are logged as TCP_DENIED in Squid's access log; the agent can reach the DIFC proxy host during startup (B5/B12) but individual in-session MCP/HTTP calls to those same hosts fail or produce Squid 403 noise buildProxyEnvironment() assembled NO_PROXY from a static list of AWF-internal hosts but did not include (a) config.topologyAttach peer hostnames or (b) config.difcProxyHost. Proxy-aware clients (Node undici, curl) honoured HTTPS_PROXY for those hosts; Squid denied them because they are not in the domain allowlist. Fixed in AWF (PR #6189, merged 2026-07-20): config.topologyAttach peer hostnames are now added to NO_PROXY/no_proxy. Fixed in AWF (PR #6438, merged 2026-07-20): config.difcProxyHost (stripped of any :port suffix) is also added to NO_PROXY. Fixed in AWF (PR #6473, merged 2026-07-21): Topology peer hostnames and difcProxyHost are also auto-added to the Squid ACL allowlist (via resolveAllowedDomains()) so that proxy-aware clients that honour HTTP(S)_PROXY but ignore NO_PROXY are still permitted by Squid. The gating mirrors buildProxyEnvironment(): only active when networkIsolation is true and the runtime is compose-based (runtimeUsesComposeAgent()); microVM backends (sbx) are excluded. Upgrade AWF to the version including #6473. Fixed in AWF (PR #6658, merged 2026-07-27): Topology-peer rules are now added to the policy manifest, closing the audit attribution gap. The block report no longer incorrectly flags awmg-mcpg as denied. Upgrade AWF to the version including #6658. Additional fix (PR #6689): isInternalAwfDomain() in src/logs/internal-domain-filter.ts filters TCP_DENIED log entries whose destination is a 172.30.0.0/24 IP or single-label hostname (Docker container names always single-label) from both the runtime ⚠️ Firewall blocked N domain(s) warning and awf logs stats/summary output. Upgrade AWF to the version including #6689. Inspect Squid access.log for TCP_DENIED entries whose target host is a topology peer or the DIFC proxy host; inside the agent container check echo $NO_PROXY — the topology peer and DIFC proxy hostname should appear; inspect Squid access.log for TCP_DENIED entries for topology peer or DIFC proxy host even after upgrading to AWF with #6189/#6438 — present on AWF missing #6473; confirm topology hostname appears in generated squid.conf as acl allowed_domains dstdomain .<topology-host> after #6473; if the Firewall Issue Dispatcher block report flags a topology peer after #6473, treat it as a reporting false positive (manifest attribution gap; tracked in #6658) #6189, #6438, #6473, #6652, #6658, #6685, #6689
B14 After upgrading to AWF including the centralized mount-policy refactor (PR #6339), the Copilot CLI agent exits immediately (exit code 1, ~0.5 s, zero stdout/stderr); no network calls are made; the issue appears on Docker and gVisor runtimes but not sbx (which moves files aside rather than overlay-masking them) PR #6339 added ~/.copilot/config.json to the credential deny list with reason "Copilot CLI persisted auth token". This is incorrect: config.json contains only experiment flags, first-launch timestamps, and plugin metadata — no token. AWF masks denied files with a read-only /dev/null bind-mount overlay. The Copilot CLI reads and atomically rewrites config.json via temp-file + rename at startup; a rename cannot target a bind-mounted path and the overlay is read-only, so the CLI crashes silently. Fixed in AWF (PR #6374, merged 2026-07-18): ~/.copilot/config.json is removed from the credential deny list; the file is bind-mounted read-write as part of the .copilot tool subdir (pre-#6339 behaviour). Upgrade to the AWF version including #6374. Workaround (older AWF): pass --keep-containers and inspect the agent container for EROFS or rename errors; downgrade AWF past #6339 if unpatched. awf ... copilot agent --version — if it exits 1 with no output in ~0.5 s, confirm AWF version includes #6374; inspect agent container logs for EROFS/rename errors on ~/.copilot/config.json #6339, #6374
B15 AWF startup fails with --network-isolation is not yet supported with --enable-host-access; happens when localhost is in the domain allowlist because the gh-aw compiler auto-enables --enable-host-access and also emits network.isolation: true + topologyAttach In src/cli.ts, --network-isolation and --enable-host-access are mutually exclusive with no overlap support. The gh-aw compiler (v0.82.x+) does not detect the conflict: it emits isolation:true + topologyAttach unconditionally while also auto-enabling --enable-host-access when localhost is in the allowlist. The compiled workflow always fails at runtime. Fixed in AWF (PR #6657, merged 2026-07-28): The mutual-exclusion guard is removed; --enable-host-access now coexists with --network-isolation in topology mode. In topology mode the agent is on an internal Docker network with no host route, so --enable-host-access only drives Squid port ACLs and the host.docker.internal hosts-file entry — no incompatible iptables changes. allowHostServicePorts (iptables-based GitHub Actions services) is still suppressed unconditionally. Upgrade AWF to version including #6657. awf --network-isolation --enable-host-access ... — on AWF versions before #6657, reproduces immediately with not yet supported error; on patched AWF the combination is accepted without error; check compiler output for both flags being emitted together when localhost is in allowlist #6651, #6657
B16 AWF rejects a --mount entry with a non-absolute host path error; the mount spec was defined as ${TERRAFORM_CLI_PATH}/terraform:... (or similar env-var pattern) but the literal string ${TERRAFORM_CLI_PATH} reaches AWF's volume validator The gh-aw compiler wraps sandbox.agent.mounts specs containing ${} references in single quotes in the generated shell invocation. Bash single quotes prevent variable substitution, so the literal variable reference string (e.g. ${TERRAFORM_CLI_PATH}) reaches AWF's volume validation instead of the resolved absolute path. Fixed in AWF (PR #6655, merged 2026-07-27): expandEnvVarsInMount() in src/parsers/volume-parsers.ts now expands ${VAR_NAME} and $VAR_NAME patterns from process.env before path validation. If the variable is undefined, AWF emits a precise error (Environment variable is not set: ${VAR_NAME}) rather than a misleading path-absoluteness failure. Upgrade AWF to the version including #6655. Inspect the generated awf invocation in the compiled lock file — if mount specs appear in single quotes containing ${...}, the bug is present. Running awf ... --mount '${VAR}/path:/dest' on AWF versions before #6655 reproduces the rejection; on patched AWF the mount spec is accepted after variable expansion. #6649, #6655
B17 In --network-isolation/topology mode, when a workflow also brings up Tailscale as a later step, every Squid CONNECT to an allowlisted API host (e.g. api.githubcopilot.com) fails with 503 TCP_TUNNEL:HIER_NONE (server field -:-, DNS never resolved); agent reports 503 Service Unavailable retries and fails. Same workflow works on host-access mode or in isolation mode without Tailscale. When Tailscale starts, it can install a policy-routing rule covering 0.0.0.0/0 (exit node / accepted subnet route). Docker bridge traffic on the isolation network follows that route through the Tailscale tunnel. Host-specific DNS servers reachable only via the original network path (Azure DHCP DNS 168.63.129.16, Tailscale Magic DNS 100.100.100.100, link-local 169.254.x.x) become unreachable once Tailscale captures the route, so Squid's DNS queries black-hole and every CONNECT fails at resolution. Fixed in AWF (PR #6705, merged 2026-07-29): new isNonPortableDns(ip) / filterForNetworkIsolation(servers, logger) in src/dns-resolver.ts strip non-portable DNS servers (Azure DHCP DNS, Tailscale Magic DNS, link-local) before generating squid.conf when config.networkIsolation is true; falls back to 8.8.8.8/8.8.4.4 if all detected servers are non-portable. Users needing specific DNS in isolation mode can still override via --dns-servers. Upgrade to AWF version including #6705. Behavior changed again in AWF (PR #7499, merged 2026-08-19): the reachability-probe/filtering approach from #6705 and #7188 was removed entirely. filterForNetworkIsolation() now preserves all detected/explicit DNS servers unchanged in network-isolation mode — including Azure DHCP DNS, Tailscale Magic DNS, and link-local resolvers — since silently substituting public DNS breaks enterprise/cloud runners where the assigned resolver is private or policy-required and public DNS itself is blocked. Fallback to DEFAULT_DNS_SERVERS now only occurs when host DNS auto-detection finds no usable non-loopback resolver, not based on resolver classification. Upgrade AWF to include #7499. On AWF including #7499, a persisting 503 TCP_TUNNEL:HIER_NONE after a Tailscale-up step is no longer routed around automatically: use --dns-servers to point Squid at a resolver reachable through the captured route, or fix the Tailscale route/Magic DNS configuration. Squid access log shows CONNECT ... 503 TCP_TUNNEL:HIER_NONE with server field -:- for an allowlisted domain, specifically after a Tailscale-up step in isolation/topology mode; inspect audit/awf-resolved-config.json for networkIsolation: true + Tailscale in the same run; check whether the DNS servers in use are Azure DHCP (168.63.129.16) or Tailscale Magic DNS (100.100.100.100) #6704, #6705, #7185, #7188, #7495, #7499
B18 Azure CLI, Azure DevOps CLI, and Azure/ADO MCP servers fail inside the AWF agent sandbox even after OIDC login completes in runner setup steps and relevant Azure/Microsoft domains are allowlisted; ~/.azure config is invisible inside the sandbox, and ADO_MCP_AUTH_TOKEN/AZURE_CONFIG_DIR are not forwarded when --enable-api-proxy is active .azure was not in the whitelisted $HOME tool subdirs (home.toolSubdirs) in the canonical sandbox mount policy — it was actually in home.forbiddenSubdirs — so Azure CLI config written by runner setup steps never reaches the sandbox; separately, AZURE_CONFIG_DIR and ADO_MCP_AUTH_TOKEN were not in the always-forwarded host env var list, so the MCP bridge could not authenticate Fixed in AWF (PR #6690, merged 2026-07-28): .azure added to home.toolSubdirs and removed from home.forbiddenSubdirs in the canonical mount policy (applies to both compose and sbx runtimes); AZURE_CONFIG_DIR and ADO_MCP_AUTH_TOKEN added to the always-forwarded env vars in src/services/agent-environment/env-passthrough.ts. Authentication caveat: #6690 deliberately scrubs Azure token cache files (msal_token_cache*, accessTokens.json, service_principal_entries.json) — a pre-AWF az login session is not inherited by the sandbox; only account metadata (tenant IDs, subscription list) is mounted. The agent must perform OIDC re-login inside the sandbox using a fresh writable AZURE_CONFIG_DIR (e.g. az login --federated-token $ARM_OIDC_TOKEN --service-principal --username $ARM_CLIENT_ID --tenant $ARM_TENANT_ID); see #6686 for the working pattern. Upgrade AWF to version including #6690 to enable this in-sandbox re-login. Inside the agent sandbox: ls -la ~/.azure — presence confirms fix; echo $AZURE_CONFIG_DIR $ADO_MCP_AUTH_TOKEN — non-empty on patched AWF when set on the runner; az account show will still fail until in-sandbox OIDC re-login is performed #6686, #6690
B19 AWF fails on a primary error (e.g. topology-peer attach failure: Failed to connect container "awmg-mcpg" to network "awf-net": ... No such container), but the last prominent diagnostic consumers see is instead a large [WARN] Could not fix squid log permissions: Error: Command failed with exit code 1: chmod -R a+rX ... Operation not permitted dump from best-effort cleanup, obscuring the real root cause preserveDirectory() in src/artifact-preservation.ts ran a direct host-side chmod -R a+rX on Squid-owned log directories during cleanup and logged the full execa error object at warn level whenever it failed with EPERM/EACCES — even though this is an expected, benign outcome on rootless runners. The existing Docker-based rootless permission repair (fixArtifactPermissionsForRootless()) already classified these as benign debug output, but the direct chmod path did not reuse that classifier. Fixed in AWF (PR #6939, merged 2026-08-04): preserveDirectory() now reuses the same benign-permission-error classifier as fixArtifactPermissionsForRootless(), demoting expected EPERM/EACCES chmod cleanup failures to concise debug output while still surfacing unexpected cleanup failures as warnings. This preserves the primary startup error as the last prominent diagnostic. Upgrade AWF to include #6939. Trigger any primary startup failure on a rootless/self-hosted runner where Squid-owned log files exist; on unpatched AWF, cleanup logs a full chmod ... Operation not permitted execa error object after the primary failure; on patched AWF this is reduced to a debug-level note github/gh-aw#50384, #6934, #6939
B20 On ubuntu-latest/GitHub-hosted or plain self-hosted runners (no Tailscale/custom routing) in --network-isolation mode, awf-cli-proxy never becomes healthy: tcp-tunnel dials ENETUNREACH 172.17.0.1:18443 against host.docker.internal, exhausting the DIFC liveness probe and failing the workflow before the agent starts awf-net is internal: true with no outbound route. Squid and api-proxy are already dual-homed onto the external bridge (awf-ext), but cli-proxy was left attached only to awf-net even though it sets extra_hosts: {'host.docker.internal': 'host-gateway'} to reach the external DIFC proxy. Without a route out, Docker's host-gateway falls back to the default bridge gateway (172.17.0.1), unreachable from the isolated network Fixed in AWF (PR #7338, merged 2026-08-14): a credential-free cli-proxy-egress relay service is created and attached to awf-ext only when the DIFC proxy target is external; the credential-bearing cli-proxy stays on awf-net and reaches the relay there. No relay is created for attached sibling DIFC proxies already on awf-net. Loopback DIFC addresses are normalized to host.docker.internal. Upgrade AWF to include #7338. Inspect generated compose for a cli-proxy-egress service attached to awf-ext with no credential env vars and a read-only filesystem; confirm cli-proxy itself remains attached only to awf-net. #7063, #7066, #7335, #7338
B21 unable to create native thread / Cannot create worker GC thread from concurrent JVM builds (javac, Android manifest merger) inside the AWF agent container; /sys/fs/cgroup exposes no pids.max/pids.current, ulimit -u reports unlimited AWF hardcoded Docker's pids_limit to 1000 with no visibility or configurability, so JVM tools can't discover or size against the real ceiling Fixed in AWF (PR #7150, merged 2026-08-09): new --pids-limit <n> CLI flag (default 1000, matches prior behavior) with container.pidsLimit config-file support, plumbed through cli-options.ts → validators/log-and-limits.ts (parsePidsLimit) → build-config.ts → services/agent-service.ts. containers/agent/entrypoint.sh adds mount_host_cgroupfs() (best-effort) to bind-mount the container's delegated /sys/fs/cgroup read-only onto /host/sys/fs/cgroup so pids.max/pids.current are visible inside chroot. Inside agent: cat /sys/fs/cgroup/pids.max — presence confirms fix; raise ceiling with --pids-limit 4000 for concurrent JVM builds #7148, #7150
B22 Strict-security (--network-isolation, no --legacy-security) workflows cannot directly reach a GitHub Actions services: container that uses a raw protocol (for example, Postgres on 5432) via --enable-host-access / --allow-host-ports Strict topology omits the agent's host.docker.internal mapping and host-access iptables bypass. A raw PostgreSQL client cannot use Squid's HTTP CONNECT protocol, so allowing the port does not create a direct service route. gh-aw also does not derive service ports when it emits AWF flags. Known unresolved: preserving or compiler-emitting --enable-host-access / --allow-host-ports is insufficient. A supported strict-topology route for raw service protocols is required in AWF, and gh-aw must derive the required service ports. Until then, run direct service clients outside strict topology (for example with --legacy-security) or use a separately verified tunnel. In strict mode, getent hosts host.docker.internal is absent and psql -h host.docker.internal -p 5432 ... cannot connect; do not treat emitted host-access flags alone as a successful probe #7149, #7152, #7132
B23 Copilot-engine workflows fail with spawn /usr/local/bin/copilot ENOENT specifically when the runner's tool-cache already has copilot-cli installed (cache hit) Two gaps contributed to the symptom. The still-open upstream gap is that gh-aw's install_copilot_cli.sh activate_cached_copilot_bin() prepends the cached dir to PATH and returns early on cache hits (skipping the wrapper install to /usr/local/bin/copilot), while the compiler-emitted harness (copilot_harness.cjs) always spawns that hardcoded absolute path. Before #7245, AWF also mounted host /usr//usr/local read-only without creating the missing hardcoded entry inside the chroot, so the harness failed unless the host symlink already existed. Fixed on the AWF side (PR #7245, merged 2026-08-11): containers/agent/entrypoint.sh adds resolve_chroot_binary_path(), ensure_usr_local_bin_shims(), and prepare_usr_local_bin_overlay(), invoked after copy_dind_runner_binary. When AWF_ENSURE_USR_LOCAL_BIN=copilot is set (auto-set for Copilot runs in tool-specific-environment.ts), AWF resolves the real copilot binary from $GITHUB_PATH, AWF_HOST_PATH, staged bin dirs, or system dirs, and creates /usr/local/bin/copilot inside the chroot via a read-only symlink-farm overlay without modifying host /usr/local/bin. Upgrade AWF to include #7245. Older AWF only: before invoking awf, use the host workaround sudo ln -sf "$(command -v copilot)" /usr/local/bin/copilot; it is unnecessary on patched AWF. The upstream installer/harness mismatch remains open in #7130. PR #7151 documents the older behavior and workaround. Confirm gh-aw took the cache-hit path (GITHUB_PATH already set). On AWF including #7245, check entrypoint logs for ensure_usr_local_bin_shims / prepare_usr_local_bin_overlay; on older AWF, absence of host /usr/local/bin/copilot reproduces the ENOENT #7130, #7147, #7151, #7245
B24 Repeated EACCES retries from the harness reading ${RUNNER_TEMP}/gh-aw/mcp-config/mcp-servers.json (or other files under ${RUNNER_TEMP}/gh-aw) with no root-cause diagnostic explaining the native-root fallback ownership mismatch; occurs when AWF is invoked as native root (no sudo, so no SUDO_UID) AWF falls back to a default sandbox uid/gid (1000:1000) when it cannot recover the original host identity from SUDO_UID. Files under ${RUNNER_TEMP}/gh-aw created by the root-run harness stay root-owned, so the sandbox identity cannot read them once they are mounted read-only into the agent container Fixed in AWF (PR #7565, merged 2026-08-20): isNativeRootWithoutSudo() (src/host-identity.ts) detects root execution without SUDO_UID and logs an explicit warning identifying the fallback sandbox identity. repairRunnerTempGhAwOwnership() (src/config-writer.ts) recursively chowns ${RUNNER_TEMP}/gh-aw to the resolved sandbox uid/gid (via chown -h -P -R, never following symlinks) before Docker Compose generation and container launch. Upgrade AWF to include #7565. Run AWF as root directly (no sudo) and check for the warning Host process is running as root with no SUDO_UID; AWF will use sandbox identity ...; compare stat -c '%U:%G' "$RUNNER_TEMP/gh-aw/mcp-config/mcp-servers.json" before/after a run — it should show the sandbox uid:gid after the fix instead of root #7564, #7565
B25 On native-root runners (root, no SUDO_UID — e.g. AWS CodeBuild), the job silently exits 0 with no output even after PR #7565's ${RUNNER_TEMP}/gh-aw ownership repair; the agent never writes any file to the checkout PR #7565 repaired ${RUNNER_TEMP}/gh-aw ownership for native-root runners but left config.containerWorkDir (the checkout, --container-workdir "$GITHUB_WORKSPACE") root-owned while the agent runs as the fallback sandbox identity (uid 1000). The workspace is writable by mount but not by ownership, so every agent write silently fails and the job exits 0 — a false green costing a full agent session per run Fixed in AWF (PR #7599, merged 2026-08-21): src/config-writer.ts adds repairContainerWorkDirOwnership(config), called from writeConfigs() alongside repairRunnerTempGhAwOwnership(); applies to config.containerWorkDir only when isNativeRootWithoutSudo() is true. Also adds isDirectoryWritableByIdentity() as a post-repair preflight (checks owner/group/other mode bits directly, since access(2) always succeeds as root) — throws with an explicit chown -R <uid>:<gid> <workdir> suggestion instead of silently proceeding. Fixing ownership also resolves git's dubious ownership error without needing safe.directory. Upgrade AWF to include #7599. Run AWF as native root (no sudo) on a runner where the checkout is root-owned; on unpatched AWF the job exits 0 with no agent writes; on patched AWF, check for either successful chown-and-proceed, or the explicit failure message Host workspace is not writable by the sandbox identity (<uid>:<gid>): <path> #7593, #7599
B26 In --network-isolation mode, gh api .../actions/artifacts/{id}/zip or gh run download fails inside the agent sandbox with error connecting to productionresultssa*.blob.core.windows.net; download fails in ~350ms Two independent, non-interacting changes: (1) gh-aw-mcpg PR github/gh-aw-mcpg#10350 stopped auto-following the GitHub 302 redirect for artifact ZIP requests, so the gh CLI (running inside cli-proxy) must follow the Location header itself; (2) cli-proxy is intentionally isolated to awf-net only (no awf-ext egress, per #7066) so it has no route to productionresultssa*.blob.core.windows.net. Since mcpg now expects the client to follow the redirect but cli-proxy cannot reach the blob storage target directly, every ZIP download fails. Fixed in AWF (PR #7635, merged 2026-08-22): cli-proxy's HTTP(S) traffic is now routed through Squid; cli-proxy remains isolated from awf-ext; Squid ACL scopes *.blob.core.windows.net access to requests originating from cli-proxy's fixed IP only (http_access allow from_cli_proxy cli_proxy_artifact_storage), preserving blocklist precedence and SSL Bump behavior. Azure Blob storage is not added to the agent's general domain allowlist. Upgrade AWF to include #7635. Reproduce with gh run download <id> or gh api .../actions/artifacts/{id}/zip inside a --network-isolation agent sandbox; on unpatched AWF this fails within ~350ms with a blob-storage connection error; on patched AWF inspect Squid access.log for an ACL entry scoping *.blob.core.windows.net to the cli-proxy source IP github/gh-aw#54371, #7615, #7635
B27 Docker Compose refuses to start AWF containers with repeated warnings: a network with name awf-net exists but was not created for project "awf-<id>", or legacy iptables mode cannot find the bridge, on a persistent self-hosted runner A prior run can leave an orphaned awf-net with the expected subnet but without the com.docker.network.bridge.name option. Legacy iptables mode reused it and then failed to find the bridge Additional fix (PR #9130, merged 2026-09-28, fixes #9121): AWF recreates an orphaned awf-net with the expected bridge options when no containers are attached, leaves occupied networks untouched, and reports an actionable docker network rm command. Upgrade AWF to include #9130. docker network inspect awf-net --format '{{json .Options}} {{json .Containers}}' — an empty options map with no com.docker.network.bridge.name and an empty containers map identifies an unoccupied orphan; do not classify a network with attached containers as safe to recreate/remove github/gh-aw#56463, #7809, #7817, #9121, #9130
B28 Custom apiProxy targets pointing at an internal/corporate LLM router (--openai-api-target, --anthropic-api-target, etc.) fail TLS verification when the upstream endpoint's certificate chains to a private or corporate CA not present in the api-proxy sidecar's trust store The containers/api-proxy Node.js sidecar had no supported way to extend its trust store for custom upstream targets; the only workarounds were disabling certificate verification (insecure) or baking a custom CA into a rebuilt image Fixed in AWF (PR #7816, merged 2026-08-28): new apiProxy.caCert config field and --api-proxy-ca-cert <path> CLI flag bind-mount the host CA file read-only into the api-proxy container at /usr/local/share/ca-certificates/awf-upstream-ca.crt and set NODE_EXTRA_CA_CERTS to that path, so Node trusts the additional CA alongside its built-in roots without disabling verification. Upgrade AWF to include #7816. Inspect generated docker-compose.yml for the api-proxy service — a read-only bind mount to /usr/local/share/ca-certificates/awf-upstream-ca.crt and NODE_EXTRA_CA_CERTS env var confirm the fix is active; reproduce the failure pre-fix with awf --openai-api-target <internal-host> --allow-domains <internal-host> -- <cmd> against an endpoint using a private-CA certificate #7807, #7816
B29 codex-engine (and similar) workflows abort with report_incomplete: "context-rebuild circuit breaker tripped" after repeatedly failing to cd into the expected workspace path (for example /home/runner/work/<repo>/<repo>: No such file or directory) --container-workdir sets the agent's starting directory to a host-style absolute path, but that path was not guaranteed to be bind-mounted inside the chroot. If it was outside the workspace mount, /tmp, system mounts, $HOME tool mounts, or an explicit --mount, entrypoint.sh silently fell back to /, causing repeated context-rebuild retries Fixed in AWF (PR #8021, merged 2026-09-02): buildContainerWorkDirMounts() in src/services/agent-volumes/workspace-mounts.ts emits an explicit <workdir>:<workdir>:rw bind mount when the configured workdir is not already reachable inside the chroot. It refuses paths inside deliberately hidden roots and warns when the host directory does not exist. Upgrade AWF to include #8021. Inspect generated docker-compose.yml for a bind mount matching --container-workdir when no other mount covers it; check startup logs for a workdir-not-found warning instead of a silent / fallback; reproduce with a workdir outside every default mount #8015, #8021
B30 AWF-sandbox workflows fail before Squid starts (for example from a bad bind-mount spec), leaving no Squid access.log; awf logs summary/awf logs stats report only "no log sources found" AWF had no mechanism to preserve startup-phase failure detail when containers never produced Squid logs, so the underlying cause was lost Fixed in AWF (PR #8023, merged 2026-09-02): AWF writes a redacted awf-startup-error.json (timestamp, phase, failure message) into the proxy logs directory on startup abort; log discovery recognizes it via AWF_LOGS_DIR and preserved /tmp/squid-logs-* discovery, and stats/summary include the diagnostic. Upgrade AWF to include #8023. After a pre-egress failure, check the preserved proxy-logs directory for awf-startup-error.json; run awf logs summary and confirm it surfaces the startup diagnostic #8014, #8023

| B31 | Under sandbox.agent.runtime: docker-sudo-iptables, a toolchain version selected via a setup action (e.g. ruby/setup-ruby choosing Ruby 3.4.8) is shadowed by the system-installed version (e.g. /usr/bin/ruby 3.2.3) inside the AWF agent container, even though --env-all/AWF_HOST_PATH capture is active | docker-sudo-iptables invoked AWF via sudo -E awf ...; sudoers' secure_path could silently overwrite the runner's $GITHUB_PATH-augmented PATH before AWF observed process.env.PATH, losing hosted-toolcache bin-dir precedence | Fixed in gh-aw (PR github/gh-aw#58625, merged 2026-09-05): privileged AWF startup preserves the caller PATH. AWF's readGitHubPathEntries()/recoverHostPaths() recovery remains defense in depth, and merged PR #8173 adds regression coverage in src/services/agent-environment/host-path-recovery.test.ts; no AWF production-code change was needed. | Confirm the setup-action hosted-toolcache bin dir remains ahead of /usr/bin in AWF_HOST_PATH under docker-sudo-iptables; if it does not, this is a regression. | github/gh-aw#58458, github/gh-aw#58625, #8141, #8173 | | B32 | A repeated/persistent-runner workflow intermittently blocks allowlisted api.github.com/github.com traffic with 403 or DNS SERVFAIL, recurring across otherwise-healthy runs | Squid's default negative_dns_ttl is 1 minute, so one transient upstream SERVFAIL is negatively cached and replayed for up to 60 seconds even after DNS recovers | Fixed in AWF (PR #8171, merged 2026-09-05): generateDnsSection() emits negative_dns_ttl 1 seconds, dns_retransmit_interval 1 seconds, and dns_timeout 10 seconds. Upgrade AWF to include #8171. | Inspect generated squid.conf for negative_dns_ttl 1 seconds, dns_retransmit_interval 1 seconds, dns_timeout 10 seconds, and dns_nameservers; correlate Squid TCP_DENIED/SERVFAIL bursts with concurrent startup or resolver load on unpatched AWF | #8168, #8171 | | B33 | [DEBUG] Could not check Squid logs: EACCES ... access.log during diagnostics, or [DEBUG] Could not preserve squid logs: chmod ... Operation not permitted during artifact preservation, although logs are intact | Squid writes logs as UID 13; the previous shutdown-time repair only changed mode bits, ran after mid-run diagnostics, and never transferred ownership to the runner before artifact preservation | Fixed in AWF (PR #8251, merged 2026-09-07): reusable fixSquidLogPermissions() now chowns to the runner UID/GID via docker exec -e and chmods; runAgentCommand() repairs permissions before checkSquidLogs(), and preserved-log chmod is reported separately after rename. Additional fix (PR #8624, merged 2026-09-16): topology startup can recreate Squid mid-run (for example, for new extra_hosts); startup preflight repairs ownership of the explicit top-level access.log, audit.jsonl, and cache.log before Squid drops to the proxy user. The scoped, best-effort repair avoids both stale unwritable logs and a new startup-abort path. Upgrade AWF to include #8624. | Trigger a blocked-domain or upstream-error diagnostic and confirm access.log is readable; after the run, ls -la <preserved-squid-logs-dir> should show runner-UID ownership and no Could not check Squid logs/Could not preserve squid logs messages; recreate Squid through topology startup and confirm docker compose up -d succeeds with pre-existing log files | #8249, #8251, #8615, #8624 | | B34 | host.docker.internal or (host.docker.internal/redacted) appears in network.allowDomains, but an Ollama or other host-side service remains unreachable from inside AWF even when its host port is allowlisted | Host-gateway trigger detection did not recognize the canonical host.docker.internal keyword or its redacted audit form, so AWF omitted the host-gateway mapping needed to route the request to the runner host | Fixed in AWF (PR #8172, merged 2026-09-05): host-gateway detection recognizes host.docker.internal and (host.docker.internal/redacted) forms and emits the required host mapping. Upgrade AWF to include #8172. | Inspect generated docker-compose.yml for the host.docker.internal:host-gateway mapping and verify getent hosts host.docker.internal plus a request to the allowlisted host service from inside the agent | #8172 | | B35 | After a workflow using tools.cache-memory completes inside AWF, the host-side validateMemoryStep (gh-aw ≥ v0.89.21) fails with EACCES: permission denied writing /tmp/gh-aw/memory-validation/cache-default.ok | AWF remaps the agent UID/GID before running the command, while /tmp is bind-mounted read-write into the chroot. Files and directories created under /tmp/gh-aw can therefore have ownership and modes that prevent the original host runner identity from writing there after the container exits | Fixed in AWF (PR #9029, merged 2026-09-26; closes #9028): the agent entrypoint sets umask 0002 for the user command and runs root-owned relax_gh_aw_shared_permissions() after normal or signal-driven command exit to chmod -R g+w the /tmp/gh-aw tree without making it world-writable. Upgrade AWF to include #9029; prefer sudo awf over invoking AWF as native root so the UID/GID remap matches the host runner user. | After a cache-memory-enabled AWF run, inspect ls -la /tmp/gh-aw/memory-validation (or another host-side-touched /tmp/gh-aw path); on affected builds, differing owner and no group-write bit reproduce the mode. On patched builds, entries show g+w with a group matching the runner user's GID and validateMemoryStep succeeds | github/gh-aw#63472, #9028, #9029 |

Category C — GHES / GHEC / ghe.com

ID Signal Root cause Fix / flag Probe Citations
C1 DR-origin PAT authenticates, then /close fails with invalid API key Data-residency Copilot token exchange must target tenant-specific endpoints Route to copilot-api.<tenant>.ghe.com; verify PAT scope on the DR tenant Check GITHUB_SERVER_URL and api-proxy routing logs #1421
C2 Copilot auth fails on *.ghe.com at startup The API proxy did not use the tenant-specific Copilot endpoint Derive copilot-api.<tenant>.ghe.com from GITHUB_SERVER_URL Inspect api-proxy and Squid logs for the target host #1315
C3 400 bad request: Authorization header is badly formatted AWF v0.27.0 assembled GHES Copilot auth headers incorrectly Upgrade AWF to v0.27.2 or newer awf --version; inspect api-proxy logs for 400s #4867
C4 none of the git remotes correspond to the GH_HOST environment variable GH_HOST leaked as localhost:18443 instead of the real enterprise host Derive GH_HOST from GITHUB_SERVER_URL even with --env-all Print GH_HOST in the agent and compare to git remote -v #1452, #1460, #1492, #1499
C5 malformed version: from gh pr list --search, gh issue list --search, or gh search prs or issues in cli-proxy gh-proxy mode or later user steps In primary gh-proxy mode, GH_HOST=localhost:18443 makes gh treat the relay as GHES and request /api/v3/meta; the proxied api.github.com response lacks installed_version, while later steps still inherit the leaked host via $GITHUB_ENV Partially mitigated in AWF (PR #9189, merged 2026-09-29, fixes #9184): cli-proxy's Unix-socket gh-http-shim.js adds installed_version: 999.0.0 to /meta only when missing, wired through GH_CONFIG_DIR/config.yml http_unix_socket; GH_CONFIG_DIR is protected. This applies to primary mode only, not enclave mode. The full leak into later user steps still requires gh-aw changes curl http://localhost:18443/api/v3/meta should include installed_version in patched primary mode; inspect $GITHUB_ENV for the remaining leak #3937, #9184, #9189
C6 Safe-outputs post-processing talks to github.com instead of GHES gh-aw emitted GH_HOST to the wrong channel for later jobs Fix the compiler / environment propagation in gh-aw Inspect $GITHUB_OUTPUT and $GITHUB_ENV for GH_HOST #1460, #1566
C7 awf-cli-proxy DIFC-proxy liveness probe loops retrying; cli-proxy logs show diagnosis=unknown (AWF < v0.27.12) or diagnosis=reachable-but-api-error (HTTP NNN) with a *.ghe.com hint (AWF ≥ v0.27.12); AWF fails to start DIFC proxy is reachable but the forwarded gh api rate_limit call returns an HTTP error because the DIFC proxy is not enterprise-host-aware on data-residency *.ghe.com tenants Partially mitigated: upgrade to AWF ≥ v0.27.12 for a targeted *.ghe.com hint and HTTP status in cli-proxy logs; root cause (DIFC proxy enterprise-host awareness) is unresolved in companion projects (github/gh-aw-mcpg#8202, github/gh-aw#41911) Check GITHUB_SERVER_URL for *.ghe.com; inspect cli-proxy logs for diagnosis=unknown or reachable-but-api-error (HTTP NNN); confirm AWF ≥ v0.27.12 for the targeted hint #5615, #5616
C8 400 bad request: Authorization header is badly formatted on GHEC (*.ghe.com) runners when COPILOT_API_TARGET=api.business.githubcopilot.com; Copilot Business calls receive Bearer instead of required token prefix. Reproduced on AWF v0.27.13 and v0.27.16; or 400 persists even after upgrading past #5872 when COPILOT_PROVIDER_API_KEY=dummy-byok-key-for-offline-mode is set by gh-aw offline mode Two distinct root causes: (1) Pre-#5872: copilotTargetRequiresGitHubTokenPrefix() checked AWF_PLATFORM_TYPE guard first. On GHEC, AWF auto-injects AWF_PLATFORM_TYPE=ghec, which short-circuited to false before querying the GITHUB_TOKEN_PREFIX_COPILOT_TARGETS catalog. (2) Post-#5872 / #6237: gh-aw offline mode sets COPILOT_PROVIDER_API_KEY=dummy-byok-key-for-offline-mode as a sentinel; AWF before #6237 treated it as a real BYOK key, which took precedence over and suppressed the GitHub-token auth path Fixed in AWF (PR #5872): catalog endpoints (api.enterprise.githubcopilot.com, api.business.githubcopilot.com) are now checked first (always token); the platform-type guard now only affects the GHES heuristic for unknown targets. Upgrade to AWF version including #5872. Additional fix (PR #6237): treats dummy-byok-key-for-offline-mode as a non-credential sentinel (same class as AWF placeholder tokens), restoring the GitHub-token auth path on Business/Enterprise targets. awf --version; inspect api-proxy logs for 400 on api.business.githubcopilot.com; confirm AWF_PLATFORM_TYPE=ghec is set; check whether COPILOT_PROVIDER_API_KEY=dummy-byok-key-for-offline-mode is present #5871, #5872, #6237
C9 400 bad request: Authorization header is badly formatted specifically on the derived GHEC data-residency Copilot target copilot-api.<tenant>.ghe.com (distinct from C8's api.business.githubcopilot.com); receives token instead of required Bearer prefix copilotTargetRequiresGitHubTokenPrefix() incorrectly classified inferred copilot-api.*.ghe.com endpoints as targets requiring the token prefix Fixed in AWF (PR #8113): derived GHEC data-residency targets use the standard Bearer prefix; token remains reserved for Enterprise and Business Copilot targets. Upgrade AWF to include #8113. Inspect api-proxy logs for 400 on copilot-api.<tenant>.ghe.com; confirm GITHUB_SERVER_URL is *.ghe.com and the target is the derived Copilot endpoint (not api.business.githubcopilot.com) #6989, #6991, #8113
C10 Fine-grained GitHub PATs (github_pat_...) sent to Copilot Business, Enterprise, and canonical GHEC (copilot-api.<tenant>.ghe.com) targets used token instead of required Bearer; canonical GHEC targets also missed GitHub-hosted /models discovery handling, and a legacy isolation placeholder could override the real GitHub credential The api-proxy Copilot adapter selected the Authorization scheme by host without classifying credential kind, did not treat copilot-api.*.ghe.com as a GitHub-hosted catalog endpoint, and did not reject placeholder-token-for-credential-isolation as a non-credential sentinel Fixed in AWF (PR #8038, merged 2026-09-02): resolved credentials are classified before scheme selection (Bearer for github_pat_ tokens, preserving token for OAuth/classic PATs and BYOK behavior); canonical GHEC /models calls receive X-GitHub-Api-Version: 2026-07-01; integration identity resolves from COPILOT_INTEGRATION_ID, GITHUB_COPILOT_INTEGRATION_ID, then the default; the legacy placeholder is rejected for inference. Upgrade AWF to include #8038. Inspect api-proxy reflection diagnostics for credential_kind, selected_scheme, inference_credential_source, and integration_id_source; confirm fine-grained PATs receive Bearer and GHEC /models calls include the API-version header #8035, #8038

Category D — Alternative runtimes and adjacent gaps

ID Signal Root cause Status Probe Citations
D1 AWF does not start or isolation is ineffective under gVisor / Kata Alternative runtimes can differ from AWF's default runtime assumptions gVisor support landed in PR #6093 (merged 2026-07-10): use --container-runtime gvisor (maps to Docker runtime runsc), or raw --container-runtime runsc which now aliases to the same gVisor capability profile (static-DNS host mappings plus non-iptables networking). Kata Containers remain an unresolved research area. `docker info grep -i runtime; run AWF with --container-runtime gvisoror rawrunsc` on a build including #6401
D2 cli-proxy fails on IPv6-enabled runners The tunnel was bound only to 127.0.0.1 Fixed by dual-stack binding Check ss -tlnp inside awf-cli-proxy #4626
D3 --enable-dind still exists after DinD removal Legacy flag cleanup is incomplete Known unresolved cleanup item `awf --help grep enable-dind`
D4 Enterprise LLM gateway needs an injected auth header API proxy lacks a user extension point for that hop Known unresolved proposal No general probe; capture the required header flow in the report #4849
D5 gVisor install step exits with HTTP 404 / download failure; runsc binary not found; AWF exits with runtime 'gvisor' is not available or compose up fails immediately after --container-runtime gvisor The gh-aw-compiler-generated gVisor install step (or smoke-test lock files) pins a specific gVisor release tag (e.g. 20250623.0); if that artifact is no longer available from storage.googleapis.com, the download returns 404 Manually update the affected generated .lock.yml install step to a currently available release (for example, 20250707.0), or upgrade to a gh-aw release containing the corresponding DefaultGVisorVersion update and recompile. PR #6143 only patched this repository's four smoke-test lock files; it did not add an AWF configuration option or release an AWF fix. ARCH=$(uname -m); curl -sI https://storage.googleapis.com/gvisor/releases/release/20250623.0/${ARCH}/runsc — HTTP 404 confirms the pinned artifact is unavailable; replace the tag with an available release and retry #6143
D6 AWF exits with Docker sbx CLI not found. Install sbx to use --container-runtime sbx., Error: sbx is not available, or spawn sbx ENOENT when --container-runtime sbx is set; agent never starts; infrastructure containers (Squid, api-proxy) start normally but no microVM is created --container-runtime sbx was added in AWF PR #6101 (merged 2026-07-11); it requires the sbx CLI to be installed on the runner. AWF's isSbxAvailable() preflight fails if sbx is absent and the run aborts before the agent starts. Install the sbx CLI on the runner before invoking AWF; see #6117 for productionization and CI runner coverage. AWF injects HTTP_PROXY/HTTPS_PROXY to route HTTP(S) egress through Squid; standard --allow-domains whitelisting applies. command -v sbx && sbx version — absence of sbx confirms; docker info — confirm Docker daemon is accessible from the runner #6101, #6117
D7 Claude Code (Bun runtime) crashes with SIGSEGV / SIGABRT under --container-runtime gvisor; the harness retries multiple times, each attempt failing with exit code 1; no output or minimal output is produced gVisor restricts the W^X (write XOR execute) memory operations required by JIT compilers. Bun uses JavaScriptCore (JSC) with JIT enabled by default; JSC generates native code via JIT, triggering SIGSEGV/SIGABRT on gVisor's restricted syscall surface AWF (PR #6276) automatically sets BUN_JSC_useJIT=0 at runtime via buildToolEnvironment() when Claude runs under gVisor — no workflow change required. For older AWF builds without #6276, pass --env BUN_JSC_useJIT=0 (or set it as a job-level env var with --env-all) as a manual fallback. Inside the AWF gVisor agent container: echo $BUN_JSC_useJIT should be 0; a crash without this flag confirms the mode via SIGSEGV/SIGABRT signal in agent logs #6260, #6261, #6276
D8 MCP tool calls (safeoutputs, github) return 403 ERR_ACCESS_DENIED under --container-runtime gvisor or raw runsc; agent completes but never writes safe outputs; smoke tests fail at "Validate safe outputs were invoked"; direct /dev/tcp connections fail with No route to host gVisor's userspace netstack is isolated from the host network namespace. The iptables DNAT RETURN bypass rule (installed by awf-iptables-init) that normally lets the agent reach the MCP gateway (172.30.0.1:8080) directly never fires under gVisor. MCP requests follow HTTP_PROXY into Squid and receive 403 ERR_ACCESS_DENIED because the MCP gateway IP is not in the domain allowlist. Fixed in AWF (PR #6401): runtimeUsesIptables() returns false for gvisor and its raw runsc alias (plus sbx); awf-iptables-init is skipped for these runtimes (AWF_SKIP_IPTABLES_INIT=1); the MCP gateway (172.30.0.1) and host.docker.internal are added to NO_PROXY so proxy-aware MCP clients connect directly. Caveat: proxy-unaware tools using raw sockets (e.g. /dev/tcp) still get No route to host under gVisor — egress requires proxy-aware clients. Inspect Squid access log for 403 on 172.30.0.1; check NO_PROXY env inside the agent container — should include 172.30.0.1 when runtime is gvisor or raw runsc on patched AWF #6401, #6326
D9 On --container-runtime sbx, credential files (~/.aws/credentials, ~/.ssh/id_rsa, ~/.docker/config.json, ~/.kube/config, ~/.azure/, ~/.gnupg/, ~/.netrc, ~/.config/gh/hosts.yml, ~/.config/gcloud/, ~/.cargo/credentials.toml, ~/.claude/.credentials.json, ~/.gemini/oauth_creds.json) are visible to the agent inside the sbx microVM AWF before PR #6336 mounted the entire host $HOME (read-write) into the sbx microVM. Unlike compose mode (empty home volume + /dev/null overlays), sbx uses virtiofs passthrough where mounts are directory-granular and file-level overlays are not expressible. Fixed in AWF (PR #6336): sbx-manager.ts mounts only whitelisted tool/cache + agent-state subdirs (HOME_TOOL_SUBDIRS + .copilot/.gemini). For whitelisted dirs that contain nested credential files (e.g. ~/.config/gh/hosts.yml, ~/.cargo/credentials.toml), scrubHomeCredentials() moves them aside to .awf-sbx-cred-backup-<pid> before sbx create and restoreHomeCredentials() restores them after teardown. Workaround (older AWF): do not use --container-runtime sbx with real credentials in $HOME. ls $HOME/.aws $HOME/.ssh $HOME/.docker 2>/dev/null inside the sbx sandbox — if dirs are visible, fix is not applied; confirm AWF version includes #6336 #6336
D10 On --container-runtime sbx, Copilot CLI installed via install_copilot_cli.sh --rootless (redirects install to ~/.local/bin) is not found even though ~/.local is mounted into the microVM; copilot: command not found or ENOENT at agent startup sbx executes agent commands via bash -lc (login shell), whose profile initialization can reset PATH and discard --env PATH=.... The binary is present in the VM but unreachable by name unless PATH is fixed after login init. Fixed in AWF (PR #6407, merged 2026-07-19): sbx wraps the executed command with export PATH="$HOME/.local/bin${PATH:+:$PATH}" after login initialization (withLocalBinOnPath() in sbx-manager.ts). Upgrade to the AWF version including #6407. Workaround (older AWF): invoke $HOME/.local/bin/copilot directly, or prefix the agent command with export PATH="$HOME/.local/bin${PATH:+:$PATH}"; .... Inside the sbx agent: which copilot or ls ~/.local/bin/copilot confirms binary presence; on patched AWF, inspect the executed command wrapper in sbx logs and verify it prepends ~/.local/bin before invoking the agent command #6407
D11 Copilot CLI agent starts under --container-runtime gvisor but exits immediately with exit code 139 (SIGSEGV) or SIGABRT (exit 1); [copilot-harness] log shows all retry attempts crashing within ~90 ms (tokenCount=0, stdout=0B); the outer bash wrapper can also segfault, often before any model or tool call is issued. Affects sandbox.agent.runtime: gvisor (compose-managed gVisor) at ~8% failure rate; identical workloads on runc and sbx are unaffected. Node.js v22 (bundled in the Copilot CLI) can trigger a V8 native assertion (StringBytes::Encode ... Assertion failed: (written) == (u16size)) during ESM module translation under gVisor's userspace netstack; gVisor's mmap/madvise emulation can return unexpected buffer contents to V8's UTF-8 decode path. The SIGSEGV/exit 139 variant can also take down the outer bash wrapper so the Copilot harness cannot retry. Root cause (gVisor ↔ Node.js v22 incompatibility) is unresolved — tracked in #6558. Mitigated in AWF (PR #6514, merged 2026-07-23): runAgentCommand() detects gVisor runtime and automatically retries the agent container once (via docker start awf-agent) when the exit code is 134 (SIGABRT) or 139 (SIGSEGV) within the first 30 s (GVISOR_STARTUP_CRASH_WINDOW_MS = 30_000); MAX_GVISOR_AGENT_RETRIES = 1. Log reattachment uses docker logs --since <restart-time> -f awf-agent to avoid replaying the crashed attempt. Non-gVisor runtimes are unaffected. Root cause remains open; upgrade AWF to include #6514 to get the retry mitigation. Agent logs show signal=SIGABRT duration=0s stdout=0B on attempt 1; with #6514, a second docker start awf-agent log block appears and usually succeeds. Without the fix, all harness retries fail identically. Confirm containerRuntime: gvisor in the resolved docker-compose.redacted.yml. #6513, #6514, #6558
D12 Copilot workflow with model: auto (or no explicit top-level model, where gh-aw v0.84.1+ emits auto) fails before the agent starts under --container-runtime gvisor or sbx; harness logs awf-reflect: fetching (apiproxy/redacted) then request failed: fetch failed, followed by 400 ... Model "auto" has no AI credits pricing and no default pricing is configured; retries fail identically with zero tokens consumed. Same workflow succeeds under default (non-isolated) AWF runtime. Isolated agent runtimes (gVisor, sbx) may not reach (apiproxy/redacted), which the harness uses to pre-resolve auto to a concrete priced model. Without that resolution, api-proxy maxAiCredits pre-flight guard (checkUnknownModelRejection in guards/ai-credits-guard.js) had no pricing for literal auto and rejected the request with HTTP 400, even though Copilot resolves auto server-side and returns priced resolved-model metadata post-response. Fixed in AWF (PR #6811, merged 2026-08-01): checkUnknownModelRejection now allows provider === 'copilot' && model.toLowerCase() === 'auto' to pass pre-flight; AI-credit accounting then uses the response's resolved model. Non-Copilot providers still reject unresolved auto. Upgrade AWF to include #6811. Workaround (older AWF): pin a concrete priced model in workflow frontmatter (for example model: claude-sonnet-4.6) to avoid catalog-based auto resolution under isolated runtimes. Confirm sandbox.agent.runtime: gvisor or sbx in resolved AWF config; check api-proxy logs for 400 ... Model "auto" has no AI credits pricing alongside harness awf-reflect: request failed: fetch failed; verify whether apiProxy.maxAiCredits is enabled (guard only fires when enabled) #6810, #6811
D13 Under --container-runtime sbx with network.verifySbxEgress/--verify-sbx-egress enabled, AWF reports Direct sbx egress reached 1.1.1.1 without proxy environment variables and aborts before the agent starts, even though Squid is healthy (squid host.docker.internal:3128 -> 200). The long-lived sbx daemon was started or restarted without DOCKER_SANDBOXES_PROXY pointing to AWF's published Squid endpoint, allowing its sandboxed egress to bypass Squid. AWF correctly detects this and fails closed. A transient Squid-startup race can produce the same probe result before the sbx daemon's proxy chaining is ready. The caller/workflow owns daemon lifecycle: start or restart sbx daemon with DOCKER_SANDBOXES_PROXY=http://host.docker.internal:3128 (or the appropriate Squid gateway) before AWF creates the sandbox. PR #8252 fixes this repository's smoke-workflow post-processing final daemon restart. Additional fix (PR #8575, merged 2026-09-15): assertSbxEgressEnforced retries a bypass hit up to 3 times, 2 seconds apart, to absorb the transient Squid-startup race; probe execution failures still throw immediately and a persistent bypass still fails closed. See docs/sbx-integration.md for the general contract. Run sbx daemon status and inspect its environment for DOCKER_SANDBOXES_PROXY; restart it without that variable and confirm AWF's direct-egress probe reaches 1.1.1.1 or a denied destination without proxy variables. If the result clears after a brief retry while Squid starts, upgrade for #8575. #8250, #8252, #8568, #8575
D14 On --container-runtime cloud-hypervisor, the agent run aborts before the engine starts with "/run/awf-cloud-hypervisor/trusted-artifacts/run-<id>/cloud-hypervisor --version" exited with code undefined; the failure propagates and kills the whole engine run even on hosts where Cloud Hypervisor is simply unsupported (no KVM) or the staged artifact is incomplete The Cloud Hypervisor secure-launcher version probe only captured exitCode; when the child was signal-terminated, never spawned, or ran against a corrupted/partial staged artifact, exitCode came back null/undefined, leaving no way to distinguish those conditions from a genuine version-check failure Fixed in AWF (PR #8622, merged 2026-09-16): the version probe reports both exitCode and signalCode; staged trusted artifacts must be nonzero size before SHA-256 hashing/execution; CloudHypervisorUnsupportedHostError falls back to Docker on unsupported runner/KVM hosts. Artifact-trust, digest, version, and configuration errors remain fail-closed. Upgrade AWF to include #8622. Further fixed in AWF (PR #8801, merged 2026-09-20): the --version probe for both the Cloud Hypervisor and virtiofsd binaries is now retried up to 3 times (250 ms apart) before failing, since the staged artifact's digest is already trusted by this point and a single failure is more likely a transient exec hiccup. Deterministic errors (ENOENT/EACCES/EISDIR/ENOTDIR/ENOEXEC) skip the retry and fail closed immediately. Upgrade AWF to include #8801. Reproduce on a host without KVM/nested virtualization using --container-runtime cloud-hypervisor; on unpatched AWF the probe reports exited with code undefined; on patched AWF confirm signal termination names signalCode, or an unsupported host warns and falls back to Docker; on AWF including #8801, a single transient probe failure no longer aborts the run — check agent/preflight logs for Cloud Hypervisor version probe failed for "<path>" (attempt N/3) warnings indicating the retry engaged, versus an immediate failure for the fail-fast error codes above #8620, #8622, #8727, #8728, #8767, #8801
D15 On --container-runtime cloud-hypervisor, the run aborts before the engine starts with Unable to execute "/run/awf-cloud-hypervisor/trusted-artifacts/run-<id>/cloud-hypervisor --version"; verify the trusted artifact exists, is executable, and is complete: code=EACCES, even though the staged binary is root-owned mode 0555 and its digest verified AWF staged the executable trusted artifacts under /run/awf-cloud-hypervisor, a runtime-state tmpfs that hardened hosts (and some runner images) mount noexec; execve() then fails with EACCES regardless of file mode. EACCES is a deterministic probe failure, so preflight fails closed immediately and the whole engine run dies Fixed in AWF (PR #8835, merged 2026-09-21): executable snapshots are staged under /var/lib/awf-cloud-hypervisor/trusted-artifacts (an exec-capable root) while /run/awf-cloud-hypervisor keeps runtime state only; the EACCES message now also reports effective uid/gid/groups, the mount options of the staged binary, and per-path-component stat/ACL data. Follow-up (PR #8866): preflight refuses to stage at all when the trusted-artifact root resolves to a noexec mount, failing with the mount point and its options instead of an opaque post-copy EACCES. Upgrade AWF to include #8835. Run findmnt -no TARGET,OPTIONS /var/lib/awf-cloud-hypervisor (and /run) on the runner and look for noexec; on unpatched AWF the probe fails with code=EACCES against a /run/... path, on patched AWF the staged path is under /var/lib/awf-cloud-hypervisor/trusted-artifacts and a noexec root is reported as trusted artifact root "<path>" is on a mount that rejects execution #8827, #8834, #8835, #8866
D16 --container-runtime nvx or cloud-hypervisor intermittently aborts after Containers started successfully with Error: NVX confinement detected a process identity or thread-set race (or the Cloud Hypervisor equivalent), while identical runs alternate between pass and fail verifyNvxConfinement and verifyCloudHypervisorConfinement snapshotted the VMM's /proc/<pid>/task thread set, awaited cgroup/namespace reads, then required the thread set and per-thread start times to match byte-for-byte. Legitimate worker-thread creation or exit during guest initialization made this TOCTOU check falsely report a confinement violation Fixed in AWF (PRs #9016 and #9017, merged 2026-09-26, fix #9012): re-verification drops threads that exit between samples (ENOENT/ESRCH) and fully checks newly appeared or recycled TIDs (Tgid, uid/gid, groups, capabilities, no_new_privs, and seccomp) before acceptance. It still fails closed on changed process start time or executable (path and device/inode for NVX), a vanished main thread, policy-violating new threads, or more than 256 threads. Error messages identify the failed invariant and observed versus expected values. Upgrade AWF to include both PRs. Run the same NVX or Cloud Hypervisor workflow several times on KVM hardware; unpatched AWF intermittently aborts at microvm-run with the race error while identical runs pass, whereas patched AWF should consistently reach guest boot despite transient worker-thread counts #9012, #9016, #9017

Error-string quick lookup

Observable Likely mode
Docker socket not found at /var/run/docker.sock with Invalid container ID format A10
chroot: failed to run command 'capsh' or capsh not found A4
AWF chroot mode requires a glibc-based daemon host A4
one-shot-token.so ... __fprintf_chk: symbol not found A5
unknown shorthand flag: 'd' in -d from docker compose up -d on ARC/DinD A14
Rootless artifact permission repair failed for .../sandbox/firewall/logs on ARC/DinD A15
[WARN] Rootless artifact permission repair failed with each attempt taking ~30 s (timeout, not an instant error), followed by cascading EACCES: permission denied on chroot-home removal B10
[WARN] Rootless artifact permission repair failed ... (exit 1) with little/no stderr detail, plus cleanup warnings around chroot-home removal and Command completed with exit code: 1 B11
FATAL: http_port: IPv6 is not available or Bungled ... [::]:3128 B3
node: command not found on self-hosted or DinD A8 or B4
none of the git remotes correspond to the GH_HOST environment variable C4
malformed version: from gh --repo C5
malformed version: from gh pr list --search or gh search in cli-proxy gh-proxy mode C5
Streaming log write failure / read-only file system under ${RUNNER_TEMP}/gh-aw on arc-dind A28
400 bad request: Authorization header is badly formatted C3 (general GHES header assembly, AWF ≤ v0.27.1); also C8 if on *.ghe.com with COPILOT_API_TARGET=api.business.githubcopilot.com
400 bad request: Authorization header is badly formatted on derived copilot-api.*.ghe.com target specifically (not api.business.githubcopilot.com) C9 (derived GHEC Copilot API target missing GitHub token prefix; fixed in #6991)
context-rebuild circuit breaker tripped together with repeated failed cd into the expected workspace path B29 (container-workdir was not bind-mounted into the chroot; fixed in #8021)
awf logs summary reports "no log sources found" after a pre-egress startup failure with no Squid access.log B30 (check the preserved logs directory for awf-startup-error.json; fixed in #8023)
A setup-action-selected toolchain version (for example ruby/setup-ruby picking Ruby 3.4.8) is shadowed by the system-default version inside the AWF agent under sandbox.agent.runtime: docker-sudo-iptables B31 (sudoers secure_path strips $GITHUB_PATH-augmented PATH before AWF sees it; already mitigated on main via recoverHostPaths() reading $GITHUB_PATH directly — not a new bug, but a diagnostic to rule out before suspecting AWF)
Recurring intermittent 403/DNS SERVFAIL blocking an allowlisted domain (for example api.github.com) across otherwise-healthy runs, with no real forbidden-domain escape B32 (Squid's default 60-second negative_dns_ttl replayed a transient upstream SERVFAIL; fixed in #8171 by setting negative_dns_ttl 1 seconds, dns_retransmit_interval 1 seconds, and dns_timeout 10 seconds)
[DEBUG] Could not check Squid logs: EACCES ... access.log mid-run, or [DEBUG] Could not preserve squid logs: chmod ... Operation not permitted during artifact preservation, even though logs are intact B33 (the previous shutdown-time repair only changed mode bits and ran after diagnostics; fixed in #8251 with reusable pre-diagnostic chown+chmod repair)
docker compose up -d fails after AWF topology recreates Squid mid-run (new extra_hosts), with stale/unwritable Squid log files B33 update (topology-restart-triggered Squid log repair fixed in #8624)
Copilot on Business/Enterprise/GHEC uses the wrong Authorization scheme specifically with a fine-grained PAT (github_pat_...) C10 (credential-kind scheme selection, GHEC model discovery, and legacy placeholder handling fixed in #8038)
EACCES: permission denied, mkdir on a /tmp/gh-aw/... path before containers start (pre-flight) B8
EACCES: permission denied writing /tmp/gh-aw/memory-validation/cache-default.ok (or a similar /tmp/gh-aw path) from a host-side gh-aw step after the AWF agent exits B35 (UID-remapped agent left shared /tmp/gh-aw paths unwritable by the host runner; fixed in #9029)
No CA certificates were loaded from the system inside AWF chroot on RHEL/Fedora/Amazon Linux B9
Error: invalid key 'build-tools' with --image-tag build-tools=sha256:... A17
EACCES / write failures from XDG-respecting tools (Flutter, etc.) writing directly under /home/runner (for example /home/runner/tool_state) under runner.topology: arc-dind A18 (XDG_CONFIG_HOME captured stale root-owned home before HOME updated to writable arc-dind path; fixed in github/gh-aw#48658)
create_pull_request fails with No patch file found on ARC/DinD despite safeoutputs reporting a successful write A19 (safeoutputs /tmp/gh-aw mount not docker-host-path-prefix-translated; fixed in #6959)
ENOENT ... /host/usr/local/bin/copilot A8
mkdirat ... : read-only file system during chroot agent startup A12
getaddrinfo EAI_AGAIN <topology-peer> with awf-cli-proxy could not connect to the external DIFC proxy, but DinD nslookup succeeds once the peer is attached B5
getaddrinfo EAI_AGAIN awmg-cli-proxy / ENOTFOUND awmg-cli-proxy on ARC/DinD, and docker run --rm alpine nslookup awmg-cli-proxy fails B12
EACCES in upload-artifact after sudo: false (--network-isolation) AWF run B6
EACCES / unlink on path containing /tmp/awf-...-chroot-home/ during AWF cleanup (not in an upload-artifact step) B7
chroot: failed to run command '/bin/sh' on glibc daemon (not musl — confirmed by ldd --version) A13
getent passwd <UID> fails or HOME=/, USER=root in chroot A6
Bind-mounted /tmp/... files are missing inside DinD containers A1
diagnosis=unknown from awf-cli-proxy DIFC probe (proxy reachable, no connection error) with GITHUB_SERVER_URL=*.ghe.com, or diagnosis=reachable-but-api-error (HTTP NNN) C7
HTTP 404 / 404 Not Found downloading runsc from storage.googleapis.com during gVisor install D5
Docker sbx CLI not found. Install sbx to use --container-runtime sbx., Error: sbx is not available, or spawn sbx ENOENT with --container-runtime sbx D6
SIGSEGV / SIGABRT signal crash with Claude Code (Bun runtime) under --container-runtime gvisor; retries all fail D7 (JSC JIT incompatible with gVisor W^X restrictions; AWF ≥ #6276 auto-injects BUN_JSC_useJIT=0; for older AWF pass --env BUN_JSC_useJIT=0)
403 ERR_ACCESS_DENIED for MCP tool calls (safeoutputs, github) to 172.30.0.1/redacted under --container-runtime gvisor or raw runsc; agent finishes but safe-output validation fails D8 (gVisor userspace netstack: iptables DNAT bypass absent, MCP gateway not in NO_PROXY; fixed in #6401)
Credential files (~/.aws, ~/.ssh, ~/.docker/config.json, ~/.kube, ~/.config/gh, ~/.cargo/credentials.toml, etc.) visible inside sbx microVM under --container-runtime sbx D9 (entire $HOME mounted into sbx microVM virtiofs; fixed in #6336 with home-whitelist + scrubHomeCredentials())
Copilot CLI exits immediately (exit 1, ~0.5 s, zero stdout/stderr) after AWF upgrade; works in sbx but not Docker/gVisor B14 (.copilot/config.json masked as credential; fixed in #6374)
copilot: command not found inside sbx microVM when binary is at ~/.local/bin/copilot D10 (bash -lc login init resets injected PATH; fixed by command wrapper export PATH="$HOME/.local/bin${PATH:+:$PATH}" in #6407)
SIGABRT / signal=SIGABRT duration=0s stdout=0B for Copilot CLI all retries under --container-runtime gvisor; or exit code 139 with Segmentation fault on bash wrapper D11 (Node.js v22 V8 ESM decode assertion under gVisor; one-shot restart mitigation in #6514; underlying Node/gVisor incompatibility unresolved in #6558)
Model "auto" has no AI credits pricing and no default pricing is configured together with awf-reflect: request failed: fetch failed under --container-runtime gvisor or sbx D12 (isolated runtime cannot reach /reflect to pre-resolve auto; AI-credits guard rejected sentinel auto; fixed in #6811)
Direct sbx egress reached 1.1.1.1 without proxy environment variables (or a similar denied-destination reach) despite Squid healthchecks passing D13 (sbx daemon restarted without DOCKER_SANDBOXES_PROXY; workaround/CI fix in #8252)
Direct sbx egress reached 1.1.1.1 without proxy environment variables that clears after a brief retry while Squid starts D13 update (transient Squid-startup race, mitigated with bounded retry in #8575)
"cloud-hypervisor --version" exited with code undefined under --container-runtime cloud-hypervisor D14 (signal-terminated, missing-binary, or corrupted-artifact failure not distinguished from a real version-check failure; fixed in #8622; further fixed with a bounded retry for transient probe failures in #8801)
Unable to execute "<path>/cloud-hypervisor --version"; verify the trusted artifact exists, is executable, and is complete: code=EACCES under --container-runtime cloud-hypervisor D15 (trusted artifacts staged on a noexec/non-exec-capable root; fixed by staging under /var/lib/awf-cloud-hypervisor/trusted-artifacts in #8835, with a pre-staging noexec check in #8866)
trusted artifact root "<path>" is on a mount that rejects execution under --container-runtime cloud-hypervisor D15 (pre-staging detection of a noexec trusted-artifact root; remount it without noexec)
NVX confinement detected a process identity or thread-set race or the Cloud Hypervisor equivalent, intermittently after Containers started successfully D16 (false positive from legitimate VMM worker-thread churn during TOCTOU re-verification; fixed in #9016/#9017)
TCP_DENIED in Squid access log for topology peer or DIFC proxy host during agent run; MCP tool calls silently fail or return connection errors in network-isolation mode B13 (topology peers and difcProxyHost missing from NO_PROXY; fixed in #6189 and #6438; if block report still flags topology peer after #6473, treat as audit/policy-manifest reporting false positive tracked in #6652 / #6658 — runtime traffic is not blocked)
⚠️ Firewall blocked N domain(s) warning lists awmgmcpg or 172.30.0.x as a blocked domain on every run, even with no actual external blocks B13 (internal MCP gateway traffic counted by log aggregator as denied; fixed in #6689 with isInternalAwfDomain() filter)
--network-isolation is not yet supported with --enable-host-access B15 (compiler auto-emits both flags when localhost in allowlist + topology; fixed in #6657)
AWF rejects --mount with "host path must be absolute" (or similar) and the path visibly contains ${VAR_NAME} unexpanded B16 (single-quote wrapping by compiler prevents shell expansion of ${} in mount specs; fixed in #6655)
503 TCP_TUNNEL:HIER_NONE (server field -:-) on an allowlisted API host in network-isolation/topology mode, specifically after a Tailscale-up step B17 (Tailscale policy-routing captures the default route, making host-specific DNS servers unreachable; the probe/filter approach from #6705/#7188 was removed in #7499 — filterForNetworkIsolation() now preserves detected resolvers, so a persisting 503 needs a --dns-servers override or Tailscale route/Magic DNS configuration instead of automatic filtering)
DNS resolution fails in --network-isolation mode specifically on GKE/ARC using NodeLocal DNSCache (resolver 169.254.20.10), even though the resolver is reachable B17 update (#7188 made filtering reachability-probed; #7499 removed filtering entirely so all detected resolvers, including 169.254.20.10, are preserved — upgrade AWF)
Azure CLI / ADO MCP auth failures with ~/.azure missing inside AWF sandbox, or AZURE_CONFIG_DIR/ADO_MCP_AUTH_TOKEN empty inside agent despite being set on the runner B18 (.azure was in home.forbiddenSubdirs and auth env vars were not forwarded; fixed in #6690 — note: pre-AWF az login is not inherited; agent must perform OIDC re-login inside the sandbox, see #6686)
[WARN] Could not fix squid log permissions: ... Operation not permitted appears as the last log after an unrelated primary AWF startup failure B19 (benign rootless chmod cleanup error not demoted, obscuring the real failure; fixed in #6939)
ENETUNREACH ... :18443 (or default bridge gateway IP) from awf-cli-proxy in --network-isolation mode B20 (fixed in #7338 via credential-free cli-proxy-egress relay attached to awf-ext; upgrade AWF)
unable to create native thread / Cannot create worker GC thread from concurrent JVM builds (javac, Android manifest merger) inside the AWF agent; /sys/fs/cgroup shows no pids.max/pids.current B21 (Docker pids_limit hardcoded to 1000 with no visibility/configurability; fixed in #7150 with --pids-limit/container.pidsLimit plus mount_host_cgroupfs())
Strict-security workflow cannot reach a GitHub Actions services: raw-protocol port (e.g. Postgres 5432) B22 (known unresolved: strict topology needs an AWF direct service route and gh-aw must derive service ports; emitted host-access flags alone are insufficient)
EACCES retries reading ${RUNNER_TEMP}/gh-aw/mcp-config/mcp-servers.json (or similar ${RUNNER_TEMP}/gh-aw paths) when AWF was invoked as native root (no sudo) B24 (root fallback to sandbox uid 1000 leaves root-owned config unreadable; fixed in #7565 with repairRunnerTempGhAwOwnership())
Job exits 0 with no agent output/writes on a native-root runner (no sudo), even after ${RUNNER_TEMP}/gh-aw ownership is fixed B25 (checkout/container-workdir left root-owned while agent runs as fallback sandbox uid; fixed in #7599)
Host workspace is not writable by the sandbox identity (<uid>:<gid>): <path> B25
error connecting to productionresultssa*.blob.core.windows.net from gh run download/artifact ZIP fetch in --network-isolation mode B26 (mcpg stopped auto-following the artifact redirect; cli-proxy had no route to Azure blob storage; fixed in #7635 with scoped Squid ACL keyed to cli-proxy's fixed IP)
invalid CapDrop: capability not supported by your kernel or not available in the current environment A22 (host capability bounding set trimmed below AWF's hardcoded cap_drop list, e.g. Talos; fixed in #7795)
error mounting "/dev/null" to .../home/.npmrc: create mountpoint ...: read-only file system (or .docker/config.json, .composer/auth.json) on arc-dind with --docker-host-path-prefix set A23 (filterAgentVolumesForSysroot() compared prefixed mount sources against unprefixed workDir/effectiveHome, so the bogus chroot-home mount wasn't dropped; fixed in #7998)
error mounting "/dev/null" to .../.npmrc: create mountpoint ...: read-only file system on arc-dind persisting even after upgrading past #7998 (A23's fix), where the credential mountpoint is missing under a declared-rw home bind backed by a genuinely read-only directory A24 (pruneUnmountableCredentialOverlays only checked declared bind mode, never real filesystem writability for rw-declared covers; fixed in #8086)
docker network connect --alias <name> awf-net <service_container> is needed for raw-protocol GitHub Actions services: containers under runner.topology: arc-dind A25 (service container must join awf-net for direct protocol access while the agent stays on the isolated awf-net; fixed in docs/workaround in #8085)
503 HIER_NONE on every Squid CONNECT on runner.topology: arc-dind deployed on OpenShift/ARO A26 (awf-net default 172.30.0.0/24 collides with OpenShift's service CIDR and CoreDNS ClusterIP 172.30.0.10; fixed in #8398 with --network-subnet)
Safe-output field (e.g. PR body) comes back empty on runner.topology: arc-dind, with no write error, and the payload was staged under /tmp/gh-aw/agent A27 (daemon-invisible staging path when RUNNER_TEMP is set; fixed in #8933)
threat-detect (or another non-engine CLI tool invoked inside the AWF sandbox) exits 127, or its --output path is unwritable/unreadable, on runner.topology: arc-dind Not an AWF defect — gh-aw codegen only auto-stages the invoking engine binary. Stage the tool onto a daemon-visible path and add an explicit :rw mount for its output directory even under an otherwise :ro parent mount (--mount /tmp/gh-aw/<tool>:/tmp/gh-aw/<tool>:rw); see docs/arc-dind.md "Staging additional CLI tools" (#8457, tracks github/gh-aw#59935)
host.docker.internal or (host.docker.internal/redacted) appears in network.allowDomains but the host service still cannot be reached from inside AWF B34 (host.docker.internal keyword handling was missing; fixed in #8172)
a network with name awf-net exists but was not created for project B27 (orphaned fixed-name awf-net from a prior run on a persistent self-hosted runner; fixed in #7817)
TLS/certificate verification failure from api-proxy against a custom --openai-api-target/--anthropic-api-target internal endpoint using a private/corporate CA B28 (api-proxy sidecar had no custom CA trust extension point; fixed in #7816 with apiProxy.caCert/--api-proxy-ca-cert)
spawn /usr/local/bin/copilot ENOENT specifically on a tool-cache hit (GITHUB_PATH already set by the installer) B23 (gh-aw's activate_cached_copilot_bin() skips the /usr/local/bin/copilot wrapper on cache hits while the compiler harness spawns that hardcoded path; AWF-side fixed via ensure_usr_local_bin_shims()/prepare_usr_local_bin_overlay() in #7245; durable upstream fix still tracked in #7130, open)
runc mountpoint creation failure for /dev/null credential overlays under /host$HOME on runner.topology: arc-dind A20
mkdir -p .../.m2 failing under set -e in agent entrypoint on arc-dind A20
mkdirat ... : read-only file system at agent container startup while a filesystem.allowWrite policy is active (not the chroot.binariesSourcePath-specific A12 case) A21
token-usage.jsonl not found / missing gh-aw.aic or gen_ai.usage.* telemetry on arc-dind A29

Known unresolved items

Flag these explicitly instead of implying there is a complete fix:

  • D1 / #3264 — Kata Containers compatibility research (gVisor resolved in #6093; raw runsc aliases to the same profile in #6401)
  • D3 / #1727 — lingering --enable-dind cleanup
  • D4 / #4849 — enterprise header injection extension point
  • C5 / #3937 — full GH_HOST leak fix still requires gh-aw changes
  • A28 / #9183 — gh-aw compiler must relocate Pi's streaming log to sandbox/agent
  • A29 / #9352 — gh-aw parse_token_usage.cjs must consume AWF_TOKEN_USAGE_LOG
  • C7 / #5615 — DIFC proxy enterprise-host awareness for *.ghe.com data-residency (root cause unresolved; tracked in github/gh-aw-mcpg#8202 and github/gh-aw#41911)
  • D11 / #6558 — gVisor + Node.js v22 V8 ESM startup crash root cause (SIGABRT StringBytes::Encode); one-shot retry mitigates (~8% failure rate) but does not prevent the crash
  • B22 / #7132 — strict topology lacks a direct raw-protocol services: route, and gh-aw does not derive service ports; compiler-emitted host-access flags alone are insufficient
  • B23 / #7130 — gh-aw's install_copilot_cli.sh/copilot_harness.cjs mismatch on tool-cache hits (cache-hit path skips the /usr/local/bin/copilot wrapper while the harness always spawns that hardcoded path) remains open upstream; AWF's own contribution to the symptom is fixed in #7245 (chroot /usr/local/bin/copilot overlay), so this item now tracks only the gh-aw-side installer/harness mismatch