Skip to content

Fix composer e2e flake on packagist connect timeouts - #1037

Merged
Mikola Lysenko (mikolalysenko) merged 1 commit into
mainfrom
ci-janitor/composer-packagist-retry
Oct 7, 2026
Merged

Mikola Lysenko (mikolalysenko) merged 1 commit into
mainfrom
ci-janitor/composer-packagist-retry

Conversation

@mikolalysenko

@mikolalysenko Mikola Lysenko (mikolalysenko) commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

The composer capstones (e2e_vendor_composer_build, e2e_redirect_composer_build) set up their fixture by running a real composer update against repo.packagist.org, once. Every CI and composer-compatibility.yml leg sets SOCKET_PATCH_COMPOSER_E2E_REQUIRED=1, so a single 10 s curl connect timeout to packagist turns skip()'s assert into a hard failure. This happens on PRs that never touch composer code:

  • job 112776608105 (2026-10-07, dependabot/cargo/rustls-0.23.45), composer 2.2.30 / php 8.3 / windows-latest, test composer_vendor_redownloads_modified_copy:
    panicked at crates\socket-patch-cli\tests\composer_e2e_common\mod.rs:91:5:
    e2e_vendor_composer_build(redownload-copy): SOCKET_PATCH_COMPOSER_E2E_REQUIRED is set but `composer update` failed (registry unreachable?):
    curl error 28 while downloading https://repo.packagist.org/packages.json: Connection timed out after 10001 milliseconds
    
  • Job 111773795177 (2026-10-05, agent/fix-yarn-berry-catalog-resolutions) failed on the same leg with the same signature.
  • Over 2026-10-03..07 the Composer workflow had 12 failed Windows jobs across 6 unrelated branches; the two above are the ones whose logs were confirmed as packagist curl-28. Each one costs a full re-run of the Composer workflow (~3 OS × 6 composer versions).

Root cause

The fixture setup is the only step in these suites that reaches the public internet (packagist, plus the GitHub zipball). composer_e2e_common::composer ran it once with no handling for transient transport errors. Composer 2.2 does not retry a curl connect timeout on packages.json. The hosted Windows runners intermittently fail to connect to packagist within composer's 10 s connect timeout.

Fix

composer_e2e_common::composer now re-runs a failed composer command only when its output has a curl error <code> while downloading <url> line where:

  • the code is connect-level (6 resolve, 7 connect, 28 timeout, 35 TLS connect, 52 empty reply, 56 recv reset), and
  • the URL host is not loopback.

It makes up to 3 attempts, with 5 s and 10 s backoff. Nothing else changes:

  • HTTP status errors, checksum refusals, resolution errors and anything against 127.0.0.1/localhost (the wiremock patch server) still fail on the first attempt. So composer_redirect_tampered_archive_fails_checksum_verification and the no-source-fallback assertions are untouched.
  • A sustained outage still fails the required leg after the third attempt.

Two self-tests pin the classifier: the exact CI line and its wrapped [TransportException] box form are transient; loopback, curl 60, HTTP 404, checksum and resolver errors are not.

Proof

All local runs use real composer 2 and real packagist.

  • Before (origin/main). I injected a one-shot blackhole proxy (HTTPS_PROXY=http://10.255.255.1:3128 for the first composer update only, through a SOCKET_PATCH_PHP_BIN wrapper). composer_vendor_redownloads_modified_copy fails with the CI signature: panicked at .../composer_e2e_common/mod.rs:91:5, curl error 28 while downloading https://repo.packagist.org/packages.json.
  • After, same injection. The run logs composer update: transport failure (curl error 28 ...); retrying (1/3) and then passes (21.9 s).
  • After, sustained blackhole. It retries 1/3 and 2/3, then fails as before (46 s). A real outage is still reported.
  • After, normal network. All 11 host capstones pass with --ignored and SOCKET_PATCH_COMPOSER_E2E_REQUIRED=1: 5 hosted, including the tampered-archive checksum test, and 6 vendored. The non-ignored tests in both binaries also pass, 17 and 11, including the 2 new self-tests.
  • rustfmt --check is clean on the touched file.
  • Clippy on the three binaries that include the module reports nothing in it. The only clippy hits are pre-existing ones in prebuilt_common/mod.rs, which this PR doesn't touch; CI's clippy job doesn't lint test targets.

Where tests run

No test was removed, moved or weakened, and no workflow changed. The composer capstones keep running in ci.yml (e2e matrix, composer 1/2/2.2) and in composer-compatibility.yml, as before.

🤖 Generated with Claude Code

https://claude.ai/code/session_011j1YKsT7s3GopU6brHft2B


Generated by Claude Code


Note

Low Risk
Test-only helper changes with bounded retries; production code and security-sensitive assertions (loopback/checksum) are unchanged.

Overview
Reduces flaky CI failures when Composer e2e fixture setup hits transient network errors talking to Packagist/GitHub.

composer() now retries up to 3 times (5s/10s backoff) only when stderr/stdout contains a connect-level curl error (codes 6, 7, 28, 35, 52, 56) for a non-loopback URL. Loopback wiremock failures, HTTP status errors, checksum failures, and resolver errors still fail immediately on the first run.

Adds remote_transport_failure to classify output, splits the old single-shot runner into composer_once, documents the behavior for SOCKET_PATCH_COMPOSER_E2E_REQUIRED, and adds two self-tests that lock in transient vs non-transient cases (including the Windows packagist curl-28 signature).

Reviewed by Cursor Bugbot for commit 8c23705. Configure here.


Generated by Claude Code

The composer capstones' fixture setup runs a real `composer update`
against repo.packagist.org once. With SOCKET_PATCH_COMPOSER_E2E_REQUIRED
set (every CI and composer-compatibility leg), a single 10 s curl
connect timeout to packagist fails the leg through skip()'s assert, on
PRs that never touch composer code (e.g. the rustls bump, job
112776608105: "curl error 28 while downloading
https://repo.packagist.org/packages.json").

Re-run a failed composer command, up to 3 attempts with 5 s/10 s
backoff, only when its output shows a connect-level curl failure
(codes 6, 7, 28, 35, 52, 56) against a non-loopback host. HTTP status
errors, checksum refusals and anything against the loopback wiremock
patch server still fail at once, so the tampered-archive and no-source
fallback assertions are unchanged, and a sustained outage still fails
the required leg.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011j1YKsT7s3GopU6brHft2B
@mikolalysenko Mikola Lysenko (mikolalysenko) added the ci-janitor Opened by the CI janitor routine (flakes, redundant tests, CI perf) label Oct 7, 2026
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

bugbot run


Generated by Claude Code

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 8c23705. Configure here.

@mikolalysenko
Mikola Lysenko (mikolalysenko) merged commit d47eab3 into main Oct 7, 2026
201 of 212 checks passed
@mikolalysenko
Mikola Lysenko (mikolalysenko) deleted the ci-janitor/composer-packagist-retry branch October 7, 2026 16:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-janitor Opened by the CI janitor routine (flakes, redundant tests, CI perf)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants