Skip to content

Fix PDM matrix flakes on PyPI/patch API transport blips - #860

Merged
Mikola Lysenko (mikolalysenko) merged 3 commits into
mainfrom
ci-janitor/pdm-transport-retry
Oct 5, 2026
Merged

Mikola Lysenko (mikolalysenko) merged 3 commits into
mainfrom
ci-janitor/pdm-transport-retry

Conversation

@mikolalysenko

@mikolalysenko Mikola Lysenko (mikolalysenko) commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

pdm-compatibility.yml is the flakiest workflow right now. From 2026-09-28 to 2026-10-05 it had 29 failed runs (ci.yml had almost no flakes in the same window: 133 runs, and nearly every red one was the PR's own bug). For the 25 failures in the native job, the log shows one random FAIL cell in an otherwise green 33–47-cell matrix, on PRs that don't touch PDM (gem, npm, yarn, uv fixes). The cell is different every time:

failed check(s) count modes / versions
rescanIdempotent 7 hosted/vendored/agent, 0.12.3 mac ×2, 2.8.2, 2.10.4, 2.22.4 ×2, 2.29.2 mac
rescanAfterRelockApplies + rollbackAfterRelockPristine 6 hosted, 0.12.3, 2.9.3 ×3, 2.22.4
appliedExactlyOne 6 agent/vendored, 0.12.3, 2.9.3, 2.20.1 ×2, 2.6.1, 2.29.2
refusalCodeReported 3 2.1.5, 2.3.4, 2.6.1
rescanAfterRelockApplies + rescanReusesWheel 2 vendored, 2.17.3, 2.20.1
rollback trio + VEX trio (same run) 1 2.8.2 mac, extras

Example runs: 37309092656 (2.17.3 space-unicode vendored), 37308344193 (2.29.2 mac extras hosted), 37306321055, 37279240870 (4 legs at once), 37225409452. Each red run costs a re-run of the ~19-leg matrix and blocks the PR's checks.

Root cause

Every failing check judges an operation that goes to production PyPI or the public patch API: a CLI scan, rollback or vex, or a pdm sync/install. Run retries a command once, after 5 s, and only when it exits non-zero. On an exhausted patch API fetch (patch_details_failed / api_batch_failed, a 503, or a 429 its retry loop gave up on) the CLI usually exits zero and reports the failure in its JSON envelope. The cell then just fails a later check. The Poetry harness had exactly this pattern and was fixed in #596. The PDM harness never got that fix. The job log also never says why a cell failed (the notes live only in the artifact), which is why these were never diagnosed.

Fix

Port backtest-poetry.py's evidence-gated, case-level retry to scripts/backtest-pdm.py:

  • record_check(..., operation=): each network-bound check names the command it judges. A failed check records transport evidence only from that command: its output if it failed, otherwise only the CLI's JSON error records (error, failed/skipped events, api_batch_failed / patch_details_failed). The patterns are the same TRANSPORT_FAILURE set Poetry uses, plus httpx errors (PDM ≥ 2 uses httpx). Lines saying "retrying" (a retry that recovered) are ignored.
  • retry_transport: re-runs a case from a fresh directory, at most 3 attempts, with 10 s/20 s backoff, only when every failed check has transport evidence. An exception is retried only if its text is a transport diagnostic and no check had already failed functionally. A functional failure is never retried.
  • Failed attempts' logs go under native-pdm/attempts/<version>-<shape>-<mode>/<n>/, now uploaded. The final row lists them in transportRetries.
  • A FAIL cell now prints each failed check's recorded note in the job log, so the next real failure names its cause without downloading the artifact.

The existing one-shot Run(retry=True) stays as is. Nothing is skipped, ignored or loosened: a persistent outage still turns the cell red after 3 attempts.

Proof

  • New PdmTransportRetryTests (7 tests, offline) cover:

    • a zero-exit CLI patch_details_failed 503 is retried from a fresh case, and the evidence is kept and saved;
    • non-zero CLI request error / exhausted 429 / httpx DNS errors are retried;
    • refusals, recovered "retrying" warnings, transport text outside error records, and checks with no operation are not retried;
    • one unexplained failed check blocks the retry;
    • a transport exception after a functional failure is not retried;
    • a persistent failure stays red after 3 attempts.

    python3 -B -m unittest discover -s scripts/tests → 237 tests OK.

  • End to end with the real harness and CLI (--versions 2.17.3 --shapes direct --modes vendored agent). The sandbox blocks patches-api.socket.dev, so the real CLI returns "error": "Network error: error sending request for url (https://patches-api.socket.dev/patch/batch) …". Both cases were classified as transport, re-run from fresh case dirs (attempts 1/3, 2/3 kept under attempts/), and stayed FAIL after the third attempt, with appliedExactlyOne: {"applied": 0, … "status": "error"} printed in the log.

  • ruff check --select F shows the same single pre-existing warning as main. The workflow YAML parses.

Where tests run

Nothing removed or moved. Every cell still runs in pdm-compatibility.yml on the same triggers. The only workflow change adds native-pdm/attempts/** to the results artifact. Required-check names are unchanged.

Also in this PR: 80538db ports #851's test-only fix for two commands::vex_consumed tests that main has failed since #605. The port is needed because those tests make coverage/test/test-release red on every PR. It becomes a no-op once #851 lands. 6b0302a addresses Bugbot's finding: the rollback and VEX checks are now judged by the commands they check.

🤖 Generated with Claude Code

https://claude.ai/code/session_01N8YeCUhdg2tKdqyyY7z3sV


Generated by Claude Code


Note

Low Risk
Changes CI/backtest harness behavior only; production CLI logic is untouched aside from test adjustments, with retries limited to transport-classified failures.

Overview
Adds evidence-gated transport retries to the PDM compatibility backtest so flaky PyPI/patch API blips no longer fail random matrix cells on an otherwise green run.

scripts/backtest-pdm.py now mirrors the Poetry harness: failed checks tie to the command they judge (record_check + operation), terminal transport patterns are detected (including zero-exit CLI JSON warnings like patch_details_failed), and retry_transport re-runs a case from a clean directory up to three times only when every failed check has transport evidence—functional failures are never retried. Failed attempts are archived under native-pdm/attempts/…, surfaced in transportRetries, uploaded by CI, and FAIL rows print per-check notes in the job log. Docs describe this behavior.

vex_consumed.rs tests are updated for post-#605 npm resolver behavior (aliases/peers found without feeding the full installed set) while still asserting alias expansion and the resolver’s copy set.

New PdmTransportRetryTests cover retry vs no-retry cases offline.

Reviewed by Cursor Bugbot for commit 80538db. Configure here.


Generated by Claude Code

The PDM matrix runs against production PyPI and the public patch API.
Over the last 7 days 25 pdm-compatibility runs failed on one random
cell each, on unrelated PRs. The version, OS, shape, mode and check
differed every time (rescanIdempotent, appliedExactlyOne,
rescanAfterRelockApplies, ...). Each check judges a CLI scan, install
or rollback. `Run` retries a command once, and only on a non-zero
exit. The CLI usually reports an exhausted patch API fetch in its
JSON while exiting zero, so the cell just fails a later check.

Port backtest-poetry.py's case-level retry (#596). A case is re-run
from a fresh directory, at most three attempts, only when every
failed check recorded transport evidence from the operation it
judged. Evidence is a failed command's request error, PyPI give-up,
patch API 5xx or exhausted 429, or the same in the CLI's JSON error
records. Functional failures are never retried, even when a later
step raises a transport error. Failed attempts' logs go under
attempts/ and are uploaded. A failing case now prints its failed
checks' notes, since the job log alone never said why.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N8YeCUhdg2tKdqyyY7z3sV
@mikolalysenko Mikola Lysenko (mikolalysenko) added the ci-janitor Opened by the CI janitor routine (flakes, redundant tests, CI perf) label Oct 5, 2026
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

bugbot run


Generated by Claude Code

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread scripts/backtest-pdm.py
Bugbot: the final hosted/vendored rollback checks, the unverifiable-
write rollback, the refused-lock VEX and the reverted-lock VEX runs
named no operation, so a transport failure there never made the case
retryable. installedBytesPatched fails together with a blipped pdm
sync and blocked the retry the same way.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N8YeCUhdg2tKdqyyY7z3sV
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

bugbot run


Generated by Claude Code

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Main is red since #605 (4646693): two commands::vex_consumed tests
assume the name-keyed resolver never returns npm-aliased copies, and
#605 taught it to find them. This fails socket-patch-cli --lib in
coverage and test on every PR. Port #851's test-only fix so this PR
can go green; it no-ops once #851 lands.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N8YeCUhdg2tKdqyyY7z3sV
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

coverage failed on 6b0302a in socket-patch-cli --lib. That failure isn't this PR's. The PR only touches the PDM harness, its tests and docs, and the failing tests are main's: commands::vex_consumed::hosted_expands_alias_only_copies and hosted_reuses_expanded_npm_copies_and_merges_alias_variants have been red since #605 (4646693). I reproduced both locally on this branch (8 passed, 2 failed).

I ported #851's test-only fix as 80538db, so this PR can go green now. It becomes a no-op once #851 lands. With it, cargo test -p socket-patch-cli --lib passes 840/840.


Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

bugbot run


Generated by Claude Code

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 80538db. Configure here.

@mikolalysenko Mikola Lysenko (mikolalysenko) added the Ready for review Agent-verified: mergeable, CI green, Bugbot clean — awaiting human review label Oct 5, 2026
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

Burn-down agent: labeled Ready for review.

  • Head: 80538db4f7e4fcf9bcbc49156ee8ecf5ccf4fb9f
  • CI: 462/462 check runs green (success/skipped) on this head; mergeable clean
  • Bugbot: reviewed 80538db, no new issues; 1 earlier thread resolved
  • Reviewer focus: CI-only: PDM matrix retry on PyPI/patch API transport errors in scripts/backtest-pdm.py and workflow

Generated by Claude Code

@mikolalysenko
Mikola Lysenko (mikolalysenko) merged commit bc2449e into main Oct 5, 2026
463 checks passed
@mikolalysenko
Mikola Lysenko (mikolalysenko) deleted the ci-janitor/pdm-transport-retry branch October 5, 2026 17:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-janitor Opened by the CI janitor routine (flakes, redundant tests, CI perf) Ready for review Agent-verified: mergeable, CI green, Bugbot clean — awaiting human review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants