Skip to content

ci: fix go-tests-short breakage, auto-rerun Spot-preempted jobs, clear rust-deny - #540

Open
jcortejoso wants to merge 3 commits into
celo-rebase-18from
jcortejoso/ci-flaky-tests
Open

jcortejoso wants to merge 3 commits into
celo-rebase-18from
jcortejoso/ci-flaky-tests

Conversation

@jcortejoso

@jcortejoso jcortejoso commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Investigation of the CI "flakiness" on celo-org/optimism (~120 pipelines, Aug 10 – Sep 30; main workflow green only ~30–40% of the time). Three changes:

1. Skip the broken Sepolia chain-upgrade subtest (cherry-pick of upstream ethereum-optimism#22933)

Since 2026-09-17 ~19:00 UTC go-tests-short fails on every branch: TestEndToEndBootstrapApplyWithUpgrade/.../upgrade_chain_v2 fails 4/4 attempts (not a flake).

  • The test forks live Sepolia head (devnet.NewForkedSepolia) and upgrades the real OP Sepolia SystemConfig.
  • The upgrade reads respectedGameType() from OP Sepolia's AnchorStateRegistry, which is now 9 (SUPER_CANNON_KONA, U20 applied). The test leaves that game disabled, so OPContractsManagerV2._assertValidFullConfig reverts with OPContractsManagerV2_InvalidGameConfigs() (0x7c165cd6).
  • Upstream hit the same break and skipped the subtest the next day (f4af487). Their proper fix (contracts-bedrock: Restore OPCM upgrade test compatibility ethereum-optimism/optimism#22939) bumps OPContractsManagerV2 to 9.0.0, so it's left for the next rebase rather than pulled into a CI fix.

2. max_auto_reruns: 2 on main and rust-ci-gate-short

Auto-reruns only retry the failed jobs. 2 retries leaves ~1% infra-only red; it's kept low because genuine failures are retried too (a real red now takes ~3 attempts to report). rust-ci is left out: it fails deterministically on rust-deny advisories.

3. Clear rust-deny (RUSTSEC advisories)

rust-deny fails on every rust-ci run: cargo-deny fetches the latest RustSec DB, so new advisories land without code changes. Cargo.lock-only bumps to the patched versions upstream develop already locks (all semver-compatible):

Crate Bump Advisory
h2 0.4.13 → 0.4.16 RUSTSEC-2026-0258
rkyv 0.8.16 → 0.8.17 RUSTSEC-2026-0233/0234/0235
ruint 1.17.2 → 1.20.0 RUSTSEC-2026-0220
rustls 0.23.38 → 0.23.45 RUSTSEC-2026-0285
imbl (→ imbl-sized-chunks 0.2.0) 7.0.0 → 7.0.2 RUSTSEC-2026-0292
anyhow 1.0.102 → 1.0.104 RUSTSEC-2026-0190
crossbeam-epoch 0.9.18 → 0.9.20 RUSTSEC-2026-0204

The other packages that moved with these (aws-lc-rs/sys, rustls-webpki, rkyv_derive, ark-* 0.6) also match upstream's lock.

RUSTSEC-2026-0253 (lru <0.18.2) is ignored in rust/deny.toml. It's informational/unsound, not a vulnerability, and has no patched 0.16.x; alloy-provider 2.0.x requires lru ^0.16. It only triggers when a key's Drop panics under catch_unwind, and our LruCache keys are u64, B256 and PreimageKey (plain data).

Locally: cargo deny --all-features check all → advisories, bans, licenses, sources ok.

Not addressed here

…mism#22933)

Cherry-pick of upstream f4af487.

TestEndToEndBootstrapApplyWithUpgrade forks live Sepolia head. Since
2026-09-17 OP Sepolia's AnchorStateRegistry respectedGameType is
SUPER_CANNON_KONA (U20 applied), which the test's game config leaves
disabled, so OPContractsManagerV2 reverts with InvalidGameConfigs
(0x7c165cd6) on every run and go-tests-short is red on all branches.

The proper upstream fix (ethereum-optimism#22939) bumps OPContractsManagerV2 to 9.0.0,
so it is left for the next rebase.
The self-hosted celo-org runners run on GKE Spot node pools. Over
Aug 10 - Sep 30, 30 of 34 infrastructure_fail jobs died ~25s before a
compute.instances.preempted event in their own pool, and ~24% of main
workflows hit at least one such failure.

max_auto_reruns reruns only the failed jobs. With 2 retries the
residual infra-only red rate is ~1%; it is kept low because genuine
failures are retried too. rust-ci is left out since it fails on
rust-deny advisories, which a rerun cannot fix.
…026-0253

rust-deny fails on every rust-ci run because cargo-deny fetches the
latest RustSec DB. Bump the affected crates (Cargo.lock only) to the
patched versions upstream develop already locks:

- h2 0.4.13 -> 0.4.16 (RUSTSEC-2026-0258)
- rkyv 0.8.16 -> 0.8.17 (RUSTSEC-2026-0233/0234/0235)
- ruint 1.17.2 -> 1.20.0 (RUSTSEC-2026-0220)
- rustls 0.23.38 -> 0.23.45 (RUSTSEC-2026-0285)
- imbl 7.0.0 -> 7.0.2, pulling imbl-sized-chunks 0.2.0 (RUSTSEC-2026-0292)
- anyhow 1.0.102 -> 1.0.104 (RUSTSEC-2026-0190)
- crossbeam-epoch 0.9.18 -> 0.9.20 (RUSTSEC-2026-0204)

Collateral bumps (aws-lc-rs/sys, rustls-webpki, rkyv_derive, ark-* 0.6)
also match upstream's lock.

RUSTSEC-2026-0253 (lru <0.18.2, informational unsound) is ignored: no
patched 0.16.x exists and alloy-provider 2.0.x requires lru ^0.16. It
needs a key whose Drop panics plus catch_unwind; our LruCache keys are
u64, B256 and PreimageKey (plain data).

`cargo deny --all-features check all`: advisories, bans, licenses,
sources ok.
@jcortejoso jcortejoso changed the title ci: fix go-tests-short breakage and auto-rerun Spot-preempted jobs ci: fix go-tests-short breakage, auto-rerun Spot-preempted jobs, clear rust-deny Oct 1, 2026
@jcortejoso
jcortejoso marked this pull request as ready for review October 1, 2026 15:47
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-01T15:51:35.171715Z fa5b606 Draft marked ready
🔒 Security Review ✅ Completed 2026-10-01T15:52:28.160575Z fa5b606 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@palango palango left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@jcortejoso nothing blocking here, and CI went green on the first attempt (59/59, no auto-reruns needed). I'd like one change before merge, plus a gap in the rerun coverage.

  1. We don't need the RUSTSEC-2026-0253 ignore, and its comment blames the wrong thing. rust/deny.toml:27-30 says alloy-provider's lru ^0.16 blocks the fix. But cargo-deny's default unsound = "workspace" scope only fails on our direct deps, and the one tripping it is our own lru = "0.16.3" at rust/Cargo.toml:616. Upstream bumped that line to 0.18.2 in ethereum-optimism#22500 and has no ignore. I checked this at the PR head: removed the ignore, set lru = "0.18.2", ran cargo update -p lru@0.16.4. alloy's lru 0.16.4 stays in the lock next to 0.18.5, and:

    • cargo deny --all-features check advisories gives advisories ok
    • cargo check -p kona-proof -p kona-client -p kona-providers-alloy -p kona-providers-local --all-targets builds with no code changes

    I'd rather take the bump than keep a graph-wide ignore with no tracking issue. Since alloy was never the cause, nobody would think to drop the ignore after the alloy bump either. (Your read of the keys is right, though. We do call LruCache::pop at rust/kona/crates/providers/providers-local/src/buffer.rs:259, but with plain-data keys, so it's safe in practice.)

  2. rust-ci is left without max_auto_reruns, and the reason given stops being true once this PR merges. The body says it "fails deterministically on rust-deny", but commit 3 fixes rust-deny. Even before that, a rerun-from-failed would only have redone the ~3 min rust-deny job on a cloud runner. rust-ci runs the same Spot-hosted celo-org/2xlarge jobs that rust-ci-gate-short now retries (rust-tests ~20 min, op-reth-integration-tests ~13 min, rust-doctest). Routing sends PRs that touch rust/ to rust-ci, so the PRs that actually change Rust are the ones that don't get retries. The rust-e2e-* workflows have the same gap, since contracts-bedrock-build runs on Spot celo-org/xlarge as a required dep. That one could be a follow-up.

A few nits:

  1. The skip comment's "real fix: upstream ethereum-optimism#22939 (OPCM v9 bump)" doesn't hold here (apply_test.go:903). Upstream needed v9 because its sequence check rejected 8.0.1 → 8.0.6. Our OPContractsManagerV2 is 7.1.20, and that check passes for anything below 8.0.0. Our revert comes from the test config leaving SUPER_CANNON_KONA disabled. On this base the subtest was also pushing older implementations onto live OP Sepolia, so it was effectively testing a downgrade. Skipping it loses nothing, but I'd fix the comment so whoever re-enables it doesn't go chasing a semver bump.
  2. The TODO points at ethereum-optimism#22934, which is already closed. Nothing fails, since the closed-issue check is off on the fork.
  3. fa5b60612f uses chore(rust):, but the repo convention is rust:. A squash merge with the PR title avoids it.
  4. "~1% infra-only red" has no stated basis, which is fine for a ballpark. Also, a real red doesn't take ~3 attempts to show: the gate goes red right away, and only the final verdict waits for all three.

What I checked and found fine: max_auto_reruns is valid at workflow level (1–5), and circleci config validate and config process both accept it next to when:. Nothing in main would double-publish (the publish workflows are all when: false), and rebase-18 has no required checks to trip over. All 19 Cargo.lock changes match upstream develop exactly, the ark-* 0.6 entries come from an optional ruint feature that never compiles, and cargo check --locked --workspace --all-targets passes. The skip is identical to upstream f4af487006 and only hits upgrade_chain_v2. I confirmed live OP Sepolia's respectedGameType() returns 9, and cast sig gives your 0x7c165cd6.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants