ci: fix go-tests-short breakage, auto-rerun Spot-preempted jobs, clear rust-deny - #540
jcortejoso wants to merge 3 commits into
Conversation
…mism#22933) Cherry-pick of upstream f4af487. TestEndToEndBootstrapApplyWithUpgrade forks live Sepolia head. Since 2026-09-17 OP Sepolia's AnchorStateRegistry respectedGameType is SUPER_CANNON_KONA (U20 applied), which the test's game config leaves disabled, so OPContractsManagerV2 reverts with InvalidGameConfigs (0x7c165cd6) on every run and go-tests-short is red on all branches. The proper upstream fix (ethereum-optimism#22939) bumps OPContractsManagerV2 to 9.0.0, so it is left for the next rebase.
The self-hosted celo-org runners run on GKE Spot node pools. Over Aug 10 - Sep 30, 30 of 34 infrastructure_fail jobs died ~25s before a compute.instances.preempted event in their own pool, and ~24% of main workflows hit at least one such failure. max_auto_reruns reruns only the failed jobs. With 2 retries the residual infra-only red rate is ~1%; it is kept low because genuine failures are retried too. rust-ci is left out since it fails on rust-deny advisories, which a rerun cannot fix.
…026-0253 rust-deny fails on every rust-ci run because cargo-deny fetches the latest RustSec DB. Bump the affected crates (Cargo.lock only) to the patched versions upstream develop already locks: - h2 0.4.13 -> 0.4.16 (RUSTSEC-2026-0258) - rkyv 0.8.16 -> 0.8.17 (RUSTSEC-2026-0233/0234/0235) - ruint 1.17.2 -> 1.20.0 (RUSTSEC-2026-0220) - rustls 0.23.38 -> 0.23.45 (RUSTSEC-2026-0285) - imbl 7.0.0 -> 7.0.2, pulling imbl-sized-chunks 0.2.0 (RUSTSEC-2026-0292) - anyhow 1.0.102 -> 1.0.104 (RUSTSEC-2026-0190) - crossbeam-epoch 0.9.18 -> 0.9.20 (RUSTSEC-2026-0204) Collateral bumps (aws-lc-rs/sys, rustls-webpki, rkyv_derive, ark-* 0.6) also match upstream's lock. RUSTSEC-2026-0253 (lru <0.18.2, informational unsound) is ignored: no patched 0.16.x exists and alloy-provider 2.0.x requires lru ^0.16. It needs a key whose Drop panics plus catch_unwind; our LruCache keys are u64, B256 and PreimageKey (plain data). `cargo deny --all-features check all`: advisories, bans, licenses, sources ok.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
palango
left a comment
There was a problem hiding this comment.
@jcortejoso nothing blocking here, and CI went green on the first attempt (59/59, no auto-reruns needed). I'd like one change before merge, plus a gap in the rerun coverage.
-
We don't need the RUSTSEC-2026-0253 ignore, and its comment blames the wrong thing.
rust/deny.toml:27-30says alloy-provider'slru ^0.16blocks the fix. But cargo-deny's defaultunsound = "workspace"scope only fails on our direct deps, and the one tripping it is our ownlru = "0.16.3"atrust/Cargo.toml:616. Upstream bumped that line to 0.18.2 in ethereum-optimism#22500 and has no ignore. I checked this at the PR head: removed the ignore, setlru = "0.18.2", rancargo update -p lru@0.16.4. alloy's lru 0.16.4 stays in the lock next to 0.18.5, and:cargo deny --all-features check advisoriesgivesadvisories okcargo check -p kona-proof -p kona-client -p kona-providers-alloy -p kona-providers-local --all-targetsbuilds with no code changes
I'd rather take the bump than keep a graph-wide ignore with no tracking issue. Since alloy was never the cause, nobody would think to drop the ignore after the alloy bump either. (Your read of the keys is right, though. We do call
LruCache::popatrust/kona/crates/providers/providers-local/src/buffer.rs:259, but with plain-data keys, so it's safe in practice.) -
rust-ciis left withoutmax_auto_reruns, and the reason given stops being true once this PR merges. The body says it "fails deterministically on rust-deny", but commit 3 fixes rust-deny. Even before that, a rerun-from-failed would only have redone the ~3 min rust-deny job on a cloud runner.rust-ciruns the same Spot-hostedcelo-org/2xlargejobs thatrust-ci-gate-shortnow retries (rust-tests~20 min,op-reth-integration-tests~13 min,rust-doctest). Routing sends PRs that touchrust/torust-ci, so the PRs that actually change Rust are the ones that don't get retries. Therust-e2e-*workflows have the same gap, sincecontracts-bedrock-buildruns on Spotcelo-org/xlargeas a required dep. That one could be a follow-up.
A few nits:
- The skip comment's "real fix: upstream ethereum-optimism#22939 (OPCM v9 bump)" doesn't hold here (
apply_test.go:903). Upstream needed v9 because its sequence check rejected 8.0.1 → 8.0.6. Our OPContractsManagerV2 is 7.1.20, and that check passes for anything below 8.0.0. Our revert comes from the test config leaving SUPER_CANNON_KONA disabled. On this base the subtest was also pushing older implementations onto live OP Sepolia, so it was effectively testing a downgrade. Skipping it loses nothing, but I'd fix the comment so whoever re-enables it doesn't go chasing a semver bump. - The TODO points at ethereum-optimism#22934, which is already closed. Nothing fails, since the closed-issue check is off on the fork.
fa5b60612fuseschore(rust):, but the repo convention isrust:. A squash merge with the PR title avoids it.- "~1% infra-only red" has no stated basis, which is fine for a ballpark. Also, a real red doesn't take ~3 attempts to show: the gate goes red right away, and only the final verdict waits for all three.
What I checked and found fine: max_auto_reruns is valid at workflow level (1–5), and circleci config validate and config process both accept it next to when:. Nothing in main would double-publish (the publish workflows are all when: false), and rebase-18 has no required checks to trip over. All 19 Cargo.lock changes match upstream develop exactly, the ark-* 0.6 entries come from an optional ruint feature that never compiles, and cargo check --locked --workspace --all-targets passes. The skip is identical to upstream f4af487006 and only hits upgrade_chain_v2. I confirmed live OP Sepolia's respectedGameType() returns 9, and cast sig gives your 0x7c165cd6.
Investigation of the CI "flakiness" on celo-org/optimism (~120 pipelines, Aug 10 – Sep 30;
mainworkflow green only ~30–40% of the time). Three changes:1. Skip the broken Sepolia chain-upgrade subtest (cherry-pick of upstream ethereum-optimism#22933)
Since 2026-09-17 ~19:00 UTC
go-tests-shortfails on every branch:TestEndToEndBootstrapApplyWithUpgrade/.../upgrade_chain_v2fails 4/4 attempts (not a flake).devnet.NewForkedSepolia) and upgrades the real OP SepoliaSystemConfig.respectedGameType()from OP Sepolia'sAnchorStateRegistry, which is now9(SUPER_CANNON_KONA, U20 applied). The test leaves that game disabled, soOPContractsManagerV2._assertValidFullConfigreverts withOPContractsManagerV2_InvalidGameConfigs()(0x7c165cd6).OPContractsManagerV2to 9.0.0, so it's left for the next rebase rather than pulled into a CI fix.2.
max_auto_reruns: 2onmainandrust-ci-gate-shortAuto-reruns only retry the failed jobs. 2 retries leaves ~1% infra-only red; it's kept low because genuine failures are retried too (a real red now takes ~3 attempts to report).
rust-ciis left out: it fails deterministically onrust-denyadvisories.3. Clear
rust-deny(RUSTSEC advisories)rust-denyfails on everyrust-cirun: cargo-deny fetches the latest RustSec DB, so new advisories land without code changes.Cargo.lock-only bumps to the patched versions upstreamdevelopalready locks (all semver-compatible):The other packages that moved with these (aws-lc-rs/sys, rustls-webpki, rkyv_derive, ark-* 0.6) also match upstream's lock.
RUSTSEC-2026-0253(lru <0.18.2) is ignored inrust/deny.toml. It's informational/unsound, not a vulnerability, and has no patched 0.16.x;alloy-provider2.0.x requireslru ^0.16. It only triggers when a key'sDroppanics undercatch_unwind, and ourLruCachekeys areu64,B256andPreimageKey(plain data).Locally:
cargo deny --all-features check all→ advisories, bans, licenses, sources ok.Not addressed here
TestBatcherAutoDA: fails its first attempt in most runs, hidden by gotestsum--rerun-fails=3. It assertsfeeRatio > 41, but the ratio decays ~2.3% per L1 block and passes only if read by block ~4.TestELSyncSafeRetractedByOffset: ~10% ofmemory-all-opn-op-gethruns, timing out during preset setup. Known upstream: flaky test:TestELSyncSafeRetractedByOffset— CrossSafe sync timeout during preset setup ethereum-optimism/optimism#20334.