Skip to content

DNM: soak the replica-recovery step against the litep2p provider-store fix - #13380

Closed
DenzelPenzel wants to merge 7 commits into
masterfrom
denzelpenzel/statement-store-v2-replica-recovery-litep2p665
Closed

DenzelPenzel wants to merge 7 commits into
masterfrom
denzelpenzel/statement-store-v2-replica-recovery-litep2p665

Conversation

@DenzelPenzel

Copy link
Copy Markdown
Contributor

Do not merge. This exists only to run the soak from #13378 on the cluster with the
litep2p fix patched in.

Builds on top of #13378 (the replica outage and recovery step) and patches litep2p to
paritytech/litep2p#666, the fix for paritytech/litep2p#665: without it a node evicted
from a key's provider list panics in MemoryStore::remove_local_provider on the next
epoch rollover, leaving the pod Running with its networking dead. Any run above 20
providers on one key hits it.

Sized at 40 nodes / 600s: above the 20-providers threshold, so the fix is actually
exercised, and short enough to finish well inside the hour the runner's k8s credentials
last (see the token expiry issue - runs past ~1h get 401 on the closing log scan).

Run it with:

gh workflow run zombienet_statement-store-soak.yml \
  --ref denzelpenzel/statement-store-v2-replica-recovery-litep2p665 \
  -f build_run_id=<a green build of this branch>

soak_nodes and soak_secs are dispatch inputs, so the size can be changed per run
without touching the branch.

Close once #13378 has been exercised.

Patches litep2p to paritytech/litep2p#666 so this branch can run the soak past the
20-providers-per-key limit where #665 panics nodes out of the network. The
KademliaEvent::PeersDiscovered arm is unrelated to the fix - the branch sits on
litep2p master, which added that event in #611, and sc-network has to compile
against it.
@DenzelPenzel
DenzelPenzel requested review from a team as code owners September 30, 2026 12:20
…ent-store-v2-replica-recovery-litep2p665

# Conflicts:
#	Cargo.lock
The orchestrator otherwise brings a whole level up at once, and the apiserver drops
the exec connections partway through: two runs here died at 35 and 39 of 43 nodes
with 'deadline has elapsed'. A 102-node run with this set spawned cleanly in ten
minutes, so the size is not the problem - the burst is.
…replica-recovery' into denzelpenzel/statement-store-v2-replica-recovery-litep2p665

# Conflicts:
#	.github/workflows/zombienet_statement-store-soak.yml
@paritytech-workflow-stopper

Copy link
Copy Markdown

All GitHub workflows were cancelled due to failure one of the required jobs.
Failed workflow url: https://github.com/paritytech/polkadot-sdk/actions/runs/36742779582
Failed job name: fmt

@bkchr bkchr closed this Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants