Skip to content

feat: grant the outage traffic advance and take it back - #66

Closed
timwu20 wants to merge 4 commits into
feat/outage-advance-capfrom
feat/outage-advance-operator
Closed

timwu20 wants to merge 4 commits into
feat/outage-advance-capfrom
feat/outage-advance-operator

Conversation

@timwu20

@timwu20 timwu20 commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Closes ChainSafe/canton-extending-mainnet#129 together with #61, which this is stacked on: #61 adds the outageAdvance cap to the registration, and this PR applies it.

While the operator's participant has seen no fresh global-synchronizer time for outageAdvanceDelay, the sync operator app sets each member's limit on its sequencer to the member's purchased total plus the registration's cap, so members keep transacting through the outage down to that floor. When time moves again, it sets each limit back to the purchased total. A member that used the advance is below zero until its purchases cover what it consumed, and the validator's top-up now buys that shortfall together with its normal top-up.

  • Outage signal. The app runs a DomainTimeAutomationService against the global synchronizer, the one the validator and SV apps run. It is a service of its own because triggers wait on their automation's domain time before each run, and these have to run exactly when it is stale. healthy from listConnectedSynchronizers was not enough: it stays true while enough sequencer subscriptions are live, even if the synchronizer is not ordering.
  • Target. The purchased total in the app's store, which cannot change while the global synchronizer is unreachable. A restarted app computes the same target, and on a synchronizer whose sequencers several operators run, apps with up-to-date stores send the same grant. The advance only raises limits.
  • Cost. Outside a switch between the two modes, the trigger does nothing per member; it walks them only until the new mode's targets are all in place.
  • Store. The operator store also ingests the RegisteredSynchronizer, which the operator observes, so the cap is at hand without Scan. The descriptor moves to version 2, so an existing store re-ingests.
  • Config. outageAdvanceDelay (5 minutes) and globalSynchronizerAlias (global) on the sync operator app.

The take-back sets any limit above the purchased total back to it, so traffic an operator granted by hand above the purchases is removed when the app starts and after an outage.

Tested in SyncOperatorTrafficIntegrationTest: bob, whose balance starts small, gets the advance when the operator's participant is disconnected from the global synchronizer, keeps it across an operator restart, transacts past his purchase and is rejected at the cap; after reconnecting, the advance is taken back, he is blocked, and one top-up covers the shortfall.

…is unreachable

The sync operator app tracks the global synchronizer's time through its
participant. After outageAdvanceDelay without a fresh time, it sets each
member's limit on its sequencer to the member's purchased total plus the
registration's outageAdvance, and sets it back to the purchased total once
time moves again. The store now also ingests the registration, which the
operator observes, so the cap is at hand without Scan.

Signed-off-by: Timothy Wu <tim.wu@chainsafe.io>
Once an operator takes back an outage advance a member used, the member's
balance is below zero. The top-up now buys that shortfall together with the
configured amount, with the funds check priced for both, instead of leaving
the member blocked until several top-ups have covered it.

SyncOperatorTrafficIntegrationTest checks the whole outage: bob gets the
advance once the operator's participant loses the global synchronizer, keeps
it across an operator restart, transacts past his purchase and is rejected
at the cap, then after reconnecting is blocked until one top-up covers it.

Signed-off-by: Timothy Wu <tim.wu@chainsafe.io>
Signed-off-by: Timothy Wu <tim.wu@chainsafe.io>
Signed-off-by: Timothy Wu <tim.wu@chainsafe.io>
@timwu20

timwu20 commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator Author

Closing this in favour of the design agreed in today's sync (see #61): the allowance comes from the operator's local config instead of the registration, and the operator applies and removes it following the runbook. That makes this PR's outage detection, synchronizer-time tracking and registration ingestion unnecessary.

A new PR will implement the agreed design, reusing parts of this one: setting each member's limit to its purchased total plus the allowance and back again, and the validator top-up buying the shortfall in one purchase. It's tracked in ChainSafe/canton-extending-mainnet#129.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant