Skip to content

feat: govern an outage traffic advance on the registration - #61

Closed
timwu20 wants to merge 3 commits into
mainfrom
feat/outage-advance-cap
Closed

timwu20 wants to merge 3 commits into
mainfrom
feat/outage-advance-cap

Conversation

@timwu20

@timwu20 timwu20 commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator

Towards ChainSafe/canton-extending-mainnet#129. This is the Daml half, so it can make the 0.10.0 cut; the operator app's advance and take-back, the top-up change and the LocalNet test follow on the 0.10.3 train.

Summary: GovernanceParameters gains outageAdvance : Int, the traffic in bytes the operator may grant each member beyond what it has purchased while its connection to the global synchronizer is down, and takes back when it returns. A member keeps transacting through a global-synchronizer outage and repays by purchase afterwards, which is FR-3's settle-on-restoration without negative balances in Canton.

It is set by SV vote on the same record as the discount, so the registration vote and the set-parameters vote carry it with no new vote action. 0 disables the advance, and the registration's ensure rejects a negative amount. The operator app applies it off-ledger; the sequencer accepts any limit, so the cap binds an honest operator, the same trust model as every traffic grant.

Upgrade compatibility. GovernanceParameters has not shipped in a release, so the field is mandatory rather than an appended Optional, as when the parameters became mandatory. The same six DARs rebuild as then.

SV UI. The registration form gains the field, prefilled with 0 and validated as a whole number of bytes, and the review step and vote details show it. The generated type requires the field, and sending a hard-coded 0 would leave no vote able to set an advance.

How it's verified

  • splice-amulet-test, splice-dso-governance-test and splice-wallet-test damlTest green, with no Daml warnings. New tests: a registration carries the advance, a set-parameters vote changes it and leaves the discount alone, and the template rejects a negative amount, with a positive control.
  • DARs, dars.lock and DarResources regenerated; apps-wallet, apps-scan, apps-syncoperator and apps-app Test/compile green.
  • SV frontend: tsc, eslint and prettier clean; the governance tests pass, including new prefill and validation tests for the field.

@moritzkiefer-da moritzkiefer-da left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for the pr! It's not quite clear to me how you expect this to be used, this gives us a parameter but that parameter is unused as of this pr. How will it get picked up to allow the outage overage?

@timwu20

timwu20 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

It's not quite clear to me how you expect this to be used, this gives us a parameter but that parameter is unused as of this pr. How will it get picked up to allow the outage overage?

Fair question. This PR is only the Daml half, nothing reads it yet. The sync operator app picks it up in a follow-up that closes (ChainSafe/canton-extending-mainnet#129):

  1. The app reads the cap from the registration, which the operator observes, so it's in the app's own store.
  2. It treats the global synchronizer as down when its participant hasn't seen a fresh time on it for a configured delay, using the domain-time ingestion the validator and SV apps already run.
  3. During the outage it sets each member's limit on the dedicated sequencer to the member's purchased total plus the cap. The total can't change while the global synchronizer is down, so a restarted app computes the same value. On a multi-node synchronizer the grant lands only once the sequencer group's threshold send it, so the advance applies only when enough operators see the outage.
  4. When time moves again it sets each limit back to the purchased total. A member that used the advance is below zero until its top-up buys the shortfall, which the follow-up also changes.

Outside an outage this costs one time fetch per polling interval, and the app walks the members only when an outage starts and ends. Canton accepts any limit, so the cap binds an honest operator app, like any grant today; on the registration it's SV-voted and public.

Is a stale synchronizer time the signal you'd use here? I went with it over healthy from listConnectedSynchronizers. That flag stays true while the participant still has enough sequencer subscriptions, even if the global synchronizer isn't ordering.

`GovernanceParameters` gains `outageAdvance : Int`, the traffic in bytes the
operator may grant each member beyond what it has purchased while its
connection to the global synchronizer is down, and takes back when it returns.
It lives on the same record as the discount, so the registration vote and the
set-parameters vote carry it with no new vote action. 0 disables it, and the
registration's `ensure` rejects a negative amount.

The record has not shipped in a release, so the field is mandatory rather than
an appended Optional, as with the switch to mandatory parameters. Tests that
built the record now update `defaultGovernanceParameters`, so the next field
does not touch them again. DARs, dars.lock and DarResources regenerated.

Towards ChainSafe/canton-extending-mainnet#129.

Signed-off-by: Timothy Wu <tim.wu@chainsafe.io>
…n [ci]

The registration form gains the advance, prefilled with 0 and validated as a
whole number of bytes, and the review step and vote details show it. The
generated type requires the field; sending a hard-coded 0 instead would leave
no vote able to set one.

Signed-off-by: Timothy Wu <tim.wu@chainsafe.io>
… [ci]

The check rejects 'global' and a bare 'member' in Daml code, so the
comment now says 'decentralized synchronizer' and 'participant'. The
DAR is rebuilt because it bundles the source; its package id is
unchanged.

Signed-off-by: Timothy Wu <tim.wu@chainsafe.io>
@moritzkiefer-da

Copy link
Copy Markdown

I see, that could work but it's unclear to me if this is the exact approach we want. A few points:

  1. It's unclear to me why this config is per synchronizer. Wouldn't you expect all synchronizers to have the same rules here?
  2. A fixed amount is inherently not gonna cover outages of sufficient length as you'll run out of traffic again.
  3. It's not super clear what the signal should be that we use here to detect an outage. Maybe synchronizer time kinda works but it's a bit weird.

I think this is a case where we are better off first sketching out the plan in the design doc and aligning there between us on what exactly the requirements are and what approach we want to pick.

@timwu20

timwu20 commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Replying to your comment. The operator side that uses this cap is now up as draft #66, with an integration test that cuts the operator off the global synchronizer and checks both the advance and the take-back.

  1. It's unclear to me why this config is per synchronizer. Wouldn't you expect all synchronizers to have the same rules here?

Per synchronizer because the operator carries the credit, and the bytes that buy a useful runway depend on its members' throughput: the same amount is minutes on a busy synchronizer and hours on a quiet one. A single network-wide rule works if it is expressed as time rather than bytes, say N minutes of a member's recent consumption, but measuring that means sampling every member's consumption continuously, which the current design avoids.

  1. A fixed amount is inherently not gonna cover outages of sufficient length as you'll run out of traffic again.

Yes, that is the bound on the operator's unpaid exposure. Covering an outage of any length means letting a member's balance run below zero and settling it afterwards, which needs negative traffic balances in Canton itself. So the Phase 2 question is how long an outage it has to survive: if a few hours is enough, a bounded advance covers it, and if not, Phase 2 is prepaid runway alone and longer outages wait for negative balances.

  1. It's not super clear what the signal should be that we use here to detect an outage. Maybe synchronizer time kinda works but it's a bit weird.

Synchronizer time is what Splice already uses to tell that an app is behind its synchronizer (DomainTimeStore), and it catches what the connection status misses: healthy stays true while the participant has enough sequencer subscriptions, even if the synchronizer isn't ordering. Its weak spot is a participant that lags while the global synchronizer is fine, which grants credit without an outage; on a multi-node synchronizer the grant threshold filters that out. If you have a better signal in mind, I'd rather use it.

I think this is a case where we are better off first sketching out the plan in the design doc and aligning there between us on what exactly the requirements are and what approach we want to pick.

Agreed. I'll write up the requirements and options in the design doc and link it here. Any cap is a Daml change, so it has to be settled by mid-October to make the 0.10.0 cut, and I'll get the write-up out this week.

@timwu20

timwu20 commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator Author

Closing this following what we agreed in today's sync: the outage allowance won't be an SV-voted parameter on the registration, so there is no Daml change and nothing for 0.10.0.

Instead, during a global-synchronizer outage the operator sets an allowance in the sync operator app's local config, following the runbook, and removes it once the global synchronizer is back. While it's set, the app raises each member's limit to its purchased total plus the allowance. Once it's removed, the app sets each limit back to the purchased total. The app also warns while the allowance is set, and when a purchase lands while it's still set, since that means the global synchronizer is reachable again.

A new PR will implement this in the sync operator app, together with the validator top-up buying a member's shortfall in one purchase. It's tracked in ChainSafe/canton-extending-mainnet#129.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants