Skip to content

docs: design for sequencing a dedicated synchronizer through a global-synchronizer outage (FR-3) - #141

Closed
timwu20 wants to merge 1 commit into
mainfrom
docs/fr-3-outage-advance
Closed

timwu20 wants to merge 1 commit into
mainfrom
docs/fr-3-outage-advance

Conversation

@timwu20

@timwu20 timwu20 commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Proposed Phase 2 answer to FR-3, to be folded into the FR-3 section of the 2026-08 Extension Traffic Manager Technical Design. It covers:

Refs #129.

…-synchronizer outage (FR-3)

Signed-off-by: Timothy Wu <tim.wu@chainsafe.io>
@timwu20
timwu20 requested a review from salindne October 1, 2026 16:57
## Open questions

1. **Per synchronizer or network-wide.** The cap is per synchronizer because the operator carries the credit and the bytes that buy a useful runway depend on members' throughput: the same amount is minutes on a busy synchronizer and hours on a quiet one. A single network-wide rule would have to be expressed as time, say N minutes of a member's recent consumption, which means sampling every member's consumption continuously, the per-member cost the design otherwise avoids.
2. **How long an outage Phase 2 must cover.** Any bounded amount runs out. Covering an outage of any length needs negative balances in Canton. If a few hours is enough for Phase 2, a bounded advance covers it. If not, Phase 2 is prepaid runway alone, the FR-3 basic model, and longer outages wait for Canton.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you allow arbitrary negative balances people can just stop purchasing traffic so that's somewhat useless. So you still need a bound or you need a really strong signal on when an outage has happened and only then allow them to go negative that you can trust which also sounds tricky.


## Options for 0.10.0

Any cap is a Daml change, so it has to be settled by mid-October to make the 0.10.0 cut on 29 October. The backend follows on 0.10.3 on 20 November.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In theory there is the option of actually having an off-ledger endpoint that the SVs update e.g. a bft read from SV scans. the downside of that is that it does introduce a dependency. So it doesn't work if you literally have zero connection to scan. But keeping scan up when the synchronizer is down is not a particularly hard problem to solve so it's at least a weaker dependency.

The advantage is that it actually can be dynamically adjusted. So if the SVs think the outage is gonna be 1h they can use a fairly small one and then extend it if it isn't resolved.

Not yet sold on this being the best option but I think we should at least consider it.

@timwu20

timwu20 commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Closing this as outdated. At the 10-05 sync we went with an operator-set outage traffic allowance instead. During an outage, the operator sets a bounded allowance in the sync operator app's local config, and removes it once the global synchronizer is back. There's no outage signal and no SV governance involved. #142 records the decision in this design note, and canton-network/splice-multi-sync#71 implements it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants