Skip to content

[P2-E5.9] Operator-set outage traffic allowance (FR-3) #129

Description

@timwu20

Context. FR-3 asks that a dedicated synchronizer keep sequencing while the global synchronizer is unavailable, with obligations settled on restoration. Canton has no negative traffic balances, and DA confirmed on 2026-10-01 that they are unlikely on current timelines. Without help, a member runs only as long as its own prepaid remainder, about one top-up (10 minutes by default), and is then refused until the global synchronizer is back.

Decision (2026-10-05). An operator-set outage traffic allowance, applied and removed by hand following the runbook. Design: #141, updated in #142. It replaces the SV-voted cap with automatic outage detection (canton-network/splice-multi-sync#61 and canton-network/splice-multi-sync#66, both closed).

Approach.

  • Applying it. During an outage, the operator adds the allowance to the sync operator app's config and restarts the app. It's a single value in bytes, applied to every member with a purchase on record. On start, the app sets each such member's limit to exactly its purchased total plus the allowance, so setting a smaller value and restarting lowers it again.
  • While it's set. Purchases that land meanwhile pay down the credit, since the reconcile trigger only raises a limit to the purchased total.
  • Removing it. Once the global synchronizer is back, the operator removes the setting and restarts. The app sets each limit back to the member's purchased total, then raises any limit a purchase ingested during the take-back left behind.
  • There's no outage detection, no on-ledger parameter and no Daml change.

Deliverable (0.10.3, the backend cut on 20 November).

  • Sync operator app:

    • the allowance setting;
    • on start, limits set to purchased total plus allowance, or back to purchased total;
    • warnings every few minutes while the allowance is set, giving the allowance, the members it covers and the total credit outstanding;
    • a warning when a purchase lands while it's set, since that means the global synchronizer is reachable and it's time to remove the allowance;
    • a warning after removal, for any member whose limit doesn't come back to its purchased total.
  • Validator top-up: when a member's remainder is negative, buy the shortfall plus a normal top-up in one purchase, with the funds check priced for both.

  • Integration test: a Splice integration test covering the whole flow, not LocalNet:

    1. Simulate the outage by disconnecting the operator's participant from the global synchronizer.
    2. Set the allowance in the operator app's config and restart the app.
    3. Check each member's limit rose to its purchased total plus the allowance, and the "allowance is set" warning is logged.
    4. Restart the app again with the allowance still set, and check no limit moved.
    5. Check a member keeps transacting during the outage past what it bought, up to the allowance, and is refused there. A second member doesn't use the allowance.
    6. Restore the global synchronizer, land a purchase for the first member, and check the "purchase landed while the allowance is set" warning is logged and the member's limit didn't move: the purchase pays down the credit.
    7. Remove the allowance from the config and restart the app.
    8. Check each limit went back down to the purchased total, and the second member keeps transacting.
    9. Check the first member is refused until its top-up buys past the shortfall in one purchase, then transacts again.

    The member's own top-up stays paused until step 9, so it doesn't pay down the credit before the take-back.

  • Unit tests:

    • a purchase ingested during the take-back is picked up by the re-raise;
    • a member with no purchase on record gets no allowance;
    • a smaller allowance lowers limits on restart;
    • the warning for a limit that doesn't come back to its purchased total;
    • the shortfall purchase's amount and funds check.
  • Runbook, in [P2-E8.3] Helm charts + operator documentation #37: the steps to set and remove the allowance, how operators coordinate on a synchronizer run by several organizations, and guidance for members' prepaid runway.

Acceptance.

  • With the allowance set, each member with a purchase on record is at its purchased total plus the allowance, never more. With it removed, each limit returns to the purchased total.
  • A member that used the allowance transacts again only once its purchases cover what it consumed, and its top-up buys that in one purchase.
  • The app warns while the allowance is set, and when a purchase lands while it's set.
  • The integration test runs the full flow above in CI.
  • The runbook steps are in the operator documentation ([P2-E8.3] Helm charts + operator documentation #37).

Not in scope.

  • Negative balances inside Canton ([P3-E7] Outage settlement #79).
  • Automatic outage detection.
  • Penalties for long outages.
  • Running a dedicated synchronizer separately for a planned period, which FR-3 marks as probably Phase 3.
  • Reporting validators' balances on reconnection, which belongs with the Phase 3 consumption reports.

Notes.

  • On a synchronizer run by several organizations, a grant lands only once the sequencer group's threshold send the same value. Every operator has to set, and later remove, the same allowance, coordinated through the runbook.
  • The allowance is the operator's own credit decision: local, not governed by SVs and not visible on-ledger. Canton accepts any limit, so an operator could always grant traffic nobody paid for; this makes it a defined, bounded procedure whose deficits are repaid by on-ledger purchases.
  • The allowance is one value for every member. A quiet member gets more headroom than it needs and a busy one may reach the allowance sooner. A per-member override (participant id to allowance, with the single value as the default) is deferred until reviewers ask for it.
  • The multi-operator case can't be tested until the four-node topology with several operators lands (test: sync operators on a four-node BFT synchronizer canton-network/splice-multi-sync#63, [P2-E8.7] Multi-node dedicated synchronizer topology (fault tolerance) #106). Testing it is a follow-up.
  • FR-3's grace period becomes a runbook choice: the operator can wait after the global synchronizer returns before removing the allowance, so active members' purchases pay down their credit first.
  • The closed feat: grant the outage traffic advance and take it back canton-network/splice-multi-sync#66 has code to reuse: setting limits to purchased total plus allowance and back, the shortfall top-up, and the integration-test scaffolding.

—
Epic: #68 · Related: #37, #79, #106, #141, #142 · Refs: FR-3 in 2026-08 Extension Traffic Manager Technical Design

Activity

  1. self-assigned this
    on Sep 23, 2026
  2. moved this from Backlog to Ready in Canton Mainnet Extension Phase 2on Sep 23, 2026
  3. moved this from Backlog to Ready in Canton Mainnet Extension Phase 2on Sep 23, 2026
  4. changed the title [-][P2-E5.9] Operator traffic advance through a global-synchronizer outage (FR-3)[/-] [+][P2-E5.9] Operator-set outage traffic allowance (FR-3)[/+] on Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:opsDeployment / operator node / configarea:scala-svSV app (Scala)area:scala-validatorValidator app (Scala)phase-2Phase 2 — One Economy (MVP)

Type

No type

Fields

Priority

None yet

Projects

Relationships

None yet

Development

No branches or pull requests

Issue actions