Commit df87548
authored
fix: prod on a dedicated CPU, queued runs no longer shown as running, watchdog no longer amplifies CPU starvation (#175)
* fix(prod): put prod on a dedicated CPU
Prod ran on `shared-cpu-1x`, whose CPU burst balance was fully drained. The
machine sat pinned at the ~6% baseline: 88% `steal` in /proc/stat, PSI cpu
`some avg10=51`, load 1.38 on 1 vCPU.
Starving the event loop that badly meant Node could not drain its DB sockets.
Every Postgres backend sat in `ClientRead` with zero lock contention, and a
plain `select 1` took 0.6-5s against a database that was otherwise healthy. The
DB watchdog reads a slow ping as a wedged pool, so it exited the process; each
restart then re-armed 109 enabled loops plus misfire catch-up, spiking CPU again
on an already-drained balance. Successive lives ran 455s, then 161s, then 100s,
so the balance never recovered. Roughly 2h of hard downtime.
A dedicated core removes the trigger entirely - measured steal after the switch
is 0.2%, load 0.11, and `/api/health/db` is a stable ~0.3s with no flapping.
The running machine was already moved with `fly machine update --vm-size
performance-1x` to stop the outage; this pins it in config so a redeploy cannot
silently put prod back on shared CPU.
Not fixed here: the watchdog still cannot distinguish a wedged pool from a
starved event loop, so it remains an outage amplifier. Tracked separately.
* fix: stop presenting queued runs as running, and stop the watchdog amplifying CPU starvation
Two independent defects surfaced by the 2026-08-10 production incident. The VM-size
fix that ended the outage is the first commit on this branch; these are the two
software problems it exposed.
1) A QUEUED run was presented as a RUNNING one.
`toRunSummary` collapsed both open phases into a single flag - `running:
r.phase === "pending" || r.phase === "running"` - and `JobSummary.running` came from
the equally phase-agnostic `hasOpenRun`. Three surfaces consumed that one flag, so a
run merely queued for a machine that was asleep or shut would render a pulsing
"Running" badge, put the loop and run pages on their 3s LIVE poll cadence, and
disable "Run once" with the tooltip "A run is already in progress".
The badge and the tooltip are simply false. The poll cadence is worse than false: a
genuinely running run is bounded by RUN_TIMEOUT_MS (~20min), so 3s is self-limiting,
but a queued run survives for DEFERRED_MAX_MS (7 days) - so every such page hammered
the server at the live rate for as long as the machine stayed away. With dozens of
deferred runs fleet-wide that is a permanent, self-inflicted load multiplier.
Split the two states at the source: `running` now means executing, `queued` means
waiting to be claimed. The fast poll is gated on `running` alone; a queued run gets
a calm 15s refresh so it still flips promptly when its machine returns. Queued
surfaces render a distinct, still state (no pulse, no elapsed clock) and name the
reason, preferring the sweep's own `progress.label` ("deferred - machine offline")
over any inference. Both open states still block a second dispatch, because the
scheduler refuses to stack two agents on one loop - only the wording changed.
2) The DB watchdog turned CPU starvation into a crash loop.
The watchdog exists to escape a wedged pool, and that job is unchanged. But it read
any slow `select 1` as a wedge. When the machine ran at ~6% of a core (88% steal),
Node could not drain its DB sockets, so the ping blew its 5s deadline against a
database that was entirely healthy - every backend idle in `ClientRead`, zero lock
contention. The watchdog exited; the restart re-armed every loop and re-fired misfire
catch-up, spiking CPU on an already-drained budget. Successive lives ran 455s, 161s,
then 100s. The recovery mechanism was the outage.
The watchdog now consults event-loop delay before blaming the database. Above the
ceiling (default 1s; a healthy server sits in single-digit ms) a failed ping is
recorded as INCONCLUSIVE: it neither trips the exit nor clears a real streak. A
restart is the right cure for a wedged pool and the wrong cure for a starved CPU.
The guard fails toward the old behavior - an unreadable lag signal, or no signal
wired at all, still blames the database - so it can never suppress a genuine wedge
exit. Tunable via LOOPANY_DB_WATCHDOG_LAG_CEILING_MS, 0 to disable.
Tests: the two suites that pinned the old collapsed behavior now pin the split;
new coverage for the adapter phase mapping, a source-level guard on the poll-cadence
and rendering coupling (that coupling is what regresses), and six watchdog cases
including the starvation regression and its fail-toward-old-behavior paths.
* fix(db): size the connection pool by the pooler's cap, not our appetite
`max: 10` against a session pooler that refuses past `pool_size: 15`
(`EMAXCONNSESSION`) left no room for a restart. A process killed without a clean
shutdown leaves its backends held until TCP keepalive reaps them while the
replacement immediately opens its own, so the real worst case is `2*max + 1` - the
extra being the prestart migrator, which shares the same budget here because
DATABASE_URL and DIRECT_DATABASE_URL both point at the session pooler. At 10 that is
21 against a cap of 15: refusals during any restart, which is what the 2026-08-10
crash loop produced (observed 14/15 occupied, new connections refused).
Steady-state demand was never the constraint. Sampling prod once a second for a
minute: 0 active connections in 56 of 60 samples, 1 in three, peak 4 - while the pool
held all 10 open the whole time, because poll traffic keeps round-robining across
them so `idle_timeout` never finds a 30s-quiet connection to reap. So the old setting
permanently occupied two thirds of the pooler's clients to serve a peak of four.
6 keeps the restart worst case at 13, stays 50% above measured peak demand, and
leaves slots for the migrator and an ops session. Queueing behind a smaller pool
costs little on a single-vCPU box, where the CPU is the actual limit.
The right value follows the pooler's cap, which is external and can change without a
deploy, so `LOOPANY_DB_POOL_MAX` overrides it. `poolOptionsFor` stays pure - it takes
the size as an argument; only `db/index.ts` reads env.
* fix: bound the watchdog's starvation guard, and restore cancel for queued runs
Review findings on this branch. Two were real defects in the preceding commit, one
of them a regression that commit introduced.
The starvation guard had no upper bound. A wedged pool can perfectly well coexist
with a busy event loop, and in that case `consecutive` never advanced, so the
watchdog never exited - quietly handing back the 2026-07-12 failure mode (~9h down,
no auto-recovery) that the watchdog exists to end. The previous commit message
claimed the guard "can never suppress a genuine wedge exit"; that was true only for
an unreadable or absent lag signal, not for sustained lag alongside a real wedge.
The guard is now an excuse, not an alibi: past `starvedCeiling` consecutive
inconclusive ticks (default 45, ~15min at the 20s cadence) the watchdog exits
anyway. Restarting is a poor cure for a starved CPU but a strictly better outcome
than staying wedged forever, and 45 ticks is far enough above the 3-failure
threshold that the ~100s crash-loop amplification stays broken.
The lag sample could also be stale. `lagMs()` was read only on the failure path, so
`monitorEventLoopDelay`'s `max` accumulated across healthy ticks and one old stall -
arming 109 loops at boot, say - could sit in the histogram indefinitely and
disqualify a much later, genuine failure. It is now read (and reset) on every tick,
so the sample always describes the interval that contained the probe.
Cancellation regressed for queued runs. Before the queued/running split, `running`
covered `pending`, so a queued run showed the stop control; the split left that call
site on `running` alone. `cancelRun` accepts both phases, and a queued run is exactly
the one worth cancelling - it can wait on an offline machine for days. Restored, and
labelled "Cancel run" rather than "Stop run" when nothing is executing.
Two queued surfaces also still carried `runPulseStyle`, the infinite live-signal
animation, and one still read "Applying your edit" for an edit that had not started.
Not a regression - the collapsed flag pulsed for pending runs before this branch too
- but it contradicted the split's whole point, so both are now still.
Also corrected two comments that overstated their case: the pool worst case is
`2*max` (the prestart migrator closes its connection before the server boots, so it
overlaps only the dead process's lingering backends), and `monitorEventLoopDelay`
does not sample while JS blocks the loop - the overdue sample lands once it resumes,
which is why reading `max` still captures the stall.
Tests: the starvation ceiling firing, that a healthy ping resets the starved streak,
that lag is read on every tick, that a stale spike cannot disqualify a later genuine
failure, that no queued surface carries the pulse, and that a queued run keeps its
cancel control. Server 893 passed, daemon 360.
* fix(env): stop a sub-1 fraction disabling a knob, and tighten the review guards
Second review round. No blocking defect in the previous commit - the four fixes
audited clean, including the case I was least sure of: alternating starved and
non-starved failures cannot stall both counters, because a starved tick preserves
`consecutive` while a non-starved one advances it, so the ordinary failure threshold
is still reached.
One real bug, introduced by the previous commit. `posIntEnv` floored AFTER its
positivity test, so `0.5` passed `n > 0` and became `0`. Every knob reads 0 as
"disabled", and for the new starved ceiling that silently turned the bounded-recovery
guarantee back off, since `starvedTicks >= 0` is true on the first starved tick.
Floor first, so a sub-1 fraction falls back to the documented default. The fix is in
the shared helper, so it covers every knob in the family, not just the new one.
Three items from the same round, all mine and all cosmetic:
- The lag-ceiling JSDoc ended up documenting the starved-ceiling function, because
the new function was inserted between the doc and its subject. Each has its own
doc now.
- The run page's confirm dialog still said "Stop this run?" under a button relabelled
"Cancel run". Both now follow the run's actual state.
- The pool test still carried the `2*max + 1` reasoning corrected in the source.
Two tests were weaker than they looked, which is worth more than the assertions they
replaced:
- The "projected fire" case in timeline.test.ts called `runToMark` twice and never
invoked `projectedMark`, so it asserted nothing about projections. It now tests
`projectedMark`, and the non-pending phases got their own case.
- The source-reading guards sliced between two `indexOf` results without checking
either. Removing a marker would yield an empty slice and make every negative
assertion pass vacuously - precisely how this kind of guard rots into a no-op. A
`between` helper now asserts both markers exist and are ordered.
Not addressed here: `cancelRun` reads a run's phase and then updates unconditionally,
so a cancel racing a poll's claim can mark a just-claimed run canceled while the
daemon still receives it, and can overwrite a run that finished in the gap. It is
pre-existing on main, the fix needs a decision about what a losing cancel should
report, and it touches lease lifetime. Raised separately.
Server 900 passed, daemon 360.1 parent 0d394c4 commit df87548
23 files changed
Lines changed: 980 additions & 54 deletions
File tree
- packages/server/src
- components
- db
- lib
- server
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
57 | 57 | | |
58 | 58 | | |
59 | 59 | | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
60 | 76 | | |
61 | | - | |
62 | | - | |
| 77 | + | |
| 78 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
70 | 70 | | |
71 | 71 | | |
72 | 72 | | |
| 73 | + | |
| 74 | + | |
73 | 75 | | |
74 | 76 | | |
75 | 77 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
110 | 110 | | |
111 | 111 | | |
112 | 112 | | |
113 | | - | |
114 | | - | |
| 113 | + | |
| 114 | + | |
| 115 | + | |
| 116 | + | |
| 117 | + | |
| 118 | + | |
115 | 119 | | |
116 | 120 | | |
117 | 121 | | |
| |||
279 | 283 | | |
280 | 284 | | |
281 | 285 | | |
| 286 | + | |
| 287 | + | |
| 288 | + | |
| 289 | + | |
| 290 | + | |
| 291 | + | |
| 292 | + | |
| 293 | + | |
| 294 | + | |
| 295 | + | |
| 296 | + | |
| 297 | + | |
| 298 | + | |
282 | 299 | | |
283 | 300 | | |
284 | | - | |
| 301 | + | |
| 302 | + | |
| 303 | + | |
285 | 304 | | |
286 | 305 | | |
287 | 306 | | |
| |||
468 | 487 | | |
469 | 488 | | |
470 | 489 | | |
471 | | - | |
| 490 | + | |
| 491 | + | |
| 492 | + | |
| 493 | + | |
472 | 494 | | |
473 | 495 | | |
474 | 496 | | |
475 | 497 | | |
476 | | - | |
477 | | - | |
478 | | - | |
| 498 | + | |
| 499 | + | |
| 500 | + | |
| 501 | + | |
| 502 | + | |
479 | 503 | | |
480 | 504 | | |
481 | 505 | | |
| |||
624 | 648 | | |
625 | 649 | | |
626 | 650 | | |
| 651 | + | |
| 652 | + | |
| 653 | + | |
627 | 654 | | |
628 | 655 | | |
629 | 656 | | |
| |||
701 | 728 | | |
702 | 729 | | |
703 | 730 | | |
| 731 | + | |
| 732 | + | |
| 733 | + | |
704 | 734 | | |
705 | 735 | | |
706 | | - | |
| 736 | + | |
707 | 737 | | |
708 | 738 | | |
709 | 739 | | |
| 740 | + | |
| 741 | + | |
| 742 | + | |
| 743 | + | |
| 744 | + | |
| 745 | + | |
| 746 | + | |
| 747 | + | |
710 | 748 | | |
711 | 749 | | |
712 | 750 | | |
| |||
910 | 948 | | |
911 | 949 | | |
912 | 950 | | |
913 | | - | |
| 951 | + | |
| 952 | + | |
| 953 | + | |
| 954 | + | |
914 | 955 | | |
915 | | - | |
916 | | - | |
| 956 | + | |
| 957 | + | |
| 958 | + | |
| 959 | + | |
| 960 | + | |
| 961 | + | |
| 962 | + | |
917 | 963 | | |
918 | 964 | | |
919 | 965 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
89 | 89 | | |
90 | 90 | | |
91 | 91 | | |
92 | | - | |
| 92 | + | |
| 93 | + | |
| 94 | + | |
93 | 95 | | |
94 | 96 | | |
95 | 97 | | |
96 | 98 | | |
97 | 99 | | |
98 | 100 | | |
99 | 101 | | |
100 | | - | |
| 102 | + | |
101 | 103 | | |
102 | | - | |
| 104 | + | |
103 | 105 | | |
104 | 106 | | |
105 | 107 | | |
| |||
204 | 206 | | |
205 | 207 | | |
206 | 208 | | |
| 209 | + | |
| 210 | + | |
| 211 | + | |
| 212 | + | |
| 213 | + | |
| 214 | + | |
| 215 | + | |
| 216 | + | |
| 217 | + | |
| 218 | + | |
| 219 | + | |
| 220 | + | |
| 221 | + | |
| 222 | + | |
| 223 | + | |
| 224 | + | |
| 225 | + | |
| 226 | + | |
| 227 | + | |
| 228 | + | |
| 229 | + | |
| 230 | + | |
| 231 | + | |
| 232 | + | |
| 233 | + | |
| 234 | + | |
| 235 | + | |
207 | 236 | | |
208 | 237 | | |
209 | 238 | | |
| |||
310 | 339 | | |
311 | 340 | | |
312 | 341 | | |
| 342 | + | |
| 343 | + | |
| 344 | + | |
313 | 345 | | |
| 346 | + | |
314 | 347 | | |
315 | | - | |
316 | | - | |
| 348 | + | |
| 349 | + | |
317 | 350 | | |
318 | | - | |
| 351 | + | |
319 | 352 | | |
320 | 353 | | |
321 | 354 | | |
| |||
328 | 361 | | |
329 | 362 | | |
330 | 363 | | |
331 | | - | |
| 364 | + | |
| 365 | + | |
| 366 | + | |
| 367 | + | |
332 | 368 | | |
333 | 369 | | |
334 | | - | |
| 370 | + | |
335 | 371 | | |
336 | 372 | | |
337 | 373 | | |
| |||
383 | 419 | | |
384 | 420 | | |
385 | 421 | | |
386 | | - | |
| 422 | + | |
| 423 | + | |
| 424 | + | |
| 425 | + | |
387 | 426 | | |
388 | | - | |
| 427 | + | |
389 | 428 | | |
390 | 429 | | |
391 | 430 | | |
| |||
397 | 436 | | |
398 | 437 | | |
399 | 438 | | |
| 439 | + | |
| 440 | + | |
| 441 | + | |
400 | 442 | | |
401 | 443 | | |
402 | 444 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
155 | 155 | | |
156 | 156 | | |
157 | 157 | | |
158 | | - | |
| 158 | + | |
| 159 | + | |
159 | 160 | | |
160 | 161 | | |
161 | 162 | | |
| |||
Lines changed: 4 additions & 2 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
54 | 54 | | |
55 | 55 | | |
56 | 56 | | |
57 | | - | |
58 | | - | |
| 57 | + | |
| 58 | + | |
| 59 | + | |
| 60 | + | |
59 | 61 | | |
60 | 62 | | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
20 | 20 | | |
21 | 21 | | |
22 | 22 | | |
| 23 | + | |
23 | 24 | | |
24 | 25 | | |
25 | 26 | | |
| |||
0 commit comments