Skip to content

Commit df87548

Browse files
authored
fix: prod on a dedicated CPU, queued runs no longer shown as running, watchdog no longer amplifies CPU starvation (#175)
* fix(prod): put prod on a dedicated CPU Prod ran on `shared-cpu-1x`, whose CPU burst balance was fully drained. The machine sat pinned at the ~6% baseline: 88% `steal` in /proc/stat, PSI cpu `some avg10=51`, load 1.38 on 1 vCPU. Starving the event loop that badly meant Node could not drain its DB sockets. Every Postgres backend sat in `ClientRead` with zero lock contention, and a plain `select 1` took 0.6-5s against a database that was otherwise healthy. The DB watchdog reads a slow ping as a wedged pool, so it exited the process; each restart then re-armed 109 enabled loops plus misfire catch-up, spiking CPU again on an already-drained balance. Successive lives ran 455s, then 161s, then 100s, so the balance never recovered. Roughly 2h of hard downtime. A dedicated core removes the trigger entirely - measured steal after the switch is 0.2%, load 0.11, and `/api/health/db` is a stable ~0.3s with no flapping. The running machine was already moved with `fly machine update --vm-size performance-1x` to stop the outage; this pins it in config so a redeploy cannot silently put prod back on shared CPU. Not fixed here: the watchdog still cannot distinguish a wedged pool from a starved event loop, so it remains an outage amplifier. Tracked separately. * fix: stop presenting queued runs as running, and stop the watchdog amplifying CPU starvation Two independent defects surfaced by the 2026-08-10 production incident. The VM-size fix that ended the outage is the first commit on this branch; these are the two software problems it exposed. 1) A QUEUED run was presented as a RUNNING one. `toRunSummary` collapsed both open phases into a single flag - `running: r.phase === "pending" || r.phase === "running"` - and `JobSummary.running` came from the equally phase-agnostic `hasOpenRun`. Three surfaces consumed that one flag, so a run merely queued for a machine that was asleep or shut would render a pulsing "Running" badge, put the loop and run pages on their 3s LIVE poll cadence, and disable "Run once" with the tooltip "A run is already in progress". The badge and the tooltip are simply false. The poll cadence is worse than false: a genuinely running run is bounded by RUN_TIMEOUT_MS (~20min), so 3s is self-limiting, but a queued run survives for DEFERRED_MAX_MS (7 days) - so every such page hammered the server at the live rate for as long as the machine stayed away. With dozens of deferred runs fleet-wide that is a permanent, self-inflicted load multiplier. Split the two states at the source: `running` now means executing, `queued` means waiting to be claimed. The fast poll is gated on `running` alone; a queued run gets a calm 15s refresh so it still flips promptly when its machine returns. Queued surfaces render a distinct, still state (no pulse, no elapsed clock) and name the reason, preferring the sweep's own `progress.label` ("deferred - machine offline") over any inference. Both open states still block a second dispatch, because the scheduler refuses to stack two agents on one loop - only the wording changed. 2) The DB watchdog turned CPU starvation into a crash loop. The watchdog exists to escape a wedged pool, and that job is unchanged. But it read any slow `select 1` as a wedge. When the machine ran at ~6% of a core (88% steal), Node could not drain its DB sockets, so the ping blew its 5s deadline against a database that was entirely healthy - every backend idle in `ClientRead`, zero lock contention. The watchdog exited; the restart re-armed every loop and re-fired misfire catch-up, spiking CPU on an already-drained budget. Successive lives ran 455s, 161s, then 100s. The recovery mechanism was the outage. The watchdog now consults event-loop delay before blaming the database. Above the ceiling (default 1s; a healthy server sits in single-digit ms) a failed ping is recorded as INCONCLUSIVE: it neither trips the exit nor clears a real streak. A restart is the right cure for a wedged pool and the wrong cure for a starved CPU. The guard fails toward the old behavior - an unreadable lag signal, or no signal wired at all, still blames the database - so it can never suppress a genuine wedge exit. Tunable via LOOPANY_DB_WATCHDOG_LAG_CEILING_MS, 0 to disable. Tests: the two suites that pinned the old collapsed behavior now pin the split; new coverage for the adapter phase mapping, a source-level guard on the poll-cadence and rendering coupling (that coupling is what regresses), and six watchdog cases including the starvation regression and its fail-toward-old-behavior paths. * fix(db): size the connection pool by the pooler's cap, not our appetite `max: 10` against a session pooler that refuses past `pool_size: 15` (`EMAXCONNSESSION`) left no room for a restart. A process killed without a clean shutdown leaves its backends held until TCP keepalive reaps them while the replacement immediately opens its own, so the real worst case is `2*max + 1` - the extra being the prestart migrator, which shares the same budget here because DATABASE_URL and DIRECT_DATABASE_URL both point at the session pooler. At 10 that is 21 against a cap of 15: refusals during any restart, which is what the 2026-08-10 crash loop produced (observed 14/15 occupied, new connections refused). Steady-state demand was never the constraint. Sampling prod once a second for a minute: 0 active connections in 56 of 60 samples, 1 in three, peak 4 - while the pool held all 10 open the whole time, because poll traffic keeps round-robining across them so `idle_timeout` never finds a 30s-quiet connection to reap. So the old setting permanently occupied two thirds of the pooler's clients to serve a peak of four. 6 keeps the restart worst case at 13, stays 50% above measured peak demand, and leaves slots for the migrator and an ops session. Queueing behind a smaller pool costs little on a single-vCPU box, where the CPU is the actual limit. The right value follows the pooler's cap, which is external and can change without a deploy, so `LOOPANY_DB_POOL_MAX` overrides it. `poolOptionsFor` stays pure - it takes the size as an argument; only `db/index.ts` reads env. * fix: bound the watchdog's starvation guard, and restore cancel for queued runs Review findings on this branch. Two were real defects in the preceding commit, one of them a regression that commit introduced. The starvation guard had no upper bound. A wedged pool can perfectly well coexist with a busy event loop, and in that case `consecutive` never advanced, so the watchdog never exited - quietly handing back the 2026-07-12 failure mode (~9h down, no auto-recovery) that the watchdog exists to end. The previous commit message claimed the guard "can never suppress a genuine wedge exit"; that was true only for an unreadable or absent lag signal, not for sustained lag alongside a real wedge. The guard is now an excuse, not an alibi: past `starvedCeiling` consecutive inconclusive ticks (default 45, ~15min at the 20s cadence) the watchdog exits anyway. Restarting is a poor cure for a starved CPU but a strictly better outcome than staying wedged forever, and 45 ticks is far enough above the 3-failure threshold that the ~100s crash-loop amplification stays broken. The lag sample could also be stale. `lagMs()` was read only on the failure path, so `monitorEventLoopDelay`'s `max` accumulated across healthy ticks and one old stall - arming 109 loops at boot, say - could sit in the histogram indefinitely and disqualify a much later, genuine failure. It is now read (and reset) on every tick, so the sample always describes the interval that contained the probe. Cancellation regressed for queued runs. Before the queued/running split, `running` covered `pending`, so a queued run showed the stop control; the split left that call site on `running` alone. `cancelRun` accepts both phases, and a queued run is exactly the one worth cancelling - it can wait on an offline machine for days. Restored, and labelled "Cancel run" rather than "Stop run" when nothing is executing. Two queued surfaces also still carried `runPulseStyle`, the infinite live-signal animation, and one still read "Applying your edit" for an edit that had not started. Not a regression - the collapsed flag pulsed for pending runs before this branch too - but it contradicted the split's whole point, so both are now still. Also corrected two comments that overstated their case: the pool worst case is `2*max` (the prestart migrator closes its connection before the server boots, so it overlaps only the dead process's lingering backends), and `monitorEventLoopDelay` does not sample while JS blocks the loop - the overdue sample lands once it resumes, which is why reading `max` still captures the stall. Tests: the starvation ceiling firing, that a healthy ping resets the starved streak, that lag is read on every tick, that a stale spike cannot disqualify a later genuine failure, that no queued surface carries the pulse, and that a queued run keeps its cancel control. Server 893 passed, daemon 360. * fix(env): stop a sub-1 fraction disabling a knob, and tighten the review guards Second review round. No blocking defect in the previous commit - the four fixes audited clean, including the case I was least sure of: alternating starved and non-starved failures cannot stall both counters, because a starved tick preserves `consecutive` while a non-starved one advances it, so the ordinary failure threshold is still reached. One real bug, introduced by the previous commit. `posIntEnv` floored AFTER its positivity test, so `0.5` passed `n > 0` and became `0`. Every knob reads 0 as "disabled", and for the new starved ceiling that silently turned the bounded-recovery guarantee back off, since `starvedTicks >= 0` is true on the first starved tick. Floor first, so a sub-1 fraction falls back to the documented default. The fix is in the shared helper, so it covers every knob in the family, not just the new one. Three items from the same round, all mine and all cosmetic: - The lag-ceiling JSDoc ended up documenting the starved-ceiling function, because the new function was inserted between the doc and its subject. Each has its own doc now. - The run page's confirm dialog still said "Stop this run?" under a button relabelled "Cancel run". Both now follow the run's actual state. - The pool test still carried the `2*max + 1` reasoning corrected in the source. Two tests were weaker than they looked, which is worth more than the assertions they replaced: - The "projected fire" case in timeline.test.ts called `runToMark` twice and never invoked `projectedMark`, so it asserted nothing about projections. It now tests `projectedMark`, and the non-pending phases got their own case. - The source-reading guards sliced between two `indexOf` results without checking either. Removing a marker would yield an empty slice and make every negative assertion pass vacuously - precisely how this kind of guard rots into a no-op. A `between` helper now asserts both markers exist and are ordered. Not addressed here: `cancelRun` reads a run's phase and then updates unconditionally, so a cancel racing a poll's claim can mark a just-claimed run canceled while the daemon still receives it, and can overwrite a run that finished in the gap. It is pre-existing on main, the fix needs a decision about what a losing cancel should report, and it touches lease lifetime. Raised separately. Server 900 passed, daemon 360.
1 parent 0d394c4 commit df87548

23 files changed

Lines changed: 980 additions & 54 deletions

‎fly.prod.toml‎

Lines changed: 18 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -57,6 +57,22 @@ primary_region = "sjc"
5757
timeout = "10s"
5858
grace_period = "30s"
5959

60+
# DEDICATED cpu, deliberately NOT a shared-cpu size - do not move it back.
61+
#
62+
# The 2026-08-10 outage: on `shared-cpu-1x` prod exhausted its CPU burst balance
63+
# and was pinned at the ~6% baseline (measured: 88% `steal` in /proc/stat, PSI
64+
# `some avg10=51`, load 1.38 on 1 vCPU). Node's event loop starved, so it could
65+
# not drain the DB sockets - every Postgres backend sat in `ClientRead` with zero
66+
# lock contention while a plain `select 1` took 0.6-5s. The DB watchdog
67+
# (server/dbWatchdog.ts) read that as a wedged pool and exited the process, and
68+
# every restart re-armed 109 enabled loops + misfire catch-up, spiking CPU again
69+
# on an already-drained balance. Each life got shorter (455s -> 161s -> 100s), so
70+
# the balance never recovered: a death spiral, ~2h of hard downtime.
71+
#
72+
# The watchdog cannot tell "pool wedged" from "event loop starved" (a known gap,
73+
# tracked separately). A dedicated core removes the trigger: steal drops to ~0,
74+
# so CPU starvation can never masquerade as a wedged pool again. Shared CPU is
75+
# also simply the wrong shape here - the scheduler must stay responsive 24/7.
6076
[[vm]]
61-
size = "shared-cpu-1x"
62-
memory = "1024mb"
77+
size = "performance-1x"
78+
memory = "2048mb"

‎packages/server/src/components/LoopCard.tsx‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -70,6 +70,8 @@ export function LoopCard({
7070
Running
7171
</Pill>
7272
)}
73+
{/* Queued is a distinct, still state - no pulse, no "Running" claim. */}
74+
{!job.running && job.queued && <Pill>Queued</Pill>}
7375
{job.graduation && <Pill>{job.graduation}</Pill>}
7476
{completed && (
7577
<Pill tone="success" dot="green">

‎packages/server/src/components/LoopDetailView.tsx‎

Lines changed: 57 additions & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -110,8 +110,12 @@ export function LoopDetailView({ id }: { id: string }) {
110110
.catch(() => {})
111111
}, [id, load])
112112

113-
// Self-poll the page (fast while a run is live), but not mid-edit (don't churn
114-
// the form) or mid-delete (the optimistic tombstone).
113+
// Self-poll the page (fast while a run is EXECUTING), but not mid-edit (don't
114+
// churn the form) or mid-delete (the optimistic tombstone). The fast cadence is
115+
// deliberately gated on `running` alone, never on a merely QUEUED run: a running
116+
// run is bounded by RUN_TIMEOUT_MS (~20min) so 3s is self-limiting, whereas a run
117+
// queued for an offline machine can sit for 7 days - polling that at 3s pinned
118+
// every such page (and the server) at the live cadence indefinitely.
115119
const running = !!detail?.summary.running
116120
useEffect(() => {
117121
if (editing || del.armed) return
@@ -279,9 +283,24 @@ export function LoopDetailView({ id }: { id: string }) {
279283
// A closed loop still working toward its goal (not yet completed).
280284
const closedActive = isClosed(s) && !completed
281285
const onMachine = detail.machine.name ? `“${detail.machine.name}”` : 'the bound machine'
286+
// A QUEUED run is waiting to be claimed, not executing. Name the reason instead
287+
// of the old pulsing "Running" badge, which claimed work was under way while the
288+
// bound machine was shut. The scheduler allows at most one open run per loop, so
289+
// any queued run in the list IS the one being described (order is irrelevant).
290+
const queuedRun = s.queued ? runs.find((r) => r.queued) : undefined
291+
const queuedLabel = !online ? (asleep ? 'Queued - machine asleep' : 'Queued - machine offline') : 'Queued'
292+
// The gateway sweep stamps the concrete reason on the run itself; prefer it over
293+
// our inference, and fall back to a plain explanation when it hasn't stamped yet.
294+
const queuedReason =
295+
queuedRun?.progress?.label ??
296+
(online
297+
? 'Waiting for the machine to pick it up'
298+
: `Waiting for ${onMachine} to reconnect - it runs as soon as it does`)
282299
// The dispatched edit run (once the poll surfaces it) drives the status card.
283300
const editRun = editDispatched ? findEditRun(runs) : undefined
284-
const editSettled = !!editRun && !editRun.running
301+
// "Settled" means the edit pass reached a terminal state - a queued edit run has
302+
// not settled either, so both open states count as still in flight here.
303+
const editSettled = !!editRun && !editRun.running && !editRun.queued
285304
const editInFlight = editDispatched && !editSettled
286305
const dismissEdit = () => {
287306
if (editRun) seenRunIds.current.add(editRun.id) // a dismissed run never re-surfaces as "the" edit
@@ -468,14 +487,19 @@ export function LoopDetailView({ id }: { id: string }) {
468487
<div className="flex flex-wrap items-center gap-2">
469488
<button
470489
className={btnPrimary}
471-
disabled={busy || !online || completed || s.running}
490+
// Both open states block a second dispatch (the scheduler refuses to stack
491+
// agents on one loop), but they are blocked for DIFFERENT reasons, so the
492+
// tooltip must not claim a run is in progress when one is merely queued.
493+
disabled={busy || !online || completed || s.running || s.queued}
472494
onClick={onRun}
473495
title={
474496
s.running
475497
? 'A run is already in progress'
476-
: completed
477-
? 'Loop completed - reopen it to run again'
478-
: offlineHint ?? (job.exec ? 'Spends credits' : undefined)
498+
: s.queued
499+
? queuedReason
500+
: completed
501+
? 'Loop completed - reopen it to run again'
502+
: offlineHint ?? (job.exec ? 'Spends credits' : undefined)
479503
}
480504
aria-label={job.exec ? 'Run once - spends credits' : 'Run once'}
481505
>
@@ -624,6 +648,9 @@ export function LoopDetailView({ id }: { id: string }) {
624648
Running
625649
</Pill>
626650
)}
651+
{!s.running && s.queued && (
652+
<Pill title={queuedReason ?? undefined}>{queuedLabel}</Pill>
653+
)}
627654
{completed ? (
628655
<Pill tone="success" dot="green">
629656
Completed
@@ -701,12 +728,23 @@ export function LoopDetailView({ id }: { id: string }) {
701728
aria-live="polite"
702729
>
703730
<div className="flex flex-wrap items-center gap-x-4 gap-y-2">
731+
{/* The pulse is the LIVE signal light - it may only appear while an agent
732+
is actually working. A waiting edit gets a still dot and says so, which
733+
is the whole point of splitting queued from running. */}
704734
{!editRun ? (
705735
<span className="inline-flex items-center gap-2.5 text-body text-secondary">
706-
<span aria-hidden className="size-1.5 shrink-0 rounded-full" style={runPulseStyle} />
736+
<span aria-hidden className="size-1.5 shrink-0 rounded-full bg-disabled" />
707737
<span className="font-medium text-primary">Edit queued</span>
708738
<span>waiting for {onMachine} to pick it up…</span>
709739
</span>
740+
) : editRun.queued ? (
741+
<span className="inline-flex min-w-0 items-center gap-2.5 text-body text-secondary">
742+
<span aria-hidden className="size-1.5 shrink-0 rounded-full bg-disabled" />
743+
<span className="shrink-0 font-medium text-primary">Edit queued</span>
744+
<span className="truncate">
745+
{editRun.progress?.label ?? `waiting for ${onMachine} to pick it up…`}
746+
</span>
747+
</span>
710748
) : editRun.running ? (
711749
<span className="inline-flex min-w-0 items-center gap-2.5 text-body text-secondary">
712750
<span aria-hidden className="size-1.5 shrink-0 rounded-full" style={runPulseStyle} />
@@ -910,10 +948,18 @@ function RunsSection({
910948
</span>
911949
</span>
912950
<span className="mt-0.5 block">
913-
{x.running && x.progress ? (
951+
{/* A queued run carries the sweep's reason in the same field
952+
("deferred - machine offline"), which is exactly what the
953+
row should say - so show the line for both open states. */}
954+
{(x.running || x.queued) && x.progress ? (
914955
<span className="inline-flex items-center gap-2 text-meta text-secondary">
915-
<span aria-hidden className="size-1.5 rounded-full" style={runPulseStyle} />
916-
<span className="text-disabled">{x.progress.step}</span>
956+
{/* Pulse only while executing; a queued row is still. */}
957+
<span
958+
aria-hidden
959+
className={`size-1.5 rounded-full${x.running ? '' : ' bg-disabled'}`}
960+
style={x.running ? runPulseStyle : undefined}
961+
/>
962+
{x.running && <span className="text-disabled">{x.progress.step}</span>}
917963
<span className="truncate">{x.progress.label}</span>
918964
</span>
919965
) : x.error ? (

‎packages/server/src/components/RunView.tsx‎

Lines changed: 52 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -89,17 +89,19 @@ function RecordedFiles({ run }: { run: RunSummary }) {
8989
function Changes({ run }: { run: RunSummary }) {
9090
const [data, setData] = useState<RunDiffResult | null>(null)
9191
useEffect(() => {
92-
if (run.running) return // snapshot is captured at finalize — nothing to diff yet
92+
// Snapshot is captured at finalize, so neither an executing NOR a queued run
93+
// has anything to diff yet.
94+
if (run.running || run.queued) return
9395
let alive = true
9496
getRunDiff({ data: { runId: run.id } })
9597
.then((d) => alive && setData(d))
9698
.catch(() => alive && setData({ hasSnapshot: false, files: [] }))
9799
return () => {
98100
alive = false
99101
}
100-
}, [run.id, run.running])
102+
}, [run.id, run.running, run.queued])
101103

102-
if (run.running)
104+
if (run.running || run.queued)
103105
return (
104106
<Card label="Changes">
105107
<div className="text-body text-disabled">File changes appear once the run finishes.</div>
@@ -204,6 +206,33 @@ function LiveActivity({ run }: { run: RunSummary }) {
204206
)
205207
}
206208

209+
/**
210+
* The QUEUED counterpart of `LiveActivity`: the run exists but no machine has
211+
* claimed it. Deliberately still, with no pulse and no elapsed clock - both would
212+
* imply work is under way. The waiting time is real and can run to days, so it is
213+
* stated plainly rather than hidden, and the sweep's own reason (`progress.label`,
214+
* e.g. "deferred - machine offline") is shown verbatim when present.
215+
*/
216+
function QueuedActivity({ run, machineName }: { run: RunSummary; machineName: string | null }) {
217+
const waiting = dur(Math.max(0, Date.now() - Date.parse(run.ts)))
218+
const target = machineName ? `“${machineName}”` : 'its machine'
219+
return (
220+
<Card label="Activity">
221+
<div className="flex items-start gap-2.5">
222+
<span aria-hidden className="mt-[5px] size-2 shrink-0 rounded-full bg-disabled" />
223+
<div className="min-w-0 flex-1">
224+
<span className="min-w-0 break-words text-body text-primary">
225+
{run.progress?.label ?? `Queued - waiting for ${target} to pick it up.`}
226+
</span>
227+
<div className="mt-1 text-meta text-disabled">
228+
{waiting ? `Queued for ${waiting} · ` : ''}it runs as soon as {target} reconnects
229+
</div>
230+
</div>
231+
</div>
232+
</Card>
233+
)
234+
}
235+
207236
// The coding agent's session id behind this run — handy for resuming that session
208237
// in your agent (e.g. `claude --resume <id>`) or feeding the auto-evolve context.
209238
// Mono + click-to-copy.
@@ -310,12 +339,16 @@ export function RunDetailView({ loopId, runId }: { loopId: string; runId: string
310339
}, [detail, run, searchDone, loopId, runId])
311340

312341
// Keep a live run streaming in (its transcript + diff settle once it finishes).
342+
// A QUEUED run still needs refreshing - it flips to running the moment its machine
343+
// polls - but at a calm cadence: it can sit queued for days, so the 3s live rate is
344+
// reserved for a run that is actually executing (and is bounded by RUN_TIMEOUT_MS).
313345
const running = !!run?.running
346+
const queued = !run?.running && !!run?.queued
314347
useEffect(() => {
315-
if (!running) return
316-
const t = setInterval(() => void poll(), 3_000)
348+
if (!running && !queued) return
349+
const t = setInterval(() => void poll(), running ? 3_000 : 15_000)
317350
return () => clearInterval(t)
318-
}, [running, poll])
351+
}, [running, queued, poll])
319352

320353
// Unconditional hook call (null sessionId while loading ⇒ renders nothing);
321354
// must sit above the early-return guards below.
@@ -328,10 +361,13 @@ export function RunDetailView({ loopId, runId }: { loopId: string; runId: string
328361

329362
async function onStop() {
330363
if (!run) return
331-
if (!confirm('Stop this run? It will be marked canceled.')) return
364+
// Match the button: a queued run is cancelled before it ever starts, so calling
365+
// that "stopping" would misdescribe what the user is doing.
366+
const verb = run.running ? 'Stop' : 'Cancel'
367+
if (!confirm(`${verb} this run? It will be marked canceled.`)) return
332368
const r = await cancelRun({ data: run.id })
333369
if (r?.error) {
334-
alert(`Stop failed: ${r.error}`)
370+
alert(`${verb} failed: ${r.error}`)
335371
return
336372
}
337373
await load()
@@ -383,9 +419,12 @@ export function RunDetailView({ loopId, runId }: { loopId: string; runId: string
383419
View the whole loop →
384420
</Link>
385421
{continueSession.button}
386-
{run.running && (
422+
{/* Cancellation must cover BOTH open states: the server accepts pending and
423+
running alike, and a queued run is precisely the one worth cancelling -
424+
it can sit waiting on an offline machine for days. */}
425+
{(run.running || run.queued) && (
387426
<button type="button" onClick={onStop} className={btnDanger}>
388-
Stop run
427+
{run.running ? 'Stop run' : 'Cancel run'}
389428
</button>
390429
)}
391430
</div>
@@ -397,6 +436,9 @@ export function RunDetailView({ loopId, runId }: { loopId: string; runId: string
397436
<div className="mt-6 grid grid-cols-1 gap-6 lg:grid-cols-[minmax(0,1fr)_minmax(300px,360px)]">
398437
<div className="flex min-w-0 flex-col gap-6">
399438
{run.running && <LiveActivity run={run} />}
439+
{!run.running && run.queued && (
440+
<QueuedActivity run={run} machineName={detail.machine.name || null} />
441+
)}
400442

401443
{run.message && (
402444
<Card label="Report">

‎packages/server/src/components/Timeline.tsx‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -155,7 +155,8 @@ export function Timeline({
155155
// pulsing RunSeg in `visible` and settles into its finished color in place once
156156
// the report lands. We only need the flag here to suppress the next-run marker
157157
// while it executes (the live block already stands in for "what's next").
158-
const running = !!job.running && atLatest
158+
// Either open state stands in for "what's next", so both suppress the marker.
159+
const running = (!!job.running || !!job.queued) && atLatest
159160

160161
// Right edge: at the live edge we show the next-run marker; otherwise the
161162
// forward "+N" pager for the newer runs currently scrolled out of view.

‎packages/server/src/components/loopGoal.regression.test.ts‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -54,7 +54,9 @@ describe('LoopDetailView closed/completed states', () => {
5454
expect(detail).toContain('completionReason')
5555
})
5656

57-
it('disables Run once while the loop is completed (until reopened) or already running', () => {
58-
expect(detail).toMatch(/disabled=\{busy \|\| !online \|\| completed \|\| s\.running\}/)
57+
it('disables Run once while the loop is completed (until reopened) or already open', () => {
58+
// Both open states block a second dispatch - the scheduler refuses to stack two
59+
// agents on one loop - so `queued` must be in the guard alongside `running`.
60+
expect(detail).toMatch(/disabled=\{busy \|\| !online \|\| completed \|\| s\.running \|\| s\.queued\}/)
5961
})
6062
})

‎packages/server/src/components/loopTimelineLanes.test.ts‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -20,6 +20,7 @@ const mark = (loopId: string, kind: TimelineMark['kind'] = 'run'): TimelineMark
2020
runId: kind === 'run' ? 'r1' : null,
2121
kind,
2222
running: false,
23+
queued: false,
2324
canceled: false,
2425
role: 'exec',
2526
outcome: kind === 'run' ? 'exec' : null,

0 commit comments

Comments
 (0)