Skip to content

chunk referencing a deleted observation stalls client import forever (dependency-safe importer never converges) #1135

Description

@cristiandavidcaminos

Bug: chunk referencing a deleted observation stalls client import forever (dependency-safe importer never gives up)

Symptom

engram sync --cloud --import --project <p> on a client fails deterministically:

engram: dependency-safe cloud import stalled after 1 pass(es); chunk <chunk-id>: apply chunk mutation 2: relation FK precondition not met: referenced observation missing; pending session dependencies: <session-id>

This blocks all subsequent sync cycles for that project on that client (in our setup a systemd timer re-runs it every 5 min; it has been red continuously since the state drift occurred).

Root cause (as observed from the hub Postgres)

Hub-side (v1.20.0), project had:

  1. cloud_chunks row whose payload contains: session upsert, observation upsert (obs-A), and a relation upsert obs-A -> obs-B.
  2. cloud_mutations history shows obs-B was upserted (twice) and later deleted.
  3. No chunk exists that re-upserts obs-B, and no chunk contains a relation delete for the stale edge.

A client that joined/synced after obs-B was deleted has no way to satisfy the relation's FK precondition at chunk-apply time. The importer treats the unsatisfiable FK as a temporary ordering problem ("stalled after N pass(es)") instead of a permanent one, so it can never converge.

Expected behavior (any of)

  • On relation apply, if the referenced observation is deleted or permanently absent, skip + record the edge as a warning (visible), not a hard stall;
  • or: hub garbage-collects relations whose endpoints were deleted when processing the delete mutation;
  • or: at minimum, distinguish "stalled, retry may help" from "unsatisfiable, needs repair" in the message, and offer an operator-side repair command (like the existing cloud upgrade repair) so clients aren't bricked.

Impact

One stale relation in one chunk freezes cloud import for that project/client indefinitely, with a message that reads like a transient condition. Workarounds are manual SQL against the hub, which shouldn't be the supported path.

Environment: hub = Podman postgres:16-alpine + engram 1.20.0 (digest-pinned), clients = Linux over Tailscale, token in env file. Chunk/observation/session ids redacted but available.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions