Skip to content

Backfill the data half of the Modified date from Edit Journals and Bulk Load Events - #2589

Merged
jh-RLI merged 2 commits into
developfrom
feature-2558-modified-backfill
Oct 2, 2026
Merged

jh-RLI merged 2 commits into
developfrom
feature-2558-modified-backfill

Conversation

@jh-RLI

@jh-RLI jh-RLI commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Closes #2558

Slice 6 of #2551. #2557 added Table.data_modified and metadata_modified, stamped from then on. This fills the data half for Tables that changed before, where a date can be recovered.

What it does

python manage.py backfill_data_modified [--apply], dry run by default.

  • The date is the later of:

    • the latest applied change in the Table's Edit Journal (_<schema>._<name>_insert / _edit / _delete). A change still waiting for Apply never reached the Main Table, so it does not count;
    • the latest successful Bulk Load Event.

    A Table with neither stays NULL and reads "–".

  • Only a NULL is filled, and that test is in the UPDATE's own WHERE. A stamp from any write since the release therefore wins, including one landing during the run, and a second run writes nothing.

    • The issue suggested "NULL or older than the recovered value". I chose the stricter rule: every write after the release stamps the application's clock, which is later than anything a journal holds, so the two rules can only differ by overwriting a real stamp.
  • Never touched: metadata_modified (no metadata save was ever timestamped) and date_updated (for older Tables it holds a date copied out of their metadata).

  • Output: the dry run counts the sources (journal only / bulk only / both / neither / already dated) and how many Tables it would date. --apply reports how many got a date, and how many were stamped meanwhile and left alone.

The OEDB is only read

  • Journals are found in the catalog first (pg_class, with the three columns the read needs), and only those that exist are read.
    • The journal names come from the table proxy, so they are clipped exactly as their creator clipped them.
    • get_sa_table is never called on a Journal Table, because it creates the journal when it is missing (oedb/utils.py, _OedbMetaTable.get_sa_table → _create_if_missing). A test pins that no journal exists after either run.
  • Every read runs in a READ ONLY transaction, one per 200 journals.
  • The read itself is an index scan on each journal's primary key: ORDER BY _id DESC LIMIT 1, because _id comes from one shared sequence. It is batched in a UNION ALL per chunk, not one query per Table.
  • Time zone: _submitted (a timestamp without time zone from the OEDB's now()) is read back AT TIME ZONE current_setting('TimeZone'), the zone it was written in.
  • A journal two Tables' names clip to (names over about 56 characters) cannot say whose rows are whose, so it is skipped and counted.
  • A Table deleted mid-run fails the run before anything is written. Run it again.

Cost

benchmarks/tables_tab/backfill_cost.py uses a throwaway database, like delete_cost.py. Measured locally: 300 Tables, 525 journals of 1,000 rows: 46 ms dry run, 74 ms --apply. A 2,000-Table run was spoiled by another session's API tests emptying the shared sandbox schema mid-run, so it is not quoted.

Deploy

One command, no migration, no internet. Run it once after migrate has applied dataedit.0056, on the host, where the OEDB is reachable. The step is in the vault's deploy checklist for the next release.

Tests

9 tests in dataedit/tests/test_backfill_data_modified.py, through call_command, against real Journal Tables in the sandbox:

  • the four source cases: journal only, bulk only, both, neither;
  • applied versus unapplied journal rows, and a failed Bulk Load Event later than a successful one;
  • the metadata half and date_updated untouched;
  • a Table already dated, whether newer or older than what the journal would give;
  • a second run, and a stamp landing during the run;
  • no journal created by reading;
  • a shared clipped journal.

🤖 Generated with Claude Code

jh-RLI and others added 2 commits October 3, 2026 01:25
`manage.py backfill_data_modified` fills Table.data_modified from the
later of the latest applied Edit Journal change and the latest
successful Bulk Load Event. Tables with neither stay NULL. Dry run by
default; --apply writes.

- Only a NULL is filled, tested in the UPDATE's own WHERE, so a stamp
  from any write since the release wins, including one landing during
  the run, and a second run writes nothing.
- The OEDB is only read: journals are found in the catalog first,
  get_sa_table (which creates a missing journal) is never called, and
  every read runs in a READ ONLY transaction, one per 200 journals. One
  transaction for all of them runs out of the server's lock table at a
  few thousand journals.
- A journal two Tables' names clip to is skipped and counted.
- metadata_modified and date_updated are never touched.

benchmarks/tables_tab/backfill_cost.py times it on a throwaway
database: 300 Tables, 525 journals of 1,000 rows, 46 ms dry run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jh-RLI
jh-RLI merged commit 7bb6307 into develop Oct 2, 2026
4 of 5 checks passed
@jh-RLI
jh-RLI deleted the feature-2558-modified-backfill branch October 2, 2026 23:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Backfill the data half of the Modified date from Edit Journals and Bulk Load Events

1 participant