Skip to content

locking: BORG_LOCK_RECHECK_DELAY sets the race recheck delay, fixes #9899 - #10496

Merged
ThomasWaldmann merged 2 commits into
borgbackup:masterfrom
ThomasWaldmann:lock-recheck-delay-9899
Oct 5, 2026
Merged

ThomasWaldmann merged 2 commits into
borgbackup:masterfrom
ThomasWaldmann:lock-recheck-delay-9899

Conversation

@ThomasWaldmann

Copy link
Copy Markdown
Member

Fixes #9899.

acquire() creates its lock object, waits for the race recheck delay (0.01s) and lists the lock objects again to detect other clients racing for the lock.

Why not make the delay slower for "cloud" repos / derive it from measured latency

With list-after-write consistency (a listing started after a write finished contains the written object), the delay is not needed for correctness at all: of two racing clients, at least the one whose lock object was created last sees the other's lock in its second listing, so they can not both get an exclusive lock (they may both back off and retry). Local filesystems, sftp, rest (borg serve --rest), AWS S3 and MinIO give that guarantee - so the 0.01s default stays.

The delay only matters for stores where a new object shows up in listings after a lag W (e.g. NFS clients caching directory listings, some cloud storages via rclone). There, a delay >= W makes the client that created its lock object last still see the other one. W is a property of the storage, not of the request latency: a far-away ssh/rest or AWS S3 repo has high latency and W = 0, while NFS has low latency and W up to the attribute cache timeout. So measuring the time to load the config object would be the wrong proxy.

Changes

  • BORG_LOCK_RECHECK_DELAY (seconds, float) overrides the race recheck delay; invalid values (not a number, negative, nan, inf) abort with an error.
  • storelocking module docstring, internals docs (storelocking section), env var docs (borg help environment + regenerated environment.rst.inc) and the FAQ document the consistency requirement and the knob.
  • Tests for parsing the env var and for acquire() waiting for the configured delay.

🤖 Generated with Claude Code

…orgbackup#9899

acquire() creates its lock object, waits for the race recheck delay and
lists the lock objects again to detect other clients racing for the lock.

With list-after-write consistency (local filesystems, sftp, rest, AWS S3,
MinIO), at least the client that created its lock object last sees the
other's lock in its second listing, so the delay is not needed for
correctness and the 0.01s default stays.

Stores that only list a new object after a lag (e.g. NFS clients caching
directory listings, some cloud storages via rclone) need a delay of at
least that lag. That lag is a property of the storage, not of the request
latency, so it is not measured, but set via BORG_LOCK_RECHECK_DELAY.
Invalid values abort with an error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov

codecov Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 89.08%. Comparing base (f3c82cd) to head (e64154b).
⚠️ Report is 9 commits behind head on master.
✅ All tests successful. No failed tests found.

Additional details and impacted files
@@            Coverage Diff             @@
##           master   #10496      +/-   ##
==========================================
+ Coverage   89.06%   89.08%   +0.01%     
==========================================
  Files         103      103              
  Lines       19685    19757      +72     
  Branches     3081     3094      +13     
==========================================
+ Hits        17533    17600      +67     
- Misses       1490     1494       +4     
- Partials      662      663       +1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

@ThomasWaldmann
ThomasWaldmann merged commit 2850460 into borgbackup:master Oct 5, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

borg2 locking: slower race_recheck_delay

1 participant