Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 13 additions & 1 deletion docs/internals/data-structures.rst
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,18 @@ cache/
``check --repair`` salvages each pack still recorded corrupt (see :ref:`packs`) and
records a pack that reads intact at the salvage as intact. Records of packs no longer
listed in packs/ are pruned when a check finishes.
checked-archives
archives check results (archive id -> timestamp, result), in the same format as
``checked-packs``. A result is ok if the check could read all metadata of the archive
and found every chunk it references in the chunks index. ``check --max-age`` skips
archives whose ok record is younger than the given age, and partial checks
(``--max-duration``) check the archives without an ok record first, then the
least-recently-checked ones. ``check --repair`` and ``check --verify-data`` check every
archive. All records are removed when a check finds a corrupt or missing pack or a
defect chunk, when it salvages a pack, and when the archives check starts with
``--repair``. No records are stored while a pack is recorded corrupt in
``checked-packs``. Records of archives no longer in archives/ (soft-deleted ones
included) are pruned when an archives check finishes.
referenced-by-archive.<hex-encoded archive ID>
what one archive references (object ID -> plaintext object size), plus the file
count and content size of that archive, in the key's store object envelope (see
Expand Down Expand Up @@ -114,7 +126,7 @@ locks/

.. _store_object_envelope:

The index fragments, the lock objects, ``checked-packs``, the
The index fragments, the lock objects, ``checked-packs``, ``checked-archives``, the
``referenced-by-archive.*`` objects and ``config/defaults`` are stored in the **store object envelope**: the repository key's ``encrypt()``,
exactly as for the metadata and data slots of the objects in a pack (see
:ref:`security_encryption`), with an empty id and an AAD of
Expand Down
5 changes: 3 additions & 2 deletions docs/internals/security.rst
Original file line number Diff line number Diff line change
Expand Up @@ -389,8 +389,9 @@ used:
client never ends up using key material of the attacker's choice.
- ``keys/<store hash>`` -- in ``repokey`` mode, the borg key(s), encrypted with the
passphrase-derived KEK (see :ref:`key_encryption`).
- ``cache/*``: ``cache/checked-packs`` (the ``borg check`` results per pack) and the
per-archive reference caches ``cache/referenced-by-archive.<hex(archive_id)>``
- ``cache/*``: ``cache/checked-packs`` and ``cache/checked-archives`` (the ``borg check``
results per pack and per archive) and the per-archive reference caches
``cache/referenced-by-archive.<hex(archive_id)>``
(written by ``borg compact`` and ``borg analyze``, they list the object ids and
plaintext sizes an archive references) are in the key's store object envelope, like
the index. The envelope binds each object to its repository and name, so the store
Expand Down
114 changes: 110 additions & 4 deletions src/borg/archive.py
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@
from .item import Item, ArchiveItem, ItemDiff
from .platform import acl_get, acl_set, set_flags, get_flags, set_times, swidth
from .hashindex import ChunkIndex, ChunkIndexEntry
from .repository import Repository, PackReader, remove_missing_pack_entries
from .repository import Repository, PackReader, ArchiveTracker, PackTracker, remove_missing_pack_entries
from .repoobj import RepoObj, object_validator

# macOS: SF_DATALESS marks dataless placeholder files (e.g. cloud files not materialized locally).
Expand Down Expand Up @@ -2349,6 +2349,12 @@ def __init__(self):
# ids of the objects with ro_type ROBJ_ARCHIVE_META that verify_data found.
# None if verify_data did not run or was interrupted.
self.archive_meta_ids = None
# the archive check results, see ArchiveTracker. None until check() loads them.
self.tracker = None
# False if this run stores no archive check results, see check().
self.keep_results = True
# True if rebuild_archives stopped at the deadline before it reached every archive.
self.stopped_at_deadline = False

def record_stored(self, results):
"""Add the pack ids in results to written_packs.
Expand Down Expand Up @@ -2392,6 +2398,8 @@ def check(
newer=None,
oldest=None,
newest=None,
max_age=0,
deadline=None,
):
"""Perform a set of checks on 'repository'

Expand All @@ -2403,6 +2411,20 @@ def check(
:param oldest/newest: only check archives older/newer than timedelta from oldest/newest archive timestamp
:param verify_data: integrity verification of data referenced by archives
:param format: format string used to describe an archive in the log output
:param max_age: seconds, 0 = check every archive. Skip the archives whose ok record (see
ArchiveTracker) is younger than max_age. Ignored with repair or verify_data, and if this run
stores no archive check results (see below).
:param deadline: time.monotonic() value at which to stop, None = no limit. The check then
analyzes the archives without an ok record first, then the least recently checked ones, and
stops between two archives. It analyzes at least one archive, as the repository check
verifies at least one pack, so that every run makes progress.

Each analyzed archive is recorded in cache/checked-archives. repair clears the records first,
because it rebuilds the index and rewrites the archives. A run clears the records and stores
none if a pack is recorded corrupt (see PackTracker), or if it runs without repair and has a
finding before it analyzes the archives (a missing pack, content the index rebuild dropped, a
verify_data or find_lost_archives error): the chunk index may name chunks that are lost or
defect then.
"""
if not isinstance(repository, Repository):
logger.error("Checking legacy repositories is not supported.")
Expand All @@ -2413,6 +2435,9 @@ def check(
self.verifying_data = verify_data
self.format = format
self.repository = repository
self.tracker = ArchiveTracker.load(repository)
if repair:
self.tracker.clear()
# A normal (non-repair) archives check trusts the in-repo index: the repository check verified
# each index object's store hash, and the index is the authoritative record of which chunks exist,
# so we do not rebuild it from the packs (reading every pack is far too slow for a routine check).
Expand Down Expand Up @@ -2474,6 +2499,11 @@ def check(
# On Ctrl-C, skip any scan not yet started; a scan already running stops at its own boundary.
if find_lost_archives and not sig_int:
self.rebuild_archives_directory()
if (self.error_found and not repair) or PackTracker.load(repository).corrupt_ids():
self.keep_results = False
self.tracker.clear()
if repair or verify_data or not self.keep_results:
max_age = 0
if not sig_int:
self.rebuild_archives(
match=match,
Expand All @@ -2484,6 +2514,8 @@ def check(
oldest=oldest,
newer=newer,
newest=newest,
max_age=max_age,
deadline=deadline,
)
# finish() writes a consistent chunk index; run it on Ctrl-C too (#9850).
self.finish()
Expand All @@ -2493,7 +2525,12 @@ def check(
else:
logger.info("Archive consistency check interrupted, no problems found so far.")
raise Error("Got Ctrl-C / SIGINT.")
if self.error_found:
if self.stopped_at_deadline:
if self.error_found:
logger.error("Archive consistency check stopped by --max-duration, problems found so far.")
else:
logger.info("Archive consistency check stopped by --max-duration, no problems found so far.")
elif self.error_found:
logger.error("Archive consistency check complete, problems found.")
else:
logger.info("Archive consistency check complete, no problems found.")
Expand Down Expand Up @@ -2822,9 +2859,37 @@ def check_archive_meta(chunk_id):
logger.info("Rebuilding missing archives directory entries completed.")

def rebuild_archives(
self, first=0, last=0, sort_by="", match=None, older=None, newer=None, oldest=None, newest=None
self,
first=0,
last=0,
sort_by="",
match=None,
older=None,
newer=None,
oldest=None,
newest=None,
max_age=0,
deadline=None,
):
"""Analyze and rebuild archives, expecting some damage and trying to make stuff consistent again."""
"""Analyze and rebuild archives, expecting some damage and trying to make stuff consistent again.

max_age, deadline: see check().
"""
tracker = self.tracker
analyzed = 0 # archives analyzed in this run
reused = 0 # archives skipped because of a recent ok record

def record_result(archive_id, found_before, *, complete=True):
"""Record the result of the archive just analyzed, then merge it into error_found.

error_found holds the findings of this archive only while it is analyzed: the loop below
resets it for each archive and passes its previous value as found_before.
An archive that was not analyzed to its end (complete=False) gets a record only if it has a
finding.
"""
if complete or self.error_found:
tracker.record(archive_id, not self.error_found)
self.error_found = self.error_found or found_before

# Missing file chunks, collected during the per-archive checks and reported grouped as
# chunk -> files -> archives after all archives were analyzed. Bounded by
Expand Down Expand Up @@ -3037,6 +3102,16 @@ def robust_item_ids():
else:
archive_infos = self.manifest.archives.list(sort_by=sort_by)
num_archives = len(archive_infos)
if deadline is not None:
# archives without an ok record first, then the least recently checked ones, so that
# repeated runs reach every archive. The sort is stable, so sort_by orders equal keys.
def check_order(info):
entry = tracker.get(info.id)
if entry is None:
return False, 0
return bool(entry.result), entry.timestamp

archive_infos.sort(key=check_order)
formatter = ArchiveFormatter(self.format, self.repository, self.manifest, self.key)

pi = ProgressIndicatorPercent(
Expand All @@ -3050,8 +3125,16 @@ def robust_item_ids():
if sig_int:
# --repair rewrites each archive as a whole, so with --repair the check stops only here.
break
if deadline is not None and analyzed and time.monotonic() >= deadline:
self.stopped_at_deadline = True
break
pi.show(i)
archive_id, archive_id_hex = info.id, bin_to_hex(info.id)
if tracker.is_recent(archive_id, max_age):
reused += 1
continue
analyzed += 1
found_before, self.error_found = self.error_found, False
try:
formatted = formatter.format_item(info, jsonline=False)
# the formatter uses defaults for keys like {comment} if it has no archive metadata.
Expand All @@ -3071,6 +3154,7 @@ def robust_item_ids():
self.manifest.archives.delete_by_id(archive_id)
else:
logger.error(f"Would delete broken archive {info.name} {archive_id_hex}.")
record_result(archive_id, found_before)
continue
cdata = self.repository.get(archive_id)
try:
Expand All @@ -3083,16 +3167,19 @@ def robust_item_ids():
self.manifest.archives.delete_by_id(archive_id)
else:
logger.error(f"Would delete broken archive {info.name} {archive_id_hex}.")
record_result(archive_id, found_before)
continue
archive = self.key.unpack_archive(data)
archive = ArchiveItem(internal_dict=archive)
if archive.version != 2:
raise Exception("Unknown archive metadata version")
items_buffer = ChunkBuffer(self.key)
items_buffer.write_chunk = add_callback
complete = True

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

complete handling is a bit weird.

guess this is a good use case for:

complete = False
for item in ...:
    ...
else:
    complete = True

for item in robust_iterator(archive):
if sig_int and not self.repair:
# without --repair the archive is only read, so the check also stops within it.
complete = False
break
if "chunks" in item:
verify_file_chunks(info.name, item)
Expand All @@ -3111,9 +3198,28 @@ def robust_item_ids():
self.create_archive_entry(info.name, new_archive_id, info.ts)
if archive_id != new_archive_id:
self.manifest.archives.delete_by_id(archive_id)
archive_id = new_archive_id
record_result(archive_id, found_before, complete=complete)
finally:
pi.finish()
report_missing_chunks()
if self.stopped_at_deadline:
# like a pack recorded corrupt, an archive recorded with a problem fails the check until
# a check finds it ok again.
unchecked = archive_infos[analyzed + reused :]
failed_ids = set(tracker.failed_ids())
failed = sum(1 for info in unchecked if info.id in failed_ids)
if failed:
self.error_found = True
logger.error(f"{failed} archive(s) recorded with problems were not checked again in this run.")
logger.info(f"Stopping the archive check after --max-duration, {len(unchecked)} archive(s) left.")
summary = f"Analyzed {analyzed} archive(s)."
if reused:
summary += f" Reused {reused} recent archive check result(s)."
logger.info(summary)
if self.keep_results:
archives = self.manifest.archives
tracker.prune(set(archives.ids()) | set(archives.ids(deleted=True)))

def verify_written_packs(self):
"""Read the object headers of the packs in written_packs and make the chunks index match them.
Expand Down
Loading
Loading