Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.
-
Updated
Sep 2, 2026 - Java
Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.
🐋 Web Archiving Integration Layer: One-Click User Instigated Preservation
Parse a Heritrix crawl.log into an XML sitemap
Single Docker container running Heritrix 3, picking up jobs from a directory.
Dockerized Web Curator Tool with Heritrix 3 and pywb
To associate your repository with the heritrix topic, visit your repo's landing page and select "manage topics."