Skip to content

Repository files navigation

pdf-crawler

The goal of pdf-crawler is to download PDF files from web pages for testing PyPDF2.

Install

pip install -r requirements.txt

Usage

It's organized in mostly isolted scripts, e.g.

python crawl.py

starts downloading PDF documents.

About

This project goal is getting a large dataset of PDF documents

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages