🐍 Scrape a web page for PDF files and download them all to your local.
-
Updated
Nov 20, 2025 - Python
🐍 Scrape a web page for PDF files and download them all to your local.
A collection of Java APIs for Xpdf - the open source library for operating on PDF files.
🔍 AI-powered web scraper & document intelligence platform — scrape any site (even bot-protected/login-walled), harvest PDF/DOCX/XLSX/CSV, and chat with them using local LLMs via Ollama. Web app · CLI · REST API. 153 tests, TDD.
Python module to scrape information from a PDF file with different data types (eg. tables, graphs) and extract the largest number it can find.
Scrape and download Treasury Bond & Bill (91/182/364-day) result PDFs from the Central Bank of Kenya (CBK). Python, Playwright, SQLite registry to avoid re-downloads. Daily runs via Windows Task Scheduler or cron.
Realiza download automático de documentos publicados no Diário Oficial dos Municípios de Santa Catarina (DOM-SC)
Automated SharePoint PDF extractor using Playwright; captures pages as images and compiles into a single PDF with persistent browser sessions and a Rich terminal UI.
A Customized workflow for scraping vehicle manual PDFs and uploading them to cloud storage.
The Web Scraping Tool is a Python script designed to simplify the process of extracting valuable data from web pages. It empowers users to effortlessly collect various types of information, including links, email addresses, social media links, author names, and phone numbers, from websites of their choice.
SRF2025 Conference PDF Data Extraction and Interactive Browsing Tool.
Extract structured content from PDFs at scale using the Olostep API — the web scraping and data extraction infrastructure used by top AI companies. Supports single and batch scraping with Markdown, HTML, JSON, and text output.
To associate your repository with the pdf-scraper topic, visit your repo's landing page and select "manage topics."