TypeScript AGPL-3.0

browsertrix-crawler

Run a high-fidelity browser-based web archiving crawler in a single Docker container

W

webrecorder

Dernière activité 28 sept. 2026
webrecorder/browsertrix-crawler

1,2 k

étoiles

152

forks

145

issues ouvertes

crawlercrawlingwaczwarcweb-archivingweb-crawlerwebrecorder

Ce README est souvent en anglais.

Browsertrix Crawler 1.x

Browsertrix Crawler is a standalone browser-based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. Browsertrix Crawler uses Puppeteer to control one or more Brave Browser browser windows in parallel. Data is captured through the Chrome Devtools Protocol (CDP) in the browser.

For information on how to use and develop Browsertrix Crawler, see the hosted Browsertrix Crawler documentation.

For information on how to build the docs locally, see the docs page.

Support

Initial support for 0.x version of Browsertrix Crawler, was provided by Kiwix. The initial functionality for Browsertrix Crawler was developed to support the zimit project in a collaboration between Webrecorder and Kiwix, and this project has been split off from Zimit into a core component of Webrecorder.

Additional support for Browsertrix Crawler, including for the development of the 0.4.x version has been provided by Portico.

License

AGPLv3 or later, see LICENSE for more details.

Projets similaires

Browsertrix is the hosted, high-fidelity, browser-based crawling service from Webrecorder designed to make web archiving easier and more accessible for all!

TypeScriptarchivingcloudkubernetes
Wwebrecorder
478 étoiles77

A High-Fidelity Web Archiving Extension for Chrome and Chromium based browsers!

TypeScriptarchivingbrowser-extensionchromium
Wwebrecorder
1,6 k étoiles112

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.

Javaheritrixjavawarc
Iinternetarchive
3,3 k étoiles795