Skip to content
85

Awesome Web Archiving

An Awesome List for getting started with web archiving

2.7k stars205 forks181 entriesLast push Sep 18, 2026 (12 days ago)License CC0-1.0

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Training/Documentation >Introductions to Web Archiving Concepts

What is a web archive?

A video from the UK Web Archive YouTube Channel

Wikipedia's List of Web Archiving Initiatives

Glossary of Archive-It and Web Archiving Terms

The Web Archiving Lifecycle Model

An attempt to incorporate the technological and programmatic arms of the web archiving into a framework that will be relevant to any organization seeking to archive content from the web. Archive-It, the web archiving service from the Internet Archive, developed the model based on its work with…

Retrieving and Archiving Information from Websites by Wael Eskandar and Brad Murray

Training/Documentation >Training Materials

IIPC and DPC Training materials: module for beginners (8 sessions)

UNT Web Archiving Course

Continuing Education to Advance Web Archiving (CEDWARC)

A Whirlwind Tour of Common Crawl's Datasets using Python

A Whirlwind Tour of Common Crawl's Datasets as a Python notebook

A Whirlwind Tour of Common Crawl's Datasets using Java

Training/Documentation >The WARC Standard

The warc-specifications

A community HTML version of the official specification and hub for new proposals.

Offical ISO 28500 WARC specification homepage

Training/Documentation >For Researchers using Web Archives

GLAM Workbench: Web Archives

See also this related blog post on 'Asking questions with web archives'.

Archives Unleashed Toolkit documentation

Tutorial for Humanities researchers about how to explore Arquivo.pt

Resources for Web Publishers

Definition of Web Archivability

This describes the ease with which web content can be preserved. (Archived version from the Stanford Libraries)

Tools & Software

Comparison of web archiving software

Awesome Website Change Monitoring

Tools & Software >Acquisition

ArchiveBox

A tool which maintains an additive archive from RSS feeds, bookmarks, and links using wget, Chrome headless, and other methods (formerly Bookmark Archiver). (In Development)

In 4 lists

archivenow

A Python library to push web resources into on-demand web archives. (Stable)

ArchiveWeb.Page

A plugin for Chrome and other Chromium based browsers that lets you interactively archive web pages, replay them, and export them as WARC & WACZ files. Also available as an Electron based desktop application.

Auto Archiver

Python script to automatically archive social media posts, videos, and images from a Google Sheets document. Read the article about Auto Archiver on bellingcat.com.

Browsertrix Crawler

A Chromium based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. (Stable)

Brozzler

A distributed web crawler (爬虫) that uses a real browser (Chrome or Chromium) to fetch pages and embedded urls and to extract links. (Stable)

Cairn

A npm package and CLI tool for saving webpages as single HTML files. It does not produce WARC or WACZ archives. (Stable)

Chronicler

Web browser with record and replay functionality. (In Development)

Community Archive

Open Twitter Database and API with tools and resources for building on archived Twitter data.

crau

A lightweight command-line tool for archiving the Web and playing archives: you just need a list of URLs. The name "crau" stems from the Brazilian pronounciation of "crawl". (Stable)

Crawl

A simple web crawler in Golang. (Stable)

crocoite

Crawl websites using headless Google Chrome/Chromium and save resources, static DOM snapshot and page screenshots to WARC files. (In Development)

DiskerNet

A non-WARC-based tool which hooks into the Chrome browser and archives everything you browse making it available for offline replay. (In Development)

F(b)arc

A commandline tool and Python library for archiving data from Facebook using the Graph API. (Stable)

freeze-dry

JavaScript library to turn page into static, self-contained HTML document; useful for browser extensions. (In Development)

grab-site

The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns. (Stable)

In 2 lists

Heritrix

An open source, extensible, web-scale, archival quality web crawler. (Stable)

In 2 lists

Heritrix Walkthrough

(In Development)

html2warc

A simple script to convert offline data into a single WARC file. (Stable)

HTTrack

An open source website copying utility. (Stable)

monolith

CLI tool to save a web page as a single HTML file. (Stable)

In 3 lists

Obelisk

Go package and CLI tool for saving web page as single HTML file. (Stable)

In 2 lists

Scoop

High-fidelity, browser-based, single-page web archiving library and CLI for witnessing the web. (Stable)

SingleFile

Browser extension for Firefox/Chrome and CLI tool to save a faithful copy of a complete page as a single HTML file. (Stable)

In 4 lists

SiteStory

A transactional archive that selectively captures and stores transactions that take place between a web client (browser) and a web server. (Stable)

Social Feed Manager

Open source software that enables users to create social media collections from Twitter, Tumblr, Flickr, and Sina Weibo public APIs. (Stable)

In 2 lists

Squidwarc

An open source, high-fidelity, page interacting archival crawler that uses Chrome or Chrome Headless directly. (In Development)

In 3 lists

StormCrawler

A collection of resources for building low-latency, scalable web crawlers on Apache Storm. (Stable)

twarc

A command line tool and Python library for archiving Twitter JSON data. (Stable)

WAIL

A graphical user interface (GUI) atop multiple web archiving tools intended to be used as an easy way for anyone to preserve and replay web pages; Python, Electron. (Stable)

Warcprox

WARC-writing MITM HTTP/S proxy. (Stable)

WARCreate

A Google Chrome extension for archiving an individual webpage or website to a WARC file. (Stable)

Warcworker

An open source, dockerized, queued, high fidelity web archiver based on Squidwarc with a simple web GUI. (Stable)

Wayback

A toolkit for snapshot webpage to Internet Archive, archive.today, IPFS and beyond. (Stable)

In 7 listsDetails

Waybackpy

Wayback Machine Save, CDX and availability API interface in Python and a command-line tool (Stable)

In 3 lists

Web2Warc

An easy-to-use and highly customizable crawler that enables anyone to create their own little Web archives (WARC/CDX). (Stable)

Web Curator Tool

Open-source workflow management for selective web archiving. (Stable)

WebMemex

Browser extension for Firefox and Chrome which lets you archive web pages you visit. (In Development)

Wget

An open source file retrieval utility that of version 1.14 supports writing warcs. (Stable)

Wget-lua

Wget with Lua extension. (Stable)

Wpull

A Wget-compatible (or remake/clone/replacement/alternative) web downloader and crawler. (Stable)

Tools & Software >Replay

InterPlanetary Wayback (ipwb)

Web Archive (WARC) indexing and replay using IPFS.

In 3 lists

OpenWayback

The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. (Stable)

PYWB

A Python 3 implementation of web archival replay tools, sometimes also known as 'Wayback Machine'. (Stable)

Reconstructive

A Service Worker module for client-side reconstruction of composite mementos by rerouting resource requests to corresponding archived copies (JavaScript).

ReplayWeb.page

A browser-based, fully client-side replay engine for both local and remote WARC & WACZ files. Also available as an Electron based desktop application. (Stable)

warc2html

Converts WARC files to static HTML suitable for browsing offline or rehosting.

Tools & Software >Search & Discovery

hyphe

A webcrawler built for research uses with a graphical user interface in order to build web corpuses made of lists of web actors and maps of links between them. (Stable)

Mink

A Google Chrome extension for querying Memento aggregators while browsing and integrating live-archived web navigation. (Stable)

PANDORÆ

A desktop research software to be plugged on a Solr endpoint to query, retrieve, normalize and visually explore web archives. (Stable)

PastPage

Browser extension for Chrome and Firefox that recovers broken or changed pages by querying the Wayback Machine and other web archives in parallel. (Stable)

playback

A toolkit for searching archived webpages from Internet Archive, archive.today, Memento and beyond. (In Development)

SecurityTrails

Web based archive for WHOIS and DNS records. REST API available free of charge.

In 6 listsDetails

Tempas v1

Temporal web archive search based on Delicious tags. (Stable)

Tempas v2

Temporal web archive search based on links and anchor texts extracted from the German web from 1996 to 2013 (results are not limited to German pages, e.g., Obama@2005-2009 in Tempas). (Stable)

webarchive-discovery

WARC and ARC full-text indexing and discovery tools, with a number of associated tools capable of using the index shown below. (Stable)

Shine

A prototype web archives exploration UI, developed with researchers as part of the Big UK Domain Data for the Arts and Humanities project. (Stable)

SolrWayback

A backend Java and frontend VUE JS project with freetext search and a build in playback engine. Require Warc files has been index with the Warc-Indexer. The web application also has a wide range of data visualization tools and data export tools that can be used on the whole webarchive. SolrWayback…

Warclight

A Project Blacklight based Rails engine that supports the discovery of web archives held in the WARC and ARC formats. (In Development)

Wasp

A fully functional prototype of a personal web archive and search system. (In Development)

Tools & Software >Utilities

ArchiveTools

Collection of tools to extract and interact with WARC files (Python).

bagnabit2warc

Convert a bag-nabit dataset stored in a ZIP into a full-content WARC.

cdx-toolkit

Library and CLI to consult cdx indexes and create WARC extractions of subsets. Abstracts away Common Crawl's unusual crawl structure. (Stable)

duckdb_warc

DuckDB extension to query WARC files. (In Development)

duckdb-web-archive-cdx

DuckDB extension to query the Internet Archive and CommonCrawl CDX APIs directly from SQL. (In Development)

Go Get Crawl

Extract web archive data using Wayback Machine and Common Crawl. (Stable)

In 2 lists

gowarcserver

BadgerDB-based capture index (CDX) and WARC record server, used to index and serve WARC files (Go).

har2warc

Convert HTTP Archive (HAR) -> Web Archive (WARC) format (Python).

In 2 lists

httpreserve.info

Service to return the status of a web page or save it to the Internet Archive. HTTPreserve includes disambiguation of well-known short link services. It returns JSON via the browser or command line via CURL using GET. Describes web sites using earliest and latest dates in the Internet Archive and…

HTTPreserve linkstat

Command line implementation of httpreserve.info to describe the status of a web page. Can be easily scripted and provides JSON output to enable querying through tools like JQ. HTTPreserve Linkstat describes current status, and earliest and latest links on archive.org. (Golang). (Stable)

Internet Archive Library

A command line tool and Python library for interacting directly with archive.org. (Python). (Stable)

httrack2warc

Convert HTTrack archives to WARC format (Java).

MementoMap

A Tool to Summarize Web Archive Holdings (Python). (In Development)

MemGator

A Memento Aggregator CLI and Server (Golang). (Stable)

node-cdxj

CDXJ file parser (Node.js). (Stable)

OutbackCDX

RocksDB-based capture index (CDX) server supporting incremental updates and compression. Can be used as backend for OpenWayback, PyWb and Heritrix. (Stable)

py-wasapi-client

Command line application to download crawls from WASAPI (Python). (Stable)

The Unarchiver

Program to extract the contents of many archive formats, inclusive of WARC, to a file system. Free variant of The Archive Browser (macOS only, Proprietary app).

In 5 listsDetails

tikalinkextract

Extract hyperlinks as a seed for web archiving from folders of document types that can be parsed by Apache Tika (Golang, Apache Tika Server). (In Development)

wasapi-downloader

Java command line application to download crawls from WASAPI. (Stable)

Warchaeology

A collection of tools for inspecting, manipulating, deduplicating and validating WARC-files. (Stable)

warcdb

A command line utility (Python) for importing WARC files into a SQLite database. (Stable)

warcbench

A tool for exploring, analyzing, transforming, recombining, and extracting data from WARC (Web ARChive) files.

warcdedupe

WARC deduplication tool (and WARC library) written in Rust. (In Development)

WARC Explorer

Browser-based inspector for WARC files and records (client-side).

warc-safe

Automatic detection of viruses and NSFW content in WARC files.

WarcPartitioner

Partition (W)ARC Files by MIME Type and Year. (Stable)

warcrefs

Web archive deduplication tools. (Stable)

webarchive-indexing

Tools for bulk indexing of WARC/ARC files on Hadoop, EMR or local file system.

wikiteam

Tools for downloading and preserving wikis. (Stable)

Tools & Software >WARC I/O Libraries

FastWARC

A high-performance WARC parsing library (Python).

HadoopConcatGz

A Splitable Hadoop InputFormat for Concatenated GZIP Files (and *.warc.gz). (Stable)

jwarc

Read and write WARC files with a type safe API (Java).

Jwat

Libraries for reading/writing/validating WARC/ARC/GZIP files (Java). (Stable)

Jwat-Tools

Tools for reading/writing/validating WARC/ARC/GZIP files (Java). (Stable)

node-warc

Parse WARC files or create WARC files using either Electron or chrome-remote-interface (Node.js). (Stable)

Sparkling

Internet Archive's Sparkling Data Processing Library. (Stable)

Unwarcit

Command line interface to unzip WARC and WACZ files (Python).

warc

A Rust library for reading and writing WARC files. (Stable)

Warcat

Tool and library for handling Web ARChive (WARC) files (Python). (Stable)

In 2 lists

Warcat-rs

Command-line tool and Rust library for handling Web ARChive (WARC) files. (In Development)

warcio

Streaming WARC/ARC library for fast web archive IO (Python). (Stable)

warctools

Library to work with ARC and WARC files (Python).

webarchive

Golang readers for ARC and WARC webarchive formats (Golang).

Tools & Software >Analysis

Archives Research Compute Hub

Web application for distributed compute analysis of Archive-It web archive collections. (Stable)

ArchiveSpark

An Apache Spark framework (not only) for Web Archives that enables easy data processing, extraction as well as derivation. (Stable)

Archives Unleashed Notebooks

Notebooks for working with web archives with the Archives Unleashed Toolkit, and derivatives generated by the Archives Unleashed Toolkit. (Stable)

Archives Unleashed Toolkit

An open-source platform for analyzing web archives with Apache Spark. (Stable)

In 2 lists

Common Crawl Columnar Index

SQL-queryable index, with CDX info plus language classification. (Stable)

Common Crawl Web Graph

A host or domain-level graph of the web, with ranking information. (Stable)

Common Crawl Jupyter notebooks

A collection of notebooks using Common Crawl's various datasets. (Stable)

Tweet Archvies Unleashed Toolkit

An open-source toolkit for analyzing line-oriented JSON Twitter archives with Apache Spark. (In Development)

Web Data Commons

Structured data extracted from Common Crawl. (Stable)

Tools & Software >Quality Assurance

Chrome Check My Links

Browser extension: a link checker with more options.

Chrome link checker

Browser extension: basic link checker.

Chrome link gopher

Browser extension: link harvester on a page.

Chrome Open Multiple URLs

Browser extension: opens multiple URLs and also extracts URLs from text.

Chrome Revolver

Browser extension: switches between browser tabs.

FlameShot

Screen capture and annotation on Ubuntu.

In 7 listsDetails

PlayOnLinux

For running Xenu and Notepad++ on Ubuntu.

PlayOnMac

For running Xenu and Notepad++ on macOS.

Windows Snipping Tool

Windows built-in for partial screen capture and annotation. On macOS you can use Command + Shift + 4 (keyboard shortcut for taking partial screen capture).

WineBottler

For running Xenu and Notepad++ on macOS.

xDoTool

Click automation on Ubuntu.

In 2 lists

Xenu

Desktop link checker for Windows.

Tools & Software >Curation

Zotero Robust Links Extension

A Zotero extension that submits to and reads from web archives. Supercedes leonkt/zotero-memento.

Community Resources >Blogs and Scholarship

IIPC Blog

Web Archiving Roundtable

Unofficial blog of the Web Archiving Roundtable of the Society of American Archivists maintained by the members of the Web Archiving Roundtable.

The Web as History

An open-source book that provides a conceptual overview to web archiving research, as well as several case studies.

WS-DL Blog

Web Science and Digital Libraries Research Group blogs about various Web archiving related topics, scholarly work, and academic trip reports.

DSHR's Blog

David Rosenthal regularly reviews and summarizes work done in the Digital Preservation field.

UK Web Archive Blog

Common Crawl Foundation Blog

rss

Community Resources >Mailing Lists

Common Crawl

IIPC

OpenWayback

WASAPI

Community Resources >Slack

IIPC Slack

Ask @netpreserve for access.

Archives Unleashed Slack

Fill out this request form for access to a researcher group of people working with web archives.

Archivers Slack

Invite yourself to a multi-disciplinary effort for archiving projects run in affiliation with EDGI and Data Together.

Common Crawl Foundation Partners

(ask greg zat commoncrawl zot org for an invite)

Community Resources >Discord

Common Crawl Foundation

Community Resources >Twitter

@commoncrawl

Official Common Crawl Foundation handle.

@NetPreserve

Official IIPC handle.

@WebSciDL

ODU Web Science and Digital Libraries Research Group.

#WebArchiving

#WebArchiveWednesday

Web Archiving Service Providers >Self-hostable, Open Source

Browsertrix

From Webrecorder, source available at https://github.com/webrecorder/browsertrix.

Conifer

From Rhizome, source available at https://github.com/Rhizome-Conifer. The hosted service was discontinued in June 2026; existing collections remain accessible.

Web Archiving Service Providers >Hosted, Closed Source

Archive-It

From the Internet Archive.

In 3 lists

Arkiwera

Hanzo

MirrorWeb

PageFreezer

Smarsh

Public Data

Common Crawl files

WARCs, CDX files, parquet url index, parquet host index, etc.

Common Crawl CDX API

Search Common Crawl's CDX URL index.

Dead-Web Index

Reachability labels (alive / blocked / dead) for the top 10 million domains, two probe arms (polite HTTP and a browser TLS fingerprint), 2026. CC BY 4.0, JSONL.

End of Term Archive

WARCs, CDX files, parquet url index.

Internet Archive Wayback

Base URL for IA's Wayback Machine.

Webrecorder US GovArchive

High-fidelity replay.

UK Government Web Archive

Main page for the UKGWA.

See category
94

Awesome OpenClaw Skills

VoltAgent/awesome-openclaw-skills

The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞

Fresh★ 53k830 entriesPushed today
92

Awesome DeepSeek Harness (DSH) Plugin

awesome-dsh-plugin/awesome-dsh-plugin

A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表

Fresh★ 17k1654 entriesPushed today
91

Awesome Guidelines

Kristories/awesome-guidelines

Programming style, best practices, and coding conventions.

Fresh★ 11k166 entriesPushed 2 days ago
90

Awesome

sindresorhus/awesome

😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]

Fresh★ 513k51 entriesPushed 28 days ago
90

Awesome Prompts

ai-boost/awesome-prompts

Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.

Fresh★ 9k288 entriesPushed yesterday
90

Awesome README

matiassingers/awesome-readme

A curated list of awesome READMEs

Fresh★ 22k143 entriesPushed yesterday