Skip to content
56

Awesome-crawler

A collection of awesome web crawler,spider in different languages

7.3k stars754 forks101 entriesLast push Jun 16, 2024 (2 years ago)License MIT

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Python

Scrapy

A fast high-level screen scraping and web crawling framework.

In 6 listsDetails

django-dynamic-scraper

Creating Scrapy scrapers via the Django admin interface.

In 2 lists

Scrapy-Redis

Redis-based components for Scrapy.

scrapy-cluster

Uses Redis and Kafka to create a distributed on demand scraping cluster.

distribute_crawler

Uses scrapy,redis, mongodb,graphite to create a distributed spider.

pyspider

A powerful spider system.

In 3 lists

CoCrawler

A versatile web crawler built using modern tools and concurrency.

cola

A distributed crawling framework.

Demiurge

PyQuery-based scraping micro-framework.

Scrapely

A pure-python HTML screen-scraping library.

feedparser

Universal feed parser.

you-get

Dumb downloader that scrapes the web.

In 5 listsDetails

MechanicalSoup

A Python library for automating interaction with websites.

portia

Visual scraping for Scrapy.

crawley

Pythonic Crawling / Scraping Framework based on Non Blocking I/O operations.

RoboBrowser

A simple, Pythonic library for browsing the web without a standalone web browser.

MSpider

A simple ,easy spider using gevent and js render.

brownant

A lightweight web data extracting framework.

PSpider

A simple spider frame in Python3.

Gain

Web crawling framework based on asyncio for everyone.

sukhoi

Minimalist and powerful Web Crawler.

spidy

The simple, easy to use command line web crawler.

newspaper

News, full-text, and article metadata extraction in Python 3

In 2 lists

aspider

An async web scraping micro-framework based on asyncio.

Java

ACHE Crawler

An easy to use web crawler for domain-specific search.

Apache Nutch

Highly extensible, highly scalable web crawler for production environment.

In 4 lists

anthelion

A plugin for Apache Nutch to crawl semantic annotations within HTML pages.

In 2 lists

Crawler4j

Simple and lightweight web crawler.

In 2 lists

JSoup

Scrapes, parses, manipulates and cleans HTML.

websphinx

Website-Specific Processors for HTML information extraction.

Open Search Server

A full set of search functions. Build your own indexing strategy. Parsers extract full-text data. The crawlers can index everything.

Gecco

A easy to use lightweight web crawler

WebCollector

Simple interfaces for crawling the Web,you can setup a multi-threaded web crawler in less than 5 minutes.

Webmagic

A scalable crawler framework.

In 3 lists

Spiderman

A scalable ,extensible, multi-threaded web crawler.

Spiderman2

A distributed web crawler framework,support js render.

Heritrix3

Extensible, web-scale, archival-quality web crawler project.

In 2 lists

SeimiCrawler

An agile, distributed crawler framework.

StormCrawler

An open source collection of resources for building low-latency, scalable web crawlers on Apache Storm

Spark-Crawler

Evolving Apache Nutch to run on Spark.

webBee

A DFS web spider.

spider-flow

A visual spider framework, it's so good that you don't need to write any code to crawl the website.

In 2 lists

Norconex Web Crawler

Norconex HTTP Collector is a full-featured web crawler (or spider) that can manipulate and store collected data into a repository of your choice (e.g. a search engine). Can be used as a stand alone application or be embedded into Java applications.

C#

ccrawler

Built in C# 3.5 version. it contains a simple extension of web content categorizer, which can separate between the web page depending on their content.

SimpleCrawler

Simple spider base on mutithreading, regluar expression.

DotnetSpider

This is a cross platfrom, ligth spider develop by C#.

Abot

C# web crawler built for speed and flexibility.

Hawk

Advanced Crawler and ETL tool written in C#/WPF.

SkyScraper

An asynchronous web scraper / web crawler using async / await and Reactive Extensions.

Infinity Crawler

A simple but powerful web crawler library in C#.

JavaScript

scraperjs

A complete and versatile web scraper.

In 2 lists

scrape-it

A Node.js scraper for humans.

In 2 lists

simplecrawler

Event driven web crawler.

In 2 lists

node-crawler

Node-crawler has clean,simple api.

In 3 lists

js-crawler

Web crawler for Node.JS, both HTTP and HTTPS are supported.

In 2 lists

webster

A reliable web crawling framework which can scrape ajax and js rendered content in a web page.

In 2 lists

x-ray

Web scraper with pagination and crawler support.

In 2 lists

node-osmosis

HTML/XML parser and web scraper for Node.js.

In 2 lists

web-scraper-chrome-extension

Web data extraction tool implemented as chrome extension.

In 2 lists

supercrawler

Define custom handlers to parse content. Obeys robots.txt, rate limits and concurrency limits.

In 2 lists

headless-chrome-crawler

Headless Chrome crawls with jQuery support

In 3 lists

Squidwarc

High fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head

In 3 lists

crawlee

A web scraping and browser automation library for Node.js that helps you build reliable crawlers. Fast.

PHP

Goutte

A screen scraping and web crawling library for PHP.

laravel-goutte

Laravel 5 Facade for Goutte.

dom-crawler

The DomCrawler component eases DOM navigation for HTML and XML documents.

QueryList

The progressive PHP crawler framework.

pspider

Parallel web crawler written in PHP.

php-spider

A configurable and extensible PHP web spider.

spatie/crawler

An easy to use, powerful crawler implemented in PHP. Can execute Javascript.

crawlzone/crawlzone

Crawlzone is a fast asynchronous internet crawling framework for PHP.

PHPScraper

PHPScraper is a scraper & crawler built for simplicity.

C++

open-source-search-engine

A distributed open source search engine and spider/crawler written in C/C++.

In 3 lists

C

httrack

Copy websites to your computer.

In 2 lists

Ruby

Nokogiri

A Rubygem providing HTML, XML, SAX, and Reader parsers with XPath and CSS selector support.

In 5 listsDetails

upton

A batteries-included framework for easy web-scraping. Just add CSS(Or do more).

In 2 lists

wombat

Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages.

In 2 lists

RubyRetriever

RubyRetriever is a Web Crawler, Scraper & File Harvester.

Spidr

Spider a site, multiple domains, certain links or infinitely.

In 3 lists

Cobweb

Web crawler with very flexible crawling options, standalone or using sidekiq.

mechanize

Automated web interaction & crawling.

In 2 lists

Rust

spider

The fastest web crawler and indexer.

crawler

A gRPC web indexer turbo charged for performance.

R

rvest

Simple web scraping for R.

In 2 lists

Erlang

ebot

A scalable, distribuited and highly configurable web cawler.

Perl

web-scraper

Web Scraping Toolkit using HTML and CSS Selectors or XPath expressions.

Go

pholcus

A distributed, high concurrency and powerful web crawler.

In 2 lists

gocrawl

Polite, slim and concurrent web crawler.

In 3 lists

fetchbot

A simple and flexible web crawler that follows the robots.txt policies and crawl delays.

go_spider

An awesome Go concurrent Crawler(spider) framework.

In 3 lists

dht

BitTorrent DHT Protocol && DHT Spider.

In 2 lists

ants-go

A open source, distributed, restful crawler engine in golang.

scrape

A simple, higher level interface for Go web scraping.

creeper

The Next Generation Crawler Framework (Go).

In 2 lists

colly

Fast and Elegant Scraping Framework for Gophers.

In 3 lists

ferret

Declarative web scraping.

In 4 lists

Dataflow kit

Extract structured data from web pages. Web sites scraping.

In 3 lists

Hakrawler

Simple, fast web crawler designed for easy, quick discovery of endpoints and assets within a web application

In 5 listsDetails

Scala

crawler

Scala DSL for web crawling.

scrala

Scala crawler(spider) framework, inspired by scrapy.

ferrit

Ferrit is a web crawler service written in Scala using Akka, Spray and Cassandra.

See category
92

Awesome Docker

veggiemonk/awesome-docker

:whale: A curated list of Docker resources and projects

Fresh★ 37k389 entriesPushed 18 days ago
92

Awesome GraphQL

chentsulin/awesome-graphql

Awesome list of GraphQL

Fresh★ 15k483 entriesPushed yesterday
91

Awesome-Kubernetes

ramitsurana/awesome-kubernetes

A curated list for awesome kubernetes sources :ship::tada:

Fresh★ 16k47 entriesPushed 8 days ago
91

Awesome Quant

wilsonfreitas/awesome-quant

A curated list of insanely awesome libraries, packages and resources for Quants (Quantitative Finance)

Fresh★ 30k678 entriesPushed today
89

Awesome Django

wsvincent/awesome-django

A curated list of awesome things related to Django

Fresh★ 11k326 entriesPushed 13 days ago
89

Awesome Terraform

shuaibiyy/awesome-tf

Curated list of resources on HashiCorp's Terraform and OpenTofu

Fresh★ 6.6k472 entriesPushed 2 days ago