Skip to content
66

Awesome Spark

A curated list of awesome Apache Spark packages and resources.

1.9k stars345 forks78 entriesLast push Feb 27, 2026 (7 months ago)License CC0-1.0

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Packages >Language Bindings

Kotlin for Apache Spark

Kotlin API bindings and extensions.

.NET for Apache Spark

.NET bindings.

sparklyr

An alternative R backend, using dplyr.

sparkle

Haskell on Apache Spark.

spark-connect-rs

Rust bindings.

spark-connect-go

Golang bindings.

spark-connect-csharp

C# bindings.

Packages >Notebooks and IDEs

almond

A scala kernel for Jupyter.

Apache Zeppelin

Web-based notebook that enables interactive data analytics with plugable backends, integrated plotting, and extensive Spark support out-of-the-box.

In 2 lists

Polynote

Polynote: an IDE-inspired polyglot notebook. It supports mixing multiple languages in one notebook, and sharing data between them seamlessly. It encourages reproducible notebooks with its immutable data model. Originating from Netflix.

In 2 lists

sparkmagic

Jupyter magics and kernels for working with remote Spark clusters, for interactively working with remote Spark clusters through Livy, in Jupyter notebooks.

Packages >General Purpose Libraries

itachi

A library that brings useful functions from modern database management systems to Apache Spark.

spark-daria

A Scala library with essential Spark functions and extensions to make you more productive.

quinn

A native PySpark implementation of spark-daria.

Apache DataFu

A library of general purpose functions and UDF's.

Joblib Apache Spark Backend

joblib backend for running tasks on Spark clusters.

Packages >SQL Data Sources

Spark XML

XML parser and writer.

Spark Cassandra Connector

Cassandra support including data source and API and support for arbitrary queries.

Mongo-Spark

Official MongoDB connector.

Packages >Storage

Delta Lake

Storage layer with ACID transactions.

In 6 listsDetails

Apache Hudi

Upserts, Deletes And Incremental Processing on Big Data..

In 4 listsDetails

Apache Iceberg

Upserts, Deletes And Incremental Processing on Big Data..

In 4 listsDetails

lakeFS

Integration with the lakeFS atomic versioned storage layer.

Packages >Bioinformatics

ADAM

Set of tools designed to analyse genomics data.

In 2 lists

Hail

Genetic analysis framework.

In 2 lists

Packages >GIS

Apache Sedona

Cluster computing system for processing large-scale spatial data.

Packages >Graph Processing

GraphFrames

Data frame based graph API.

neo4j-spark-connector

Bolt protocol based, Neo4j Connector with RDD, DataFrame and GraphX / GraphFrames support.

In 2 lists

Packages >Machine Learning Extension

Apache SystemML

Declarative machine learning framework on top of Spark.

Mahout Spark Bindings

[status unknown] - linear algebra DSL and optimizer with R-like syntax.

KeystoneML

Type safe machine learning pipelines with RDDs.

JPMML-Spark

PMML transformer library for Spark ML.

ModelDB

A system to manage machine learning models for spark.ml and scikit-learn .

Sparkling Water

H2O interoperability layer.

In 2 lists

BigDL

Distributed Deep Learning library.

MLeap

Execution engine and serialization format which supports deployment of o.a.s.ml models without dependency on SparkSession.

In 3 lists

Microsoft ML for Apache Spark

A distributed ml library with support for LightGBM, Vowpal Wabbit, OpenCV, Deep Learning, Cognitive Services, and Model Deployment.

In 2 lists

MLflow

Machine learning orchestration platform.

Packages >Middleware

Livy

REST server with extensive language support (Python, R, Scala), ability to maintain interactive sessions and object sharing.

spark-jobserver

Simple Spark as a Service which supports objects sharing using so called named objects. JVM only.

Apache Toree

IPython protocol based middleware for interactive applications.

Apache Kyuubi

A distributed multi-tenant JDBC server for large-scale data processing and analytics, built on top of Apache Spark.

Packages >Monitoring

Data Mechanics Delight

Cross-platform monitoring tool (Spark UI / Spark History Server replacement).

In 3 lists

Packages >Utilities

sparkly

Helpers & syntactic sugar for PySpark.

Flintrock

A command-line tool for launching Spark clusters on EC2.

Optimus

Data Cleansing and Exploration utilities with the goal of simplifying data cleaning.

Packages >Natural Language Processing

spark-nlp

Natural language processing library built on top of Apache Spark ML.

In 4 listsDetails

Packages >Streaming

Apache Bahir

Collection of the streaming connectors excluded from Spark 2.0 (Akka, MQTT, Twitter. ZeroMQ).

Packages >Interfaces

Apache Beam

Unified data processing engine supporting both batch and streaming applications. Apache Spark is one of the supported execution environments.

In 6 listsDetails

Koalas

Pandas DataFrame API on top of Apache Spark.

In 3 lists

Packages >Data quality

deequ

Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.

In 3 lists

python-deequ

Python API for Deequ.

Packages >Testing

spark-testing-base

Collection of base test classes.

spark-fast-tests

A lightweight and fast testing framework.

chispa

PySpark test helpers with beautiful error messages.

Packages >Web Archives

Archives Unleashed Toolkit

Open-source toolkit for analyzing web archives.

In 2 lists

Packages >Workflow Management

Cromwell

Workflow management system with Spark backend.

In 3 lists

Resources >Books

Learning Spark, 2nd Edition

Introduction to Spark API with Spark 3.0 covered. Good source of knowledge about basic concepts.

Advanced Analytics with Spark

Useful collection of Spark processing patterns. Accompanying GitHub repository: sryza/aas.

Mastering Apache Spark

Interesting compilation of notes by Jacek Laskowski. Focused on different aspects of Spark internals.

Spark in Action

New book in the Manning's "in action" family with +400 pages. Starts gently, step-by-step and covers large number of topics. Free excerpt on how to setup Eclipse for Spark application development and how to bootstrap a new application using the provided Maven Archetype. You can find the…

In 3 lists

Resources >Papers

Large-Scale Intelligent Microservices

Microsoft paper that presents an Apache Spark-based micro-service orchestration framework that extends database operations to include web service primitives.

Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing

Paper introducing a core distributed memory abstraction.

Spark SQL: Relational Data Processing in Spark

Paper introducing relational underpinnings, code generation and Catalyst optimizer.

Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark

Structured Streaming is a new high-level streaming API, it is a declarative API based on automatically incrementalizing a static relational query.

Resources >MOOCS

Data Science and Engineering with Apache Spark (edX XSeries)

Series of five courses (Introduction to Apache Spark, Distributed Machine Learning with Apache Spark, Big Data Analysis with Apache Spark, Advanced Apache Spark for Data Science and Data Engineering, Advanced Distributed Machine Learning with Apache Spark) covering different aspects of software…

Big Data Analysis with Scala and Spark (Coursera)

Scala oriented introductory course. Part of Functional Programming in Scala Specialization.

Resources >Workshops

AMP Camp

Periodical training event organized by the UC Berkeley AMPLab. A source of useful exercise and recorded workshops covering different tools from the Berkeley Data Analytics Stack.

Resources >Projects Using Spark

Oryx 2

Lambda architecture platform built on Apache Spark and Apache Kafka with specialization for real-time large scale machine learning.

In 5 listsDetails

Photon ML

A machine learning library supporting classical Generalized Mixed Model and Generalized Additive Mixed Effect Model.

PredictionIO

Machine Learning server for developers and data scientists to build and deploy predictive applications in a fraction of the time.

In 2 lists

Crossdata

Data integration platform with extended DataSource API and multi-user environment.

Resources >Docker Images

apache/spark

Apache Spark Official Docker images.

jupyter/docker-stacks/pyspark-notebook

PySpark with Jupyter Notebook and Mesos client.

In 3 lists

sequenceiq/docker-spark

Yarn images from SequenceIQ.

datamechanics/spark

An easy to setup Docker image for Apache Spark from Data Mechanics.

Resources >Miscellaneous

Spark with Scala Gitter channel

"A place to discuss and ask questions about using Scala for Spark programming" started by @deanwampler.

Apache Spark User List

and Apache Spark Developers List - Mailing lists dedicated to usage questions and development topics respectively.

See category
87

Awesome Data Engineering

igorbarinov/awesome-data-engineering

A curated list of data engineering tools for software developers

Fresh★ 9.1k313 entriesPushed 22 days ago
84

Awesome Big Data

oxnr/awesome-bigdata

A curated list of awesome big data frameworks, ressources and other awesomeness.

Active★ 15k645 entriesPushed 2 months ago
84

Awesome Public Datasets

awesomedata/awesome-public-datasets

A topic-centric list of HQ open datasets.

Fresh★ 79k3 entriesPushed today
83

Awesome Ada

ohenley/awesome-ada

A curated list of awesome resources related to the Ada and SPARK programming language

Fresh★ 869418 entriesPushed 7 days ago
80

Awesome Network Analysis

briatte/awesome-network-analysis

A curated list of awesome network analysis resources.

Fresh★ 4.1k717 entriesPushed 1 month ago
78

Awesome Public Real-Time Datasets and Sources

bytewax/awesome-public-real-time-datasets

A list of publicly available datasets with real-time data maintained by the team at bytewax.io

Active★ 2.9k91 entriesPushed 2 months ago