Kotlin for Apache Spark
Kotlin API bindings and extensions.
A curated list of awesome Apache Spark packages and resources.
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Kotlin API bindings and extensions.
.NET bindings.
An alternative R backend, using dplyr.
Haskell on Apache Spark.
Rust bindings.
Golang bindings.
C# bindings.
A scala kernel for Jupyter.
Web-based notebook that enables interactive data analytics with plugable backends, integrated plotting, and extensive Spark support out-of-the-box.
Polynote: an IDE-inspired polyglot notebook. It supports mixing multiple languages in one notebook, and sharing data between them seamlessly. It encourages reproducible notebooks with its immutable data model. Originating from Netflix.
Jupyter magics and kernels for working with remote Spark clusters, for interactively working with remote Spark clusters through Livy, in Jupyter notebooks.
A library that brings useful functions from modern database management systems to Apache Spark.
A Scala library with essential Spark functions and extensions to make you more productive.
A native PySpark implementation of spark-daria.
A library of general purpose functions and UDF's.
joblib backend for running tasks on Spark clusters.
XML parser and writer.
Cassandra support including data source and API and support for arbitrary queries.
Official MongoDB connector.
Storage layer with ACID transactions.
Upserts, Deletes And Incremental Processing on Big Data..
Upserts, Deletes And Incremental Processing on Big Data..
Integration with the lakeFS atomic versioned storage layer.
Cluster computing system for processing large-scale spatial data.
Data frame based graph API.
Bolt protocol based, Neo4j Connector with RDD, DataFrame and GraphX / GraphFrames support.
Declarative machine learning framework on top of Spark.
[status unknown] - linear algebra DSL and optimizer with R-like syntax.
Type safe machine learning pipelines with RDDs.
PMML transformer library for Spark ML.
A system to manage machine learning models for spark.ml and scikit-learn .
H2O interoperability layer.
Distributed Deep Learning library.
Execution engine and serialization format which supports deployment of o.a.s.ml models without dependency on SparkSession.
A distributed ml library with support for LightGBM, Vowpal Wabbit, OpenCV, Deep Learning, Cognitive Services, and Model Deployment.
Machine learning orchestration platform.
REST server with extensive language support (Python, R, Scala), ability to maintain interactive sessions and object sharing.
Simple Spark as a Service which supports objects sharing using so called named objects. JVM only.
IPython protocol based middleware for interactive applications.
A distributed multi-tenant JDBC server for large-scale data processing and analytics, built on top of Apache Spark.
Cross-platform monitoring tool (Spark UI / Spark History Server replacement).
Collection of the streaming connectors excluded from Spark 2.0 (Akka, MQTT, Twitter. ZeroMQ).
Unified data processing engine supporting both batch and streaming applications. Apache Spark is one of the supported execution environments.
Pandas DataFrame API on top of Apache Spark.
Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.
Python API for Deequ.
Collection of base test classes.
A lightweight and fast testing framework.
PySpark test helpers with beautiful error messages.
Open-source toolkit for analyzing web archives.
Workflow management system with Spark backend.
Introduction to Spark API with Spark 3.0 covered. Good source of knowledge about basic concepts.
Useful collection of Spark processing patterns. Accompanying GitHub repository: sryza/aas.
Interesting compilation of notes by Jacek Laskowski. Focused on different aspects of Spark internals.
New book in the Manning's "in action" family with +400 pages. Starts gently, step-by-step and covers large number of topics. Free excerpt on how to setup Eclipse for Spark application development and how to bootstrap a new application using the provided Maven Archetype. You can find the…
Microsoft paper that presents an Apache Spark-based micro-service orchestration framework that extends database operations to include web service primitives.
Paper introducing a core distributed memory abstraction.
Paper introducing relational underpinnings, code generation and Catalyst optimizer.
Structured Streaming is a new high-level streaming API, it is a declarative API based on automatically incrementalizing a static relational query.
Series of five courses (Introduction to Apache Spark, Distributed Machine Learning with Apache Spark, Big Data Analysis with Apache Spark, Advanced Apache Spark for Data Science and Data Engineering, Advanced Distributed Machine Learning with Apache Spark) covering different aspects of software…
Scala oriented introductory course. Part of Functional Programming in Scala Specialization.
Periodical training event organized by the UC Berkeley AMPLab. A source of useful exercise and recorded workshops covering different tools from the Berkeley Data Analytics Stack.
Lambda architecture platform built on Apache Spark and Apache Kafka with specialization for real-time large scale machine learning.
A machine learning library supporting classical Generalized Mixed Model and Generalized Additive Mixed Effect Model.
Machine Learning server for developers and data scientists to build and deploy predictive applications in a fraction of the time.
Data integration platform with extended DataSource API and multi-user environment.
Apache Spark Official Docker images.
PySpark with Jupyter Notebook and Mesos client.
Yarn images from SequenceIQ.
An easy to setup Docker image for Apache Spark from Data Mechanics.
"A place to discuss and ask questions about using Scala for Spark programming" started by @deanwampler.
and Apache Spark Developers List - Mailing lists dedicated to usage questions and development topics respectively.
igorbarinov/awesome-data-engineering
A curated list of data engineering tools for software developers
oxnr/awesome-bigdata
A curated list of awesome big data frameworks, ressources and other awesomeness.
awesomedata/awesome-public-datasets
A topic-centric list of HQ open datasets.
ohenley/awesome-ada
A curated list of awesome resources related to the Ada and SPARK programming language
briatte/awesome-network-analysis
A curated list of awesome network analysis resources.
bytewax/awesome-public-real-time-datasets
A list of publicly available datasets with real-time data maintained by the team at bytewax.io