Skip to content

Entry

Apache Spark

Appears in 7 awesome lists

"a fast and general-purpose cluster computing system. It provides high-level APIs in Scala, Java, and Python that make parallel jobs easy to write, and an optimized engine that supports general computation graphs. It also supports a rich set of higher-level tools including Shark (Hive on Spark),…

Open spark.apache.org

Found in these lists

Awesome Data Engineering

Section: Batch Processing · A multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.

FreshScore 87

AWESOME DATA SCIENCE

Section: Miscellaneous Tools · Lightning-fast cluster computing

FreshScore 92

Awesome ETL

Section: Big Data (Hadoop Stack) · "a fast and general-purpose cluster computing system. It provides high-level APIs in Scala, Java, and Python that make parallel jobs easy to write, and an optimized engine that supports general computation graphs. It also supports a rich set of higher-level tools including Shark (Hive on Spark),…

ActiveScore 72

Awesome MLOps

Section: Data Processing · Unified analytics engine for large-scale data processing.

FreshScore 80

Table of Contents

Section: SQL/NoSQL Tools and Databases · is a unified analytics engine for big data processing, with built-in modules for streaming, SQL, machine learning and graph processing.

SlowScore 61

Awesome Stacks

Section: Streaming Analytics with Kafka, Spark, and Cassandra · 🛠 - 🐙 - Fast and general engine for large-scale data processing.

StaleScore 54

Table of Contents

Section: ML Frameworks, Libraries, and Tools · is a unified analytics engine for large-scale data processing. It provides high-level APIs in Scala, Java, Python, and R, and an optimized engine that supports general computation graphs for data analysis. It also supports a rich set of higher-level tools including Spark SQL for SQL and…

SlowScore 51

Opik

Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…

In 16 listsDetails

Apache Airflow

"Use airflow to author workflows as directed acyclic graphs (DAGs) of tasks. The airflow scheduler executes your tasks on an array of workers while following the specified dependencies. Rich command line utilities make performing complex surgeries on DAGs a snap. The rich user interface makes it…

In 13 listsDetails

Ray

A fast and simple framework for building and running distributed applications. Ray is packaged with RLlib, a scalable reinforcement learning library, and Tune, a scalable hyperparameter tuning library. ray.io

In 13 listsDetails

Gradio

Build and share delightful machine learning apps, all in Python. The de facto standard for creating interactive ML demos with automatic UI generation from function signatures. Powers thousands of Hugging Face Spaces.

In 11 listsDetails

Prefect

Workflow management system that makes it easy to take your data pipelines and add semantics like retries, logging, dynamic mapping, caching, failure notifications, and more.

In 11 listsDetails

MindsDB

MindsDB is an Explainable AutoML framework for developers. With MindsDB you can build, train and use state of the art ML models in as simple as one line of code.

In 11 listsDetails

Polars

| Rust, Python | - Polars is a blazingly fast DataFrames library implemented in Rust using Apache Arrow Columnar Format as memory model.

In 10 listsDetails

Streamlit

The fastest way to build and share data apps. Transform Python scripts into beautiful web applications with minimal code. Widely used for ML model demos, data visualization, and internal tools.

In 10 listsDetails