Skip to content

Entry

Apache Arrow

Appears in 5 awesome lists

Language-agnostic columnar in-memory format for fast data interchange, including the Arrow IPC format and Flight RPC for moving data between systems.

Open github.comapache/arrow

Found in these lists

Anything About Game

Section: Serialization

FreshScore 85

Awesome Data Analysis

Section: Tools · Universal columnar format and multi-language toolbox for fast data interchange.

FreshScore 80

Awesome Integration

Section: Data Formats · Language-agnostic columnar in-memory format for fast data interchange, including the Arrow IPC format and Flight RPC for moving data between systems.

FreshScore 82

Awesome Production Machine Learning

Section: Data Storage Optimisation · In-memory columnar representation of data compatible with Pandas, Hadoop-based systems, etc..

FreshScore 92

awesome-cpp

Section: Other · Apache Arrow is the universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics

FreshScore 79

Luigi

Python module for building complex pipelines of batch jobs. Handles dependency resolution, workflow management, visualization, and Hadoop integration. Built at Spotify and battle-tested in production. Apache 2.0 licensed.

In 15 listsDetails

Apache Airflow

"Use airflow to author workflows as directed acyclic graphs (DAGs) of tasks. The airflow scheduler executes your tasks on an array of workers while following the specified dependencies. Rich command line utilities make performing complex surgeries on DAGs a snap. The rich user interface makes it…

In 13 listsDetails

Dagster

Cloud-native orchestration platform for developing and maintaining data assets including ML models. Declarative programming model with integrated lineage and observability. Apache 2.0 licensed.

In 12 listsDetails

Prefect

Workflow management system that makes it easy to take your data pipelines and add semantics like retries, logging, dynamic mapping, caching, failure notifications, and more.

In 11 listsDetails

FlatBuffers

An efficient cross-platform serialization library from Google that allows direct access to serialized data without parsing or unpacking.

In 9 listsDetails

protobuf

A language-neutral and platform-neutral serialization mechanism that is designed to be highly efficient and extensible. It supports rich data types and is widely used in distributed systems, such as gRPC and Apache Kafka.

In 9 listsDetails

Apache Spark

Unified analytics engine for large-scale data processing. In-memory cluster computing with high-level APIs in Python, Scala, Java, and R. Powers MLlib for distributed machine learning and Structured Streaming for real-time data. Apache 2.0 licensed.

In 8 listsDetails

Apache Thrift

. [Delphi] Lightweight, language-independent software stack for point-to-point RPC implementation. Thrift provides clean abstractions and implementations for data transport, data serialization, and application level processing. The code generation system takes a simple definition language as input…

In 8 listsDetails