Awesome Big Data
Section: Applications · Substation is a cloud native data pipeline and transformation toolkit written in Go.
Entry
Appears in 5 awesome lists
Substation is a cloud native data pipeline and transformation toolkit written in Go.
Section: Applications · Substation is a cloud native data pipeline and transformation toolkit written in Go.
Section: Batch Processing · A cloud native data pipeline and transformation toolkit written in Go.
Section: Extract, transform, load (ETL) · Substation is a cloud native data pipeline and transformation toolkit written in Go.
Section: Monitoring / Logging · Substation is a cloud native data pipeline and transformation toolkit written in Go.
Section: Detection, Alerting and Automation Platforms · A cloud native data pipeline and transformation toolkit for security teams.
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG. Features 350+ connectors with always-in-sync data from SharePoint, Google Drive, S3, Kafka, PostgreSQL and more. BSL 1.1 license (becomes Apache 2.0 after 4 years).
"a fast and general-purpose cluster computing system. It provides high-level APIs in Scala, Java, and Python that make parallel jobs easy to write, and an optimized engine that supports general computation graphs. It also supports a rich set of higher-level tools including Shark (Hive on Spark),…
Interactive personal genome analysis toolkit using Claude Code and Python. Parses raw genotyping data from consumer DNA services and analyzes SNPs across 17 categories including health risks, pharmacogenomics, ancestry, and nutrition, with a terminal-style HTML dashboard.
Open source engineering platform to debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards. (Source Code)
Open source mobile & web analytics, push notifications and crash reporting platform, based on Node.js, MongoDB and Linux.
Run and schedule dbt-style SQL transformations without Airflow. Adds data ingestion (50+ sources) and built-in data quality to the transformation layer. Open-source CLI or managed Bruin Cloud for teams who want dbt Cloud-like experience with ingestion included.
is an open-source web application that allows you to create and share documents that contain live code, equations, visualizations and narrative text. Jupyter is used widely in industries that do data cleaning and transformation, numerical simulation, statistical modeling, data visualization, data…