Awesome Data Engineering
Section: Batch Processing · Apache Spark's API for graphs and graph-parallel computation.
Entry
Appears in 4 awesome lists
Apache Spark module to perform graph-related parallel computation.
Section: Batch Processing · Apache Spark's API for graphs and graph-parallel computation.
Section: Graph Computing Frameworks · Apache Spark's API for graphs and graph-parallel computation
Section: Software · Apache Spark module to perform graph-related parallel computation.
Section: ↑ Infographics and Data Visualization
"a fast and general-purpose cluster computing system. It provides high-level APIs in Scala, Java, and Python that make parallel jobs easy to write, and an optimized engine that supports general computation graphs. It also supports a rich set of higher-level tools including Shark (Hive on Spark),…
Interactive personal genome analysis toolkit using Claude Code and Python. Parses raw genotyping data from consumer DNA services and analyzes SNPs across 17 categories including health risks, pharmacogenomics, ancestry, and nutrition, with a terminal-style HTML dashboard.
Leading visualization and exploration software for all kinds of graphs and networks.
Substation is a cloud native data pipeline and transformation toolkit written in Go.
Paid technical-computing system for symbolic and numerical mathematics, visualization, programming, and data analysis.
A listener that streams your spark events logs to delight, a free and improved spark UI
Open source leader in AI with a mission to democratize AI for everyone.
Data engineering and classic ML toolkit with batch processing, type coercion, and 7 algorithms in pure Go with zero dependencies.