Skip to content
78

Awesome High Performance Computing

A curated list of awesome high performance computing resources

1.3k stars141 forks736 entriesLast push Sep 15, 2026 (14 days ago)License none

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

General Info >Most Recent List of the Top500 Supercomputers

Top500 (Jun. 2026)

HPCG Top500 (Jun. 2026)

Green500 (Jun. 2026)

io500

General Info >History

History of Supercomputing (Wikipedia)

History of Parallel Computing (Wikipedia)

In 2 lists

History of the Top500 (Wikipedia)

History of LLNL Computing

The Supermen: The Story of Seymour Cray ... (1997)

Unmatched - 50 Years of Supercomputing (2023)

Trends in HPC for AI workloads

AI predictions up to 2027

Software

alpaka

The alpaka library is a header-only C++17 abstraction library for accelerator development

async-rdma

A framework for writing RDMA applications with high-level abstraction and asynchronous APIs

CANN

Compute Architecture for Neural Networks for Huawei Ascend GPUs

CAF

An Open Source Implementation of the Actor Model in C++

In 3 lists

Chapel

A Programming Language for Productive Parallel Computing on Large-scale Systems

In 2 lists

Charm++

Parallel Programming with Migratable Objects

Cilk Plus

C/C++ Extension for Data and Task Parallelism

Codon

high-performance Python compiler that compiles Python code to native machine code without any runtime overhead

In 4 listsDetails

CUDA

High performance NVIDIA GPU acceleration

In 2 lists

CUDA-oxide

custom rustc backend for compiling GPU kernels in pure Rust

dask

Dask provides advanced parallelism for analytics, enabling performance at scale for the tools you love

In 3 lists

DeepSpeed

An easy-to-use deep learning optimization software suite that enables unprecedented scale and speed for Deep Learning Training and Inference

In 7 listsDetails

DeterminedAI

Distributed deep learning

Dispenso

Meta/facebook C++ Task Library

In 2 lists

FastFlow

High-performance Parallel Patterns in C++

Galois

A C++ Library to Ease Parallel Programming with Irregular Parallelism

Halide

A language for fast, portable computation on images and tensors

In 2 lists

Heteroflow

Concurrent CPU-GPU Task Programming using Modern C++

highway

Performance portable SIMD intrinsics

In 2 lists

HIP

HIP is a C++ Runtime API and Kernel Language for AMD/Nvidia GPU

HPC-X

Nvidia implementation of MPI

HPX

A C++ Standard Library for Concurrency and Parallelism

In 2 lists

Horovod

Distributed deep learning training framework for TensorFlow, Keras, PyTorch, and Apache MXNet

In 4 lists

ISPC

An open-source compiler for high-performance SIMD programming on the CPU and GPU

Intel ISPC

SPMD compiler

Intel TBB

Threading Building Blocks

In 2 lists

joblib

Data-flow programming for performance (python)

Kompute

The general purpose GPU compute framework for cross vendor graphics cards (AMD, Qualcomm, NVIDIA & friends)

In 3 lists

Kokkos

A C++ Programming Model for Writing Performance Portable Applications on HPC platforms

In 3 lists

Kubeflow MPI Operator

MPI Operator for Kubeflow

Legate

Nvidia replacement for numpy based on Legion

Legion

Distributed heterogeneous programming library

MAGMA

Next generation linear algebra (LA) GPU accelerated libraries

Merlin

A distributed task queuing system, designed to allow complex HPC workflows to scale to large numbers of simulations

Metal

Apple's GPU API

Microsoft MPI

Microsoft's implementation of MPI

MOGSLib

User defined schedulers

mpi4jax

Zero-copy mpi for jax arrays

mpi4py

Python bindings for MPI

MPI

OpenMPI implementation of the Message passing interface

In 3 lists

MPI

MPICH implementation of the Message passing interface

In 2 lists

MPI Standardization Forum

Forum for MPI standardization

MPAVICH

Implementation of MPI

In 2 lists

NCCL

The NVIDIA Collective Communication Library for multi-GPU and multi-node communication

NVSHMEM

GPU-accelerated implementation of the OpenSHMEM programming model developed by NVIDIA

cuNumeric

GPU drop-in for numpy

stdpar

GPU accelerated C++ from NVIDIA

numba

A JIT compiler that translates a subset of Python into fast machine code

In 2 lists

oneAPI

A unified, multiarchitecture, multi-vendor programming model

OpenACC

"OpenMP for GPUs"

OpenCilk

MIT continuation of Cilk Plus

OpenMP

Multi-platform Shared-memory Parallel Programming in C/C++ and Fortran

In 5 listsDetails

OpenSHMEM

OpenSHMEM is a one-sided, PGAS-based parallel programming model enabling direct remote memory access for high-performance computing

PVM

Parallel Virtual Machine: A predecessor to MPI for distributed computing

PMIX

Standard for process management

Pollux

Message Passing Cloud orchestrator

Pyfi

Distributed flow and computation system

Pyper

concurrent python made simple

RAJA

Architecture and programming model portability for HPC applications

RaftLib

A C++ Library for Enabling Stream and Dataflow Parallel Computation

In 2 lists

ray

Scale AI and Python workloads from reinforcement learning to deep learning

ROCM

First open-source software development platform for HPC/Hyperscale-class GPU computing

RS MPI

Rust bindings for MPI

Scalix

Data parallel computing framework

Simgrid

Simulate cluster/HPC environments

SkelCL

A Skeleton Library for Heterogeneous Systems

STAPL

Standard Template Adaptive Parallel Programming Library in C++

STLab

High-level Constructs for Implementing Multicore Algorithms with Minimized Contention

SYCL

C++ Abstraction layer for heterogeneous devices

Taichi

Parallel programming language for high-performance numerical computations in Python

In 2 lists

Taskflow

A Modern C++ Parallel Task Programming Library

In 6 listsDetails

The Open Community Runtime

Specification for Asynchronous Many Task systems

Transwarp

A Header-only C++ Library for Task Concurrency

In 4 listsDetails

Triton

Triton is a language and compiler for parallel programming

Tuplex

Blazing fast python data science

UCX

Optimized production proven-communication framework

In 2 lists

Zluda

Run unmodified CUDA applications with near-native performance on Intel AMD GPUs.

In 3 lists

HyperQueue

HyperQueue is a tool designed to simplify execution of large workflows (task graphs) on HPC clusters.

In 2 lists

cpuid

A software instruction available on Intel, AMD, and other processors that can be used to determine processor type and features.

cpuid instruction note

A detailed note on the CPUID instruction used for processor identification.

cpufetch

A simple yet fancy CPU architecture fetching tool.

In 3 lists

gpufetch

A tool similar to cpufetch, but for fetching GPU architecture.

intel cpuinfo

Intel tool providing information about the characteristics of Intel CPUs.

Likwid

Provides all information about the supercomputer/cluster.

In 3 lists

LIKWID.jl

Julia wrapper for LIKWID.

openmpi hwloc

Portable Hardware Locality (hwloc) software project.

PRK - Parallel Research Kernels

A collection of kernels for parallel programming research.

ClusterVisor

Cluster management tool by Advanced Clustering.

BeeGFS

A parallel file system designed for performance-critical environments.

Bluebanquise

An open-source cluster management tool.

In 2 lists

NVIDIA Base Command Manager (formerly Bright Cluster Manager)

Software for deploying and managing HPC and AI server clusters.

In 2 lists

Ceph

An open-source distributed storage system.

In 3 lists

DeepOps

Nvidia's GPU infrastructure and automation tools for Kubernetes and Slurm clusters.

In 2 lists

E4S - The Extreme Scale HPC Scientific Stack

A collection of open-source software packages for HPC environments.

Easybuild

A package manager for HPC/supercomputers.

EESSI

A shared stack of scientific software installations.

Flux framework

A framework for high-performance computing clusters.

fpsync

A tool for fast parallel data transfer using fpart and rsync.

GPFS

A high-performance parallel file system developed by IBM.

Guix

A package manager for HPC/supercomputers.

Intel DAOS

A software-defined scale-out object store for HPC applications.

LSF

A batch system for HPC and distributed computing environments.

Lmod

A Lua-based module system for software environment management on HPC systems.

In 3 lists

Lustre Parallel File System

A high-performance distributed filesystem for large-scale cluster computing.

In 3 lists

moosefs

A fault-tolerant, highly available, distributed file system.

In 3 lists

MrPackMod

TACC Alternative to Easybuild/Spack

OKA

Analytics and reporting tool for HPC schedulers: helps administrators and users to understand how compute resources are used, and optimize their usage.

Open Cluster Scheduler

A scalable HPC/AI workload manager based on SGE.

OpenHPC

A community-led set of HPC components.

OpenOnDemand

A web portal for accessing supercomputing resources.

In 2 lists

OpenPBS

A software for workload management and job scheduling.

In 2 lists

OpenXdMod

A tool for managing high-performance computing resources.

RADIUSS

Rapid Application Development via an Institutional Universal Software Stack.

rocks

An open-source Linux cluster distribution.

In 2 lists

Ruse

A tool for managing software environments in HPC clusters.

SGE

A resource management software for large clusters of computers.

Slurm

A cluster management and job scheduling system for Linux clusters.

Spectrum LSF

Workload management platform and job scheduler for distributed high performance computing (HPC)

In 2 lists

Spack

A package manager for HPC/supercomputers.

In 4 listsDetails

sstack

A tool to install multiple software stacks such as Spack, EasyBuild, and Conda.

Starfish

Unstructured data management and metadata solution for files and objects.

Warewulf

An operating system provisioning system and cluster management tool.

Velda

A modern cluster management and job scheduler, with personalizable dev-containers and scale-to-cloud capabilities.

In 2 lists

xCat

A distributed computing management and provisioning tool.

In 2 lists

XDMoD

An open-source tool for managing high-performance computing resources.

Globus Connect

A fast data transfer tool between supercomputers.

Slurm Web

Open source web dashboard for Slurm HPC clusters.

s9s

TUI for SLURM cluster management

lazyslurm

A lazygit-style terminal UI for Slurm. Monitor jobs, tail logs, and inspect nodes and partitions.

In 2 lists

slurm-why

Plain-English diagnoses for confusing Slurm job states (OOM kills, pending-reason codes, impossible requests) instead of decoding raw sacct/squeue/scontrol output by hand.

Kitten

A lightweight kernel designed for high-performance computing. It focuses on providing low noise and predictable performance for HPC applications.

McKernel

A hybrid kernel that combines Linux and a lightweight kernel designed to provide high performance for HPC applications.

mOS

A specialized operating system for high-performance computing, designed to support large-scale, manycore processors.

Apache Airflow

A platform to programmatically author, schedule, and monitor workflows.

In 9 listsDetails

Apptainer (formerly Singularity)

Container platform designed for scientific and high-performance computing (HPC) environments.

arbiter2

Monitors and protects interactive nodes with cgroups.

Charliecloud

Lightweight container solution for high-performance computing (HPC).

Docker

A set of platform as a service products that use OS-level virtualization to deliver software in packages called containers.

In 18 listsDetails

genv

GPU Environment Management for managing and scheduling GPU resources.

Grafana

Open-source platform for monitoring and observability, visualizing metrics.

In 7 listsDetails

grpc

A high-performance, open-source universal RPC framework.

In 4 lists

HPC Rocket

Allows submitting Slurm jobs in Continuous Integration (CI) pipelines.

HTCondor

An open-source high-throughput computing software framework.

Jacamar-ci

CI/CD tool designed for HPC and scientific computing workflows.

Kubernetes

An open-source system for automating deployment, scaling, and management of containerized applications.

In 10 listsDetails

nextflow

A workflow framework to deploy data-driven computational pipelines.

In 3 lists

perun

Energy monitor for HPC systems, focusing on performance and energy efficiency.

In 2 lists

Prefect

A workflow management system, designed for modern infrastructure and powered by the open-source Prefect Core workflow engine.

In 3 lists

Prometheus

An open-source monitoring system with a dimensional data model, flexible query language, efficient time series database and modern alerting approach.

In 13 listsDetails

redun

Workflow engine that emphasizes simplicity, reliability, and scalability.

In 3 lists

remora

Tool for monitoring and reporting the performance of batch jobs on HPC systems.

ruptime

A utility for monitoring the status of computational jobs and systems.

In 2 lists

server-spy

A monitoring tool for multi-run experiments on shared servers. It tracks congestion (PSI, scheduler wait) and shows how much specific runs got slowed down by other users' jobs, helping researchers identify skewed experiment comparisons.

slop

"top"-like command for slurm clusters

Slurmvision slurm dashboard

A dashboard for monitoring and managing Slurm jobs.

slurm docker cluster

A Slurm cluster implemented using Docker containers, for development and testing.

snakemake

A workflow management system that reduces the complexity of creating reproducible and scalable data analyses.

In 2 lists

srunx

Python toolkit for managing SLURM jobs and workflows. Provides a CLI (sbatch / squeue / scancel parity), a FastAPI Web UI, an MCP server for natural-language control from Claude, YAML workflows with parameter sweeps and inter-job dependencies, GPU resource monitoring, SSH-based remote execution…

Stui slurm dashboard for the terminal

A terminal-based UI for managing and monitoring Slurm clusters.

Vaex

A Python library for lazy Out-of-Core DataFrames (similar to Pandas), to visualize and explore big tabular datasets.

In 9 listsDetails

HPC Client GUI

A cross-platform desktop client for SSH, SFTP, remote file management, Slurm job submission and monitoring, terminal workflows, and optional X11 forwarding. Runs locally and requires no server-side installation. Source-available for noncommercial use.

ddt

A powerful debugger designed for developers to solve complex problems on multi-threaded and multi-process environments in HPC.

marmot MPI checker

A tool for detecting and reporting issues in MPI (Message Passing Interface) applications.

python debugging tools

A collection of tools for debugging Python applications, including pdb and other utilities.

seer modern gui for gdb

A graphical user interface for GDB, aiming to improve the debugging experience with modern features and visuals.

In 2 lists

Summary of C/C++ debugging tools

An overview of various debugging tools available for C/C++ applications, focusing on HPC environments.

totalview

A comprehensive source code analysis and debugging tool designed for complex software running on HPC systems, supporting a wide range of languages and architectures.

demonspawn

A framework for automated execution of benchmarks and simulations, designed for HPC environments.

Google benchmark

A microbenchmark support library for C++ that tracks performance over time.

In 4 listsDetails

HPL benchmark

The High Performance Linpack Benchmark for measuring floating-point computing power of systems.

kerncraft

A tool for analytical modeling of loop performance and cache behavior on HPC systems.

NASA parallel benchmark suite

A set of benchmarks designed to evaluate the performance of parallel supercomputers.

papi

Provides standard APIs for accessing hardware performance counters available on modern microprocessors.

scalasca

A software tool that supports performance analysis of large-scale parallel applications.

scalene

A high-performance, high-precision CPU, GPU, and memory profiler for Python.

In 6 listsDetails

Summary of code performance analysis tools

An overview of tools for analyzing HPC application performance.

Summary of profiling tools

A comprehensive list of profiling tools for performance analysis in HPC.

tau

TAU (Tuning and Analysis Utilities) is a profiling and tracing toolkit for performance analysis of parallel programs.

In 2 lists

The Bandwidth Benchmark

A tool for measuring memory bandwidth across various CPUs and systems.

vampir

A tool for detailed analysis of MPI program executions by visualizing their event traces.

bytehound memory profiler

A detailed memory profiler for tracking down memory issues and leaks.

In 3 lists

Flamegraphs

Visualization tool for profiling software, allowing quick identification of performance bottlenecks.

flameox

Profiling and optimization toolkit for agents that captures and compares evidence from native services, PyTorch, GPU kernels, and inference workloads.

In 2 lists

fio

Flexible I/O tester for benchmarking and stress/hardware verification.

IBM Spectrum Scale Key Performance Indicators (KPI)

Provides key performance indicators for IBM Spectrum Scale, aiding in performance tuning and monitoring.

Ior

A parallel file system I/O benchmarking tool used widely in HPC for testing storage systems.

ngstress

A versatile tool for stressing various subsystems of a computer to find hardware faults or to benchmark performance.

In 2 lists

Hotspot

The Linux perf GUI for in-depth performance analysis and visualization of software behavior.

In 3 lists

HPC Challenge Benchmark Suite

benchmark suite that measures a range memory access patterns across CPU/GPU nodes.

mixbench

A benchmark suite designed to evaluate CPUs and GPUs across different compute and memory operations.

pmu-tools (toplev)

Performance monitoring tools for modern Intel CPUs, offering detailed insights into hardware and application performance.

SPEC CPU Benchmark

A benchmark suite designed to provide a comparative measure of compute-intensive performance across the widest practical range of hardware.

STREAM Memory Bandwidth Benchmark

Measures sustainable memory bandwidth and the corresponding computation rate for simple vector kernels.

Intel MPI benchmarks

A set of benchmarks designed to measure the performance and scalability of MPI implementations on Intel architectures.

Ohio state MPI benchmarks

A comprehensive suite of benchmarks for evaluating MPI performance across a variety of message passing patterns and communication protocols.

In 2 lists

hpctoolkit

An integrated suite of tools for measurement and analysis of program performance on computers ranging from desktops to supercomputers.

core-to-core-latency

A diagnostic tool designed to measure and report the latency between CPU cores, aiding in the optimization of parallel computing tasks.

speedscope

An interactive, web-based viewer for performance profiles of software. It supports various formats and provides a flamegraph visualization to identify hot paths efficiently.

In 2 lists

Differential Flamegraphs

A visualization technique developed by Brendan Gregg that highlights differences between performance profiles, making it easier to spot performance regressions or improvements.

Hyperfine

A command-line benchmarking tool that provides a simple and user-friendly means to compare the performance of commands, featuring statistical analysis across multiple runs.

In 6 listsDetails

Openfoam HPC benchmark

A benchmarking suite for evaluating the High Performance Computing capabilities of OpenFOAM, an open-source CFD software, under various computational loads.

fio flexible I/O tester

A versatile tool for I/O workload simulation and benchmarking, capable of testing a wide array of storage and filesystem configurations.

vftrace

A tracing tool specifically designed for the NEC SX-Aurora TSUBASA Vector Engine, enabling detailed performance analysis of vectorized code.

tinymembench

A simple memory benchmark tool, focusing on benchmarking memory bandwidth and latency with minimal dependencies, suitable for various platforms.

Geekbench

Cross platform benchmarking tool

Empirical Roofline Tool (ERT)

Create empirical roofline plots, alternative to intel vtune for any machine

Roofline Visualizer for ERT

Visualizer for ERT

Caliper

A Performance Analysis Toolbox in a Library

KDiskMark

Benchmarking Tool For SSD/HDD Drives

OpenBenchmarking

Open benchmarks on a variety of algorithms and hardware

Phoronix Test Suite

Benchmarking suite for Linux

Palanteer Python/C++ Profiler

Profiler for both Python and C++

In 3 lists

HECBioSim Benchmarks

The HECBioSim benchmark suite consists of a set of simple benchmarks for a number of popular Molecular Dynamics (MD) engines

JUPITER Benchmark Suite

The JUPITER Benchmark Suite incorporates 16 applications from various domains. It was designed for and used in the procurement of JUPITER, the first European exascale supercomputer.

OpenCL Cross-Hardware PhysX Benchmarks

Benchmark across all types of GPU, CPU

MLPerf benchmarks

Flopper.io

Comparison of GPUs

ADIOS2

The Adaptable IO System version 2, designed for flexible and efficient I/O for scientific data, supporting a wide range of HPC simulations.

Amira

A powerful, multifaceted 3D software platform for visualizing, manipulating, and understanding Life Science and bio-medical data coming from all types of sources.

hdf5

The Hierarchical Data Format version 5 (HDF5), is an open source file format that supports large, complex, heterogeneous data.

In 2 lists

paraview

An open-source, multi-platform data analysis and visualization application.

In 4 lists

Scientific Visualization Wiki

A comprehensive guide to the field of scientific visualization, detailing techniques, tools, and applications.

the yt project

An open-source, Python-based package for analyzing and visualizing volumetric data.

In 2 lists

vedo

A lightweight and powerful python module for scientific analysis and visualization of 3D objects and point clouds based on VTK.

In 2 lists

visit

An Open Source, interactive, scalable, visualization, animation and analysis tool.

WebDataset

library for writing I/O pipelines for large datasets.

petsc

Parallel solution of scientific applications modeled by PDEs. (C, 2-clause BSD, GitLab)

In 2 lists

ginkgo

High-performance manycore linear algebra library, focus on sparse systems. (C++, BSD, GitHub)

In 2 lists

GSL

is a numerical library for C and C++ programmers. It is free software under the GNU General Public License. The library provides a wide range of mathematical routines such as random number generators, special functions and least-squares fitting. There are over 1000 functions in total with an…

In 5 listsDetails

Scalapack

trilinos

tnl project

RunMat

MATLAB-syntax runtime with automatic CPU/GPU execution and fused array math kernels.

In 5 listsDetails

mimalloc memory allocator

mimalloc is a compact general purpose allocator with excellent performance.

In 4 listsDetails

jemalloc memory allocator

General purpose malloc(3) implementation that emphasizes fragmentation avoidance and scalable concurrency support. [BSD] website

In 4 listsDetails

tcmalloc memory allocator

Google's fast, multi-threaded malloc implementation. [Apache-2.0] website

In 4 listsDetails

Horde memory allocator

Fast, Scalable, and Memory-efficient Malloc for Linux, Windows, and Mac. [Apache-2.0] website

In 3 lists

Software utilization at UK National Supercomputing Service, ARCHER2

SIMD Info

Comparison of cluster software

List of cluster management software

Hardware >Vendors (Work in Progress)

Dell

Lenovo

PureStorage

Supermicro

VDURA

Hardware >Interconnects/Topology

Ethernet

Infiniband

Network topologies

Battle of the infinibands - Omnipath vs Infiniband

Mellanox infiniband cluster config

RoCE - RDMA Over Converged Ethernet

Slingshot interconnect

CXL - Compute Express Link

Infiniband Essentials

NVlink

List of lan-based interconnect bit rates

List of internet-based interconnect bit rates

Hardware >CPU

Wikichip

Microarchitecture of Intel/AMD CPUs

Apple M1

Apple M2

Apple M2 Teardown

Apply M1/M2 AMX

Apple M3

List of Intel processors

List of Intel micro architectures

Comparison of Intel processors

Comparison of Apple processors

List of AMD processors

List of AMD CPU micro architectures

Comparison of AMD architectures

Hardware >GPU

Inside NVIDIA GPUs: Anatomy of high performance matmul kernels

by Aleksa Gordić

In 2 lists

Gpu Architecture Analysis

A trip through the Graphics Pipeline

🟪 - A somewhat-dated dive into a typical graphics pipeline, intended for those with some exposure to graphics APIs such as OpenGL or Direct3D 11.

In 4 lists

A100 Whitepaper

MIG

Gentle Intro to GPU Inner Workings

In 2 lists

AMD Instinct GPUs

AMD GPU ROCm Support and OS Compatibility

List of AMD GPUs

Comparison of CUDA architectures

Tales of the M1 GPU

List of Intel GPUs

Performance of DGX Cluster

Cuda Ontology

Hardware >TPU/Tensor Cores

Google TPU

TPU Wiki

NVIDIA Tensor Cores

Hardware >Many integrated core processor (MIC)

Xeon Phi

Hardware >Cloud

Awesome Cloud HPC

Official NVIDIA Vendors

AWS HPC

Azure HPC

rescale

vast.ai

hetzner - cheap servers incl. 80-core ARM

In 3 lists

Ampere ARM cloud-native processors

Scaleway

Chameleon Cloud

Lambda Labs

NVIDIA Brev

Runpod

The use of Microsoft Azure for high performance cloud computing – A case study

AWS Cluster in the cloud

AWS Parallel Cluster

AWS HPC Workshop

An Empirical Study of Containerized MPI and GUI Application on HPC in the Cloud

Unlocking Python’s Cores: Hardware Usage and Energy Implications of Removing the GIL

Hardware >Custom/FPGA/ASIC/APU

OpenPiton

Parallela

AMD APU

Hardware >Certification

Intel Cluster Ready

Hardware >Student Opportunities / Workshops

Supercomputing Conference Student Opportunities

SCC Student cluster competition

Winter Classic Invitational

Linux Cluster Institute

Hardware >Other/Wikis

Supercomputer

most generally refers to the practice of aggregating computing power in a way that delivers much higher performance than one could get out of a typical desktop computer or workstation in order to solve large problems in science, engineering, or business.

In 2 lists

Supercomputer architecture

Beowulf cluster

Computer cluster

Comparison of Intel processors

Comparison of Apple processors

Comparison of AMD architectures

Comparison of CUDA architectures

Cache

TPU Wiki

IPMI

FRU

Disk Arrays

RAID

Cray

Digital Signal Processors

Vector Processor

Transputer

People

Jack Dongarra - 2021 Turing Award - LINPACK, BLAS, LAPACK, MPI

Bill Gropp - 2010 IEEE TCSC Medal for Excellence in Scalable Computing

David Bader - built the first Linux supercomputer

Thomas Sterling - "Father of Beowulf clusters", ParalleX/HPX

Seymour Cray - Inventor of the Cray Supercomputer

Larry Smarr - HPC Application Pioneer

Donald Becker - Beowulf cluster software, Gordon Bell Prize Winner

HPCWire Class of 2025

HPCWire Class of 2024

Resources

Distributed AI Systems

Fuheng Wu 2026

Supercomputers for Linux SysAdmins

Vladimir Zhumatiy 2025

Grokking concurrency

Kirill Bobrov 2024

High Performance Computing in Biomimetics Modeling, Architecture and Applications

2024

Programming Massively Parallel Processors 4th Edition 2023

Scaling Python with Dask

Holden Karau, Mika Kimmins 2023

Unmatched - 50 Years of Supercomputing (2023)

HPC, Big Data, AI Convergence Towards Exascale: Challenge and Vision

2022

MPI with Python (free)

Gregor von Laszewski, Fidel Leal 2022

Parallel and High Performance Computing

Robert Robey, Yuliana Zamora 2021

In 2 lists

High Performance Parallel Runtimes

Michael Klemm, Jim Cownie 2021

Introduction to High Performance Scientific Computing

Victor Eijkhout 2021

Parallel Programming for Science and Engineering

Victor EIjkhout 2021

Parallel Programming for Science and Engineering - HTML Version

Victor Eijkhout 2021

The Art of Writing Efficient Programs: An Advanced Programmer's Guide to Efficient Hardware Utilization and Compiler…

Fedor Pikus 2021

Computer Organization and Design

2020

C++ High Performance

Björn Andrist, Viktor Sehr 2020

Data Parallel C++ Mastering DPC++ for Programming of Heterogeneous Systems using C++ and SYCL

2020

Systems Performance - Brendan Gregg

2020

Performance Analysis and Tuning on Modern CPUs

Denis Bakhvalov 2020

The OpenMP Common Core: Making OpenMP Simple Again

2019

The Student Supercomputer Challenge Guide

2018

High Performance Computing: Modern Systems and Practices

Thomas Sterling, Maciej Brodowicz, Matthew Anderson 2017

Introduction to Parallel Computing

Zbigniew J. Czech 2017

Practical guide to bare metal C++

Alex Robenko 2017

Raspberry Pi Supercomputing and Scientific Programming

2017

Parallel Computing: Theory and Practice

Umut A. Acar 2016

In 2 lists

OpenMP Examples - openmp.org

2016

Problem-solving in High Performance Computing: A Situational Awareness Approach with Linux

Igor Ljubuncic 2015

Power and Performance_ Software Analysis and Optimization

Jim Kukunas 2015

High Performance Python

Micha Gorelick, Ian Ozsvald 2014

Gropp books on MPI

2014

C++ Concurrency in Action: Practical Multithreading

Anthony Williams 2012

The Art of Multiprocessor Programming

Maurice Herlihy 2012

Computer Architecture: A Quantitative Approach

2011

Introduction to High Performance Computing for Scientists and Engineers

Hager 2010

The Little Book of Semaphores

Allen B. Downey 2008

In 2 lists

Software Optimization Cookbook

2005

Building Clustered Linux Systems

Robert W. Lucke 2004

Introduction to parallel computing

Ananth Grama 2003

Parallel Programming with MPI

Peter Pacheco 1997

The Supermen: The Story of Seymour Cray ... (1997)

Programming with POSIX threads

David Butenhof 1997

Free Modern HPC Books by Victor Eijkhout

by Victor Eijkhout

In 2 lists

Algorithms for Modern Hardware

Sergey Slotin

In 2 lists

Optimizing software in C++

Agner Fog

Optimizing subroutines in assembly code

Agner Fog

Microarchitecture of Intel/AMD CPUs

The Rust Performance Book

E-Zines on Bash, Linux, Perf, etc - Julia Evans

Latest books on OpemMP - openmp.org

Is Parallel Programming Hard, And, If So, What Can You Do About It? - Paul E. McKenney

The Engineers Guide to C++

HPC Carpentry

Teaching basic skills for high-performance computing.

In 2 lists

Berkeley: Applications of Parallel Computers

Detailed course on HPC

CS6290 High-performance Computer Architecture

Milos Prvulovic and Catherine Gamboa at George Tech

Udacity High Performance Computing

Parallel Numerical Algorithms

Vanderbilt - Intro to HPC

Illinois - Intro to HPC

Creator of PyCuda

Archer1 Courses

TACC tutorials

Livermore training materials

Xsede training materials

Parallel Computation Math

Introduction to High-Performance and Parallel Computing - Coursera

Foundations of HPC 2020/2021

Principles of Distributed Computing

High Performance Visualization

Temple course on building/maintaining a cluster

Nvidia Deep Learning Course

In 2 lists

Coursera GPU Programming Specialization

Coursera Fundamentals of Parallelism on Intel Architecture

Archer2 Shared Memory Programming with OpenMP

Archer2 Message-Passing Programming with MPI

HetSys 2022 Course

Edukamu Introduction to Supercomputing

Heterogeneous Parallel Programming by S K

Supercomputing in plain english

Cornell workshop

Carpentries Incubator HPC Intro

UL HPC School

Introduction to High-Performance Parallel Distributed Computing using Chapel, UPC++ and Coarray Fortran

Performance Engineering off Software Systems (MIT-OCW)

Introduction to Parallel Computing (CMSC 498X/818X)

Infiniband Essentials

Performance Ninja Optimization Course

This is an online course where you can learn and master the skill of low-level performance analysis and tuning.

In 3 lists

HPC Administration Virtual Residency 2024

Programming Parallel Computers

High Performace Machine Learning - Columbia University

HPC.NRW: Various tutorials about Linux, GPU programming, performance tools, and research data management

Introduction to HPC - IIT Bombay

Aalto University course CS-E4580 Programming Parallel Computers

by Jukka Suomela, Samuli Laine, and Jaakko Lehtinen

In 2 lists

HLRS Stuttgart HPC Courses

Rookie HPC Guide

RedHat High Performance Computing 101

Foundations of Multithreaded, Parallel, and Distributed Programming

Building pipelines using slurm dependencies

Writing slurm scripts in python,r and bash

Xsede new user tutorials

Improving Performance with SIMD intrinsics

Want speed? Pass by value

Introduction to low level bit hacks

How to write fast numerical code: An Introduction

Lecture notes on Loop optimizations

A practical approach to code optimization

Software optimization manuals

Guide into OpenMP: Easy multithreading programming for C++

An Introduction to the Partitioned Global Address Space (PGAS) Programming Model

Jax in 2022

C++ Benchmarking for beginners

Mapping MPI ranks to multiple cuda GPU

Oak Ridge National Lab Tutorials

How to perform large scale data processing in bioinformatics

Step by step SGEMM in OpenCL

Frontier User Guide

Allocating large blocks of memory in bare-metal C programming

Hashmap benchmarks 2022

LLNL HPC Tutorials

The dirty secret of high performance computing

Multiple GPUs with pytorch

Brendan Gregg on Linux Performance

Automatic Slurm build scripts

Memory bandwith NapkinMath

Avoiding Instruction Cache Misses

Multi-GPU Programming with Standard Parallel C++

EuroCC National Competence Center Sweden (ENCCS) HPC tutorials

LLNL hpc tutorials

python.org Python Performance Tips

HPC toolset tutorial (cluster management)

OpenMP tutorials

CUDA best practices guide

Understanding CPU Architecture And Performance Using LIKWID

32 OpenMP Traps For C++ Developers

Best practices for running jobs on a HPC cluster

Glossary of HPC related terms

Setting the record straight: What is HPC?

Atomic operations and contention

A concurrency cost hiearchy

Energy vs. performance: Introducing the Z-plot

Georg Hager on visualizing energy-to-solution vs. performance to analyze core count and clock speed tradeoffs.

hpc-wiki.info - Tutorials and articles for HPC users, developers, administrators and specific HPC systems

How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklog

An Introduction to Parallel Programming with MPI and Python

Performance Hints

Best practices for machine learning with HPC

How to pick the right hardware for AI - Gigabyte - Part 1

A practitioner's guide to testing and running large GPU clusters for training generative AI models

AWS HPC Workshop

Hardware Acceleration of LLMs: A comprehensive survey and comparison

The Utralscale Playbook - Training LLMs on GPU Clusters

In 2 lists

Interactive and Urgent HPC Challenges (2024)

The Landscape of Exascale Research: A Data-Driven Literature Analysis (2020)

The Landscape of Parallel Computing Research: A View from Berkeley

Programming for Exascale Computers - Will Gropp, Marc Snir

On the Memory Underutilization: Exploring Disaggregated Memory on HPC Systems (2020)

Advances in Parallel & Distributed Processing, and Applications (conference proceedings)

Designing Heterogeneous Systems: Large Scale Architectural Exploration Via Simulation

Reinventing High Performance Computing: Challenges and Opportunities (2022)

Challenges in Heterogeneous HPC White Paper (2022)

An Evolutionary Technical & Conceptual Review on High Performance Computing Systems (Dec 2021)

New Horizons for High-Performance Computing (2022)

CConfidential High-Performance Computing in the Public Cloud

Containerisation for High Performance Computing Systems: Survey and Prospects

Heterogeneous Computing Systems (2023)

Myths and Legends in High-Performance Computing

Energy-Aware Scheduling for High-Performance Computing Systems: A Survey

Ultimate Physical limits to computation - Seth Lloyd

Myths and Legends in High-Performance Computing

Abstract Machine Models and Proxy Architectures for Exascale Computing, 2014, Sandia National Laboratories and…

Some thoughts on the environmental impact of High Performance Computing

A Research Retrospective on AMD's Exascale Computing Journey

InsideHPC

insideHPC is a global publication recognized for its comprehensive and insightful coverage of the HPC-AI community, linking vendors, end-users and HPC strategists.

In 2 lists

HPCWire

Since 1987 covering the fastest computers in the world and the people who run them.

In 2 lists

NextPlatform

Datacenter Dynamics

Admin Magazine HPC

Toms hardware

Tech Radar

Phoronix

In 2 lists

The Register

This week in HPC

Preparing Applications for Aurora in the Exascale Era

Slurm podcast

HPCPodcast

Join Shahin Khan and Doug Black as they discuss Supercomputing technologies and the applications, markets, and policies that shape them.

In 2 lists

Developer Stories - The path to a career in high performance computing is not always equitable or clear.

Developer Stories - HPCToolkit

Argonne lectures on Extreme Scale Computing 2022

Argonne supercomputer tour

Containers in HPC - what they fix and what they break

HPC Tech Shorts

CppCon

zap: - The C++ conference.

In 2 lists

Create a clustering server

Argonne national lab

Oak Ridge National Lab

Concurrency in C++20 and Beyond

A. Williams

Is Parallel Programming still Hard?

P. McKenney, M. Michael, and M. Wong at CppCon 2017

The Speed of Concurrency: Is Lock-free Faster?

Fedor G Pikus in CppCon 2016

Expressing Parallelism in C++ with Threading Building Blocks

Mike Voss at Intel Webinar 2018

A Work-stealing Runtime for Rust

Aaron Todd in Air Mozilla 2017

C++11/14/17 atomics and memory model: Before the story consumes you

Michael Wong in CppCon 2015

The C++ Memory Model

Valentin Ziegler at C++ Meeting 2014

Sharcnet HPC

Low Latency C++ for fun and profit

scalane python profiler

Kokkos lectures

EasyBuild Tech Talk I - The ABCs of Open MPI, part 1 (by Jeff Squyres & Ralph Castain)

The Spack 2022 Roadmap

A Not So Simple Matter of Software | Talk by Turing Award Winner Prof. Jack Dongarra

Vectorization/SIMD intrinsics

New Silicon for Supercomputers: A Guide for Software Engineers

How to write the perfect hash table

FosDem 2024 HPC Big Data Conference videos

Bright Computing Cluster Management Technical Overview

What is HPC? An introduction by Canonical

Slurm job schedular basics

EasyBuild Tech Talk I - The ABCs of Open MPI, part 1 (by Jeff Squyres & Ralph Castain)

Warewulf HPC Youtube Channel

Scott Meyers: Cpu Caches and Why You Car

Task based Parallelism and why it's awesome

Pedro Gonnet

Tuning Slurm Scheduling for Optimal Responsiveness and Utilization

Parallel Programming Models Overview (2020)

Comparative Analysis of Kokkos and Sycl (Jeff Hammond)

Hybrid OpenMP/MPI Programming

Designs, Lessons and Advice from Building Large Distributed Systems - Jeff Dean (Google)

(Dean)

In 2 lists

Practical Debugging and Performance Engineering

Optimizing a Math Expression Parser in Rust

by Ricardo Pallás

In 2 lists

Performance Hints of the Week

Resources for learning about HPC networks and storage r/HPC

Slurm for dummies guide

Build a cluster under 50k

Build a Beowulf cluster

Build a Raspberry Pi Cluster

Puget Systems

Lambda Labs

Titan computers

Detailed reddit discussion on setting up a small cluster

Tiny titan - build a really cool pi supercomputer

Turing PI - mini PI cluster off the shelf

Building an Intel HPC cluster with OpenHPC

Reddit r/HPC post on building clusters

Build a virtual cluster with PelicanHPC

Building a High-performance Computing Cluster Using FreeBSD

Supermicro GPU racks

VirtualOrfeo - Virtual HPC Cluster

Is there a reason to build a raspberry pi clluster

Building a NVIDIA Jetson Cluster

Building your own HPC using ebay parts

Magic Castle - Terraform modules to replicate HPC in cloud

FireHPC

r/hpc

r/homelab

In 3 lists

r/slurm

HPC University Careers search

HPC wire career site

HPC wire job postings

HPC certification

HPC SysAdmin Jobs (reddit)

The United States Research Software Engineer Association

NCSA Internship

AI and Future HPC Job Prospect

HPC sys admin career (reddit)

ETP4HPC

The SIGHPC Systems Professionals

1024 Cores

Dmitry Vyukov

The Black Art of Concurrency

Internal Pointers

Cluster Monkey

Johnathon Dursi

Arm Vendor HPC blog

HPC Notes

Brendan Gregg Performance Blog

Performance engineering blog

Concurrency Freaks

Servers@Home

Dr.Bandwith Blog

Johnny's Software Lab

by Ivica Bogosavljević

In 2 lists

Daniel Lemire Blog

Gigabyte HPC Blog

IEEE Transactions on Parallel and Distributed Systems (TPDS)

Journal of Parallel and Distributed Computing

ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP)

ACM Symposium on Parallel Algorithms and Architectures (SPAA)

SC conference (SC)

The International Conference for High Performance Computing, Networking, Storage, and Analysis.

In 2 lists

IEEE International Parallel and Distributed Processing Symposium (IPDPS)

International Conference on Parallel Processing (ICPP)

IEEE High Performance Extreme Computing Conference (HPEC)

FosDem

Annual 2-day gathering of F/OSS developers in Brussels which sometimes has a "Lua devroom".

In 4 lists

Energy HPC Conference

International Workshop on OpenMP

HPC Social Discord server

HPC Social slack group

HPC Social

Beowulf Mailing List

Society of Research Software Engineering

Women In HPC

HPC Hallway

The High Performance Computing Special Interest Group

SigHPC

Top500

HPE HPC

HPC Wire

Rookie HPC

HPC_Guru

Jeff Hammond

Advanced Clustering

Redline Performance

R systems

Reddit Entry Level HPC interview help

Reddit HPC Admin Interview help

Prace

Xsede

Compute Canada

Riken CSS

Pawsey

International Data Corporation

A global provider of market intelligence and advisory services.

In 2 lists

List of Federally funded research and development centers

The HPC.NRW competence network of North-Rhine-Westphalia

CINECA Summer HPC School for Heterogeneous Computing

CSC Summer School in High-Performance Computing

IHPCSS Summer school

finding a supercomputer to use for research

Amdahl's Law

HPC Wiki

FLOPS

Computational complexity of math operations

Many Task Computing

High Throughput Computing

Parallel Virtual Machine

OSI Model

Workflow management

Compute Canada Documentation

Network Interface Controller (NIC)

In 2 lists

Just in time compilation

List of distributed computing projects

Computer cluster

Quasi-opportunistic supercomputing

Limits of Computation

Bremermann's Limit

Concurrency patterns

History of Parallel Computing (Wikipedia)

In 2 lists

Server Management

Advanced Parallel Programming in C++

Tools for scientific computing

Quantum Computing for High Performance Computing

Benchmarking data science: Twelve ways to lie with statistics and performance on parallel computers.

Establishing the IO500 Benchmark

NVIDIA High Performance Computing articles

Let's write a superoptimizer

Why I think C++ is still a desirable coding platform compared to Rust

The State of Fortran (arxiv paper 2022)

50 years later, is two phase locking still the best

Estimating your memory bandwith

libsc - Supercomputing library

xbyak jit assembler

A JIT assembler for x86/x64 architectures supporting the latest instruction set extensions such as AVX10.2 and ACE

In 3 lists

cpufetch

A simple yet fancy CPU architecture fetching tool.

In 3 lists

RRZE-HPC

Argonne Github

Argonne Leadership Computing Facility

Oak Ridge National Lab Github

Compute Canada

HPCInfo by Jeff Hammond

Texas Advanced Computing Center (TACC) Github

LANL HPC Github

Rust in HPC

University of Buffalo - Center for Computational Research

Center for High Performance Computing - University of Utah

Top500 Supercomputer Data Analysis

Rust programming language in the high-performance computing environment

Exascale Project

Pocket HPC Survival Guide

Overview of all linear algebra packages

Latency numbers

In 2 lists

Nvidia HPC benchmarks

Intel Intrinsics Guide

AWS Cloud calculator

In 2 lists

Quickly benchmark C++ functions

Quick C++ Benchmarks.

In 2 lists

LLNL Software repository

Boinc - volunteer computing projects

Prace Training Events

Nice discussion on FlameGraph profiling

Nice discussion on parts of a supercomputer on reddit

Technical Report on C++ performance

BOINC Compute for science

Open-source cross-platform volunteer computing network connecting computers of volunteers/citizen scientists to scientific research projects. Used by CERN and universities around the world.

In 2 lists

Count prime numbers using MPI

How to build your LEGO Scafell Pike Supercomputer

Deadlock empire - practice concurrency

Sad Server - practice linux server management

Vim Adventures

Learning Vim while playing a game.

In 2 lists

Other Curated Lists

Awesome Cloud HPC

Parallel Computing Guide

Awesome Parallel Computing

Princeton resources on OpenMP

Awesome HPC

Sig HPC Education

Fortran Codes On Github

Fortran Tools

See category
94

Awesome OpenClaw Skills

VoltAgent/awesome-openclaw-skills

The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞

Fresh★ 53k830 entriesPushed today
92

Awesome DeepSeek Harness (DSH) Plugin

awesome-dsh-plugin/awesome-dsh-plugin

A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表

Fresh★ 17k1654 entriesPushed today
91

Awesome Guidelines

Kristories/awesome-guidelines

Programming style, best practices, and coding conventions.

Fresh★ 11k166 entriesPushed 2 days ago
90

Awesome

sindresorhus/awesome

😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]

Fresh★ 513k51 entriesPushed 28 days ago
90

Awesome Prompts

ai-boost/awesome-prompts

Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.

Fresh★ 9k288 entriesPushed yesterday
90

Awesome README

matiassingers/awesome-readme

A curated list of awesome READMEs

Fresh★ 22k143 entriesPushed yesterday