Skip to content

Entry

GPUStack

Appears in 5 awesome lists

GPU cluster manager that orchestrates inference engines like vLLM and SGLang. Automated engine selection, parameter optimization, and distributed multi-GPU deployment for high-performance AI workloads.

Open github.comgpustack/gpustack

Found in these lists

awesome-ChatGPT-repositories

Section: Langchain · Manage GPU clusters for running AI models

FreshScore 87

Awesome LLMOps

Section: LLMOps · An open-source GPU cluster manager for running and managing LLMs

ActiveScore 75

Awesome local LLM

Section: Inference engines · simple, scalable AI model deployment on GPU clusters

FreshScore 87

Awesome Open Source AI

Section: 3. Inference Engines & Serving · GPU cluster manager that orchestrates inference engines like vLLM and SGLang. Automated engine selection, parameter optimization, and distributed multi-GPU deployment for high-performance AI workloads.

FreshScore 89

Awesome Production Machine Learning

Section: Computation and Communication Optimisation · GPUStack is an open-source GPU cluster manager for running AI models.

FreshScore 92

LangChain

Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs

In 20 listsDetails

LiteLLM

Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…

In 16 listsDetails

Opik

Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…

In 16 listsDetails

LlamaIndex

(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.

In 14 listsDetails

Dify

February 2026 release making human oversight a native workflow primitive: suspend execution at critical decision points, expose review-and-edit UI mid-flow, and route subsequent execution based on human action (approve/reject/escalate). Demonstrates how HITL transitions from bolt-on approval gates…

In 14 listsDetails

Haystack

Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…

In 13 listsDetails

Ollama

Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile

In 12 listsDetails

vLLM

State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.

In 11 listsDetails