Awesome LLMOps
Section: ML Compiler · Accessible large language models via k-bit quantization for PyTorch.
Entry
Appears in 4 awesome lists
Bitsandbytes library is a lightweight Python wrapper around CUDA custom functions, in particular 8-bit optimizers, matrix multiplication (LLM.int8()), and 8 & 4-bit quantization functions.
Section: ML Compiler · Accessible large language models via k-bit quantization for PyTorch.
Section: Quantization and training toolkits · Low-bit linear layers and quantized optimizers, including implementations used by LLM.int8() and QLoRA.
Section: 3. Inference Engines & Serving · 8-bit and 4-bit optimizers + quantization.
Section: Computation and Communication Optimisation · Bitsandbytes library is a lightweight Python wrapper around CUDA custom functions, in particular 8-bit optimizers, matrix multiplication (LLM.int8()), and 8 & 4-bit quantization functions.
Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile
State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.
Pure C/C++ inference engine with GGUF format support. The gold standard for CPU/GPU/Apple Silicon on-device running. Includes llama-server for OpenAI-compatible API. Now at 100K+ stars.
Production-grade platform for running any open-source LLMs as OpenAI-compatible API endpoints. Supports 50+ models with built-in streaming, batching, and auto-acceleration. Apache 2.0 licensed.
(MPL-2.0) allows specifying JSON schemas using regular expressions or Pydantic models for constrained decoding. Its high-performance runtime accelerates JSON decoding.
🟢🟠 — Apache-licensed AI and MCP gateway with multi-provider routing, virtual-key access controls, budgets, rate limits, MCP aggregation, OAuth, automatic fallbacks, and load balancing. (Maxim) — note: model-provider credentials and proxied request/response data are sensitive; restrict and…;…
Python-free Rust inference server with OpenAI API compatibility. Supports GGUF and SafeTensors formats with hot model swap, auto-discovery, and single binary deployment for zero-dependency inference. Apache 2.0 licensed.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference…