Skip to content

Entry

bitsandbytes

Appears in 4 awesome lists

Bitsandbytes library is a lightweight Python wrapper around CUDA custom functions, in particular 8-bit optimizers, matrix multiplication (LLM.int8()), and 8 & 4-bit quantization functions.

Open github.combitsandbytes-foundation/bitsandbytes

Found in these lists

Awesome LLMOps

Section: ML Compiler · Accessible large language models via k-bit quantization for PyTorch.

ActiveScore 75

Awesome Model Quantization

Section: Quantization and training toolkits · Low-bit linear layers and quantized optimizers, including implementations used by LLM.int8() and QLoRA.

FreshScore 83

Awesome Open Source AI

Section: 3. Inference Engines & Serving · 8-bit and 4-bit optimizers + quantization.

FreshScore 89

Awesome Production Machine Learning

Section: Computation and Communication Optimisation · Bitsandbytes library is a lightweight Python wrapper around CUDA custom functions, in particular 8-bit optimizers, matrix multiplication (LLM.int8()), and 8 & 4-bit quantization functions.

FreshScore 92

Ollama

Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile

In 12 listsDetails

vLLM

State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.

In 11 listsDetails

llama.cpp

Pure C/C++ inference engine with GGUF format support. The gold standard for CPU/GPU/Apple Silicon on-device running. Includes llama-server for OpenAI-compatible API. Now at 100K+ stars.

In 10 listsDetails

OpenLLM

Production-grade platform for running any open-source LLMs as OpenAI-compatible API endpoints. Supports 50+ models with built-in streaming, batching, and auto-acceleration. Apache 2.0 licensed.

In 10 listsDetails

SGLang

(MPL-2.0) allows specifying JSON schemas using regular expressions or Pydantic models for constrained decoding. Its high-performance runtime accelerates JSON decoding.

In 9 listsDetails

Bifrost

🟢🟠 — Apache-licensed AI and MCP gateway with multi-provider routing, virtual-key access controls, budgets, rate limits, MCP aggregation, OAuth, automatic fallbacks, and load balancing. (Maxim) — note: model-provider credentials and proxied request/response data are sensitive; restrict and…;…

In 8 listsDetails

shimmy

Python-free Rust inference server with OpenAI API compatibility. Supports GGUF and SafeTensors formats with hot model swap, auto-discovery, and single binary deployment for zero-dependency inference. Apache 2.0 licensed.

In 7 listsDetails

TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference…

In 7 listsDetails