Skip to content

Entry

BitNet

Appears in 5 awesome lists

Official inference framework for 1-bit LLMs (BitNet b1.58). Enables running large models on CPU with minimal memory footprint. Features custom kernels for ternary weight quantization and efficient matmul operations. MIT licensed.

Open github.commicrosoft/bitnet

Found in these lists

Awesome local LLM

Section: Inference engines · official inference framework for 1-bit LLMs

FreshScore 87

Awesome Model Quantization

Section: Inference and hardware · Inference framework for supported native low-bit BitNet models.

FreshScore 83

Awesome Open Source AI

Section: 3. Inference Engines & Serving · Official inference framework for 1-bit LLMs (BitNet b1.58). Enables running large models on CPU with minimal memory footprint. Features custom kernels for ternary weight quantization and efficient matmul operations. MIT licensed.

FreshScore 89

Awesome Generative AI

Section: Developer tools · Official inference framework for 1-bit LLMs, by Microsoft. #opensource

FreshScore 90

awesome-cpp

Section: Other · Official inference framework for 1-bit LLMs

FreshScore 79

Ollama

Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile

In 12 listsDetails

vLLM

State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.

In 11 listsDetails

llama.cpp

Pure C/C++ inference engine with GGUF format support. The gold standard for CPU/GPU/Apple Silicon on-device running. Includes llama-server for OpenAI-compatible API. Now at 100K+ stars.

In 10 listsDetails

OpenLLM

Production-grade platform for running any open-source LLMs as OpenAI-compatible API endpoints. Supports 50+ models with built-in streaming, batching, and auto-acceleration. Apache 2.0 licensed.

In 10 listsDetails

SGLang

(MPL-2.0) allows specifying JSON schemas using regular expressions or Pydantic models for constrained decoding. Its high-performance runtime accelerates JSON decoding.

In 9 listsDetails

Bifrost

🟢🟠 — Apache-licensed AI and MCP gateway with multi-provider routing, virtual-key access controls, budgets, rate limits, MCP aggregation, OAuth, automatic fallbacks, and load balancing. (Maxim) — note: model-provider credentials and proxied request/response data are sensitive; restrict and…;…

In 8 listsDetails

shimmy

Python-free Rust inference server with OpenAI API compatibility. Supports GGUF and SafeTensors formats with hot model swap, auto-discovery, and single binary deployment for zero-dependency inference. Apache 2.0 licensed.

In 7 listsDetails

TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference…

In 7 listsDetails