Skip to content

Entry

llama.cpp

Appears in 10 awesome lists

Pure C/C++ inference engine with GGUF format support. The gold standard for CPU/GPU/Apple Silicon on-device running. Includes llama-server for OpenAI-compatible API. Now at 100K+ stars.

Open github.comggml-org/llama.cpp

Found in these lists

Awesome Ai Agents 2026

Section: Local LLM Runners · C/C++ inference. CPU, GPU, Apple Silicon. Foundation of local AI.

ActiveScore 74

Awesome Stars

Section: C++ · LLM inference in C/C++

SlowScore 58

Awesome local LLM

Section: Inference engines · LLM inference in C/C++

FreshScore 87

Awesome Model Quantization

Section: Inference and hardware · Local LLM inference with GGUF models and multiple quantization formats.

FreshScore 83

Awesome Open Source AI

Section: 3. Inference Engines & Serving · Pure C/C++ inference engine with GGUF format support. The gold standard for CPU/GPU/Apple Silicon on-device running. Includes llama-server for OpenAI-compatible API. Now at 100K+ stars.

FreshScore 89

Awesome Production Machine Learning

Section: Deployment and Serving · llama.cpp is an open source software library that performs inference on various large language models such as Llama.

FreshScore 92

Awesome Privacy

Section: Artificial Intelligence · Inference of Facebook's LLaMA model in pure C/C++ so it can run locally on a CPU.

ActiveScore 84

Awesome Generative AI

Section: Developer tools · Inference of Meta's LLaMA model (and others) in pure C/C++. #opensource

FreshScore 90

awesome-c

Section: LLM and Inference · LLM inference in C/C++

FreshScore 80

awesome-cpp

Section: LLM and Inference · LLM inference in C/C++

FreshScore 79

LocalAI

robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…

In 14 listsDetails

Ollama

Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile

In 12 listsDetails

vLLM

State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.

In 11 listsDetails

LM Studio

LM Studio offers a platform for running various local LLMs like LLaMa, Falcon, MPT, and others offline, featuring a Chat UI, OpenAI-compatible server, and model downloads from Hugging Face, with support for Mac, Windows, and Linux, emphasizing privacy and no data collection, free for personal use…

In 11 listsDetails

gpt4all

is an ecosystem of open-source chatbots trained on a massive collections of clean assistant data including code, stories and dialogue based on LLaMa.

In 11 listsDetails

SGLang

(MPL-2.0) allows specifying JSON schemas using regular expressions or Pydantic models for constrained decoding. Its high-performance runtime accelerates JSON decoding.

In 9 listsDetails

Jan

Jan is an open-source, development-stage ChatGPT alternative that operates fully offline on diverse hardware platforms, supporting universal architectures from PCs to multi-GPU clusters github | github profile

In 8 listsDetails

btop

Resource monitor that shows usage and stats for processor, memory, disks, network, and processes. C++ version and continuation of bashtop and bpytop.

In 9 listsDetails