Awesome Ai Agents 2026
Section: Local LLM Runners · C/C++ inference. CPU, GPU, Apple Silicon. Foundation of local AI.
Entry
Appears in 10 awesome lists
Pure C/C++ inference engine with GGUF format support. The gold standard for CPU/GPU/Apple Silicon on-device running. Includes llama-server for OpenAI-compatible API. Now at 100K+ stars.
Section: Local LLM Runners · C/C++ inference. CPU, GPU, Apple Silicon. Foundation of local AI.
Section: C++ · LLM inference in C/C++
Section: Inference engines · LLM inference in C/C++
Section: Inference and hardware · Local LLM inference with GGUF models and multiple quantization formats.
Section: 3. Inference Engines & Serving · Pure C/C++ inference engine with GGUF format support. The gold standard for CPU/GPU/Apple Silicon on-device running. Includes llama-server for OpenAI-compatible API. Now at 100K+ stars.
Section: Deployment and Serving · llama.cpp is an open source software library that performs inference on various large language models such as Llama.
Section: Artificial Intelligence · Inference of Facebook's LLaMA model in pure C/C++ so it can run locally on a CPU.
Section: Developer tools · Inference of Meta's LLaMA model (and others) in pure C/C++. #opensource
Section: LLM and Inference · LLM inference in C/C++
Section: LLM and Inference · LLM inference in C/C++
robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…
Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile
State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.
LM Studio offers a platform for running various local LLMs like LLaMa, Falcon, MPT, and others offline, featuring a Chat UI, OpenAI-compatible server, and model downloads from Hugging Face, with support for Mac, Windows, and Linux, emphasizing privacy and no data collection, free for personal use…
is an ecosystem of open-source chatbots trained on a massive collections of clean assistant data including code, stories and dialogue based on LLaMa.
(MPL-2.0) allows specifying JSON schemas using regular expressions or Pydantic models for constrained decoding. Its high-performance runtime accelerates JSON decoding.