Skip to content

Entry

llama.cpp

Appears in 7 awesome lists

Ports for inferencing LLaMA in C/C++ running on CPUs, supports alpaca, gpt4all, etc.

Open github.comggerganov/llama.cpp

Found in these lists

awesome-ChatGPT-repositories

Section: Langchain · Port of Facebook's LLaMA model in C/C++

FreshScore 87

Awesome Generative AI

Section: Running LLMs Locally · Port of Facebook's LLaMA model in C/C++

SlowScore 62

Awesome LLMOps

Section: Large Model Serving · Port of Facebook's LLaMA model in C/C++

ActiveScore 75

awesome-nlp

Section: Efficient and Small Language Models · portable quantized inference.

FreshScore 90

Table of Contents

Section: Other LLaMA-derived projects: · Ports for inferencing LLaMA in C/C++ running on CPUs, supports alpaca, gpt4all, etc.

StaleScore 52

Awesome Transformer & Transfer Learning in NLP

Section: Other · Port of Facebook's LLaMA model in C/C++.

StaleScore 52

Table of Contents

Section: LLMs · is a Port of Facebook's LLaMA model in C/C++.

SlowScore 61

LangChain

Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs

In 20 listsDetails

LiteLLM

Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…

In 16 listsDetails

LlamaIndex

(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.

In 14 listsDetails

Ollama

Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile

In 12 listsDetails

vLLM

State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.

In 11 listsDetails

Flowise

Flowise simplifies the creation of applications leveraging large language models (LLMs) by providing a drag-and-drop interface for customizing AI workflows, offering easy installation, Docker support, development tools, and documentation for integrating various functionalities such as…

In 12 listsDetails

Open WebUI

Open WebUI is an extensible, feature-rich, and user-friendly self-hosted AI platform designed to operate entirely offline. It supports various LLM runners like Ollama and OpenAI-compatible APIs, with built-in inference engine for RAG, making it a powerful AI deployment solution.

In 11 listsDetails

LM Studio

LM Studio offers a platform for running various local LLMs like LLaMa, Falcon, MPT, and others offline, featuring a Chat UI, OpenAI-compatible server, and model downloads from Hugging Face, with support for Mac, Windows, and Linux, emphasizing privacy and no data collection, free for personal use…

In 11 listsDetails