Skip to content

Entry

promptfoo

Appears in 11 awesome lists

Test your prompts, models, RAGs. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality. LLM evals for OpenAI/Azure GPT, Anthropic Claude, VertexAI Gemini, Ollama, Local & private models like Mistral/Mixtral/Llama with CI/CD

Open github.compromptfoo/promptfoo

Found in these lists

Awesome AI Coding Tools

Section: AI Frameworks and SDKs · Open-source tool for testing, evaluating, and red-teaming LLM prompts and applications.

FreshScore 85

Awesome AI Security Tools

Section: Scanners, Evals & Guardrails · 🟢 — LLM eval + red-teaming/pentesting CLI with 50+ attack plugins (MIT). Note: OpenAI announced an acquisition agreement in March 2026; remains MIT-licensed — track governance. · updated 2026-08-17)

FreshScore 85

awesome-ChatGPT-repositories

Section: Prompts · Test your prompts, models, RAGs. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality. LLM evals for OpenAI/Azure GPT, Anthropic Claude, VertexAI Gemini, Ollama, Local & private models like Mistral/Mixtral/Llama with CI/CD

FreshScore 87

Awesome Harness Engineering

Section: Verification & CI Integration · YAML-driven LLM testing framework with LLM-as-judge, assertion DSL, and native CI integration. The most practical tool for adding agent output regression tests to a PR pipeline without writing a test harness from scratch.

FreshScore 88

Awesome LangChain

Section: Other LLM Frameworks · Test your prompts. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality.

FreshScore 90

Awesome Learning

Section: AI

FreshScore 82

Awesome Machine Learning

Section: Tools · Open-source LLM evaluation and red teaming framework. Test prompts, models, agents, and RAG pipelines. Run adversarial attacks (jailbreaks, prompt injection) and integrate security testing into CI/CD.

FreshScore 93

Awesome Open Source AI

Section: 8. MLOps / LLMOps & Production · Open-source LLM evaluation and red teaming framework. Test prompts, agents, and RAGs with automated security vulnerability scanning, side-by-side model comparison, and CI/CD integration. Now part of OpenAI. MIT licensed.

FreshScore 89

Awesome Production Machine Learning

Section: Evaluation and Monitoring · LLM red teaming and evaluation framework for testing jailbreaks, prompt injection, and other vulnerabilities with CI/CD integration.

FreshScore 92

Awesome Prompts

Section: Eval & Testing · Test-driven prompt engineering: regression tests, red teaming, model comparison, CI/CD integration. Acquired by OpenAI (Mar 2026) — remains open source.

FreshScore 90

Awesome Testing

Section: AI & LLM Testing · Open-source framework for testing and red teaming LLM applications. Compare prompts, test RAG architectures, run multi-turn adversarial attacks, and catch security vulnerabilities with CI/CD integration.

FreshScore 86

LiteLLM

Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…

In 16 listsDetails

Semantic Kernel

Semantic Kernel is an SDK that integrates Large Language Models (LLMs) like OpenAI, Azure OpenAI, and Hugging Face with conventional programming languages like C#, Python, and Java. Semantic Kernel achieves this by allowing you to define plugins that can be chained together in just a few lines of…

In 11 listsDetails

Prompt Engineering Guide

🐙 Guides, papers, lecture, notebooks and resources for prompt engineering

In 11 listsDetails

Phoenix

Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…

In 10 listsDetails

Langfuse

The most widely adopted self-hostable LLM observability platform: traces every agent step, manages prompt versions, and runs evals in one tool. Preferred over cloud-only alternatives when data residency or cost control is a constraint.

In 10 listsDetails

DocsGPT

Private AI platform for building intelligent agents and assistants with enterprise search. Features Agent Builder, deep research tools, multi-format document analysis, and multi-model support. MIT licensed.

In 10 listsDetails

DemoGPT

DemoGPT enables you to create quick demos by just using prompt. It applies ToT approach on Langchain documentation tree.

In 9 listsDetails

Helicone

Open-source LLM observability proxy (YC W23) with the largest open-source pricing database (300+ models). One-line proxy integration provides cost tracking, token monitoring, session tracing, and prompt versioning across providers. The AI Gateway component handles request routing and caching with…

In 8 listsDetails