Skip to content

Entry

OrcaPromptVault

Appears in 4 awesome lists

Open corpus of the instructions and tool schemas shipping AI agents are actually sent: 119 artifacts from 43 products, 44 of them recorded off the wire with the command that reproduces each, every file marked captured or reported. AGPL-3.0.

Open github.comcontinuum-ai-corp/orcapromptvault

Found in these lists

Awesome Gemini CLI

Section: Prompts · Dated archive of the system prompts and tool schemas shipped AI agents send, including the one Gemini CLI itself puts on the wire.

FreshScore 82

Awesome-ChatGPT

Section: Prompts · 119 production system prompts and tool schemas from 43 shipping AI products, 44 of them recorded off the wire with the command that reproduces each; every file marked as recorded or as reported by the model.

FreshScore 84

Awesome Open Source AI

Section: 9. Evaluation, Benchmarks & Datasets · Open corpus of the instructions and tool schemas shipping AI agents are actually sent: 119 artifacts from 43 products, 44 of them recorded off the wire with the command that reproduces each, every file marked captured or reported. AGPL-3.0.

FreshScore 89

favorite-link

Section: September 30, 2026

FreshScore 85

DeepEval

The most complete open-source LLM/agent eval framework: 20+ built-in metrics (hallucination, answer relevancy, RAGAs, tool correctness), pytest integration, and a CI-friendly runner. Removes the need to hand-roll eval infrastructure when you need structured, repeatable agent quality gates.

In 9 listsDetails

OpenAI Evals

Framework for evaluating LLMs and LLM systems with an open-source registry of 100+ community-contributed benchmarks. MIT licensed.

In 8 listsDetails

LM Evaluation Harness

Language Model Evaluation Harness is a framework to test generative language models on a large number of different evaluation tasks.

In 7 listsDetails

HuggingFace Datasets

The largest hub of ready-to-use NLP datasets for ML models with fast, easy-to-use and efficient data manipulation tools.

In 6 listsDetails

AutoRAG

RAG AutoML tool for automatically finding optimal RAG pipelines. Evaluates and optimizes retrieval-augmented generation with AutoML-style automation for your own data and use-case. Apache 2.0 licensed.

In 5 listsDetails

OpenCompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

In 4 listsDetails

VLMEvalKit

Open-source evaluation toolkit for large multi-modality models (LMMs). Supports 220+ LMMs and 80+ benchmarks including MMMU, MathVista, and ChartQA. Powers the OpenVLM Leaderboard. Apache 2.0 licensed.

In 4 listsDetails

Evaluate

Evaluate machine learning models (huggingface). pycm - Multi-class confusion matrix. pandas_ml - Confusion matrix. Plotting learning curve: link. yellowbrick - Learning curve. pyroc - Receiver Operating Characteristic (ROC) curves.

In 4 listsDetails