awesome-nlp
Section: Libraries · reference implementations for NLP metrics.
Entry
Appears in 4 awesome lists
Evaluate machine learning models (huggingface). pycm - Multi-class confusion matrix. pandas_ml - Confusion matrix. Plotting learning curve: link. yellowbrick - Learning curve. pyroc - Receiver Operating Characteristic (ROC) curves.
Section: Libraries · reference implementations for NLP metrics.
Section: 9. Evaluation, Benchmarks & Datasets · Standardized evaluation metrics.
Section: Evaluation and Monitoring · Evaluate is a library that makes evaluating and comparing models and reporting their performance easier and more standardized.
Section: General · Evaluate machine learning models (huggingface). pycm - Multi-class confusion matrix. pandas_ml - Confusion matrix. Plotting learning curve: link. yellowbrick - Learning curve. pyroc - Receiver Operating Characteristic (ROC) curves.
Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…
(formerly known as pytorch-transformers and pytorch-pretrained-bert) provides state-of-the-art general-purpose architectures (BERT, GPT-2, RoBERTa, XLM, DistilBert, XLNet, CTRL...) for Natural Language Understanding (NLU) and Natural Language Generation (NLG) with over 32+ pretrained models in…
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…
Test your prompts, models, RAGs. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality. LLM evals for OpenAI/Azure GPT, Anthropic Claude, VertexAI Gemini, Ollama, Local & private models like Mistral/Mixtral/Llama with CI/CD
A curated list of resources dedicated to Natural Language Processing and text processing for Ruby.
Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…
The most widely adopted self-hostable LLM observability platform: traces every agent step, manages prompt versions, and runs evals in one tool. Preferred over cloud-only alternatives when data residency or cost control is a constraint.
The most complete open-source LLM/agent eval framework: 20+ built-in metrics (hallucination, answer relevancy, RAGAs, tool correctness), pytest integration, and a CI-friendly runner. Removes the need to hand-roll eval infrastructure when you need structured, repeatable agent quality gates.