Awesome Generative AI Data Scientist
Section: LLMOps · Open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM Observability all in one place.
Entry
Appears in 6 awesome lists
Agenta provides end-to-end tools for the entire LLMOps workflow: building (LLM playground, evaluation), deploying (prompt and configuration management), and (LLM observability and tracing).
Section: LLMOps · Open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM Observability all in one place.
Section: LLMOps · The LLMOps platform to build robust LLM apps. Easily experiment and evaluate different prompts, models, and workflows to build robust apps.
Section: Testing, Evaluation and Observability · an open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place
Section: 8. MLOps / LLMOps & Production · Open-source LLMOps platform combining prompt playground, prompt management, LLM evaluation, and observability.
Section: Deployment and Serving · Agenta provides end-to-end tools for the entire LLMOps workflow: building (LLM playground, evaluation), deploying (prompt and configuration management), and (LLM observability and tracing).
Section: Developer tools · An open-source end-to-end LLMOps platform for prompt engineering, evaluation, and deployment. #opensource
Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…
Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…
February 2026 release making human oversight a native workflow primitive: suspend execution at critical decision points, expose review-and-edit UI mid-flow, and route subsequent execution based on human action (approve/reject/escalate). Demonstrates how HITL transitions from bolt-on approval gates…
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…
Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…
The most widely adopted self-hostable LLM observability platform: traces every agent step, manages prompt versions, and runs evals in one tool. Preferred over cloud-only alternatives when data residency or cost control is a constraint.