Run agent evals and private benchmarks in sandboxed cloud environments — compare Claude Code, Codex, Cursor and Copilot on the same real tasks.
Evaluation
Measuring whether models, agents and prompts actually work — scores, traces and test harnesses.
11 apps, 7 skills and 6 MCP servers tagged Evaluation.
Apps
Microsoft's open-source optimizer that trains an agent's SKILL.md like a model parameter — epochs, learning-rate budgets and a held-out validation gate, with frozen weights.
Continuous automated QA for voice and chat AI agents, from pre-launch simulation to live call monitoring.
Audit-grade fact-checking API for AI products — eight models across five stages, adversarial debate, and a full citation trail behind every verdict.
Reads your production chat and voice agent conversations to surface the silent failures, frustration loops and policy breaches that offline evals never catch.
Agent observability and evaluation that scores every production run for quality, drift and risk — and can pause or block a risky action before it executes.
The standard experiment tracker for machine learning, now with tracing and evaluation for LLM applications.
Evaluation-first AI development — measure whether a prompt or model change actually improved anything.
Open-source LLM observability — trace, evaluate, and improve AI applications with production data.
The framework and observability platform most teams use to build, debug, and ship LLM agents.
An open-source platform for building production LLM applications — visual workflows, RAG, and agents in one place.
Skills
Diagnoses wrong gradients in differentiable NVIDIA Warp programs by measuring first — comparing autodiff against finite differences on a shrunk reproduction before proposing any fix.
Tencent's skill-evolution harness: it rewrites a whole skill folder — SKILL.md, scripts and references together — and lands every decision as a real Git issue, PR and wiki entry you can review.
Skill: Agnost AI Integration
by Agnost AI
Walks a coding agent through wiring Agnost AI conversation analytics into a Python or TypeScript app — inspecting existing OpenTelemetry first and only adding an SDK when the traces are not usable.
Skill: NVIDIA NeMo Data Designer
by NVIDIA
NVIDIA's official skill for building synthetic datasets with NeMo Data Designer — describe the data you want and the agent writes a declarative generation pipeline, column by column.
Skill: Doubt-Driven Development
by Addy Osmani
Sends every non-trivial decision to a fresh-context adversarial reviewer before it stands — for production, security-sensitive or irreversible work.
Skill: MCP Builder
by Anthropic
Anthropic's guide for building high-quality MCP servers end to end — research, tool design, implementation in TypeScript or Python, testing, and a 10-question evaluation suite.
Skill: Skill Creator
by Anthropic
Create, edit, and optimize Agent Skills — scaffold new skills, refine existing ones, run evals, and tune descriptions for reliable triggering.
MCP servers
The official W&B MCP server: query experiment runs, Weave LLM traces, artifacts and registries in natural language, and write findings back as a W&B report.
Testing and evaluation workbench for MCP server authors — inspect tools and prompts, debug OAuth step by step, and score behaviour across 16 client configurations before shipping.
Official MCP server for Arize Phoenix: read traces and spans, manage prompt versions, and run and inspect datasets and experiments from an AI assistant.
MCP: Braintrust MCP Server
by Braintrust
Query Braintrust evals and production logs in SQL from your editor — compare experiments against a baseline, debug a bad trace, and hand a teammate a permalink.
MCP: Langfuse
by Langfuse
Langfuse's own MCP server for LLM observability — read traces and observations, manage prompt versions, run datasets and evaluators, and query cost and latency metrics from inside your agent.
MCP: Everything
by Model Context Protocol
A reference/test server that exercises every MCP feature — prompts, resources, tools, and sampling — for building and debugging clients.
Related tags
Tags that appear alongside this one, ranked by how often.