Skip to content

Evaluation

Measuring whether models, agents and prompts actually work — scores, traces and test harnesses.

11 apps, 7 skills and 6 MCP servers tagged Evaluation.

Apps

All apps
Featured

Run agent evals and private benchmarks in sandboxed cloud environments — compare Claude Code, Codex, Cursor and Copilot on the same real tasks.

Coding & DevelopmentFreemium

App: SkillOpt

Microsoft

Featured

Microsoft's open-source optimizer that trains an agent's SKILL.md like a model parameter — epochs, learning-rate budgets and a held-out validation gate, with frozen weights.

Coding & DevelopmentFree

App: Cekura

Tatva Labs Inc.

Featured

Continuous automated QA for voice and chat AI agents, from pre-launch simulation to live call monitoring.

Coding & DevelopmentFreemium

App: Lenz

Lenz IO

Audit-grade fact-checking API for AI products — eight models across five stages, adversarial debate, and a full citation trail behind every verdict.

Coding & DevelopmentFreemium

Reads your production chat and voice agent conversations to surface the silent failures, frustration loops and policy breaches that offline evals never catch.

Data & AnalyticsFreemium

App: Prefactor

Prefactor Pty Ltd

Agent observability and evaluation that scores every production run for quality, drift and risk — and can pause or block a risky action before it executes.

Coding & DevelopmentFreemium

The standard experiment tracker for machine learning, now with tracing and evaluation for LLM applications.

Data & AnalyticsFreemium

Evaluation-first AI development — measure whether a prompt or model change actually improved anything.

Data & AnalyticsFreemium

Open-source LLM observability — trace, evaluate, and improve AI applications with production data.

Data & AnalyticsFreemium

The framework and observability platform most teams use to build, debug, and ship LLM agents.

Coding & DevelopmentFreemium

App: Dify

LangGenius

An open-source platform for building production LLM applications — visual workflows, RAG, and agents in one place.

Automation & WorkflowsFreemium

Skills

All skills

Diagnoses wrong gradients in differentiable NVIDIA Warp programs by measuring first — comparing autodiff against finite differences on a shrunk reproduction before proposing any fix.

3 views

Skill: SkillHone

by Tencent

Featured

Tencent's skill-evolution harness: it rewrites a whole skill folder — SKILL.md, scripts and references together — and lands every decision as a real Git issue, PR and wiki entry you can review.

39 views

Walks a coding agent through wiring Agnost AI conversation analytics into a Python or TypeScript app — inspecting existing OpenTelemetry first and only adding an SDK when the traces are not usable.

11 views

NVIDIA's official skill for building synthetic datasets with NeMo Data Designer — describe the data you want and the agent writes a declarative generation pipeline, column by column.

5 views

Sends every non-trivial decision to a fresh-context adversarial reviewer before it stands — for production, security-sensitive or irreversible work.

4 views

Skill: MCP Builder

by Anthropic

Anthropic's guide for building high-quality MCP servers end to end — research, tool design, implementation in TypeScript or Python, testing, and a 10-question evaluation suite.

3 views

Create, edit, and optimize Agent Skills — scaffold new skills, refine existing ones, run evals, and tune descriptions for reliable triggering.

4 views 1 copies

MCP servers

All MCP servers
New

The official W&B MCP server: query experiment runs, Weave LLM traces, artifacts and registries in natural language, and write findings back as a W&B report.

New

Testing and evaluation workbench for MCP server authors — inspect tools and prompts, debug OAuth step by step, and score behaviour across 16 client configurations before shipping.

Official MCP server for Arize Phoenix: read traces and spans, manage prompt versions, and run and inspect datasets and experiments from an AI assistant.

MCP: Langfuse

by Langfuse

Langfuse's own MCP server for LLM observability — read traces and observations, manage prompt versions, run datasets and evaluators, and query cost and latency metrics from inside your agent.

MCP: Everything

by Model Context Protocol

A reference/test server that exercises every MCP feature — prompts, resources, tools, and sampling — for building and debugging clients.

Related tags

Tags that appear alongside this one, ranked by how often.

All tags