Skip to content

Evaluation

Measuring whether models, agents and prompts actually work — scores, traces and test harnesses.

9 apps, 2 skills and 2 MCP servers tagged Evaluation.

Apps

All apps

Run agent evals and private benchmarks in sandboxed cloud environments — compare Claude Code, Codex, Cursor and Copilot on the same real tasks.

Coding & DevelopmentFreemium

App: SkillOpt

Microsoft

New

Microsoft's open-source optimizer that trains an agent's SKILL.md like a model parameter — epochs, learning-rate budgets and a held-out validation gate, with frozen weights.

Coding & DevelopmentFree

App: Cekura

Tatva Labs Inc.

Featured

Continuous automated QA for voice and chat AI agents, from pre-launch simulation to live call monitoring.

Coding & DevelopmentFreemium

App: Dify

LangGenius

An open-source platform for building production LLM applications — visual workflows, RAG, and agents in one place.

Automation & WorkflowsFreemium

The framework and observability platform most teams use to build, debug, and ship LLM agents.

Coding & DevelopmentFreemium

Open-source LLM observability — trace, evaluate, and improve AI applications with production data.

Data & AnalyticsFreemium

Evaluation-first AI development — measure whether a prompt or model change actually improved anything.

Data & AnalyticsFreemium

The standard experiment tracker for machine learning, now with tracing and evaluation for LLM applications.

Data & AnalyticsFreemium

App: Prefactor

Prefactor Pty Ltd

Agent observability and evaluation that scores every production run for quality, drift and risk — and can pause or block a risky action before it executes.

Coding & DevelopmentFreemium

Skills

All skills

Skill: SkillHone

by Tencent

New

Tencent's skill-evolution harness: it rewrites a whole skill folder — SKILL.md, scripts and references together — and lands every decision as a real Git issue, PR and wiki entry you can review.

4 views

Skill: MCP Builder

by Anthropic

Anthropic's guide for building high-quality MCP servers end to end — research, tool design, implementation in TypeScript or Python, testing, and a 10-question evaluation suite.

1 views

MCP servers

All MCP servers

MCP: Langfuse

by Langfuse

New

Langfuse's own MCP server for LLM observability — read traces and observations, manage prompt versions, run datasets and evaluators, and query cost and latency metrics from inside your agent.

Related tags

Tags that appear alongside this one, ranked by how often.

All tags