Skip to content

Observability

Understanding running systems from the outside via metrics, logs, traces and dashboards.

14 apps, 15 skills and 28 MCP servers tagged Observability.

Apps

All apps
Featured

Warehouse-native product analytics that joins AI agent traces to user behaviour, so you can see which tool calls actually convert.

Data & AnalyticsFreemium

Session replay, error tracking and vision-AI agents in one SDK — agents watch real user sessions on a schedule, file the bugs they find and post to Slack.

Data & AnalyticsSubscription
Featured

An AI on-call agent that drops read-only probes into running production code, capturing live variable snapshots and stack context without a redeploy or a restart.

Coding & DevelopmentFreemium

An OpenTelemetry-native control plane for AI agents that enforces policy mid-run instead of paging you after the incident, with cost attribution and EU AI Act evidence built in.

Data & AnalyticsFreemium

Reads your production chat and voice agent conversations to surface the silent failures, frustration loops and policy breaches that offline evals never catch.

Data & AnalyticsFreemium

Real-time mobile observability with unsampled on-device logs, session replay and AI agents that triage and investigate app issues autonomously.

Coding & DevelopmentFreemium

A governed platform for running AI agents in production, with Omni — an automated forward-deployed engineer that builds and maintains them for you.

Automation & WorkflowsFreemium

Kube-DC.cloud is a fixed-price Kubernetes cloud for developers - with VMs, containers, managed databases, observability, backups, and unlimited traffic included.

Coding & DevelopmentPaid

App: Prefactor

Prefactor Pty Ltd

Agent observability and evaluation that scores every production run for quality, drift and risk — and can pause or block a risky action before it executes.

Coding & DevelopmentFreemium

App: Sim

Sim, Inc.

Open-source AI workspace where teams build, deploy and monitor agents on one canvas — visually or in code, across 1,000+ integrations.

Automation & WorkflowsFreemium

The standard experiment tracker for machine learning, now with tracing and evaluation for LLM applications.

Data & AnalyticsFreemium

Evaluation-first AI development — measure whether a prompt or model change actually improved anything.

Data & AnalyticsFreemium

Open-source LLM observability — trace, evaluate, and improve AI applications with production data.

Data & AnalyticsFreemium

The framework and observability platform most teams use to build, debug, and ship LLM agents.

Coding & DevelopmentFreemium

Skills

All skills

Walks a coding agent through wiring Agnost AI conversation analytics into a Python or TypeScript app — inspecting existing OpenTelemetry first and only adding an SDK when the traces are not usable.

11 views

Sentry's official router skill: detects your language or framework, then loads the matching Sentry SDK setup guide for error monitoring, tracing and session replay.

6 views

Design a staged LaunchDarkly rollout — traffic steps, monitoring windows, regression thresholds and automatic rollback — and start it from your coding agent.

6 views

Skill: OpenSearch Trace Analytics

by OpenSearch Project

Investigate OpenTelemetry traces stored in OpenSearch — slow and error spans, reconstructed trace trees, service maps, plus agent invocation and token-usage analysis.

4 views

Skill: OpenSearch Log Analytics

by OpenSearch Project

Query and analyse logs held in OpenSearch using PPL and Query DSL — error-pattern discovery, error-rate tracking and anomaly detection, driven from natural language.

6 views

Search Datadog logs from an agent and keep the bill under control — query syntax, exclusion filters, log-based metrics, archives and PII scrubbing.

7 views

Create and audit Datadog monitors from an agent via the pup CLI — with alerting rules that stop the usual flapping, unscoped and unowned monitors.

14 views

Elastic's official skill for writing and running ES|QL - the piped Elasticsearch query language - covering schema discovery, query patterns, error handling and result formats.

6 views

Qdrant's meta-skill: instead of shipping static docs, it loads the current official Qdrant skill tree live from skills.qdrant.tech and diagnoses from that.

2 views

Grafana's official PromQL skill: write, validate, and optimise Prometheus queries, and hunt down the label that blew up your cardinality.

12 views

Vercel's official optimization skill: collects real production metrics first, then returns ranked, citation-checked cost and performance fixes.

3 views

MCP servers

All MCP servers
New

The official W&B MCP server: query experiment runs, Weave LLM traces, artifacts and registries in natural language, and write findings back as a W&B report.

Featured

InfluxData's official MCP server for InfluxDB 3 — schema discovery, bounded SQL and InfluxQL queries, line-protocol writes and token administration across Core, Enterprise and Cloud.

Featured

Datadog's official managed MCP server — query logs, metrics, traces, incidents and dashboards from an AI agent, inside your existing Datadog permissions.

MCP: Grafana

by Grafana Labs

Featured

Official Grafana MCP server: search and edit dashboards, query Prometheus, Loki and Pyroscope, and work with alerts, incidents and on-call from an agent.

MCP: Dagster+ MCP Server

by Dagster Labs

New

Dagster Labs' official remote MCP server for Dagster+: inspect assets, runs, run logs and Insights metrics, and launch runs from an AI session.

Official MCP server for Arize Phoenix: read traces and spans, manage prompt versions, and run and inspect datasets and experiments from an AI assistant.

MCP: SmartBear MCP

by SmartBear

New

SmartBear's official MCP server, putting BugSnag error monitoring, Reflect and QMetry test management, Zephyr, PactFlow contract testing and the Swagger API Hub behind one agent connection.

MCP: Rootly

by Rootly

Rootly's official MCP server for incident management and on-call — declare and update incidents, page the right responder, and pull timelines and retrospectives into an agent session.

MCP: Prefect

by Prefect

Prefect's official read-only MCP server for inspecting flow runs, deployments and automations on Prefect Cloud or self-hosted, with a docs proxy for current CLI and SDK guidance.

MCP: Railway

by Railway

Railway's official MCP server, bundled with the Railway CLI — 40+ tools for projects, services, deployments, variables, domains and observability, with destructive actions gated behind explicit confirmation.

MCP: PlanetScale

by PlanetScale

PlanetScale's official hosted MCP server — browse organizations, databases and branches, read schema, run guarded SQL, and pull Insights and schema recommendations.

Buildkite's official MCP server: query pipelines, builds, jobs, logs, test runs and flaky-test data from an agent, over a hosted OAuth endpoint or a self-run binary.

MCP: Langfuse

by Langfuse

Langfuse's own MCP server for LLM observability — read traces and observations, manage prompt versions, run datasets and evaluators, and query cost and latency metrics from inside your agent.

MCP: PagerDuty MCP

by PagerDuty

PagerDuty's official MCP server with 80+ tools for incidents, schedules, services and escalation policies — read-only by default, with write actions behind an explicit flag.

MCP: Cloudflare

by Cloudflare

Cloudflare's official remote MCP servers — query analytics, manage Workers and bindings, read docs, and debug from your agent.

MCP: Sentry

by Sentry

Sentry's official MCP server — investigate errors and issues, inspect events, and use Seer for root-cause analysis.

Related tags

Tags that appear alongside this one, ranked by how often.

All tags