Session replay, error tracking and vision-AI agents in one SDK — agents watch real user sessions on a schedule, file the bugs they find and post to Slack.
Monitoring
Watching systems in production — metrics, health checks, monitors and the alerts they raise.
12 apps, 10 skills and 13 MCP servers tagged Monitoring.
Apps
Agentic IDE that runs coding agents, self-healing QA tests and production monitoring in one workspace so fixes reach a PR automatically.
An AI on-call agent that drops read-only probes into running production code, capturing live variable snapshots and stack context without a redeploy or a restart.
Continuous automated QA for voice and chat AI agents, from pre-launch simulation to live call monitoring.
Finds every AI agent an enterprise is running across AWS, Databricks, Google, Microsoft, Salesforce and Snowflake, then measures and risk-ranks them.
Open-source, OpenTelemetry-native observability platform for logs, metrics and traces — with first-class tracing for LLM calls and agent runs.
An OpenTelemetry-native control plane for AI agents that enforces policy mid-run instead of paging you after the incident, with cost attribution and EU AI Act evidence built in.
Reads your production chat and voice agent conversations to surface the silent failures, frustration loops and policy breaches that offline evals never catch.
A shared DevOps workspace where AI agents investigate incidents and operate AWS, GCP and Kubernetes under IAM scoping, approval gates and a full audit trail.
Free, open-source macOS menu bar app that shows live status for every parallel Codex task, so you can see which agent needs you without reopening Codex.
Watch and approve Claude Code, Codex and OpenCode sessions running on your Mac from an iPhone, with live transcripts and push notifications.
Agent observability and evaluation that scores every production run for quality, drift and risk — and can pause or block a risky action before it executes.
Skills
Skill: Datadog Browser RUM Instrumentation
by Datadog
Official Datadog skill that adds or repairs Browser RUM in a web app — detects framework, router, bundler and existing setup, then wires init once and verifies the build still passes.
Skill: Datadog Pup CLI
by Datadog
Official Datadog skill for the pup CLI — OAuth2 login with token refresh, plus a command map covering logs, monitors, traces, metrics, incidents, SLOs, on-call, security signals and audit logs.
Skill: Grafana Alloy Collector Configs
by Grafana Labs
Grafana's official Alloy skill: write one collector config that carries metrics, logs, traces and profiles, validate it locally, and prove samples are actually leaving the box.
Skill: Web Performance Audit
by Cloudflare
Audits Core Web Vitals through the Chrome DevTools MCP server — LCP, INP, CLS plus FCP, TBT and Speed Index — and traces each one back to the render-blocking resource or layout shift causing it.
Skill: Sentry AI Agent Monitoring Setup
by Sentry
Sentry's official skill for instrumenting LLM and agent code — detects your installed AI SDK and wires up traces for model calls, tool use and token spend.
Skill: Sentry SDK Setup
by Sentry
Sentry's official router skill: detects your language or framework, then loads the matching Sentry SDK setup guide for error monitoring, tracing and session replay.
Skill: LaunchDarkly Guarded Rollout
by LaunchDarkly
Design a staged LaunchDarkly rollout — traffic steps, monitoring windows, regression thresholds and automatic rollback — and start it from your coding agent.
Skill: bitdrift Critical User Journey Monitoring
by bitdrift
Stands up end-to-end monitoring for a business-critical flow — path discovery, conversion funnel, completion SLO, step-duration alerts and a dashboard.
Skill: Datadog Monitors
by Datadog
Create and audit Datadog monitors from an agent via the pup CLI — with alerting rules that stop the usual flapping, unscoped and unowned monitors.
Skill: PromQL Query Patterns
by Grafana Labs
Grafana's official PromQL skill: write, validate, and optimise Prometheus queries, and hunt down the label that blew up your cardinality.
MCP servers
ClickHouse's observability MCP server — investigate logs, traces and metrics with semantic tools rather than raw SQL, hosted or self-hosted.
The official W&B MCP server: query experiment runs, Weave LLM traces, artifacts and registries in natural language, and write findings back as a W&B report.
Nutanix's official open-source MCP server: 1,000+ Prism Central V4 API operations across 19 namespaces, read-only by default and audit-logged.
Datadog's official managed MCP server — query logs, metrics, traces, incidents and dashboards from an AI agent, inside your existing Datadog permissions.
Official Grafana MCP server: search and edit dashboards, query Prometheus, Loki and Pyroscope, and work with alerts, incidents and on-call from an agent.
Astronomer's MCP server for Apache Airflow: 30+ tools over the Airflow REST API, with consolidated helpers that explore a DAG, diagnose a failed run or summarise system health in a single call. Works with Airflow 2.x and 3.x.
Checkly's hosted MCP server gives an agent read access to your synthetic monitoring account — check health, failing results, test sessions, Rocky AI root-cause analyses, incidents and status pages — over streamable HTTP with OAuth.
MCP: Memfault MCP Server
by Memfault (Nordic Semiconductor)
Remote MCP server for Memfault (Nordic Semiconductor) exposing 18 read-only tools over device fleet telemetry — device vitals, timeseries, crash issues, traces, metrics and software versions.
MCP: Axiom
by Axiom
Axiom's official MCP server — query event data with APL, read dataset schemas, and create or edit dashboards, monitors and notifiers from an agent.
MCP: SmartBear MCP
by SmartBear
SmartBear's official MCP server, putting BugSnag error monitoring, Reflect and QMetry test management, Zephyr, PactFlow contract testing and the Swagger API Hub behind one agent connection.
MCP: Dynatrace Managed MCP Server
by Dynatrace
MCP server for self-hosted Dynatrace Managed deployments — problems, vulnerabilities, entities, SLOs, logs and metrics from one or many clusters, in stdio or remote mode.
Official Google server exposing Google Cloud Observability to agents — querying logs, metrics and traces from a Google Cloud project without leaving the assistant.
MCP: Sentry
by Sentry
Sentry's official MCP server — investigate errors and issues, inspect events, and use Seer for root-cause analysis.
Related tags
Tags that appear alongside this one, ranked by how often.
