Warehouse-native product analytics that joins AI agent traces to user behaviour, so you can see which tool calls actually convert.
Observability
Understanding running systems from the outside via metrics, logs, traces and dashboards.
17 apps, 23 skills and 30 MCP servers tagged Observability.
Apps
Session replay, error tracking and vision-AI agents in one SDK — agents watch real user sessions on a schedule, file the bugs they find and post to Slack.
An AI on-call agent that drops read-only probes into running production code, capturing live variable snapshots and stack context without a redeploy or a restart.
Finds every AI agent an enterprise is running across AWS, Databricks, Google, Microsoft, Salesforce and Snowflake, then measures and risk-ranks them.
AI engineering platform for tracing, evaluating and improving LLM apps and agents — with Phoenix, its open-source local-first counterpart, for teams that want to start without a vendor.
Open-source, OpenTelemetry-native observability platform for logs, metrics and traces — with first-class tracing for LLM calls and agent runs.
An OpenTelemetry-native control plane for AI agents that enforces policy mid-run instead of paging you after the incident, with cost attribution and EU AI Act evidence built in.
Reads your production chat and voice agent conversations to surface the silent failures, frustration loops and policy breaches that offline evals never catch.
Real-time mobile observability with unsampled on-device logs, session replay and AI agents that triage and investigate app issues autonomously.
A governed platform for running AI agents in production, with Omni — an automated forward-deployed engineer that builds and maintains them for you.
Kube-DC.cloud is a fixed-price Kubernetes cloud for developers - with VMs, containers, managed databases, observability, backups, and unlimited traffic included.
Agent observability and evaluation that scores every production run for quality, drift and risk — and can pause or block a risky action before it executes.
Open-source AI workspace where teams build, deploy and monitor agents on one canvas — visually or in code, across 1,000+ integrations.
The standard experiment tracker for machine learning, now with tracing and evaluation for LLM applications.
Evaluation-first AI development — measure whether a prompt or model change actually improved anything.
Open-source LLM observability — trace, evaluate, and improve AI applications with production data.
The framework and observability platform most teams use to build, debug, and ship LLM agents.
Skills
Production review rules for Cloudflare Workers — compatibility dates, observability wiring, and the runtime anti-patterns that only bite at the edge.
Skill: Datadog Triage Flaky Test
by Datadog
Official Datadog skill for investigating one flaky test — pulls its history and failure pattern, categorises the root cause and recommends fixing, quarantining or escalating.
Skill: Datadog Unblock PR
by Datadog
Official Datadog skill that triages a failing PR pipeline, attributing every CI failure as flaky, infra or genuine regression and proposing a targeted fix.
Skill: Datadog Browser RUM Instrumentation
by Datadog
Official Datadog skill that adds or repairs Browser RUM in a web app — detects framework, router, bundler and existing setup, then wires init once and verifies the build still passes.
Skill: Datadog Pup CLI
by Datadog
Official Datadog skill for the pup CLI — OAuth2 login with token refresh, plus a command map covering logs, monitors, traces, metrics, incidents, SLOs, on-call, security signals and audit logs.
Skill: Grafana Alloy Collector Configs
by Grafana Labs
Grafana's official Alloy skill: write one collector config that carries metrics, logs, traces and profiles, validate it locally, and prove samples are actually leaving the box.
Skill: Grafana Tempo & TraceQL
by Grafana Labs
Grafana's official Tempo skill: stand up a tracing backend that needs nothing but object storage, then write TraceQL that finds the slow span instead of scrolling for it.
Skill: Grafana Loki & LogQL
by Grafana Labs
Grafana's official Loki skill: write LogQL that returns rows instead of timeouts, ship logs through Alloy, and reason about a store that indexes labels rather than log text.
Skill: Agnost AI Integration
by Agnost AI
Walks a coding agent through wiring Agnost AI conversation analytics into a Python or TypeScript app — inspecting existing OpenTelemetry first and only adding an SDK when the traces are not usable.
Skill: Sentry AI Agent Monitoring Setup
by Sentry
Sentry's official skill for instrumenting LLM and agent code — detects your installed AI SDK and wires up traces for model calls, tool use and token spend.
Skill: Sentry SDK Setup
by Sentry
Sentry's official router skill: detects your language or framework, then loads the matching Sentry SDK setup guide for error monitoring, tracing and session replay.
Skill: LaunchDarkly Guarded Rollout
by LaunchDarkly
Design a staged LaunchDarkly rollout — traffic steps, monitoring windows, regression thresholds and automatic rollback — and start it from your coding agent.
Skill: bitdrift Critical User Journey Monitoring
by bitdrift
Stands up end-to-end monitoring for a business-critical flow — path discovery, conversion funnel, completion SLO, step-duration alerts and a dashboard.
Skill: bitdrift Instrumentation
by bitdrift
Adds the bitdrift Capture SDK to an iOS, Android or React Native app and wires up logging, network monitoring, screen views and TTI correctly.
Skill: Elastic LLM and Agentic Observability
by Elastic
Elastic's official skill for monitoring LLM and agent workloads — token cost, latency, response quality and workflow orchestration, queried with ES|QL.
Skill: OpenSearch Trace Analytics
by OpenSearch Project
Investigate OpenTelemetry traces stored in OpenSearch — slow and error spans, reconstructed trace trees, service maps, plus agent invocation and token-usage analysis.
Skill: OpenSearch Log Analytics
by OpenSearch Project
Query and analyse logs held in OpenSearch using PPL and Query DSL — error-pattern discovery, error-rate tracking and anomaly detection, driven from natural language.
Skill: Datadog Logs
by Datadog
Search Datadog logs from an agent and keep the bill under control — query syntax, exclusion filters, log-based metrics, archives and PII scrubbing.
Skill: Datadog Monitors
by Datadog
Create and audit Datadog monitors from an agent via the pup CLI — with alerting rules that stop the usual flapping, unscoped and unowned monitors.
Skill: Elasticsearch ES|QL
by Elastic
Elastic's official skill for writing and running ES|QL - the piped Elasticsearch query language - covering schema discovery, query patterns, error handling and result formats.
Skill: Qdrant Advisor
by Qdrant
Qdrant's meta-skill: instead of shipping static docs, it loads the current official Qdrant skill tree live from skills.qdrant.tech and diagnoses from that.
Skill: PromQL Query Patterns
by Grafana Labs
Grafana's official PromQL skill: write, validate, and optimise Prometheus queries, and hunt down the label that blew up your cardinality.
Skill: Vercel Optimize
by Vercel
Vercel's official optimization skill: collects real production metrics first, then returns ranked, citation-checked cost and performance fixes.
MCP servers
Query SigNoz metrics, traces, logs, alerts and dashboards in natural language — and create alerts, dashboards and saved views back, with deep links into the UI.
ClickHouse's observability MCP server — investigate logs, traces and metrics with semantic tools rather than raw SQL, hosted or self-hosted.
The official W&B MCP server: query experiment runs, Weave LLM traces, artifacts and registries in natural language, and write findings back as a W&B report.
InfluxData's official MCP server for InfluxDB 3 — schema discovery, bounded SQL and InfluxQL queries, line-protocol writes and token administration across Core, Enterprise and Cloud.
Datadog's official managed MCP server — query logs, metrics, traces, incidents and dashboards from an AI agent, inside your existing Datadog permissions.
Official Grafana MCP server: search and edit dashboards, query Prometheus, Loki and Pyroscope, and work with alerts, incidents and on-call from an agent.
Checkly's hosted MCP server gives an agent read access to your synthetic monitoring account — check health, failing results, test sessions, Rocky AI root-cause analyses, incidents and status pages — over streamable HTTP with OAuth.
MCP: Memfault MCP Server
by Memfault (Nordic Semiconductor)
Remote MCP server for Memfault (Nordic Semiconductor) exposing 18 read-only tools over device fleet telemetry — device vitals, timeseries, crash issues, traces, metrics and software versions.
MCP: Splunk MCP Server
by Splunk
Splunk's own MCP server, hosted inside your Splunk deployment, letting agents write SPL from natural language and run searches under existing RBAC.
MCP: Axiom
by Axiom
Axiom's official MCP server — query event data with APL, read dataset schemas, and create or edit dashboards, monitors and notifiers from an agent.
MCP: incident.io
by incident.io
incident.io's official hosted MCP server — reach incidents, alerts, on-call schedules, the catalog and workflows from Claude, Cursor or any MCP client over OAuth.
MCP: Inngest MCP Server
by Inngest
Inngest's official MCP servers for Cloud and the local Dev Server — inspect functions, replay runs, send test events and query Insights from a coding agent.
MCP: Dagster+ MCP Server
by Dagster Labs
Dagster Labs' official remote MCP server for Dagster+: inspect assets, runs, run logs and Insights metrics, and launch runs from an AI session.
MCP: Arize Phoenix MCP Server
by Arize AI
Official MCP server for Arize Phoenix: read traces and spans, manage prompt versions, and run and inspect datasets and experiments from an AI assistant.
MCP: SmartBear MCP
by SmartBear
SmartBear's official MCP server, putting BugSnag error monitoring, Reflect and QMetry test management, Zephyr, PactFlow contract testing and the Swagger API Hub behind one agent connection.
MCP: Dynatrace Managed MCP Server
by Dynatrace
MCP server for self-hosted Dynatrace Managed deployments — problems, vulnerabilities, entities, SLOs, logs and metrics from one or many clusters, in stdio or remote mode.
MCP: Rootly
by Rootly
Rootly's official MCP server for incident management and on-call — declare and update incidents, page the right responder, and pull timelines and retrospectives into an agent session.
MCP: Prefect
by Prefect
Prefect's official read-only MCP server for inspecting flow runs, deployments and automations on Prefect Cloud or self-hosted, with a docs proxy for current CLI and SDK guidance.
MCP: PlanetScale
by PlanetScale
PlanetScale's official hosted MCP server — browse organizations, databases and branches, read schema, run guarded SQL, and pull Insights and schema recommendations.
MCP: Buildkite MCP Server
by Buildkite
Buildkite's official MCP server: query pipelines, builds, jobs, logs, test runs and flaky-test data from an agent, over a hosted OAuth endpoint or a self-run binary.
MCP: Braintrust MCP Server
by Braintrust
Query Braintrust evals and production logs in SQL from your editor — compare experiments against a baseline, debug a bad trace, and hand a teammate a permalink.
MCP: Langfuse
by Langfuse
Langfuse's own MCP server for LLM observability — read traces and observations, manage prompt versions, run datasets and evaluators, and query cost and latency metrics from inside your agent.
MCP: New Relic MCP Server
by New Relic
New Relic's official hosted MCP server — query telemetry with NRQL or natural language, investigate alerts and incidents, and assess deployment impact from your agent.
MCP: Fastly MCP Server
by Fastly
Fastly's official MCP server — manage CDN services, deploy Compute, read real-time metrics and troubleshoot edge configuration through natural language.
Official Google server exposing Google Cloud Observability to agents — querying logs, metrics and traces from a Google Cloud project without leaving the assistant.
MCP: Honeycomb MCP Server
by Honeycomb
Honeycomb's hosted MCP server — query traces, metrics, logs and events, run BubbleUp, and read Boards, Triggers and SLOs from an AI agent.
MCP: PagerDuty MCP
by PagerDuty
PagerDuty's official MCP server with 80+ tools for incidents, schedules, services and escalation policies — read-only by default, with write actions behind an explicit flag.
MCP: Elasticsearch
by Elastic
Elastic's official MCP server — explore indexes, inspect mappings, and query Elasticsearch conversationally.
MCP: Cloudflare
by Cloudflare
Cloudflare's catalogue of managed remote MCP servers — docs, Workers bindings and builds, observability, Radar, Browser Run, AI Gateway, CASB and more, each on its own OAuth endpoint.
MCP: Sentry
by Sentry
Sentry's official MCP server — investigate errors and issues, inspect events, and use Seer for root-cause analysis.
Related tags
Tags that appear alongside this one, ranked by how often.

