Warehouse-native product analytics that joins AI agent traces to user behaviour, so you can see which tool calls actually convert.
Observability
Understanding running systems from the outside via metrics, logs, traces and dashboards.
11 apps, 9 skills and 11 MCP servers tagged Observability.
Apps
An AI on-call agent that drops read-only probes into running production code, capturing live variable snapshots and stack context without a redeploy or a restart.
Session replay, error tracking and vision-AI agents in one SDK — agents watch real user sessions on a schedule, file the bugs they find and post to Slack.
The framework and observability platform most teams use to build, debug, and ship LLM agents.
Open-source LLM observability — trace, evaluate, and improve AI applications with production data.
Evaluation-first AI development — measure whether a prompt or model change actually improved anything.
The standard experiment tracker for machine learning, now with tracing and evaluation for LLM applications.
Open-source AI workspace where teams build, deploy and monitor agents on one canvas — visually or in code, across 1,000+ integrations.
Agent observability and evaluation that scores every production run for quality, drift and risk — and can pause or block a risky action before it executes.
Kube-DC.cloud is a fixed-price Kubernetes cloud for developers - with VMs, containers, managed databases, observability, backups, and unlimited traffic included.
A governed platform for running AI agents in production, with Omni — an automated forward-deployed engineer that builds and maintains them for you.
Skills
Elastic's official skill for monitoring LLM and agent workloads — token cost, latency, response quality and workflow orchestration, queried with ES|QL.
Investigate OpenTelemetry traces stored in OpenSearch — slow and error spans, reconstructed trace trees, service maps, plus agent invocation and token-usage analysis.
Query and analyse logs held in OpenSearch using PPL and Query DSL — error-pattern discovery, error-rate tracking and anomaly detection, driven from natural language.
Search Datadog logs from an agent and keep the bill under control — query syntax, exclusion filters, log-based metrics, archives and PII scrubbing.
Create and audit Datadog monitors from an agent via the pup CLI — with alerting rules that stop the usual flapping, unscoped and unowned monitors.
Elastic's official skill for writing and running ES|QL - the piped Elasticsearch query language - covering schema discovery, query patterns, error handling and result formats.
Skill: Qdrant Advisor
by Qdrant
Qdrant's meta-skill: instead of shipping static docs, it loads the current official Qdrant skill tree live from skills.qdrant.tech and diagnoses from that.
Skill: PromQL Query Patterns
by Grafana Labs
Grafana's official PromQL skill: write, validate, and optimise Prometheus queries, and hunt down the label that blew up your cardinality.
Skill: Vercel Optimize
by Vercel
Vercel's official optimization skill: collects real production metrics first, then returns ranked, citation-checked cost and performance fixes.
MCP servers
Datadog's official managed MCP server — query logs, metrics, traces, incidents and dashboards from an AI agent, inside your existing Datadog permissions.
Official Grafana MCP server: search and edit dashboards, query Prometheus, Loki and Pyroscope, and work with alerts, incidents and on-call from an agent.
Buildkite's official MCP server: query pipelines, builds, jobs, logs, test runs and flaky-test data from an agent, over a hosted OAuth endpoint or a self-run binary.
Query Braintrust evals and production logs in SQL from your editor — compare experiments against a baseline, debug a bad trace, and hand a teammate a permalink.
Langfuse's own MCP server for LLM observability — read traces and observations, manage prompt versions, run datasets and evaluators, and query cost and latency metrics from inside your agent.
New Relic's official hosted MCP server — query telemetry with NRQL or natural language, investigate alerts and incidents, and assess deployment impact from your agent.
Fastly's official MCP server — manage CDN services, deploy Compute, read real-time metrics and troubleshoot edge configuration through natural language.
Official Google server exposing Google Cloud Observability to agents — querying logs, metrics and traces from a Google Cloud project without leaving the assistant.
Honeycomb's hosted MCP server — query traces, metrics, logs and events, run BubbleUp, and read Boards, Triggers and SLOs from an AI agent.
PagerDuty's official MCP server with 80+ tools for incidents, schedules, services and escalation policies — read-only by default, with write actions behind an explicit flag.
MCP: Elasticsearch
by Elastic
Elastic's official MCP server — explore indexes, inspect mappings, and query Elasticsearch conversationally.
Related tags
Tags that appear alongside this one, ranked by how often.

