An autonomous AI software engineer that plans, writes, tests, and ships code in its own cloud development environment.
Testing
Verifying software behaves — unit, integration, end-to-end and cross-browser test suites.
14 apps, 13 skills and 15 MCP servers tagged Testing.
Apps
Run agent evals and private benchmarks in sandboxed cloud environments — compare Claude Code, Codex, Cursor and Copilot on the same real tasks.
AI code review that builds and runs your app on every pull request, then attaches a failing test, video and repro steps to each bug it finds.
Open-source testing and evals platform for MCP servers — trace tool calls, debug OAuth, and catch cross-client regressions in CI.
Agentic IDE that runs coding agents, self-healing QA tests and production monitoring in one workspace so fixes reach a PR automatically.
Continuous automated QA for voice and chat AI agents, from pre-launch simulation to live call monitoring.
Autonomous QA agents that generate, run and self-heal Playwright end-to-end and API tests on every pull request, as code you own.
AI review agents that QA a website against your own checklist, then route the findings for client approval.
Cloud coding agents that work in isolated sandboxes, run your tests, preview the app in a browser and open a pull request you can review like any other.
Describe a browser or mobile test in plain English and run it from your terminal — Kane CLI drives real Chrome, no selectors or test scripts required.
AWS's spec-driven agentic IDE, CLI and web workspace: prompts become requirements, design docs and tracked tasks before any code is written.
Evaluation-first AI development — measure whether a prompt or model change actually improved anything.
Code integrity tooling — AI review, test generation, and an agentic IDE assistant grounded in your repository.
AI Assistant integrates deeply with JetBrains IDEs.
Skills
Microsoft's official Playwright skill — drives a real browser from the command line using accessibility snapshots and element refs, and plans, generates and heals Playwright tests.
Never claim work is complete, fixed, or passing without running the check and reading the output first.
Enforce a real red-green-refactor loop when implementing features or fixes, so behaviour is specified before it is written.
A disciplined debugging process — reproduce, isolate, hypothesise, and verify — used before proposing any fix.
A protocol for measuring whether agents can discover and use your tool without being told about it — WITH/WITHOUT trials, binary scoring, and lift as the only metric. The subject under test is your docs, not the model.
Skill: Datadog Triage Flaky Test
by Datadog
Official Datadog skill for investigating one flaky test — pulls its history and failure pattern, categorises the root cause and recommends fixing, quarantining or escalating.
Skill: Datadog Unblock PR
by Datadog
Official Datadog skill that triages a failing PR pipeline, attributing every CI failure as flaky, infra or genuine regression and proposing a targeted fix.
Skill: k6 Load Test Authoring
by Grafana Labs
Grafana's official k6 skill: generate a load, stress, spike or soak test from a plain-language brief, then actually run the script before handing it over.
Skill: Robotics Agent Skills
by arpitg1304
Ten SKILL.md modules that give coding agents real ROS engineering judgement — ROS 1 and ROS 2, design patterns, perception, testing, Docker, bringup and robot security.
Skill: Playwright Test Automation
by TestMu AI
TestMu AI's official skill for generating production-grade Playwright suites in TypeScript, JavaScript, Python, Java or C#, locally or on a 3,000-combination browser cloud.
Skill: Semgrep Rule Creator
by Trail of Bits
Trail of Bits' skill for writing production-quality Semgrep rules — pattern design, taint mode for data-flow bugs, and mandatory test-driven validation.
Skill: Cypress Author
by Cypress.io
Cypress' official skill for writing, updating and repairing E2E and component tests — and it triggers on "write tests" even when nobody says Cypress.
Skill: Web App Testing
by Anthropic
Interact with and test local web applications using Playwright — verify frontend behavior, debug UI, capture screenshots, and read browser logs.
MCP servers
Qase test management over MCP: 36 task-oriented tools for cases, runs, results, defects and QQL search, hosted at mcp.qase.io or self-run with your own API token.
Give a coding agent a live feedback loop on real mobile, TV and desktop apps — inspect, tap, type and capture evidence on iOS, Android and HarmonyOS over MCP or a CLI.
Postman's official MCP server — reach your workspaces, collections, specs and environments from Claude Code, Cursor, VS Code or Copilot, in four tool tiers.
Platform-agnostic MCP server that drives iOS and Android apps on simulators, emulators and real devices through the accessibility tree rather than screenshots.
MCP: Burp Suite MCP Server
by PortSwigger
PortSwigger's official Burp Suite extension exposing proxy history, Repeater and scanning to AI agents over MCP.
MCP: MATLAB MCP Server
by MathWorks
MathWorks' official MCP server: let Claude Code, Copilot or Codex run MATLAB code, lint it and execute unit tests on your local install.
MCP: MCPJam Inspector
by MCPJam
Testing and evaluation workbench for MCP server authors — inspect tools and prompts, debug OAuth step by step, and score behaviour across 16 client configurations before shipping.
MCP: SmartBear MCP
by SmartBear
SmartBear's official MCP server, putting BugSnag error monitoring, Reflect and QMetry test management, Zephyr, PactFlow contract testing and the Swagger API Hub behind one agent connection.
MCP: FetchSandbox MCP
by FetchSandbox
A sandbox for third-party API integrations: agents replay Stripe, GitHub, Twilio and 50+ other workflows — webhooks, retries and failure cases included — without touching the real API.
MCP: Plaid Sandbox MCP Server
by Plaid
Plaid's official sandbox MCP server: search Plaid docs, mint sandbox access tokens, generate mock financial data and fire test webhooks while building an integration.
MCP: CircleCI MCP Server
by CircleCI
CircleCI's official MCP server: pull build failure logs, hunt flaky tests, validate config, trigger and rerun pipelines, and roll back - in natural language from your IDE.
MCP: BrowserStack
by BrowserStack
Official BrowserStack MCP server: run manual and automated tests on real browsers and devices, then debug and fix failing tests from your IDE in plain English.
MCP: Chrome DevTools MCP
by Google Chrome DevTools team
Google's official MCP server for driving and inspecting a real Chrome browser — automation, debugging, and performance traces.
MCP: Playwright
by Microsoft
Browser automation for agents — navigate, click, type, and read pages using Playwright's accessibility tree, no screenshots required.
MCP: Everything
by Model Context Protocol
A reference/test server that exercises every MCP feature — prompts, resources, tools, and sampling — for building and debugging clients.
Related tags
Tags that appear alongside this one, ranked by how often.

