AI usability testing, run by your coding agent
Point Claude Code or Cursor at your product and it signs up, works a real persona's job cold, screenshots every point of friction, and hands back a severity-ranked review you can share or act on. Free, over MCP, nothing to install.
This page is for anyone who wants to catch usability problems before real users do, using the AI coding agent they already have. See the full use cases for visual product feedback hub for the same tool applied to bug reports, design reviews, and client feedback.
What AI usability testing is
AI usability testing is pointing an AI coding agent at your product and having it act as a first-time user: it takes on a realistic persona, tries to complete a real job, and reports every point where a real person would hesitate, misread, or get stuck. Instead of you clicking through your own product (where you already know where everything is), the agent enters cold, the way a new user actually would, and the friction it hits is the friction they will hit.
With CobaltCapture the agent connects over MCP, drives a real browser, signs itself up, and hands back the findings as a shareable review with annotated screenshots — not a chat message that scrolls away. It is the difference between "the AI told me my onboarding is confusing" and "here is the exact screen, with the exact step circled, ranked by severity, ready to paste into a ticket."
How it works
- Connect once. Add CobaltCapture's MCP server to Claude Code, Cursor, or Codex with a single command. The developer setup has the exact line for each.
- Pick a study, or describe one. Invoke a ready-made playbook — a first-time-user usability pass, a specific-flow walkthrough like signup or checkout, a positioning audit — or just tell the agent the job you want tested. You do not have to write the method; the playbook carries it.
- The agent runs it cold. It proposes a persona and a job, signs up with a disposable inbox, drives the browser as that user, and screenshots every point of confusion. It never enters a password you did not set, completes a payment, or solves a CAPTCHA — it drives up to any such wall and reports it.
- You get a review. A durable URL with the annotated screenshots, the findings ranked by severity, and — when you run it from your own repo — a root-cause hypothesis with the file and line to fix. Share it, export it to a GitHub issue, or let a coding agent read it back over MCP and start fixing.
Why the review beats a chat summary
The output is the whole point, and it is where AI usability testing with a real artifact pulls away from asking an agent to "review my site" in chat:
- It shows, it does not describe. The annotated screenshot is burned into the image, so the next agent reads the arrow you drew, not a sentence about it.
- It is a work item, not a report. Export to a GitHub issue or to markdown a coding agent acts on directly. The finding travels to the fix.
- It remembers. Re-run the same study next week and the review leads with what got fixed, what is new, and what still hurts. A chat summary forgets; a durable review lets you watch the product improve. That history is the thing an ephemeral agent structurally cannot give you.
- It reaches people who were not in the chat. A clean stakeholder link for a PM or client, exports for a doc, all from the same run.
AI usability testing vs the alternatives
| Captures + annotates the screen | Runs unattended in minutes | Output an agent can act on | Tracks improvement over time | Cost | |
|---|---|---|---|---|---|
| CobaltCapture + your agent | Yes (real browser + markup) | Yes | Yes (review, GitHub, MCP) | Yes (re-run and diff) | Free |
| Traditional usability test | Manual | No (days, recruiting) | No (a PDF report) | Rarely | High |
| Ask an agent in chat | No (describes only) | Yes | No (dies in scrollback) | No | — |
Traditional testing is still the right call for deep, emotional, or accessibility-specific insight with real humans. AI usability testing is what you run on every build in between, so a human session is spent on the hard questions instead of catching the obvious broken button.
Where it fits
- Before a launch. Run a first-time-user pass on staging and fix the friction before anyone signs up.
- On every meaningful build. Cheap enough to run continuously, so onboarding regressions get caught the day they ship.
- Handing off to a developer or a coding agent. The review is already in the format a coding agent reads, so "find the problem" and "fix the problem" become one loop.
Get started
Connect the MCP server to your agent, then ask it to run a usability pass on your product. First run in under a few minutes, free, no signup to try.
Frequently asked questions
What is AI usability testing?
AI usability testing is having an AI coding agent act as a first-time user of your product: it takes on a realistic persona, tries to complete a real job (sign up, create the thing, get to first value), and reports every point where a real person would hesitate, misread, or get stuck. With CobaltCapture the agent connects over MCP, drives a real browser, and delivers the findings as a shareable review with annotated screenshots, not just a chat message that disappears.
How is it different from a traditional usability test?
A traditional usability test recruits people, schedules sessions, and takes days to a report. AI usability testing runs in minutes, unattended, as often as you like, so you can check every build instead of once a quarter. It does not replace testing with real humans for deep, emotional, or accessibility-specific insight, but it catches the large majority of first-time-user friction long before you spend a human session on it.
Why use a durable review instead of just reading the agent's summary in chat?
A chat summary dies in scrollback, describes screens instead of showing them, and forgets everything next session. The review is a durable URL: it carries the annotated screenshots as real images the next agent can read, it exports to a GitHub issue or markdown a coding agent can act on directly, and because it persists you can re-run the same study later and see what got fixed and what regressed. History is the thing an ephemeral chat cannot give you.
Does the agent need a login to test the product?
It makes its own. CobaltCapture provides a disposable email inbox, so the agent signs up with a throwaway address and monitors the confirmation email itself, no test account to hand it. For safety it never enters a password you did not set, never completes a payment, and stops at any irreversible action, capturing the handoff as a finding instead.
Can it test my own app in development?
Yes, and that is the best case. When you run it from your own repo, the agent reads your code to understand how to start the app and to explain findings with file and line references a coding agent can fix in one pass. Running against your own build also means no bot-detection or rate limits to fight, so the run is faster and cleaner than testing a site you do not control.
Which agents does it work with?
Any MCP-aware coding agent. Claude Code, Cursor, and Codex connect with a single command; other clients add CobaltCapture as an HTTP MCP server. Once connected, you invoke a ready-made study (a first-time-user usability pass, a specific-flow walkthrough, a positioning audit) or describe your own, and the agent runs it.
Capture your first review.
About a minute from open tab to a shareable URL your agent can ingest.
Start capturing