Product comparison study, run by an AI that has to pick
Name two or three products and what matters to you. The AI works out what decides the category, proposes weighted criteria you can edit in a sentence, captures each maker's page, researches every cell with a dated source, and comes back with a matrix that names the edge on every row, the differences that actually matter, and an overall lean with what would flip it. Software tools or physical goods. No account, no hands-on use, and it says so.
This page is for anyone choosing between two or three products and tired of comparison articles that end in a shrug. It is one of the ready-made studies in the use cases for visual product feedback hub, alongside the pricing page teardown, whose buyer math it reuses for software.
What a comparison study is
A comparison study answers one buying question, which of these should this person choose and why, with evidence they can check and a lean they can disagree with. It is not a feature list. Every row of the matrix names which product has the edge and by how much, the differences section keeps only the rows that would change the decision, and the study ends with an overall lean, its strength, the rows that drove it, and what would flip it.
The lean is the part most comparisons refuse to give. It is not bias. With the criteria and weights shown, it is arithmetic the reader can redo with different weights, and the flip conditions are the exceptions written down.
How the AI runs it
- The thinking phase. It identifies the products precisely, works out what decides this category, and proposes about ten weighted criteria plus the two or three questions whose answers would change the weights. It stops once to show you that list. You edit it in a sentence or say go. If you say you are early and don't know where to start, it switches to a mode that walks the questions in order before it leans.
- Capture, one product at a time. Each maker's page on desktop, at least one on a phone, with the price or stock line on screen as the receipt. It looks at every image that comes back and retakes any that a cookie wall or a bot check ruined. Prices behind toggles are read from the page's markup or a second capture, and the review says which.
- Research every cell. Independent reviews, dealers, review sites, forums, and company news, each figure with a link and a date. Maker claims and independent findings are marked differently. A cell it cannot verify says so.
- The matrix, with an edge on every row. Then the critical differences, worst first, and the list of things people argue about that do not decide this comparison.
- The lean. Which product, how strongly, driven by which rows, and what would flip it. It may only cite rows that are in the matrix; if it finds itself leaning for a reason it has not written down, it adds the row first.
Each product's analysis sits on its first screenshot, in three blocks: what the maker's page says, what independent sources say, and who it is for.
What comes back
A review at a shareable link with the short version on top: the lean in two sentences and what would flip it. Then the matrix with its edge column, dated. Then the not-separating factors, how the study was run, and a closing note listing everything it could not verify with the full source list grouped by type.
Two worked examples
- Software: Grasshopper Signup vs SignUpGenius. Pricing at three buyer sizes, features that matter for sign-up sheets, company status, reviews. The free plans gate different things (ads on one, response visibility on the other), which turns out to be the whole decision.
- Physical goods: Rolex Datejust 36 vs Oyster Perpetual 36 vs Omega Aqua Terra 38. List price against what you would actually pay and wait, date or no date, discretion, resale. The sticker order and the real order invert, and the study says which to buy for a first watch and when that changes.
Running it
From Claude chat, no code. Add Cobalt Capture as a connector once, then ask:
Compare Linear, Jira, and Asana for a team of ten, with a Cobalt review
Add what matters to you in the same message if you know it. It reads public pages only, so there is no account and nothing to approve beyond the one criteria check.
From a coding agent. In Claude Code, /mcp__cobalt__comparison_study "Linear vs Jira vs Asana". The studies page has the paste-ready block for Cursor and Codex.
When to run it
- Before buying a tool for a team, when the shortlist is down to two or three and the pricing pages disagree with each other.
- Before a personal purchase where the forums argue about specs that don't decide anything.
- Against your own product and its two nearest competitors, to see the decision the way a buyer does. Say that you make one of them.
- After a vendor changes pricing. A saved review is a baseline; the re-run leads with what changed.
Frequently asked questions
What is a product comparison study?
A structured comparison of two or three products for one buying decision. An AI first works out what decides the category and proposes weighted criteria, then fills a matrix where every cell has a dated source, names which product has the edge on each row, lists only the differences that would change the decision, and gives an overall lean with the conditions that would reverse it. It is delivered as a shareable review with the makers' pages captured as screenshots.
Does it actually recommend one?
Yes. A comparison that ends in 'if you want X pick A, if you want Y pick B' has handed the work back to you, so the study is required to lean: which product, for the default buyer, how strongly, and which two or three rows drove it. That is not an opinion. It is the arithmetic of the rows you can see at the weights you approved, and the 'what would flip it' section covers the exceptions.
How does it decide what matters?
It starts with a thinking phase. For software it begins from real price at three buyer sizes, the features that matter for the job, whether the product is actively developed and the company stable, integrations and lock-in, support and trust, and independent verdicts. For physical goods it begins from what you would actually pay and where you can actually buy it, the specs that matter in the category, build, ownership cost, resale, and independent reviews. It cuts that to about ten rows that could separate these products, weights them, and shows you the list once before it runs. You edit in plain words or say go.
Where do the numbers come from?
Each maker's own page is captured as a screenshot, so prices and stock lines are dated to the capture. Independent figures come from dealers, review sites, forums, and company news via web search, each with a link and a date. A cell it cannot verify says 'not verified' and why; it is never filled from memory. Prices hidden behind toggles are read from the page's markup or a second capture, and the review says which.
Does it use the products?
No. It reads public pages and public sources only, creates no accounts, and submits nothing. The review states that, and its interface claims are attributed to reviewers rather than presented as hands-on. If you want the product driven, the usability pass is the study that signs up and uses it.
Can I run it from Claude chat?
Yes. Add Cobalt Capture as a connector once, then ask in plain words: 'Compare Linear, Jira, and Asana for a team of ten, with a Cobalt review.' Cobalt renders the pages for the assistant and the assistant does the searching. From Claude Code, Cursor, or Codex the same study runs with the agent driving a real browser, which also lets it flip pricing toggles itself.
What if I make one of the products?
Say so, or the study will find out. It states the relationship in the review once, applies the identical method to both products, and still gives the lean. The Grasshopper Signup vs SignUpGenius example below was requested by Grasshopper's maker and leans Grasshopper for small buyers on the evidence, with the cases where SignUpGenius wins spelled out.
Capture your first review.
About a minute from open tab to a shareable URL your agent can ingest.
Start capturing