The 12 prompt test: how to check whether ChatGPT, Perplexity and Gemini know your product exists
Six unbranded prompts test whether an assistant names you to a buyer who has never heard of you. Six branded prompts test whether it describes you correctly when it does. Run all twelve in a fresh signed-out session in each engine, mark what the answer body says, and read the pattern rather than any single result.
In one sentence
The 12 prompt test is a manual, repeatable check of whether answer engines discover your product unprompted and describe it accurately when named, run as six unbranded prompts and six branded ones across each engine separately.
What the twelve prompt test tells you, and what it does not
The test tells you how a handful of assistants describe your category and your product on the day you run it, from the position of a buyer with no history with you. It is a snapshot across engines rather than a metric, and treating it as one is where most people go wrong.
Answer engines are non-deterministic. Ask the same question twice in one session and you can get two different shortlists, with a competitor appearing in one and vanishing from the other. That variance is why a single absence is close to worthless as evidence, and why this test is built around repetition across engines rather than depth within one. Missing from the same unbranded prompt in ChatGPT, Perplexity and Gemini is a finding; missing once in Gemini and present everywhere else teaches you only that models vary.
It cannot tell you how often buyers ask these questions, and it gives you no share of anything, since there is no denominator. What it does give you, in about forty minutes, is a defensible answer to whether you exist inside the answer layer at all. The AI Visibility Check generates all twelve prompts with your details filled in and prints the scoring sheet. Everything here works without it; the placeholders are below if you prefer to write them by hand.
Two different things are being measured, and they must be kept apart
Prompts one to six measure discovery: whether an assistant names you to somebody who does not know you exist. Prompts seven to twelve measure accuracy: whether it describes you correctly and currently once handed your name. Separate failures with separate causes, and merging them into one score destroys the diagnostic value.
Discovery is the commercially interesting half. A buyer typing a category question is in a shortlist-forming moment, and the products named in that answer get evaluated. Your own site does not decide this alone. Muck Rack's Generative Pulse series puts earned media at 84 per cent of AI citations, and Ahrefs, across roughly 75,000 brands, found brand web mentions correlate with AI Overview visibility around three times as strongly as backlinks. That second figure is correlational, but it points where the discovery prompts point: products other people write about get named.
Accuracy is the half that surprises founders. Assistants regularly describe products with pricing from two rounds ago, features that were removed, or a positioning the company walked away from. The cause is usually old third-party coverage winning the retrieval slot over your current page, so the fix has less to do with getting mentioned more than with making current facts about yourself easy to find and hard to misread.
The twelve prompts
Substitute your own details for the bracketed placeholders first. Use the category words buyers use rather than the phrase on your homepage, because the two are rarely the same and it is the buyer wording that gets typed in.
Discovery, unbranded. Do not mention your product name anywhere in these sessions.
- What are the best [your category] options available right now? Give me a shortlist with a sentence on each.
- I am evaluating [your category]. What are the main alternatives to [competitor], including newer or smaller ones?
- How do I [the problem in a buyer's words]? What tools do people actually recommend for this?
- Compare [competitor] and [second competitor] for someone choosing [your category]. What else should be on my shortlist that I might not have heard of?
- What is the best [your category] for a small team with a limited budget, and why?
- Who are the newer or less well known players in [your category], and what do they do differently?
Accuracy, branded. Start a fresh session before you begin these.
- What is [your product]? Who is it for and what does it do?
- How much does [your product] cost, and what is included at each tier?
- What do users say about [your product]? What are the common complaints?
- How does [your product] compare with [competitor]? Which suits which kind of buyer?
- Is [your product] a good choice for someone who needs to [the problem in a buyer's words]?
- What are the main limitations or downsides of [your product]?
Prompts two, four and six deserve the closest attention, because they are where a newer product can realistically win. A model answering the generic best-in-category question reaches for incumbents, while one asked for alternatives or lesser known players has permission to surface something smaller, a gap you can fill with the right comparison and alternatives content.
How to run it properly
Use a fresh temporary chat, signed out wherever the engine allows it, and run each engine separately. The principle: you are reproducing what a stranger sees, and a modern assistant's convenience features work against that.
Memory, chat history and personalisation bias answers towards things you have discussed before, which for a founder testing their own product is close to a guaranteed false positive. Where an engine has no signed-out mode, use its temporary chat and disable memory. Run the discovery prompts before your product name has appeared anywhere in the session; once the name is in the context window you are testing only whether the model can repeat you.
Do not lead the model. Resist the follow-ups that feel natural, particularly "what about [your product]?", since an assistant asked directly nearly always says something flattering about a product it barely knows.
Score the answer body, not the source list. A citation panel records that a crawler fetched your page while assembling the answer; the buyer reads the prose. If your name is in the footnotes and a competitor is in the paragraph, the competitor won.
The scoring sheet
Use two scales, one per half of the test, and mark every engine on every prompt. Copy this table, add a column for any other engine your buyers use, and keep the sheet, since most of the value lies in comparing it against last quarter.
| # | Type | ChatGPT | Perplexity | Gemini |
|---|---|---|---|---|
| 1 | Discovery | |||
| 2 | Discovery | |||
| 3 | Discovery | |||
| 4 | Discovery | |||
| 5 | Discovery | |||
| 6 | Discovery | |||
| 7 | Accuracy | |||
| 8 | Accuracy | |||
| 9 | Accuracy | |||
| 10 | Accuracy | |||
| 11 | Accuracy | |||
| 12 | Accuracy |
On the discovery prompts, mark each cell Named, Mentioned or Absent. Named means you appear in the recommended set with a description attached. Mentioned means your name is in the prose but outside the recommendation. Absent covers everything else, the citation-only case included.
On the accuracy prompts, mark each cell Correct, Partly correct or Wrong, under a hard rule: any factual error puts the whole answer in Wrong. Partly correct covers answers that are incomplete but contain nothing false, such as one omitting your top tier. An answer that gets your positioning right and your price wrong is worse than useless, because the buyer cannot know which half to distrust and acts on the wrong number.
Reading your results: four patterns
Absent on one to six, across every engine
The normal starting position, and not a verdict on the product. Nothing outside your own site describes you in terms an engine can associate with your category, which is the default state of a product nobody else writes about. The work that changes it is off-site: roundups, review sites, community threads, marketplace listings, coverage discussing your category rather than your company, which is the substance of answer engine optimisation for a product launch.
Known on seven to twelve, absent on one to six
The engines describe you accurately when handed your name but never volunteer it. This is a positioning and comparison-content problem rather than a visibility one, and it usually means your material explains what you do without stating which category you are in or who you sit against. Nothing connects you to the category in buyer words, so a category question cannot retrieve you. Naming competitors on your own site, publishing comparison pages and getting into third-party lists that rank for that question are what move it.
Present on one to six, wrong on seven to twelve
The most commercially damaging pattern in the set, and the one founders least expect. You are being recommended with incorrect information attached, so the recommendation works against you: a buyer told your entry tier costs three times what it does will not enquire to check. Fix the source material, since the model repeats what it read. Put current pricing on a crawlable page rather than behind a form, date it, add structured data, and get stale third-party pages corrected where the publisher will oblige.
Strong across both halves
Rare, and usually the product of a long-established brand or eighteen months of deliberate work. Retest quarterly rather than relaxing, because the position decays with nothing going wrong on your side. Roughly half of AI-cited pages were updated within the previous thirteen weeks, so a page you leave alone drifts out of the citation pool while competitors publish. A strong result in April can be mediocre by August, with no announcement that it happened.
The one thing to check before concluding anything
Open your own robots.txt and confirm you are not blocking the retrieval crawlers, because if you are, every result above is already explained and no content work will help. It takes two minutes and is the commonest cause of a blank sheet.
Check OAI-SearchBot, which fetches pages for ChatGPT answers, PerplexityBot, Claude-SearchBot, and Bingbot, which still underpins Copilot. These are retrieval agents, distinct from training crawlers such as GPTBot and ClaudeBot. You can disallow training while allowing retrieval, which for most companies is the sensible configuration. Blocking both because a blog post said to opt out of AI is how products end up absent from answers.
Google-Extended is where most articles on this subject get it wrong. It controls whether your content is used for Gemini training and for grounding Gemini responses. It does not govern AI Overviews and it does not govern AI Mode. Both are built on Google's ordinary search index and ride on standard Googlebot access, so disallowing Google-Extended removes you from neither and allowing it puts you into neither. If you want out of AI Overviews, the only levers are the nosnippet, max-snippet and data-nosnippet directives, which also strip your featured snippets and rich results, a price almost nobody should pay. If you want in, the requirement is that Googlebot can crawl and index the page normally.
Check also for a staging rule that reached production. A blanket Disallow: / left over from a pre-launch environment produces the same blank sheet as a deliberate block.
How often to re-run it, and what to do with the results
Quarterly, recorded in the same sheet each time, so you end up with a trend rather than a feeling. One run tells you where you stand; four across a year tell you whether the work between them changed anything at all.
Add an off-cycle run about six weeks after any launch, rebrand or pricing change, since those are the events that make existing coverage of you wrong. Resist re-running whenever you feel uneasy, because the variance between runs hands you a different answer often enough to keep you busy and none the wiser.
Turn the sheet into work with a simple split. Discovery failures are off-site: mentions, listings, comparison content, community presence, anything that puts your category and your name in one document you did not write. Accuracy failures are on-site and source-correction: current pricing on a crawlable page, structured data, and corrections pushed to whichever third-party page the model is evidently reading. With a small team, do not attempt both in one quarter.
How AI visibility sits alongside everything else that decides a launch outcome runs through the AI visibility section, and it is one of eight weighted dimensions in the launch readiness framework. For the whole picture rather than this slice, the launch readiness assessment scores all eight in about seven minutes.
Questions people ask
How many prompts do I need to run before the result means anything?
All twelve, in at least three engines, which is thirty six answers. Assistants are non-deterministic, so one absence tells you very little and could be reversed by running the same prompt again five minutes later. A product that is missing from the same discovery prompt in ChatGPT, Perplexity and Gemini is genuinely missing.
Does appearing in the source list count as being visible?
No, and this is the most common scoring mistake. A citation panel means a crawler fetched your page while assembling the answer; it does not mean the buyer reading the answer will ever see your name. Score only what the answer body says in prose, because that is the only part most people read.
Should I be signed out when I run the test?
Yes wherever the engine permits it, and in a temporary chat where it does not. Memory, saved preferences and prior conversations all bias the answer towards products you have discussed before, which for a founder means your own. The result you want is the one a stranger with no history gets.
Does blocking GPTBot stop me appearing in ChatGPT answers?
Not directly, because GPTBot and the retrieval crawler are separate agents. GPTBot gathers training data, while OAI-SearchBot fetches pages to answer live queries, and blocking one does not block the other. If you want to stay out of training but remain citable, allow the retrieval crawler and disallow the training one.
Does Google-Extended control whether I show up in AI Overviews?
No. Google-Extended governs whether your content is used for Gemini training and grounding, and it has no effect on AI Overviews or AI Mode, both of which rely on ordinary Googlebot access and standard indexing. Many articles state otherwise. If you want out of AI Overviews the only levers are the nosnippet family of directives, which also cost you featured snippets.
How often should I re-run the test?
Quarterly for a stable product, and again about six weeks after any launch, rebrand or pricing change. Recording the results in the same sheet each time turns a set of impressions into a trend you can act on. Ad hoc re-running when you feel anxious produces noise rather than information.
Put a number on it
Score your own launch across all forty checks
Free, about seven minutes, and no email needed to see the result.
Read next
Nobody can find your launch: answer engine optimisation for new products
A launch can be executed perfectly and still be absent from AI answers. How to get retrieved, named and described correctly by answer engines before launch day.
AI visibilityComparison and alternatives pages: the launch asset answer engines quote most
Comparison and alternatives pages get cited because they already answer the question in the shape it was asked. How to build them, and what makes them fail.
Launch guidesThe launch readiness framework: eight dimensions and forty checks
Launch readiness scored across eight weighted dimensions and forty checks. The full framework, the weights, the five result bands, and how to act on your score.
Product Launch Blog is an EbizIndia publication. This article does not pitch anything; the disclosure sits here instead, and in the footer, on every page.