Home/ AI visibility

AI crawler access, schema and entities for a product nobody has heard of yet

Answer engines cite what they can fetch, parse and resolve, in that order. A new product usually fails at the first layer without knowing it, spends its effort on the second, and never reaches the third, which is the one that decides whether being cited is possible at all.

In one sentence

Answer engine readiness for a new product has three layers in fixed order: access, meaning retrieval crawlers can fetch the page; structure, meaning the page states facts a machine can extract; and entity resolution, meaning the brand name maps to something the engine can identify and corroborate elsewhere.

There are three layers between a new product and being named in an answer, and they only work in order. An engine has to be able to fetch your pages, then to parse them into facts, then to resolve your name to an entity it can distinguish from every other string and find corroborated somewhere other than your own site. Almost every founder spends their effort on the middle layer, because structured data is the part with documentation and validators. The first layer is where new products silently fail, and the third is the one that actually decides whether citation is possible at all.

This is the practical order of work for a product with no history, no coverage and no established brand. It is deliberately unglamorous.

Layer one: can the crawlers fetch you at all?

Start here every time, because everything downstream is void if the answer is no, and the answer is no more often than people expect. Three places block retrieval, and only one of them is the file everybody checks.

robots.txt, and the trap in how groups work

A robots.txt group named for a specific user agent replaces the wildcard group for that crawler rather than adding to it. This produces a failure that looks like diligence. A site writes a sensible wildcard group with a few Disallow lines, then adds a friendly block of per bot Allow lines for the AI crawlers, and in doing so removes every one of those Disallow rules for exactly the crawlers it was trying to welcome. The reverse mistake is more damaging: a named group with a bare Disallow, copied from a template written when blocking AI crawlers was fashionable, quietly removes the product from the answer layer entirely.

The safe pattern for a site that wants to be read is one group listing every agent you care about, sharing one set of rules. Then read your own file rather than assuming, because the file that is live is frequently not the file somebody remembers writing.

The layer above your server, which robots.txt cannot speak for

A permissive robots.txt is a request. A firewall rule is enforcement, and enforcement wins. CDN and WAF bot protection blocks retrieval crawlers regardless of what your file says, and the settings are often on by default or switched on by whoever set up the domain. Cloudflare's AI bot blocking and its aggressive bot fight modes are the usual culprits, and they are easy to enable in a security review without anybody connecting them to marketing. Managed hosting bot rules do the same thing more quietly.

The test costs a minute: request your own homepage with a retrieval crawler user agent and see what comes back. A challenge page, a 403, or a JavaScript interstitial means you are blocked, whatever the file says.

JavaScript rendering, the invisible one

Googlebot renders JavaScript. Most retrieval crawlers used by assistants do not; they read the HTML your server returned. A single page application that ships an empty shell and paints its content client side is therefore perfectly indexable in ordinary search and close to blank in the answer layer. Nothing warns you about this, because your rankings look fine.

Fetch your homepage with scripts disabled and read what is actually in the response. If your value proposition, your pricing and your product description are not there as text, the fix is server rendering or pre-rendering for the pages that describe what you sell, and it outranks every other item on this page in importance.

CrawlerWhat it doesEffect of blocking it
OAI-SearchBotRetrieval for ChatGPT search resultsCannot be cited in ChatGPT search answers
ChatGPT-UserFetches a page when a user's question requires itCannot be read when a user asks about you directly
PerplexityBotIndexing for PerplexityAbsent from Perplexity answers and citations
Claude-SearchBotRetrieval for Claude's searchCannot be surfaced in Claude answers
GooglebotOrdinary search index, which AI Overviews and AI Mode are built onAbsent from search and from Google's AI surfaces alike
Google-ExtendedControl token for Gemini training and groundingNo effect on AI Overviews or AI Mode. Read that row twice.
GPTBotTraining data collectionExcluded from training corpora. A separate decision from retrieval, and not one that governs citation.

Layer two: can a machine turn your pages into facts?

Once a page can be fetched, the question is whether the things buyers ask about are stated plainly enough to be extracted and attributed. Two failures dominate, and neither is about markup.

The first is putting the answer somewhere unreachable. Pricing behind a demo request form is the classic: when somebody asks an assistant what your product costs, there is nothing to retrieve, so the assistant either says it is not publicly available, which reads as evasive, or reconstructs a number from an old third party article, which is worse. The same applies to what the product does and who it is for. If a fact is commercially important enough that buyers ask it, it belongs in crawlable prose.

The second is writing that never commits. Marketing copy built from abstractions gives a machine nothing to extract, because there is no proposition to lift. Compare "the modern platform for ambitious teams" with "project management for construction subcontractors with fifteen to fifty people, priced per project rather than per seat". Only one of those can be turned into a sentence in an answer.

The structured data worth publishing

Schema does not make you visible. It removes ambiguity for a machine that has already fetched you, which is a real but modest benefit, and it is cheap enough to be obviously worth doing.

  • Organization with a stable @id and sameAs pointing at every profile that is definitively you. This is the piece that matters most, because it is what lets the entity be referenced across pages rather than re-guessed on each one.
  • SoftwareApplication or Product for the thing itself, with offers and price where a price exists.
  • Article on editorial pages with a real author reference rather than a name string, and an honest dateModified.
  • FAQPage and HowTo where they genuinely describe the content. Be clear eyed about why: Google stopped showing FAQ rich results for most sites and removed HowTo rich results altogether, so the benefit now is machine readability rather than anything you will see in a search result. Publish them for that reason, and do not let anyone sell them to you as a ranking feature.

One discipline is worth more than any of the types: keep dateModified honest. Changing a date without changing content is detectable and is treated as exactly what it is.

Layer three: does your name resolve to anything?

This is the layer new products skip, and it is the one that decides whether the other two ever pay off.

Before an engine can decide whether to trust a claim about your product, it has to work out what your product is. For an established brand this is trivial. For a two month old name it is genuinely hard, and it fails in two specific ways. Either the name matches nothing at all, in which case there is nothing to attach your pages to, or it collides with something that already exists, in which case whatever the engine says about you may be about a company in a different industry with the same name. Both outcomes look identical from your side, which is silence.

What resolves a name is corroboration from sources that are not you. This is why the ordering matters: perfect markup describing an entity nothing else in the world mentions is a claim, not a fact. The nodes that do the work for a new product are unglamorous and mostly free.

  • A Wikidata item, if the product genuinely meets the notability bar. This is the most directly useful identifier, because it is machine readable, widely ingested, and explicitly designed for entity resolution.
  • A company profile on the directories your category actually uses, plus a LinkedIn company page. These are boring, they are indexed, and they corroborate that the organisation exists.
  • Consistent naming, everywhere, with no variants. One spelling, one capitalisation, one form. "Acme", "Acme.io" and "Acme Software" used interchangeably are three strings to a machine, and splitting your own evidence across three of them is a self inflicted wound.
  • An About page that states what the thing is in one unambiguous sentence, naming the category in the words buyers use. This sentence gets lifted more often than any other on a site, so write it as the sentence you want quoted.
  • Third party mentions, linked or not. Being named in a round-up, a comparison, a newsletter or a community thread is the currency here. Ahrefs, looking across roughly 75,000 brands, found brand web mentions correlate with AI Overview visibility around three times as strongly as backlinks do. That is correlational and worth holding loosely, but it points the same direction as everything else in this layer: what other people write about you is doing the heavy lifting.

If your name collides with something bigger

Disambiguation is a real piece of work and worth doing early, while there is little to unpick. Always pair the name with the category in your own copy, so the phrase that appears across the web is "Acme, the invoicing tool for freelancers" rather than "Acme" alone. Make the sameAs graph unambiguous. And accept that if the collision is with a genuinely large entity, the honest answer may be that the name is a long term liability, which is far cheaper to conclude in month two than in year three.

The order of work, and what to expect

Do them in this sequence, because each one is void without the one before it.

  1. Access. Fetch your own site as a crawler, with JavaScript off. Fix robots.txt, the firewall layer, and rendering. This is days of work at most and everything depends on it.
  2. Structure. Put price, purpose, audience and alternatives into crawlable prose. Add Organization schema with a stable identifier. A week.
  3. Entity. Create the off-site nodes, fix naming consistency, write the one sentence. Then earn mentions, which is the slow part and never stops.

Expect the timeline to be measured in months rather than weeks. Indexation of a new domain is quick; being named in answers requires corroboration to accumulate, and no amount of technical work compresses that. What the technical work does is make sure that when the coverage arrives, it can be fetched, parsed and attached to you, which is the difference between a mention that compounds and one that evaporates.

Run the readiness check with your URL if you want the first layer answered from evidence rather than from memory. It reads your robots.txt, checks whether your homepage is readable without JavaScript, and tells you what it found, which is usually the fastest way to discover that layer one was the problem all along.

Questions people ask

Does blocking Google-Extended remove me from AI Overviews?

No. Google-Extended governs whether your content is used for Gemini training and grounding. AI Overviews and AI Mode are built on the ordinary search index and rely on normal Googlebot access, so blocking Google-Extended removes you from neither and allowing it puts you into neither. Most articles on this get it wrong, and teams end up believing they have opted out of something they have not.

Do AI crawlers run JavaScript?

Mostly not. Googlebot renders JavaScript, which is why a client rendered site can still rank in ordinary search. The retrieval crawlers used by assistants generally fetch the raw HTML response and read what is there. A page whose content only appears after scripts run can therefore be indexed by Google and effectively empty to an assistant, which is the most common invisible failure for modern product sites.

Is schema markup enough on its own to get cited?

No. Schema helps a machine parse facts it can already fetch, and it does nothing if retrieval is blocked or if no external source corroborates what you claim. It is worth doing because it is cheap and removes ambiguity, but a product with excellent structured data and no third party coverage is still a product nothing else in the world confirms the existence of.

How long before a new product starts appearing in AI answers?

Longer than search indexing and for different reasons. Indexation of a new domain often happens within days. Being named in answers depends on accumulating corroborating sources, which is a matter of months rather than weeks. Anyone promising presence in AI answers within a few weeks of launch is describing something they cannot control.

Put a number on it

Score your own launch across all forty checks

Free, about seven minutes, and no email needed to see the result.

Read next

Product Launch Blog is an EbizIndia publication. This article does not pitch anything; the disclosure sits here instead, and in the footer, on every page.