How to Get Your Website Cited by ChatGPT, Perplexity and Google AI Overviews
Published • Last verified • 8 min read
By Clunky AI editors
Practical website readiness guidance for founders using AI builders. We do not invent customer results, scan data or proof.
About the editorsThe short version: AI assistants cite sites they can find, read and trust. Most websites, and many AI-built ones, fail at the first step: their important content is missing from the initial HTML or blocked by crawler and CDN rules. Fix crawlability, publish specific material worth citing, keep your business identity consistent across the web, and measure whether the same buyer questions produce citations over time. Everything else sold under “AI SEO” is either a repackaging of those fundamentals or a claim that needs evidence.
A growing share of customers now ask ChatGPT, Perplexity or Google's AI features questions they once typed into a search box. When the answer names three companies and yours is not one of them, you did not lose a blue-link ranking; you were never in the answer set. Here is how to improve the odds without pretending anyone can guarantee placement.
How AI assistants decide what to cite
Two separate mechanisms matter, and conflating them is the source of much bad advice:
Training knowledge. What a model absorbed about you from data available before or during its training. You influence this slowly through consistent presence on the open web. It is not an outcome you can schedule for this quarter.
Retrieval. When an assistant answers a current or source-sensitive question, it may search or fetch live web pages and cite what it used. ChatGPT search, Perplexity and Google's AI features all surface links, although their exact systems and ranking methods differ. Retrieval is actionable now because crawlability, indexation, useful content and corroboration all affect whether your page can enter the candidate set.
The nine steps below focus on retrieval.
Step 1: Let the AI crawlers in
Check your robots.txt and CDN or firewall rules. The retrieval and user-directed bots to know are OAI-SearchBot for OpenAI, PerplexityBot and Perplexity-User for Perplexity, Claude-SearchBot and Claude-User for Anthropic, and Googlebot for Google Search.
Blocking a training crawler such as GPTBot or ClaudeBot is a separate choice about model training. Blocking search or user-directed retrieval bots can reduce your chance of being fetched for live answers. OpenAI explicitly separates OAI-SearchBot visibility from GPTBot training controls. Make those decisions separately, and check whether a broad CDN “block AI bots” switch has combined them for you.
Step 2: Serve content bots can read without relying on JavaScript
If your site renders important text only in the browser, a crawler may receive little more than an app shell. This is a common failure mode on sites built with Lovable, Bolt, v0 and other React-first tools. The test and the fix are in our Lovable SEO guide; the rendering section applies to every builder, not just Lovable.
Do not confuse “the page works in Chrome” with “the server response contains the page”. Fetch the live HTML and look for the visible headline, body copy, title, description and canonical URL.
Step 3: Write useful answers, not machine bait
Assistants need passages they can understand in context. Pages that are easy to retrieve and cite usually have a clear shape:
- Direct opening paragraphs. Answer the section question before elaborating.
- Descriptive headings. Use language a reader would recognise, including real questions where they fit naturally.
- Self-contained sections. A passage should still make sense when surfaced away from the rest of the article.
- Specifics. Numbers, dates, named tools, concrete steps and explicit caveats give an assistant something verifiable to use.
This is good editorial structure, not a special incantation for AI. Google's current guidance explicitly says you do not need to rewrite content into tiny chunks or use a special style for generative search.
Step 4: Publish something only you can publish
Original data, your own measurements, documented case results, a tested process or a defensible first-hand opinion gives other writers and answer engines a reason to cite you rather than the source you summarised. If you have proprietary data, publish the methodology with the number. One transparent statistic can be more useful than ten thousand words of paraphrase.
Do not invent corpus statistics to sound authoritative. A smaller sample with a visible method is stronger than a large number nobody can audit.
Step 5: Keep your entity consistent
Models and search systems need to resolve your organisation into a recognisable entity: a name, category, location where relevant, and a set of products or services. Use the same business name and plain-English description across your site, business profiles, directories and social bios. Add an About page that states what you are, who you serve and where you operate.
If your homepage says “reimagining digital experiences”, a machine and a buyer both struggle to categorise you. “We scan AI-built websites for accessibility, search and trust failures” is less glamorous and more useful.
Step 6: Mark it up and date it honestly
Structured data such as Organization and Article can remove ambiguity and make pages eligible for established search features when the markup matches visible content. It is not a special AI-ranking layer, and Google says there is no dedicated schema required for AI Overviews or AI Mode.
Visible dates help readers judge freshness. Add “last verified” only when someone has actually rechecked the claims and links; changing the date without changing the work is not freshness.
Step 7: Earn real corroboration
Third-party comparison pages, specialist directories, customer reviews and community discussions can corroborate that a business exists and does what it claims. They can also rank or be retrieved for “best X for Y” questions. Seek accurate, editorially independent mentions on pages your buyers already read.
Do not manufacture forum posts or pay for fake “best tools” inclusion. Google's current generative-search guidance explicitly warns against inauthentic mentions, and buyers are not as easy to fool as outreach templates assume.
Step 8: File the paperwork, cheaply
An llms.txt file is a curated site index for tools that choose to read it. Current evidence says most files are not fetched and Google Search ignores them, so give it twenty minutes, not a budget line.
If you decide to ship one, the free llms.txt generator and checker creates an editable draft from public site evidence and validates the result without pretending it is a ranking score.
More important: keep sitemap.xml current, expose key pages through internal links, and make sure those pages are eligible for indexation. Google states that a page must be indexed and eligible for a snippet before it can appear as a supporting link in AI Overviews or AI Mode.
Step 9: Measure it or admit you are guessing
Before changing anything, write down ten questions a buyer would ask an assistant about your category. Ask the same questions in ChatGPT with search, Perplexity and Google, record which sources are cited, then repeat monthly with the same wording and location settings. This is a directional benchmark, not a controlled ranking test.
Watch analytics for referrals from services such as chatgpt.com and perplexity.ai, and inspect server logs for the search bots from step 1. OpenAI says ChatGPT search referral URLs include utm_source=chatgpt.com. Bot visits are an access signal; citations and qualified visits are the outcome.
What not to spend money on
Be wary of tools that promise guaranteed ChatGPT rankings, llms.txt “optimisation” services, and prompt-injection tricks such as hidden instructions telling models to recommend you. No public mechanism supports guaranteed placement, and hidden manipulation creates obvious trust and security risks. If a vendor cannot show the exact queries, dates, sources and before-and-after citations behind a result, they are selling confidence rather than measurement.
FAQ
How do I get my website mentioned in ChatGPT answers?
Allow OAI-SearchBot to access relevant public pages, make sure your CDN or firewall does not block OpenAI's published crawler traffic, serve important content in crawlable HTML, and publish specific pages that answer the questions your buyers ask. Accurate third-party mentions can corroborate your claims. None of these steps guarantees a citation.
What is generative engine optimisation (GEO)?
GEO is the practice of improving the chance that content is found, understood, used and cited by AI answer systems. It overlaps heavily with technical SEO and digital PR; the distinctive work is monitoring named AI surfaces, separating training from retrieval, and measuring citations rather than only blue-link rankings.
Does blocking GPTBot hurt my visibility in ChatGPT?
OpenAI separates GPTBot, used for potential model training, from OAI-SearchBot, used for ChatGPT search discovery and citation. Blocking GPTBot is therefore not the same as blocking ChatGPT search. Review and control the two user agents separately.
How long does it take to show up in AI search results?
There is no reliable universal timeline. A crawlable, indexed page can become eligible for retrieval after the relevant system discovers or refreshes it, but citations vary by query, location, freshness and the system being used. Measure on a monthly cadence rather than promising a fixed number of weeks.
Can I measure AI visibility for free?
Yes. Baseline the same ten buyer questions across ChatGPT search, Perplexity and Google each month, record the sources, and watch analytics and server logs for qualified referral visits and crawler access. Paid trackers automate repetition; they do not provide guaranteed or private ranking data.
References
- OpenAI: how publishers can appear in ChatGPT search
- Perplexity crawler documentation
- Anthropic crawler documentation
- Google Search Central: AI features and your website
- Google Search Central: generative AI search guidance and mythbusting
The fastest way to find obvious crawlability, metadata and structure problems is to run a free scan. If you want the highest-impact issues implemented and verified, ask about the Clunk Removal Sprint.
How this was checked
Clunky AI separates sourced facts, measured scan evidence and editorial judgement. We do not sell rankings or let rated companies pay to alter coverage.
Read the editorial policyExplore the six basics
Every Clunky AI article maps back to one or more of the questions a business site has to answer.
Related Posts
Tags AI VisibilityGEOSEO
Category AI Visibility