Skip to main content
Haney Strategy

The note · August 26, 2026

How AI Search Actually Works

Google hands your buyer a list. An answer engine hands them an answer. This is what happens in between, in plain English, and the five checks that decide whether your company is in it.

Ranking and being cited are two different jobs. Google ranks pages and returns a list, while an answer engine breaks your buyer's question into a fan of smaller questions, runs them at once, and assembles one answer from whatever it can read and trust. That is why your page can sit at the top of a Google search and still not show up in an AI search, usually because of a crawler setting or a page structure nobody has looked at. Five questions decide it: can they read you, do they know who and where you are, do you answer the question, does the world vouch for you, and can you prove it is working.

Website, Organic Search & AI SearchJim Haney18 min read
A machined copper stem branching at a single junction into five arms, each holding a blank cream card, representing one search question fanning out into many simultaneous queries.

From my desk, August 26.

Someone in my network called me yesterday. He owns a small business, he had been reading about AI search, and he wanted one thing explained: what is actually different between ranking on Google and getting picked up by ChatGPT, and what does he have to do differently.

Some version of that question has come at me repeatedly over the past two weeks. Different rooms, different sizes of company, same thing underneath. Sometimes it arrives as “is AEO just SEO with a new label.” Sometimes as “we rank fine, so why does the AI never mention us.” Once it arrived as a genuinely good argument that this is search in its infancy all over again, so the toolkit from fifteen years ago should still work.

I have been answering it one call at a time. Writing it down is a better use of everyone's afternoon. So this is the fundamentals, in plain English. If you do not know what a crawler is, you will finish knowing exactly what to ask whoever runs your website. If you do, there are still two or three things in here worth your time.

Google hands you a list. An answer engine hands you an answer.

When your buyer searches Google, Google's job is to hand them a list. Ten links, ranked. Your job as a business is to sit high enough on it that somebody clicks. That has been the game since the beginning, and the entire search industry is built around it.

When your buyer asks ChatGPT, Claude, Perplexity, or Google's own AI Mode, the job is different. They are not asking for a list. They are asking for an answer, and the engine's job is to write one, in sentences, with a few sources named underneath.

Those two jobs run on much of the same machinery. Both have to reach your website, read it, and decide whether to trust it. That is why real search discipline still matters and is not going anywhere. But the finish lines are different. Ranking puts you on a list. Being cited puts you inside the answer. You can win one and lose the other.

That is measurable, and three different groups have measured it. None of them found what most marketing plans still assume.

When an AI engine writes an answer and names its sources, a lot of those sources are not on the first page of Google at all. AI Overviews, the written answer Google now puts above the blue links, are the easiest place to watch it happen.

Who measured itAhrefs, March 2026What they found38% of cited pages were also ranking in Google's top ten. Eight months earlier it was about 76%.What they sellAI visibility tools
Who measured itBrightEdge, nine industriesWhat they found16.7% of citations came from top-ten results.What they sellAI search tools
Who measured itOriginality.AI, November 2025What they foundAbout half of citations came from pages outside the top 100 entirely.What they sellAI detection

All three of those companies sell tools in this category, so do not lean hard on any single number. What they agree on is the part that matters: the list Google ranks and the answer an AI writes are not the same list.

Do not take either of the easy lessons here. Ranking still matters, and the work that earns it is most of the work that makes you quotable. It just does not buy you a mention on its own.

Meanwhile the click keeps getting rarer. SparkToro's analysis of Similarweb's US panel found 68 percent of US Google searches between January and April 2026 ended without anyone clicking anything. And getting named in an answer does appear to pay: Seer Interactive found brands cited in AI Overviews earned about 120 percent more clicks per time they appeared. Seer sells services in this category, and that is a pattern they observed rather than a proven cause.

68%

of US Google searches end with no click

38%

of AI Overview citations rank in Google's top 10

+120%

more clicks for brands that got cited

Fewer clicks to win, a shrinking overlap between ranking and being quoted, and a real premium for the companies that make it into the answer. That is the case for treating this as its own job.

The five questions

Every engine runs its own retrieval logic, with different indexes, different crawlers, and different rules about who gets cited. Google's two AI surfaces, the AI Overview above the blue links and the fully conversational AI Mode, both run on Google's own search index, so the technical groundwork you lay for Google carries straight over to them. That is not the same as saying a good ranking earns you a mention, as the numbers above make plain. ChatGPT, Claude, and Perplexity each run their own stacks, and I walked through how the platforms diverge in an earlier note.

Underneath the differences, the same five things decide whether any of them can use you. I take them one at a time below: what each one means, what is going on underneath it, and how to tell when yours is broken.

The five checks, in order

  1. Can they read you?Whether a machine can reach your pages, load them, and turn them into text it is allowed to use.
  2. Do they know who and where you are?Whether every engine resolves you to one company, in one place, selling one clear set of things.
  3. Do you answer the question?Whether your pages answer what a buyer actually asks, in a shape an engine can lift.
  4. Does the world vouch for you?Whether the record outside your website agrees with the claims inside it.
  5. Can you prove it is working?Whether you have a baseline, a repeatable test, and a line back to pipeline.

Work them in order. Checks one and two are the foundation, and everything below is worth very little until they are true. The one exception is the baseline in check five: take that before you change anything, because you cannot prove movement without a before.

Can they read you?

Two words do most of the work here, and most people use them as if they mean the same thing. Crawling is the visit: a program arrives at your page and reads what is there. Indexing is the filing: the engine decides the page is worth keeping and puts it in the library it searches later. A page can be crawled and never indexed. A page that was never indexed cannot be quoted, no matter how good it is.

That is the most common version of the problem I get called about. The page ranks on Google, so everyone assumes the work is done, but the engine your buyer is actually using never filed it. That is not a writing problem. It is plumbing, and it is invisible from the marketing seat.

The plumbing starts with robots.txt, a plain text file that lives at yourcompany.com/robots.txt. You can go look at yours right now. It is a note taped to the door listing which automated visitors are welcome and which are not. It is a request rather than a lock. Most declared crawlers honor it, but not always, and the exceptions are public: in August 2025 Cloudflare accused Perplexity of fetching pages from sites that had blocked it by using undeclared user agents, and removed it from its verified bot list. Perplexity denied it. Treat robots.txt as the front door you control, and remember that an actual block lives on the server, not in a text file.

Here is where it goes wrong, and this is the part almost nobody knows. Every major AI company gives you two separate controls: one governs whether they can train on you, the other governs whether they can find you and quote you in an answer. Different settings, different consequences, and they are not even the same kind of thing from one company to the next.

CompanyOpenAIControls trainingGPTBot, a crawlerControls whether you can be quotedOAI-SearchBot, a crawler
CompanyAnthropicControls trainingClaudeBot, a crawlerControls whether you can be quotedClaude-SearchBot, a crawler
CompanyPerplexityControls trainingDoes not train on crawled contentControls whether you can be quotedPerplexityBot, a crawler
CompanyGoogleControls trainingGoogle-Extended, a token, not a crawlerControls whether you can be quotedGooglebot, a crawler
CompanyAppleControls trainingApplebot-Extended, a token, not a crawlerControls whether you can be quotedApplebot, a crawler

In OpenAI's documentation, OAI-SearchBot is the one “used to surface websites in search results in ChatGPT's search features.” Anthropic splits it the same way, and Perplexity states plainly that PerplexityBot “is not used to crawl content for AI foundation models.”

The last two rows work differently, and the difference is worth understanding before you touch anything. Google and Apple each run one crawler and pair it with a control token that does no crawling at all. Google's documentation says Google-Extended “doesn't have a separate HTTP request user agent string” and that the token is “used in a control capacity.” It governs whether content Google already crawled can train Gemini, and Google states it has no effect on your inclusion in Google Search. Apple says the same about Applebot-Extended: pages that disallow it can still appear in search results.

Microsoft is different again. There is no Copilot crawler to block. Bing controls AI usage with nocache and noarchive tags on the page itself, and neither one removes you from Bing's ordinary results. So if your plan is “check the robots.txt file,” that plan does not cover Copilot at all.

Back in 2023 a lot of companies blocked GPTBot. That was a defensible call and for some businesses it still is. What most of them did not realize is that the block did nothing to their eligibility to be recommended inside ChatGPT, because a different crawler handles that. Block the training control and you have opted out of future training, though nothing recalls what a model already learned or what reached it through somebody else's dataset. Block the search crawler and you have opted out of the answer itself. Same file, two lines apart, completely different consequences.

Every crawler and control named here was checked against the company's own documentation on August 26, 2026. This is the fastest-moving corner of the whole subject, so if you are reading this some months later, check again before you act on it.

Then there is the second trap, and it is the cleanest explanation I know for “we rank well and still get nothing.” Google's own documentation says that to be eligible to appear as a supporting link in AI Overviews or AI Mode, “a page must be indexed and eligible to be shown in Google Search with a snippet.” A snippet is the short preview of text under a search result, and there are standard settings that suppress it or cap its length. Some publishers use them deliberately. Plenty of other sites have them switched on because of a plugin default nobody has looked at since.

If a page cannot show a snippet, it cannot be a supporting link. It can still rank. It just cannot be quoted.

You can rank at the top of Google and be invisible inside the answer, and the reason is usually a setting nobody knew was there, not the quality of your writing.

One more piece. Google's guidance, in its own words: “Making sure that important content is available in textual form.” Many modern sites build their pages in the browser after loading and put the most persuasive material inside images, sliders, and interactive components. People see it. A machine may not. If your value proposition lives only inside a graphic, assume it is not being read.

The tell: you rank respectably on Google, no AI engine ever names you, and nobody at your company can tell you what your robots.txt file says.

Do they know who and where you are?

An engine cannot recommend you until it decides who you are. It is trying to resolve every mention of your name into one thing: one company, in one place, selling one clear set of services. Marketers call that an entity. Think of it as the engine's index card about your business.

Now count how many versions of you exist. The website says one name and the invoices say a slightly different one. There are two location profiles because somebody set up a second after the move. The services page and the LinkedIn page describe the business differently. A directory listing still carries the old address. None of that is a crisis alone. Together they make the engine unsure, and an unsure engine reaches for a competitor it is sure about.

The technical layer underneath is structured data, sometimes called schema: a small block of code stating plainly what the page is about. This is a company, this is what it sells, this is where it operates. Google's own words are that it provides “explicit clues about the meaning of a page,” and that Google uses it to gather information “about the people, books, or companies that are included in the markup.”

Now the correction, because somebody may be selling you the wrong version of this. There is a proposed file called llms.txt being pitched as the thing you must add to be visible in AI search. On Google's documentation page about AI features, in writing: “You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add.”

That is Google saying the file does nothing for Google. As far as I can find, no other engine has published a claim either way, so here is the fair version: it is a cheap file, it may turn out to help somewhere, and it is not a plan. If it is the centerpiece of what somebody is proposing to you, ask what else is in there.

The tell: ask three AI engines what your company does and where it operates, and you get three meaningfully different answers.

Do you answer the question?

Here is the part that surprises people who have run search programs for years. Your buyer asks one question. The engine does not go looking for one answer.

Google calls it query fan-out, and describes it this way: AI Mode “breaks down your question into subtopics and issues a multitude of queries simultaneously on your behalf.” It splits the question into the smaller questions somebody would have to answer to answer the big one, runs them at the same time, and assembles a single response from what comes back.

What your buyer typed

Who should handle IT for a 30-person law firm in Nashville?

What the engine ran, all at once

  • managed IT providers Nashville small business
  • law firm IT compliance requirements Tennessee
  • average cost of managed IT for 30 employees
  • best rated IT providers Nashville reviews
  • do law firms need a specialized IT provider
  • managed IT versus in-house for a small firm
One question in. A fan of searches out. The company that shows up across the most of them is the one that gets named.

Google has never published how many of those it runs. Outside estimates range from eight to more than twenty depending on the question, and I would not build a plan on any particular number. The count is not the point. The shape is.

Because the shape changes what a good page looks like. If your site has one page built around the phrase you want to rank for, you have entered one branch of the fan and forfeited the rest. The company that gets named has real material sitting under several branches at once: what it costs, who it is for, what the alternatives are, what your industry specifically requires.

And the engine does not read your page the way a person does. It is looking for the paragraph that answers the question. It pulls passages, not pages. So a genuinely useful answer buried in the fourth paragraph under a heading that says “Our Approach” is worth less than the same answer under a heading that says what it is, in the words a buyer would use.

That is why a frequently asked questions section helps, and why bolting one on does not fix this by itself. FAQs work because they are shaped like the thing the engine is hunting for: a real question, followed by a direct answer. That shape belongs throughout your site, not only in a box at the bottom.

The tell: your best pages are built around two or three phrases somebody picked years ago, and nobody has written down the questions buyers actually ask on a first call.

Does the world vouch for you?

This is the check people like least, because you cannot fix it by editing your own website.

An engine assembling an answer weighs what other credible sources say about you more heavily than what you say about yourself. Your website is the claim. Reviews, trade press, directories, association listings, partner pages, and earned coverage are the corroboration. When the two agree, the engine gets confident. When your website is the only place a claim appears, it stays a claim.

There is an obvious counter here, and it is a good one: if the engines are young, why not run the old playbook and manufacture the corroboration? Publish your own rankings, seed the lists, appear everywhere.

The honest answer is that the cheap version of this got priced out a long time ago, at least on the search side, and the closing of it is documented. Google now defines link spam as “creating links to or from a site primarily for the purpose of manipulating search rankings,” with sites that do it liable to “rank lower in results or not appear in results at all.” I would not assume every answer engine inherited that machinery wholesale, because they run different stacks and some of them lean heavily on forums and aggregators. What I would assume is that manufactured volume is the most contested and least durable thing you can buy, and that a genuine third-party record is the part nobody can take away from you later.

Which is good news for a company that does good work and has been quiet about it. The gap between the quality of your work and the visibility of your record is the most fixable gap on this list. It is just not fixable this week.

The tell: look your company up and everything on the first two pages is something you published or paid for. That is a thin record, and every engine is reading the same thin record.

Can you prove it is working?

Here is the measurement problem, and it almost always arrives in the same shape. With Google it is easy. Before the work you look yourself up and you are on page 25. After the work you look again and you are on page two. You know it moved. With an answer engine there is no page 25 to be on. There is one answer, and it either mentions you or it does not.

And the test almost everybody runs first is the wrong one. You open ChatGPT, you type your own company name, and there you are. That is not a result. That is a mirror. You are signed in, the model has your history, and you put the answer inside the question. Ask again tomorrow and it will be even more certain, because now it knows you care about that company.

Then be patient in a specific way. The settings in check one can register within a crawl or two. Everything downstream of other people, your content depth and your off-site record, moves over months, the same way organic search always has. The number that matters at the end is not how many times you got mentioned. It is whether the mentions turned into conversations, which means somebody has to ask new inquiries where they heard about you and write the answer down.

The tell: your last report on this was a screenshot of a chat window, taken while signed in.

Six moves, in order

  1. Take a baseline this week, before you change anything. Four engines, real buyer questions, no company names, recorded with the date. You cannot prove movement without the before.
  2. Open yourcompany.com/robots.txt and read it. It is a public file, so anyone can. Find out which crawlers you block and whether anybody chose that on purpose. Do not edit it yourself: one wrong line in that file can remove you from Google entirely, and some blocks are there deliberately for legal or security reasons. Take what you find to whoever runs the site.
  3. Check whether your key pages can show a snippet. If snippet suppression is set on pages that matter, you are ineligible to be quoted by Google's AI features no matter how well you rank.
  4. Make one true index card. One company name, one address, one description of what you sell, and make the website, the location profiles, the directories, and the social profiles all say it.
  5. Write down the twenty questions buyers actually ask you, then count how many have a real answer on your site under a heading in their words. Fix the biggest gaps first.
  6. Name three places your record should exist and does not. A trade publication, an association directory, a review site your buyers actually read. This is the slowest item on the list and the only one a competitor cannot copy off your website.

None of that requires a new vendor or a rebuild. It does require real hours from somebody senior enough to make the call and technical enough to read the file, which is the combination most companies are missing, and the reason this sits undone in so many good businesses.

Everything in this note, plus the full crawler directory across six engines, a plain-English glossary, the robots.txt cheat sheet, and a printable version of the five checks, is in The AI Search Field Reference. No form, no email, no gate. Take it to whoever runs your website.

The part I would not skip

These five questions are not a framework I invented for a note. They are the plain-English version of the standard I score against when a company brings me in to fix this, which is what the AI Search Power-Up is: I grade the site, ship the fixes myself, then re-score against the original baseline so you can see what actually moved.

You do not need me for the first four moves on that list, though. You need somebody to open the file and look. Most companies find the answer is embarrassing and cheap, which is the best outcome available.

If you want to walk through what came back when you looked, book an introductory call and bring the screenshots. That is a good hour.

Ranked is not the same as read. Go find out where you actually stand.

All signal. No noise.

Frequently asked questions

What is answer engine optimization, and is it different from SEO?

Answer engine optimization, or AEO, is the work of making your company usable by AI systems that write answers instead of returning lists. It overlaps heavily with SEO, because both depend on a machine reaching your site, reading it, and trusting it. The difference is the finish line. SEO is about earning a position on a ranked list. AEO is about being understood, selected, and quoted inside a written answer, which is decided by different mechanics and measured differently.

Why does my site rank well on Google but never get mentioned by ChatGPT?

Usually one of three reasons. The engine's crawler cannot reach your pages, because your robots.txt file blocks it, often a block nobody remembers making. Or the page is not eligible to be quoted, because a snippet setting prevents it. Or the content is not in a shape the engine can lift, because the answer to the buyer's question is buried inside a page built around a keyword. Ranking and citation are separate jobs now. In Ahrefs' March 2026 study of 863,000 sets of search results, only about 38 percent of the pages cited in Google's AI Overviews were also ranking in the top ten for that question, which tells you how much of what these engines quote is coming from somewhere other than page one.

Do I need an llms.txt file?

Google says in its own documentation that you do not need new machine readable files, AI text files, or special structured data to appear in AI Overviews or AI Mode. It costs almost nothing to publish one and I would not tell anyone to remove theirs. Google has said it does nothing for Google, and as far as I can find no other engine has published a claim either way, so treat it as a cheap file rather than a plan. Spend the effort on crawler access, content structure, and your off-site record instead.

Should I block AI crawlers from my website?

For most companies, no. The upside of being quotable is larger than the downside of being trained on, and the two are separately controllable anyway. If you do want to opt out of training, it is worth knowing that every major AI company gives you separate controls for separate purposes. One collects material to train models. A different one finds and cites you when somebody asks a question. Blocking the training crawler keeps your material out of model training. Blocking the search crawler removes you from the answers your buyers are reading. They are different lines in the same file, and most companies that blocked one in 2023 did not know the other existed.

How do I test whether AI engines recommend my company?

Not while signed in. The model remembers what you have asked before, so typing your own company name into an account you use daily produces a flattering result that means nothing. Get out of your own history, which works differently on each engine: ChatGPT has a temporary chat, Perplexity and Google's AI Mode run in a private window, and Claude has no signed-out mode at all, so turn memory off in settings and start a new conversation. Then ask the question the way a buyer would ask it, without naming your company, and write down what each one says with the date. That is your baseline, and you need it before any work starts.

How long before this shows results?

Plan on a couple of months for real movement, the same way organic search has always worked. Technical fixes such as crawler access and snippet eligibility can register quickly once the engines re-crawl. Content depth and your off-site record take longer because they depend on other people, and that is the part nobody can shortcut for you.

Share

Next note · In the works

Want this thinking applied to your business?

Signal Notes can sharpen the thinking. A strategy call turns it into a plan.