Blog · 5 min read

How AI Engines Choose Which Businesses to Cite

By the HowsMyGEO team · Published · Updated

When an AI engine answers a question, it typically cites a small number of sources — far fewer than a page of search results. Getting into that set is the whole game of GEO. Engines don't publish their selection criteria, but the observable mechanics — what crawlers can read, what answers quote, what the engines' own documentation asks for — point to a consistent set of page-level signals.

What has to be true before you can be cited at all?

Two prerequisites, and most losing sites fail here rather than at content quality. First, the engine's crawler must be allowed in — a robots.txt block is absolute. Second, the content must exist in the HTML the crawler receives: most AI crawlers don't execute JavaScript, so client-rendered content effectively isn't there. Every signal below only matters after these two pass.

What makes a passage quotable?

  • It answers the question directly. Engines favour passages that resolve a query in one lift: a question-shaped heading followed by a plain answer, a definition sentence, an FAQ pair. Content buried in narrative has to be reconstructed; content shaped like an answer gets used as one.
  • It contains checkable specifics. Numbers, dates, prices, named entities. A generated answer built on specifics is more defensible, so engines prefer sources that provide them. "Fixed fees from £89/month" is citable; "affordable pricing" is not.
  • It is attributable. A named author, an organisation with an About page, Organization and Person schema. Engines carrying a citation are lending your page their credibility; anonymous, undated pages are a risk they don't need to take.
  • It is dated. Machine-readable and visible dates let an engine judge whether your claim is current. Undateable content gets discounted, especially for anything time-sensitive — prices, regulations, product features.
  • It is structured. Clean heading hierarchy, lists, structured data. Not because engines reward tidiness morally, but because extraction is parsing, and parseable pages yield more usable passages.

Does brand size decide citations?

Partly, and it's worth being honest about: famous brands appear in answers because the models already know them from training data, page signals or not. The flip side is that this is where smaller businesses have the most to gain — for queries where no brand dominates ("best fixed-fee accountant for UK contractors"), the citation goes to whoever made their page most readable, specific and attributable. Those queries are won on the signals above, and they are exactly the high-intent queries that convert.

How do you find out which signals you're missing?

Every signal in this post is externally checkable, which means it can be audited. Our free scan runs 21 of these checks — crawler access, extractability, answer-shape, factual density, entity signals, freshness — and the methodology page publishes every weight, so you can see exactly why you scored what you scored.

Where does your site stand?

Run the free 21-check scan — SEO score, GEO score and your grade in about a minute.

Get my SEO + GEO score

Related reading