Person using MacBook Pro — illustrating How to Rank in ChatGPT Search: What We Learned Testing Real Queries

How to Rank in ChatGPT Search: What We Learned Testing Real Queries

“How do we rank in ChatGPT?” is now a standing question in most marketing meetings, and it’s usually answered with speculation. The honest position is that nobody outside the labs knows the selection logic, but the behaviour is observable, and observation is a legitimate substitute for documentation.

What follows is a method for testing ChatGPT search systematically, and the kinds of patterns that reliably show up when you do. We’re deliberately not attaching percentages to any of it. The value is in the mechanism and the method, not in numbers that would be unreproducible next month.

Why testing beats theorising

ChatGPT doesn’t expose a ranking report. There’s no position, no impression count, and no explanation of why one source was quoted and another ignored. The only instrument available is the answer itself.

That makes testing an evidence-gathering exercise rather than a measurement exercise. You’re not trying to find your rank; you’re trying to characterise how the system treats your category so you can decide where to invest. Done properly, it answers questions like which competitors anchor the category, which of their pages get pulled repeatedly, and which questions have no settled answer yet.

How to design a test set that tells you something

Most teams test badly, in a predictable way: they ask ten flattering questions once, see themselves mentioned twice, and conclude things are fine. A usable test set has structure.

  1. Cover the question types separately. Definitional (“what is X”), procedural (“how do we do X”), comparative (“X versus Y”), selection (“best provider for X”), and problem-symptom (“our X keeps failing when we Y”). These behave differently and should be scored separately.
  2. Write prompts the way buyers speak. Full sentences, with context and constraints attached, not keyword strings.
  3. Include unbranded questions. Branded prompts almost always surface you and tell you little.
  4. Repeat each prompt on separate days, in fresh sessions. Outputs are non-deterministic; a single run is an anecdote.
  5. Vary phrasing deliberately. Ask the same underlying question three ways and see whether the source set moves.
  6. Record structure, not just presence. Note which URLs were cited, in what order, and how each source was characterised.

Keep the raw answers. The competitor citation pattern in your archive is often more useful six months later than the original scoring was.

What patterns typically emerge

When you run this kind of test set across engines, a set of recurring behaviours tends to appear. These are qualitative observations, not measurements, and you should verify them in your own category rather than take them on faith.

The same handful of sources anchor a category

Across many phrasings of the same underlying question, you usually see a small, stable set of domains recur, with a longer tail that swaps in and out. Getting into the tail is achievable with a strong page; displacing an anchor takes sustained work and independent corroboration.

Cited passages are usually short and self-contained

When you trace a citation back to the source page, the material being used is typically a compact definitional sentence, a short list, or a direct answer sitting immediately under a heading. It’s rarely the middle of a long argumentative paragraph. This is the strongest practical signal from testing: structure influences whether a page can be used at all.

Phrasing changes the source set more than expected

Reword a question from an expert framing to a layperson framing and the sources often shift substantially. That implies the system is matching on the question as asked rather than on a normalised topic, which is why literal, question-shaped headings work.

Comparison and “alternatives” queries favour neutral pages

For prompts that ask which option is better, sources that treat multiple options evenhandedly appear more often than vendor pages arguing for one. A vendor page that fairly describes when a competitor is the better fit is unusually citable, because it is safer for a synthesis system to rely on.

Recency matters more for some questions than others

Fast-moving topics tend to surface recently updated pages; stable definitional questions tend to surface older, well-established ones. A visible last-updated date and genuinely refreshed content help mainly in the first case.

Engines disagree, and the disagreement is informative

Run the same prompt through ChatGPT, Perplexity, and Google AI Overviews and you’ll often get overlapping but distinct source sets. Where an engine cites you and another does not, the difference usually traces to how each one retrieves: index-driven engines respond faster to new content, while others lean more on established authority signals.

Turning observations into changes

Testing is only useful if it changes what you publish. Three conversions are worth making routinely.

  • Gap list. Questions where you’re absent and no anchor source dominates. These are your highest-value content targets.
  • Rewrite list. Pages that rank in classic search for a tested question but are never quoted. These almost always need structural work rather than more words.
  • Correction list. Answers where you are mentioned but described inaccurately. These are entity-facts problems, fixed on your About page, schema, and third-party profiles rather than in blog content.

Score the pages on your rewrite list with the AI citability scorer before and after, and keep the test set running on a schedule with the ARIA citation tracker so you can see whether the changes moved anything.

What testing cannot tell you

Be clear about the limits. You cannot infer the ranking algorithm from outputs, you cannot attribute a change with confidence when the model itself may have been updated between runs, and you cannot generalise from your category to another. Non-determinism means you are always reading a distribution, not a result.

Treat the whole exercise the way you would treat qualitative market research: a disciplined way to reduce uncertainty and prioritise, not a measurement system that produces a number you can defend to two decimal places.

If you want this run continuously and interpreted against your competitive set, our AI visibility practice does exactly that, or start a project to scope it.

Frequently asked questions

Is there a ranking position in ChatGPT search?

No. There’s no stable position to track, only whether you’re mentioned, whether a specific URL is cited, and how you’re characterised. Any tool reporting a precise ChatGPT rank is modelling, not measuring.

Why do I get different answers to the same question?

Model outputs are non-deterministic, and retrieval results change as indexes update. This is why repeated runs across days matter far more than any single test.

How many prompts should a test set contain?

Twenty to forty covering all the major question types is enough to reveal patterns while remaining maintainable. Depth of repetition matters more than breadth of coverage.

Should I test logged in or logged out?

Use a fresh, logged-out or memory-disabled session wherever possible. Personalisation and conversation history contaminate results and make your own brand look more prominent than it is.

Enjoyed this?

Get the next one in your inbox.

Practical insights — no fluff, straight to your inbox.

Or follow us on LinkedIn:

Follow StrategyPeeps

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *