How these studies are run
Published before any data was collected, so the method cannot be adjusted after seeing the results.
In short
This page describes the standing methodology behind every study in this section. It is deliberately published first, and separately, so that the method is fixed before results exist and cannot be quietly reshaped to fit them.
Why the method comes first
Almost every published claim about AI search visibility is unfalsifiable. It describes an outcome without describing how it was measured, which means nobody can check it and nobody can reproduce it. A number with no method behind it is a marketing asset, not a finding.
So the order here is fixed: the method is published, then data is collected, then results are reported against the method that was published. If a study needs a method change partway through, the change is recorded with its date and reason, and the earlier data is reported separately rather than silently merged.
This costs speed. It is the only way the results are worth anything.
How prompts are chosen
Prompts are written as a customer would actually type them, not as a marketer would like them phrased. "Who builds websites for restaurants on Long Island" is a real query. "Best AI-powered Long Island web design agency near me 2026" is not; it is a keyword string wearing a question mark.
The prompt set is fixed in writing before the first run and does not change between intervals. Adding a prompt mid-study creates a new series rather than extending the old one, because a set that grows over time cannot be compared against itself.
Prompts are grouped by intent: branded, where the business name is in the prompt; category, where it is not; local, where a place name is; and comparison, where the user is choosing between named options. These behave very differently and averaging across them hides more than it shows.
Controls, and the ones that are not fully controllable
Every run is unpersonalized: logged out, no chat history, fresh session, US locale. Personalization is the largest uncontrolled variable in this kind of measurement and a logged-in run measures a single account rather than a system.
Geography is set explicitly where the platform allows it, and recorded as uncontrolled where it does not. A local query answered from an unknown inferred location is a different measurement from one answered for a stated location.
Model version is recorded where the platform exposes it. Where it does not, that is recorded as unknown rather than assumed stable, because these systems change under a fixed product name and a result from one week is not necessarily comparable to the next.
Repetition matters more here than in traditional search. The same prompt can return materially different answers on consecutive runs, so a single observation is an anecdote. Each prompt is run multiple times per interval and the variation itself is reported, not smoothed away.
What gets recorded
For each run: the date, the platform, the product or model where exposed, the exact prompt, whether any business was named, the order in which businesses appeared, every cited URL, and whether the answer was grounded in live retrieval or answered without it where that is observable.
Citations are recorded as full URLs, not domains, because the specific page matters and because a domain-level count conflates a homepage with a deep reference page.
Absence is recorded as data. A prompt returning no local business at all is a result, and a common one, and studies that only record hits systematically overstate how often these systems recommend anyone.
What this kind of study cannot establish
It cannot establish causation. If a business appears in more answers after a change, the study can report the correlation and the interval, and cannot attribute it to the change. Too much moves at once: the platform, the index, the competitors, and the model itself.
It cannot generalise from one site. A finding on one domain is a case report. Where the sample is one, the page says so in the same breath as the number.
It cannot see inside the system. None of these platforms publishes its source-selection criteria. Any account of why a particular source was chosen is inference, and is labelled as inference here rather than presented as mechanism.
Which claims here are ours, and which are the platforms’
This page mixes two kinds of statement and they deserve different treatment. Most of it is original methodology: what we do, in what order, and why. Those need no external citation, because they are not claims about the world. They are a description of a process, and the process is the thing being published.
A smaller number are claims about how these platforms behave, and those do need support. The load-bearing one is that a product name is not a stable measurement target. OpenAI publishes a deprecation schedule and states that it regularly retires older models, which is why this study records model version where a platform exposes it and records it as unknown where it does not, rather than treating a fixed name as a fixed system.
Two further statements are observations rather than documented behaviour, and are labelled that way wherever they appear. The first is that repeated identical prompts can return materially different answers, which we have seen directly and report as run variance rather than assert as a property of the system. The second is that personalization materially changes answers, which is why every run is unpersonalized. We are not aware of a platform that publishes the size of either effect, and we do not estimate it.
The distinction matters because a methodology page that cites nothing looks either careless or like it has nothing worth citing. This one had that problem until a documentation review found it, which is recorded here rather than quietly corrected.
Reproducibility, and the honest limit on it
Prompts, dates, platforms and raw counts are published so a reader can run the same set and compare. That is real reproducibility and it is worth having.
It is also bounded in a way that a laboratory study is not. These systems are non-deterministic, they are updated without notice, and they do not version their behaviour publicly. Someone repeating this set in three months may get different answers because the system changed, not because the method failed. That is a property of the subject, not a defect in the study, but it does mean results are a snapshot with a date attached rather than a stable measurement.
Nothing is backfilled. An interval that was not measured is recorded as missed. Reconstructing a data point after the fact, from memory or from a later observation, would make the whole series untrustworthy for the sake of one cell in a table.
What this study cannot tell you
Published alongside the method, not buried after the results, because a study that hides its limits is advertising.
Sample sizes are small. These are manual studies run by a small agency, not instrumented data collection at scale, and small samples on a non-deterministic system carry wide uncertainty.
Web search is used as a proxy for the retrieval layer in some measurements. It is a useful proxy and it is not the same instrument as querying an assistant directly. Where a proxy was used, the study says so.
Assistant answers must be captured by a human, so intervals are days rather than continuous, and a change occurring between intervals is invisible.
There is no control site. Comparing a changed site against itself over time cannot separate the effect of the change from the effect of the platform changing underneath it.
Where this came from
Every factual claim above traces to one of these. Each entry says what it supports and the date it was read, because platform documentation changes without notice.
-
Model deprecations — OpenAI
That models are retired and replaced on a published schedule, which is why a product name is not a stable measurement target and why model version is recorded as unknown rather than assumed constant when a platform does not expose it.
Primary source · read 2026-09-03