Do online reviews affect which businesses AI assistants recommend?
The answer everyone repeats has a traceable evidence chain, and the chain does not reach the claim.
Do online reviews affect which businesses AI assistants recommend?
Nobody has demonstrated that reviews cause an AI assistant to recommend a business. The claim traces back to a study that measured Google Local Pack visibility, not AI recommendations, which was then restated as an AI finding. The strongest study specific to AI was commissioned by a review platform, is correlational by its own statement, and had its headline number reported as 75 percent when the study's own tier comparison was 53.5 percent. The one piece of research that inspected ChatGPT's internal response data found no sorting by rating at all. Reviews plainly matter for customers and for classic local search. That is a separate question with a much better answer.
Follow the chain backwards
Start with what a local business owner is told: that reviews are a ranking signal for AI assistants, sometimes with a specific threshold attached, such as a minimum review count and star average. Then ask where any of that came from.
The upstream study is Yext's analysis of 8.7 million Google search results across 2,500 US ZIP codes and six verticals, with more than 200 structured data points per location. It found that active engagement with reviews, in both volume and responsiveness, was the most consistent driver of visibility.
The dependent variable in that work is Google Local Pack visibility. It is not an AI recommendation. Somewhere downstream the outcome variable was swapped and nobody flagged the substitution.
That substitution is the load-bearing move in the whole genre. A well-run study of one system was relabelled as evidence about a different system, and everything built on top inherits the error.
The strongest AI-specific study, and who paid for it
Seer Interactive analysed 804,491 AI responses covering 1,926 brands, 15,783 prompts, four platforms and eight verticals in March 2026. Brands were tiered from no review profile through to an optimized one, and matched on domain authority. Median citation rate rose from 1 percent at the lowest tier to 53.5 percent at the first tier with a profile, with only about six further points of lift across the two higher tiers.
That shape is interesting and it is a threshold effect, not a dose-response curve. Having a profile at all is associated with most of the difference. Adding more on top of that is associated with very little.
Two facts about this study are almost never carried with it. It was commissioned by Trustpilot, and it measures the effect of having a Trustpilot profile. And the press release led with a rise from 1 percent to 75 percent, using a subgroup figure for brands with active profiles and 80 or more reviews, rather than the 53.5 percent tier comparison the study itself reports. The 75 is the number that propagated.
Seer states in the study that this is correlation rather than causation and names marketing spend and audience size as uncontrolled confounders. Those caveats are in the original and are absent from nearly every restatement.
None of this makes the work worthless. It makes it a threshold-shaped correlational finding, produced for an interested party, about national brands rather than local businesses, with a headline chosen for a press release.
The evidence pointing the other way
When Turrado inspected ChatGPT's internal business data structure rather than the rendered text, he found no sorting by rating, no alphabetical sorting, and no sorting by opening hours. The ordering looked like provider response order with a hard cut above twenty venues.
That is small, dated and Spanish, and it is also the only work anyone has published that examined the mechanism instead of the output. A single mechanistic observation outweighs a stack of confident assertions that examined nothing.
Sterling Sky tracked AI local packs across 322 markets and reported counts: they appeared on roughly 7 percent of tracked keywords, contained 5,943 unique businesses against 18,330 in traditional three-packs, and typically showed one or two businesses instead of three. Crucially, Sterling Sky explicitly declines to say why, and reports no testing of reviews, profile optimization or schema. That refusal is the most credible thing published in this area.
Google, for its part, mentions keeping business information current only inside general search hygiene, and states that no additional optimizations are necessary for its AI features.
| Source | Measured | Supports the claim that reviews drive AI recommendations? |
|---|---|---|
| Yext, 8.7M results, 2,500 ZIP codes | Google Local Pack visibility | No. Different outcome variable entirely |
| Seer Interactive, 804,491 responses | Association between having a review profile and citation rate | Correlational, commissioned by a review platform, national brands |
| Turrado, internal response data | Ordering inside ChatGPT local results | Points against. No rating-based sorting found |
| Sterling Sky, 322 markets | How often AI local packs appear and how many businesses they hold | Makes no causal claim and explicitly tests none |
| Vendor threshold figures, e.g. 30 reviews at 4.3 stars | Nothing | No study, no method, no source located |
The numbers that are simply invented
These circulate widely and none of them has a locatable primary source. A working baseline of thirty reviews at 4.3 stars, rising to a hundred in competitive markets. A pace of two to four new reviews per week outperforming a burst. A weighted table giving review quality and sentiment 20 percent of ChatGPT's decision and location 15 percent. Conversion rates up to 300 percent higher for AI-referred customers. A complete profile increasing AI recommendations by 3.4 times.
The weighted table deserves particular attention. It assigns percentage weights to ranking factors for a system whose owner has never published any ranking factors. There is no dataset it could have come from.
A close relative of this pattern is real but misread. One vendor page reports that ChatGPT referenced reviews in 58 percent of responses and Perplexity in 100 percent, and concludes that reviews are therefore a direct ranking signal. Referencing a review page inside an answer is not evidence that reviews determined which business was chosen. Those are different events and only one of them was observed.
The largest expert-opinion dataset in local search is a survey of 47 practitioners scoring 187 factors, and its publisher states plainly that none of those experts has special access to the algorithm. In 2026 they were also asked to score AI search impact, and the author flagged that question as ambiguous himself. Downstream this becomes hard weights presented as measurement.
What to actually do about reviews
Collect them and respond to them. The justification is that customers read them, that they are associated with Local Pack visibility in a large well-run study, and that a business with recent reviews is easier for anyone to verify. Those reasons are sufficient and they are honest.
What should not be promised is an AI outcome. If a vendor invoices for review generation on the grounds that it will get you named by ChatGPT, ask which study, on which platform, with what sample. The answers do not exist yet.
Consumer behaviour data also argues against treating an AI mention as the finish line. BrightLocal found that 43 percent of people who started a search with an AI tool went on to check Google anyway, and only 18 percent felt ready to contact a business without verifying it first. Whatever the assistant says, the reviews are getting read afterwards by a human.
What is actually established, and how
Sorted by how strong the evidence is, not by how convenient it is.
| Claim | Basis |
|---|---|
| Yext measured review engagement as the most consistent driver of Google Local Pack visibility, across 8.7 million results in 2,500 ZIP codes. | We measured this |
| Seer Interactive found median citation rate rising from 1 percent without a review profile to 53.5 percent with one, across 804,491 responses. | We measured this |
| That study was commissioned by Trustpilot and states it is correlation rather than causation, with uncontrolled confounders named. | Documented by the platform |
| Inspection of ChatGPT internal local response data found no rating-based, alphabetical or hours-based sorting. | We measured this |
| AI local packs appeared on roughly 7 percent of tracked keywords and held about a third as many unique businesses as traditional three-packs. | We measured this |
| 43 percent of consumers who began with an AI tool went on to check Google, and 18 percent felt ready to make contact without verifying. | We measured this |
| Whether a review profile causes an AI assistant to recommend a local business. | Not publicly documented |
What nobody can currently tell you
Stated because the alternative is implying a certainty that does not exist.
No controlled experiment has ever tested whether a local business's review profile changes its odds of being recommended by any AI assistant. Not one, on any platform, by anyone.
The closest study measured national brands with profiles on the platform that paid for it, so its transfer to a single-location trade business is untested.
Whether a threshold exists, and where it would sit, is unestablished. The one threshold-shaped result is correlational and specific to one review platform.
Whether review recency or response rate matters to any AI system is asserted constantly and has never been measured.
Whether an AI recommendation produces a customer at all. Nobody has connected one to a transaction.
What people get wrong about this
-
Reviews are a confirmed ranking factor for AI assistants.
What is actually the case
No platform publishes ranking factors for AI answers, and the study most often cited for this measured Google Local Pack visibility instead. The substitution of one outcome variable for another happened downstream and was never flagged.
-
Brands with reviews get cited 75 percent of the time versus 1 percent without.
What is actually the case
The study's own tier comparison is 1 percent to 53.5 percent. The 75 figure comes from a subgroup with 80 or more reviews and appeared in a press release from the company that commissioned the research.
-
You need roughly 30 reviews at 4.3 stars to be considered.
What is actually the case
No study produced that threshold. It has no locatable source, and it appears alongside other invented figures such as weighted percentage tables for systems that publish no weights.
-
Perplexity cites review sites in every answer, so reviews decide the recommendation.
What is actually the case
Citing a review page and selecting a business are separate events. Only the first one was observed, and one study of citation behaviour suggests sources are often gathered to support a choice already made.
How to check this on your own site
You should not have to take our word for any of it.
- Pick two competitors in your category with clearly different review profiles and ask an assistant for recommendations twenty times, logged out. Record how often each appears. Wide run-to-run variation will show up before any review effect does, which is itself the finding.
- When a business is named, check whether its review pages were cited or whether the citation went somewhere else entirely. The gap between the two tells you which event you are actually observing.
- Before acting on any review threshold you read, search for its primary source. If the trail ends at another blog post, treat the number as fiction.
Questions we get asked constantly
So should I stop asking customers for reviews?
No. Reviews have a solid evidence base for Google Local Pack visibility and an obvious one for human decision-making. The narrow point is that no evidence supports paying specifically for an AI recommendation outcome.
Does replying to reviews help with AI?
Unmeasured. Review responsiveness was part of what Yext associated with Local Pack visibility, and nobody has tested it against AI recommendation outcomes on any platform.
Why is this stated so confidently everywhere else?
Because it is a monetizable recommendation and the underlying study sounds close enough. A large, well-run study of the wrong outcome variable is more persuasive than an honest statement that the question is open.
Where this came from
Every factual claim above traces to one of these. Each entry says what it supports and the date it was read, because platform documentation changes without notice.
-
Best practices will only take you so far — Yext
Review engagement as the most consistent driver of Google Local Pack visibility across 8.7 million results and 2,500 ZIP codes. The outcome variable is the Local Pack, not AI answers.
Primary source · read 2026-09-03
www.yext.com/research/articles/best-practices-will-only-take-you-so-far/
-
Seer Interactive study of reviews and AI brand presence — Seer Interactive
Tiered citation rates from 1 percent to 53.5 percent across 804,491 responses, the threshold shape, the Trustpilot commission, and the stated correlational limits.
Primary source · read 2026-09-03
-
Brands that build trust through reviews increase AI citations from 1% to 75% — Trustpilot press release via PR Newswire
The headline 75 percent figure, its subgroup basis, and the simultaneous announcement of products built on the finding.
-
SEO local para ChatGPT — Natzir Turrado
Direct inspection of ChatGPT internal local response data finding no rating-based, alphabetical or hours-based sorting.
Primary source · read 2026-09-03
natzir.com/posicionamiento-buscadores/seo-local-para-chatgpt/
-
The state of local SEO in 2026 — Sterling Sky, Joy Hawkins
AI local pack prevalence at roughly 7 percent of tracked keywords, 5,943 unique businesses against 18,330 in three-packs, and the explicit refusal to attribute causes.
Primary source · read 2026-09-03
-
Consumer search behavior across channels — BrightLocal
43 percent of AI-starters went on to check Google, and only 18 percent felt ready to contact an AI-recommended business without verifying.
Primary source · read 2026-09-03
www.brightlocal.com/research/consumer-search-behavior-channels/
-
Local search ranking factors — Whitespark
The survey basis: 47 practitioners scoring 187 factors, with the publisher stating none has special access to the algorithm.
Primary source · read 2026-09-03
-
GEO myths and lies — Search Engine Land, Philipp Goetza
Independent critique of the evidentiary quality of AI search claims, including the absence of solid data behind commonly sold local tactics.
Secondary source · read 2026-09-03