Insights · Genrise Study

Ratings, Reviews, and Rufus: What 610,000 AI Recommendations Reveal About Visibility.

A Genrise study of 610,096 Amazon Rufus recommendations across eight product categories and three marketplaces, examining what customer ratings and reviews actually do when an AI shopping assistant decides which products to surface — and, just as importantly, where they stop working.

By Devansh Verma12 min read
A Genrise study of 610,096 Amazon Rufus recommendations across eight product categories and three marketplaces, examining what customer ratings and reviews actually do when an AI shopping assistant decides which products to surface — and, just as importantly, where they stop working.

Every brand operating on the digital shelf has internalized the same instinct: win the reviews, win the shelf. Accumulate rating, accumulate review volume, and visibility follows. For two decades of keyword-driven search, that instinct was broadly correct — review signals fed the ranking algorithm, and more was better.

Amazon Rufus does not behave that way. We know because we watched it make 610,096 recommendations over sixty days, captured 99.6% of the product universe it was choosing from, and then read the pages of the products it passed over — the category best-sellers, heavy with reviews, that Rufus never once surfaced.

The finding that frames everything else: the best-reviewed products are frequently not the ones Rufus recommends. Ratings and reviews still matter — but not in the way most brands assume, and not for the job most brands assign them. This is what the data shows they actually do.

The paradox in one number

Start with the comparison that should not be possible if reviews were the driver.

1,114
Median reviews on best-sellers Rufus never showed
Against 648 across all recommended products.
23.6%
Of never-shown best-sellers answered the shopper's question
Versus 88.8% for the most-recommended products.
4.5 ★
Median rating in both groups
Same rating, more reviews in the invisible group.

Across the categories we studied, the best-selling products that Rufus never recommended carried a median of 1,114 reviews. The products Rufus did recommend carried a median of 648 — across every tier of recommendation, from the ones it showed constantly to the ones it showed once. The never-shown best-sellers had more social proof than the products that got surfaced, at the same 4.5-star median rating. And yet one group appeared in Rufus recommendations and the other never did.

Held side by side, the two groups diverge on one axis and one axis only. The never-shown best-sellers answered the shopper's underlying question 23.6% of the time. The products Rufus recommended most often answered it 88.8% of the time. Same rating, more reviews in the invisible group — and a nearly four-fold gap in whether the page actually addressed what the shopper was asking.

If review mass determined visibility, the most-reviewed products would win. They don't. Which raises the real question: if ratings and reviews aren't the deciding factor, what job are they doing? The data gives a precise answer, and it is a different job for each.

Rating is a threshold, not a lever

The first thing to disentangle is rating from review volume, because brands routinely treat them as one signal. When separated in the data, rating turns out to do very little work.

Products Rufus kept recommending and products it dropped had the same median rating — 4.5 stars against 4.5 stars. Rating did not distinguish the winners from the losers. In a model that isolates rating from review volume across the full product universe, the effect of rating on how often a product appears is a precise statistical null — it neither helps nor hurts once volume is accounted for.

That does not mean rating is irrelevant. It means rating behaves as a gate rather than a dial. Products below 4.0 stars are penalized: they make up 8.7% of recommended products but only 2.6% of actual appearances. The floor is real. But once a product clears roughly 4.0 stars, additional rating stops buying visibility.

The clearest evidence is what happens at the very top. Appearance likelihood rises with rating up to a peak in the 4.7-to-4.8-star band — where products appear 1.89 times more than their share of the catalog would predict — and then collapses to 0.22 times at 4.9 to 5.0 stars. A near-perfect rating is, counterintuitively, a disadvantage. The reason is structural: products sitting at 4.9 or 5.0 stars are overwhelmingly low-review products whose ratings simply haven't been tested at volume, and those are exactly the products Rufus is most cautious about surfacing.

The rising curve up to that peak is not a contradiction of the threshold finding — it is the reason the threshold finding needs a model to see. That raw curve is confounded with review volume: higher-rated products in the mid-4-star range also tend to carry more reviews, and the appearance lift is riding on the review volume, not the rating. When the two signals are separated statistically — holding review volume constant — the independent effect of rating flattens to the null described above. The lift curve shows what the two signals do together; the model shows that, on its own, rating past the floor does almost nothing.

The practical read is uncomfortable for anyone running a rating-improvement program as a visibility play. Clearing 4.0 stars matters. Moving from 4.6 to 4.8 does not appear to move visibility. And chasing a perfect score can signal the very thinness that keeps a product out of the recommendation set. Rating is table stakes, not leverage.

What review volume actually buys: staying power

Review volume is a different story — but not the story brands expect. It does real, measurable work, and that work is durability.

The products still being recommended at the end of our window held a median of 1,457 reviews, against 443 for the products that had dropped out. Review mass predicts whether a product persists in the recommendation set roughly twice as strongly as rating does. So far, this looks like confirmation of the conventional wisdom.

The nuance surfaces when you measure how long products survive. We tracked more than 104,000 individual stretches of time that products spent inside Rufus carousels and split them by review volume. The pattern is stark:

  • Products in the bottom third by review volume lasted a median of 4 days in the carousel.
  • Products in the middle third lasted 8 days.
  • Products in the top third lasted 19 days.
1.89×
Appearance lift at 4.7–4.8 stars
Collapsing to 0.22 times at 4.9 to 5.0 stars.
4 / 8 / 19
Median days in carousel by review-volume third
Bottom / middle / top — a 4.8-fold shelf-life gap.
1.6×
Review mass carried by the slot-one product
Past the first slot, the curve is nearly flat.

A 4.8-fold gap in shelf-life between the most-reviewed and least-reviewed products. Review mass buys endurance — once a product is being shown, a deep review base is what keeps it there run after run, while thinly-reviewed products flicker in and then fall out within days.

But endurance is not the same as climbing. When we looked at where products sat within a carousel — the order Rufus presents them in — the top slot was systematically review-heavy: the product in slot one carried roughly 1.6 times the review mass of products further down. Past that first slot, though, the curve is nearly flat, and a product's own review count did not predict the position it occupied. Review mass helps a product hold the top slot once it has it; it does not ladder a product up the carousel. What review volume governs is whether you stay in the set — not how high you sit inside it.

This is the pivot the whole picture turns on. Reviews are a retention layer. They keep a product on the shelf once it has earned a place there. What they conspicuously do not do — as the paradox in the opening section showed — is earn that place to begin with.

Where reviews stop working

Return now to the never-shown best-sellers, because the two sections above resolve the paradox.

Those products had cleared the rating gate — 4.5 stars, comfortably above the floor. They had the review volume that, in any other product, would have bought weeks of shelf-life. They had, in other words, everything the conventional model says should produce visibility. And Rufus never showed them.

The reason is that ratings and reviews are both necessary and neither is sufficient. Rating earns eligibility. Review volume earns durability. But there is a prior question that both leave unanswered: does the product page actually address what the shopper is asking? For the never-shown best-sellers, the answer was usually no — they resolved the shopper's question 23.6% of the time. Their social proof was intact and their answer was absent, and the absence is what kept them off the shelf.

This is the ceiling on what reviews can do. A product cannot review its way into a recommendation set if its page does not answer the question that triggers the recommendation. The heaviest review base in the category does not compensate for a page that doesn't speak to shopper intent. Social proof is the floor a product stands on. It is not the thing that gets the product chosen.

Social proof is the floor a product stands on. It is not the thing that gets the product chosen.

The quiet job of review content

There is one more layer to the review story, and it is the one most brands overlook entirely — because it treats reviews not as a score but as content.

On an Amazon product page, the customer reviews section is not just a star count. It is a block of text, written in the shopper's own language, that sits on the page alongside the title, bullets, and description. And Rufus reads it. In our data, on the pages that answered the shopper's question at all, the reviews section was part of that answer 55% of the time (across every page judged, it carried the answer 42% of the time) — less often than the title or bullets, but frequently, and with a distinctive property: it carries the shopper's own phrasing more faithfully than most brand-authored copy does.

The consequence shows up in the numbers. When the reviews section participated in answering a question, the product earned more visibility than it did on the strength of title and bullets alone. Reviews are quietly pulling weight that brands rarely credit them for — and the mechanism is not the star rating attached to them, but the substance of what they say.

Two honest caveats belong here. First, this is a floor, not a ceiling: Rufus can only read the handful of reviews Amazon renders directly on the page — a small fraction of a popular product's full review corpus — so the true contribution of review content is almost certainly larger than we could measure. Second, this is the layer brands control least. You cannot author your reviews. But you can influence what shoppers have to write about — by making sure the questions they care about are answerable in the first place, so their reviews reinforce the intent your page is already built around rather than filling a vacuum the page left open.

Review content, in other words, is most valuable when it echoes a page that was already doing its job. It is a multiplier on good content, not a substitute for it.

The pattern holds across categories and markets

A finding drawn from one category or one marketplace is a curiosity. This one repeats everywhere we looked.

Across eight product categories — spanning oral and digestive health, skincare and beauty, pet care, confectionery, snacking and cereal, and packaged meal solutions — the same structure held: rating as a threshold, review volume as a persistence signal, and content answering as the deciding factor for entry. The categories differed in how often their pages answered shopper questions, from a high of 87% in oral and digestive health down to 52% in packaged meal solutions — a 35-point spread within a single assistant. But the shape of the mechanism did not change from one category to the next.

The marketplace comparison is more striking still, because review norms differ so enormously across regions — a product considered heavily reviewed in one market would look thin in another. If raw review volume drove visibility, markets with very different review baselines would behave like different systems. They don't. The gate-persistence-content structure repeats across all three marketplaces we observed — the review thresholds shift with local norms, but the roles that rating, review volume, and content play stay constant. (A caution on reading too much into the non-US figures: outside the US, our coverage in this study rests on a single category portfolio per market, so we treat the cross-market comparison as evidence that the structure travels, not as a reliable benchmark of what review volume any given market requires.)

That consistency is what tells us this is not a category quirk or a US artifact. It is how the assistant behaves.

What this means for your shelf

The temptation reading a study like this is to convert it into a checklist. We will resist that, because the honest conclusion is structural, not prescriptive.

Ratings and reviews are the floor, not the ceiling. A product needs to clear the rating gate to be eligible and needs review depth to stay durable once it is being shown. Both are real, both are necessary, and both are things most established brands already have. What they do not do — what no amount of additional review volume will do — is win the recommendation in the first place. That is decided by whether the page answers the question the shopper actually asked, a layer that sits upstream of reviews and that the products with the heaviest review bases in our study most often failed.

For a brand with strong review equity, this is either a warning or an opening depending on where its content stands. The warning: your review advantage is not protecting you in AI-assisted shopping the way it protected you in keyword search, and best-sellers are being passed over for exactly this reason. The opening: the layer that decides visibility is the one you have the most direct control over, and it is where most of your category is underinvested.

This kind of analysis — reading not just what an assistant recommends but what it passes over, across categories and markets, at the scale where the pattern becomes undeniable — is the work Genrise does continuously for enterprise consumer brands. The always-on system monitors how AI shopping surfaces are behaving against a brand's catalog, identifies where pages are and are not answering the questions shoppers are asking, and keeps the digital shelf aligned as the surfaces shift. Rufus visibility is one measurable outcome of that work.

What this study deliberately does not do is tell any single brand what to change. That conversation is specific to a portfolio, a category, a retailer footprint, and a content starting point. The data here is the evidentiary backbone. What to do about it is the next conversation.

Request a Demo

See how Genrise scores your catalog against the questions shoppers are actually asking AI assistants.

How we know this

610,096
Rufus recommendation events observed
Over a sixty-day window.
22,824
Distinct products
99.6% of the universe Rufus was choosing from captured.
2,404
Never-shown best-sellers evaluated
Across 235 categories, plus 832 shopper questions tracked.

This study is not a sampled audit. Over a sixty-day window, we observed 610,096 Rufus recommendation events across 22,824 distinct products, capturing the product pages behind 99.6% of the universe Rufus was choosing from. We tracked 832 shopper questions, each run a median of 26 times, across three marketplaces — the US, India, and the UK — and eight anonymized product categories.

Two design choices give the findings their weight.

First, we read the products Rufus did not recommend. Almost every study of AI shopping visibility examines only the winners — the products that got surfaced. That approach cannot tell you what separates a winner from a loser, because it never observes the losers. We identified, for each recommended product, the leaf category it competes in, pulled Amazon's public Best Sellers list for that category, and captured and evaluated 2,404 best-selling products that Rufus never once recommended — across 235 categories. The never-shown best-sellers are the control group that makes the central finding possible.

Second, we separated the signals statistically rather than eyeballing them. Because rating, review volume, and content quality tend to move together, a raw correlation cannot tell you which one is doing the work. The rating and review conclusions above come from models that isolate each signal while holding the others constant, and from a survival analysis of more than 104,000 individual carousel appearances. That is how we can say rating is a threshold once volume is controlled, and that volume governs duration rather than position — claims a simpler analysis would blur together.

A note on what this study is and is not. It observes a single AI assistant over a defined window; the system is updated continuously, and Rufus in six months may weight signals differently than the Rufus we observed. The reviews layer, as noted, is a lower bound, limited to the reviews Amazon renders on-page. And these are the mechanics of how ratings and reviews relate to visibility — the details of how Genrise evaluates whether a page answers a shopper's question, and how that feeds the AI Shelf Readiness Index, sit inside our platform and our client work. The purpose of this piece is to share what the data reveals about the shelf, not the instrument that reads it.

Frequently asked questions

See it in action

See what your pages answer —
and what they leave open.

Get a tailored walkthrough of Genrise — how it scores every PDP against the questions shoppers are actually asking, where the gaps live, and how an always-on system closes them at full catalog scale.