Recommendation widgets often look productive because popular products attract clicks. Yet a system can increase short-term engagement while repeatedly exposing the same narrow set of items, hiding new or long-tail products, recommending unavailable variants, and concentrating demand where stock is already constrained.
What we see in ecommerce analysis is this: relevance alone does not describe recommendation quality. Merchants also need coverage, diversity, repetition, availability, margin, and customer-level outcomes. The objective is not random variety. It is useful discovery without allowing popularity bias to become the entire merchandising strategy.

Table of Contents
- Keyword decision and intent
- Instrument recommendation exposure
- Measure diversity with relevance
- Diagnose popularity and repetition
- Test business guardrails
- EcomToolkit point of view
Keyword decision and intent
- Primary keyword: ecommerce recommendation diversity analytics
- Secondary keywords: recommendation catalog coverage, popularity bias, product exposure concentration, recommendation repetition rate
- Search intent: evaluate whether recommendation systems create useful, commercially healthy discovery
- Funnel stage: mid funnel
- Page type: merchandising analytics guide
Shopify’s Search & Discovery documentation distinguishes complementary and related product recommendations and lets merchants customize recommendations (Shopify Search & Discovery). That distinction matters analytically: substitutes, complements, recently viewed items, and personalized rankings solve different customer jobs and should not share one benchmark.
Instrument recommendation exposure
Capture request ID, session and consent-safe customer key, timestamp, page context, source product, widget type, algorithm and rule version, candidate set, ranked items, displayed positions, viewport visibility, click, add-to-cart, purchase, quantity, net revenue, margin, return, and inventory state at exposure time.
Log candidates as well as displayed items where practical. Without the candidate set, analysts cannot tell whether low diversity came from the model, merchandising rules, inventory filters, or the final interface.
| Statistic | Calculation | Decision supported |
|---|---|---|
| catalog coverage | unique recommended products / eligible catalog | exposure breadth |
| exposure concentration | share of impressions held by top 1% or 10% of items | popularity dominance |
| intra-list diversity | dissimilarity among items in one widget | shopper choice breadth |
| session repetition rate | repeated item exposures / recommendation exposures | fatigue |
| availability rate | in-stock displayed items / displayed items | recommendation validity |
| recommendation CTR | recommendation clicks / viewable impressions | immediate relevance |
| assisted margin per 1,000 views | attributed contribution / viewable impressions × 1,000 | commercial quality |
Use viewable impressions, not render events, when the widget may sit below the fold. Keep last-click revenue separate from assisted outcomes.
Measure diversity with relevance
Segment by widget purpose, page type, device, customer state, category, price band, brand, inventory depth, traffic source, season, and model version. A complementary-products widget should be evaluated on attach behavior; an alternatives widget should be evaluated on successful product continuation and substitution.
Diversity can be measured across three levels: within one list, across a customer’s session, and across the total catalog. A widget can look varied within each list yet repeatedly show the same pool to everyone. Combine all three views.
| Pattern | Likely interpretation | Response |
|---|---|---|
| CTR rises, coverage falls | popular items dominate | test relevance-preserving diversity |
| coverage rises, conversion falls | variety became weak relevance | refine candidates and context |
| strong clicks, poor margin | discount or low-margin bias | add commercial guardrails |
| repeated out-of-stock exposure | stale availability filter | shorten inventory freshness SLA |
| new items receive no exposure | cold-start failure | create controlled exploration pool |
Diagnose popularity and repetition
Plot exposure share against product sales rank and inventory. Compare recommendation exposure with organic catalog demand. If the top products receive far more recommendation share than their natural demand share, the widget may be amplifying popularity rather than helping discovery.
An anonymous merchant can see higher recommendation revenue after simplifying every widget to bestsellers. The result may still be negative if customers would have found those products anyway while high-margin compatible items lose exposure. Holdout groups and incremental margin are more useful than attributed widget revenue alone.

Test business guardrails
Run experiments with explicit guardrails: relevance, viewable CTR, add-to-cart, conversion, contribution margin, return rate, availability, exposure concentration, new-product coverage, and page latency. A more complex ranking model that delays the page can erase merchandising gains.
Define exclusion and boost rules with owners and expiry dates. Manual rules often accumulate until the model is no longer making the decision. Track the percentage of exposures changed by rules and the incremental result of each rule family.
Pair this guide with search query analytics and assortment productivity analytics. Merchandising should own discovery goals, data science should own ranking quality, engineering should own latency and logging, and finance should validate contribution margin.
Create a recommendation scorecard by job
Do not publish a single blended dashboard. Maintain separate scorecards for alternatives, complements, replenishment, recently viewed, personalized discovery, and editorial modules. Each should state its customer job, eligible catalog, success event, attribution window, and guardrails. Blending them rewards whichever widget receives the most traffic, not the one that best performs its role.
Review exposure concentration weekly. List products gaining or losing the most share, items repeatedly recommended without engagement, high-demand items exposed beyond available depth, and new products receiving no meaningful test traffic. Merchandisers need examples alongside distribution statistics because a healthy aggregate can hide nonsensical pairs.
For cold-start items, reserve a controlled exploration budget rather than inserting them everywhere. Define eligibility using product completeness, imagery, availability, price, compliance, and merchandising approval. Measure whether exploration produces incremental discovery without damaging relevance or page speed.
Also test the interface. A model can return a diverse list that becomes repetitive after responsive truncation, deduplication across modules, or client-side filtering. Log the final viewable order on the shopper’s device. Server candidates alone cannot describe the experience customers actually received.
EcomToolkit point of view
A recommendation engine should help shoppers discover the right next product, not repeatedly congratulate the catalog’s winners. Measure relevance and incrementality, but keep coverage, availability, margin, and repetition visible so the system serves the whole commercial strategy.