Skip to content
Skip article header Engineering

Popularity Bias in Recommender Systems

Popularity bias makes a recommender show already popular items more often than users' tastes justify, and retraining on the resulting clicks makes it worse. This guide computes popularity lift, exposure Gini, coverage, novelty and miscalibration from serving logs, ranks the fixes by engineering cost and covers what the DSA asks of recommender systems.

Updated 18 min read 21 views

Technically reviewed by Victor Sineglazov, D.Sc.

Three smartphones side by side on a desk, each showing a shopping feed whose top row of product tiles is identical while the rows below differ.
Skip key takeaways

In short: popularity bias in recommender systems is the tendency of a model to show already popular items more often than users' own tastes justify, so a small head of the catalog absorbs most impressions. It starts with skewed interaction data and grows through a feedback loop when the system retrains on clicks its own exposure produced. The fix is to measure exposure from serving logs, judge each reading against the model's own baseline after every retrain and correct with serving-time re-ranking first, reserved exploration slots second and training-time corrections such as propensity weighting last.

This guide assumes a recommender already in production. Model selection and cold-start basics are covered in our recommender system development guide.

What Is Popularity Bias in Recommender Systems?

Data-side bias

Two different things go by this name. One is a property of the data. Item popularity is usually counted as interactions (ratings, clicks or purchases), and those counts are skewed for honest reasons: some products are better or cheaper, and some are heavily advertised. Klimashevskaia, Jannach, Elahi and Trattner put it this way in their survey: "We refer to such pre-existing, commonly skewed distributions regarding the popularity of items as the natural bias in the data."

Exposure-side bias

The other is a property of what the system shows. The same survey notes that "Most commonly, popularity bias is considered a characteristic of the recommendations that are shown (exposed) to users." Its worry is that "a serious problem of recommender systems is that they might reinforce these pre-existing distributions" rather than merely reflect them. Chen et al., following Abdollahpouri and Mansoury, give the over-proportional version: "Popular items are recommended even more frequently than their popularity would warrant". This exposure-side definition is the one engineers can act on, because it compares two measurable distributions: what users interacted with and what the system served.

Where the long tail starts

Long-tail items are the many items with few interactions each, and where the tail starts is a convention. The popularity bias survey records that one group of works, Abdollahpouri and colleagues among them, treats the top 20% of items by interactions as popular, while others put the head at the top 10% or even 1% of items. It warns that "the threshold popularity value used to divide the items into head and tail is an important factor and can substantially impact the outcome of the experiments". It also found "no agreed-upon definition of what represents popularity bias has emerged so far", so write your own definitions down before you measure anything. Pick one head cut-off, record it next to every dashboard and keep it fixed between releases. Tail share, long-tail coverage and any grouping of users by how mainstream their histories are all depend on it.

How Does the Feedback Loop Amplify Popularity?

Our recommender system development page makes the short version of this point: relying on click-through rate alone creates a popularity bias feedback loop. A model trained on skewed interactions scores head items higher, those items get more impressions and clicks, and the next training run reads those clicks as stronger preference for the head.

Chaney, Stewart and Engelhardt name the problem in their RecSys 2018 paper: "These systems are often evaluated or trained with data from users already exposed to algorithmic recommendations; this creates a pernicious feedback loop." Schnabel et al. state the statistical side: "Most data for evaluating and training recommender systems is subject to selection biases, either through self-selection by the users or through the actions of the recommendation system itself."

Missing interactions are therefore not negative feedback. As Chen et al. explain, "Exposure bias happens as users are only exposed to a part of specific items so that unobserved interactions do not always represent negative preference." A tail item with zero impressions has told you nothing, yet most training pipelines treat it like an item users rejected.

Damage is also uneven. In an offline simulation, Mansoury et al. found that the loop amplified popularity bias and reduced aggregate diversity, and wrote that "we show that the impact of feedback loop is generally stronger for the users who belong to the minority group". In that paper the minority group is the female users of the MovieLens 1M dataset, a demographic split rather than a taste-based one.

Measuring Popularity Bias From Serving Logs

A data analyst holding a laptop in a product sample room while a merchandiser pulls a box from behind the front rows of a shelf.

In production the better source for these metrics is the impression log: one row per served slot with the request, the user or session, the position, the item, the model version and whether the slot came from the model, a fallback or an exploration policy. Join it to item popularity counted over the training window the model saw, and to a catalog snapshot that includes items with zero impressions. Without those zero rows, coverage and Gini look better than they are.

No cited source gives an absolute threshold for any of these metrics, so the table defines a bad reading relative to the previous model version, the popularity of the users' own profiles and the trend across retrains.

Six exposure metrics

Metric What it measures How to compute it from serving logs What a bad reading looks like Typical failure mode
Average recommendation popularity (ARP) Mean popularity of served items Average the training-window interaction count of the items in each list, then average across users Above the average popularity of the same users' histories, or rising across retrains while profiles stay flat Head items fill most slots regardless of history
Popularity lift Relative gap between item popularity in lists and in profiles, per user group Group users by how mainstream their history is, then compare mean item popularity in lists with that in profiles as (lists minus profiles) over profiles Positive and growing, largest for the niche-taste group Niche users receive mainstream lists
Exposure Gini Inequality of impressions across the catalog Count impressions per item over a fixed window, including zero-impression items, then compute the Gini index Rising release over release, or a widening gap over the Gini of training interactions computed on the same item set A few items absorb most impressions
Catalog and long-tail coverage Share of the catalog served at least once, and tail share within lists Distinct served items over active catalog items. For the tail, the average share of slots below a fixed head cut-off Coverage falling while the catalog grows, or a tail share in lists well below the tail share in profiles Tail items never get the impressions they need to earn interactions
Novelty Distance of served items from the popular head Mean of minus log of each served item's share of training interactions Dropping against the previous version, or below the novelty of users' own histories Lists repeat what users already know
Miscalibration Gap between the category mix of a user's history and of their list Per user, category shares of history and of served slots, compared by KL divergence or Hellinger distance, then averaged per group Rising across retrains, concentrated in users with mixed or niche histories A multi-category user gets lists from the most popular category only

Where the definitions come from

Abdollahpouri, Burke and Mobasher's FLAIRS 2019 paper uses the long-tail share from their own earlier work, where "this metric measures the average percentage of long tail items in the recommended lists", alongside average recommendation popularity. Popularity lift comes from Abdollahpouri, Mansoury, Burke and Mobasher, a measure "which quantifies the difference between average item popularity in input (user profile) and output (recommendation list) for an algorithm". A positive value means the algorithm amplified popularity, and a negative one means the lists are less concentrated on popular items than the profiles.

For exposure inequality, the Klimashevskaia survey gives the reading rule: "The Gini index expresses the inequality of the distribution, with values closer to 1 indicating a high inequality (range: 0-1)." Compute it over logged impressions, not training interactions, or you are measuring demand rather than exposure.

Miscalibration has two published distances that should not be merged. Steck's calibrated recommendations work (RecSys 2018) uses KL divergence. As restated by Corrêa da Silva and Jannach, "Steck relied on the Kullback-Leibler (KL) divergence, as this measure possesses three useful properties for calibration". KL is undefined when a list has zero share in a category the history contains, so implementations add a small smoothing term. The popularity lift paper uses the Hellinger distance instead, and its abstract links the two problems: "the more a group is affected by the algorithmic popularity bias, the more their recommendations are miscalibrated". Pick one distance, record which and never compare readings across them.

Two open-source libraries implement most of these metrics. RecBole's metric reference states that "GiniIndex presents the diversity of the recommendation items. It is used to measure the inequality of a distribution." Its TailPercentage metric treats as long tail the least-interacted share of items set by tail_ratio, which defaults to 0.1, so only the bottom 10% of items count as tail (a value above 1 instead means an interaction-count threshold). Abdollahpouri and colleagues call the top 20% the head, which makes their tail the remaining 80%, so matching them means a tail_ratio of 0.8, not 0.2. The Microsoft Recommenders evaluation docs say "Novelty is computed as the minus logarithm of (number of interactions with item / total number of interactions)." RecBole's default does not match the 20% head convention, so set the cut-off as described under where the long tail starts instead of inheriting one.

Why Can Offline and Online Readings Diverge?

An offline improvement in popularity metrics is a hypothesis about live behavior, and the evidence base is thin. Among the mitigation studies the popularity bias survey reviewed, only four reported user studies, all from 2015 or earlier, and "a single work was found which examined popularity bias effects in a field test".

Readings can diverge for structural reasons. Offline splits are products of the old policy's exposure, so a model that ranks tail items higher may be scored down for recommending items nobody had the chance to click. Online, users react to the new lists and shift the very popularity counts the metrics depend on. Field evidence on how recommenders shift popularity in general is also scarce. As summarized in the same survey, a field study by Lee and Hosanagar found that introducing a recommender lowered sales diversity while absolute sales of long-tail items still rose, so the direction you see depends on what you measure.

Business value is a separate question. Discussing recommenders that shifted consumption away from the most popular items, Jannach and Jugovac warn that "such a shift in the consumption distribution does not necessarily mean that there is more business value". Compute the same exposure metrics on both arms of an A/B test from the same impression logs, decide on business metrics and use exposure metrics as guardrails. On a marketplace both arms share popularity counts and inventory, so interference between the arms can bias arm-level exposure differences. Where the stakes justify it, confirm with a longer or cluster-randomized test.

A Mitigation Ladder: From Re-Ranking to Retraining

For a production team the useful order for mitigation is by engineering cost, starting with steps that leave the trained model untouched: re-rank what the model returns, then reserve exploration slots and only then change training.

Serving-time re-ranking

Post-processing works on the candidate list the model already returns, which the survey describes as "re-ranking the items in a way that less popular items are brought to the front of the list". The simplest variant is a popularity penalty: subtract a weighted popularity term from each score before sorting, tune the weight offline and confirm it online.

Maximal marginal relevance is the classic diversity re-ranker. In Carbonell and Goldstein's SIGIR 1998 paper, "a document has high marginal relevance if it is both relevant to the query and contains minimal similarity to previously selected documents". MMR targets redundancy, so it reduces popularity concentration only when popular items also resemble each other.

Abdollahpouri, Burke and Mobasher adapted xQuAD, a search diversification method by Santos, Macdonald and Ounis, to popularity. Their abstract says "Our approach is a post-processing step that can be applied to the output of any recommender system." In the adapted objective, the paper explains that "the second term promotes diversity between two different categories of items (i.e. short head and long tail)". The gains reported are offline.

Calibration re-ranks toward each user's own category mix rather than toward the tail in general. The calibration survey states the goal: "the properties of the items that are suggested to users should match the distribution of their individual past preferences". It does not push tail items at users who genuinely prefer the head.

Reserved exploration slots

Exploration is the only rung that creates evidence for items the model has never seen succeed, and it still leaves the trained model alone. Reserve a share of slots, for example one position per list, for low-impression items and new arrivals, and fill them with a bandit policy. Our recommender system development page describes one e-commerce marketplace case where Thompson sampling bandits were used to explore long-tail products. Log the selection probability of every exploration slot. Epsilon-greedy and softmax policies give it exactly. Thompson sampling has no closed form for it, so estimate it, for example by Monte Carlo resampling of the posterior at serving time, or log the posterior parameters behind each decision so it can be estimated later. Those probabilities are propensities for the exploration traffic only, and the training-time corrections on the next rung can use them.

Training-time corrections

Inverse propensity scoring reweights each observed interaction by the inverse of the probability that it was observed, which counters the selection bias Schnabel et al. describe. In practice, the survey notes, "The propensity score is often based on the popularity of the items". Propensities for tail items can be tiny, so clip their inverses or a handful of rare interactions will dominate the loss. The clipping threshold is tuned offline and recorded with each model release, so a later shift in Gini can be traced to it.

Which rung to try first

The failing reading points to the rung. This mapping is our engineering judgment, not a published procedure.

Failing reading First rung to try Why
Miscalibration rising against users' own category mix Calibrated re-ranking Moves each list toward the user's own category mix
ARP above the profiles of every user group Popularity penalty at serving time The lift appears in every group, so one global weight is the cheapest first try
Popularity lift concentrated in the niche-taste group Personalized re-ranking such as the xQuAD adaptation Weights the long-tail bonus by each user's interest in the tail, per Abdollahpouri, Burke and Mobasher
Near-zero tail coverage or many zero-impression items Reserved exploration slots Re-ranking only reorders scored candidates, and an unseen item has no evidence
Exposure Gini creeping up across retrains with serving-time fixes in place Training-time correction The skew enters through training data, which propensity weighting targets

Stop climbing when the guardrail metric recovers to its baseline and the business metric in the A/B test does not degrade against control. If the business metric drops, step back a rung instead of adding another.

Cold Start, New Arrivals and Inventory

Popularity-based fallbacks are a legitimate default for anonymous sessions and new users, and our sibling guide recommends a popularity baseline when interaction history is thin. The risk is scale: when many sessions are anonymous, the fallback quietly becomes the main policy and the most concentrated one. Cap the share of impressions it may serve, report that share on the exposure dashboard and compute every metric in the table separately for fallback and model slots.

E-commerce adds constraints a media catalog lacks. New arrivals have no interactions on day one, so they need exploration slots or content-based scoring to get any exposure. An out-of-stock bestseller still carries its popularity, so filter on availability before re-ranking, or the penalty is spent on items nobody can buy. Sponsored placements are a separate exposure channel: tag them in the impression log and keep them out of organic metrics, or paid exposure will mask organic concentration.

Marketplaces have a second audience for these numbers: sellers. Sum impressions per seller and compute the same Gini and coverage metrics to see whether a few large sellers absorb most exposure on the supply side. Abdollahpouri, Burke and Mobasher cite earlier work in which "xQuAD was used to make a fair representation of items from different item providers".

What Should You Monitor After Each Retrain?

Popularity bias moves fastest at a retrain, when the new model learns from exposure the old one created. Treat exposure metrics as release gates next to accuracy.

  • Before promotion, score a shadow sample of live requests with the candidate and compute the six table metrics against the production model on the same requests.
  • Escalate when a metric leaves its tolerance band, set from its spread over several past retrains.
  • After promotion, track the same metrics daily on impression logs, split by user group and by slot source (model, fallback, exploration and sponsored).
  • Record the tolerance bands, the popularity-window length and the miscalibration distance with each release, change them only deliberately and keep the head cut-off fixed as set under where the long tail starts.

A Gini that rises a little on every retrain is a feedback loop in progress even if no release crosses a tolerance. These checks belong in the same pipeline stage that gates accuracy and latency, which our MLOps services page describes.

What Does the DSA Require of Recommender Systems?

The Digital Services Act defines a recommender system in Article 3(s) as "a fully or partially automated system used by an online platform to suggest in its online interface specific information to recipients of the service or prioritise that information". The obligations attach to online platforms, not to every website that recommends something, so check whether you are an online platform. Marketplaces usually are.

Article 27 transparency

Under Article 27, providers of online platforms that use recommender systems "shall set out in their terms and conditions, in plain and intelligible language, the main parameters used in their recommender systems", together with any options recipients have to modify or influence them. Article 27(2) adds: "The main parameters referred to in paragraph 1 shall explain why certain information is suggested to the recipient of the service." They must cover "the criteria which are most significant in determining the information suggested to the recipient of the service" and the reasons for the relative importance of those parameters.

Article 27 sits in the section that Article 19 opens, and that section "shall not apply to providers of online platforms that qualify as micro or small enterprises". The exclusion runs on for 12 months after a company outgrows that status and never covers a provider designated as a very large online platform.

Article 38 and the non-profiling option

Very large online platforms and search engines carry an extra duty: Article 38 says they "shall provide at least one option for each of their recommender systems which is not based on profiling", with profiling as the GDPR defines it. The EDPB Guidelines 3/2025 on the interplay between the DSA and the GDPR, version 2.0, adopted on 17 September 2026, say providers of very large platforms and search engines should present both options equally on first use and "should not nudge recipients of the service to select the option for a recommender system that is based on profiling". The profiling-based option may be used only after the user has chosen it, and while the non-profiling option is active, "the provider of the online platform cannot lawfully continue to collect and process personal data to profile the user for the purposes of future recommendations".

Why a non-profiling feed still needs monitoring

A common non-profiling option is a feed ranked by popularity, overall or within a category. Such a feed shows every user in scope the same head items, so it needs the same exposure metrics as the main feed. Monitor it on the same dashboard and consider non-personal diversity controls such as freshness or category quotas. A popularity penalty or an exploration share also changes why an item is suggested, so whoever drafts the Article 27 text should know about it.

How Pharos Production Helps

Popularity bias is cheapest to control when impression logging, exposure metrics and release gates are designed together with the model. Our recommender system development team builds that measurement layer, the re-ranking and exploration policies on top of it and the post-retrain checks. The same page states that every recommendation sprint includes fairness and bias auditing before production deployment. Our e-commerce software development practice connects the recommender to inventory, new arrivals and sponsored placements so each exposure channel is measured on its own.

Sources: Klimashevskaia, Jannach, Elahi and Trattner, A Survey on Popularity Bias in Recommender Systems; Chen et al., Bias and Debias in Recommender System: A Survey and Future Directions; Abdollahpouri, Burke and Mobasher, Managing Popularity Bias in Recommender Systems with Personalized Re-ranking; Abdollahpouri, Mansoury, Burke and Mobasher, RecSys 2020 paper on popularity bias, calibration and fairness; Steck, Calibrated Recommendations (RecSys 2018); Corrêa da Silva and Jannach, Calibrated Recommendations: Survey and Future Directions; Santos, Macdonald and Ounis, Exploiting Query Reformulations for Web Search Result Diversification (WWW 2010); Carbonell and Goldstein, The Use of MMR, Diversity-Based Reranking (SIGIR 1998); Chaney, Stewart and Engelhardt, How Algorithmic Confounding in Recommendation Systems Increases Homogeneity; Mansoury et al., Feedback Loop and Bias Amplification in Recommender Systems; Schnabel et al., Recommendations as Treatments; Jannach and Jugovac, Measuring the Business Value of Recommender Systems; RecBole evaluator metrics; Microsoft Recommenders evaluation; Regulation (EU) 2022/2065 (Digital Services Act); European Data Protection Board, Guidelines 3/2025 on the interplay between the DSA and the GDPR, version 2.0. Engineering guidance, not legal advice.

FAQ

Last updated:

Quick answers to common questions about custom software development, pricing, process and technology.

  • Is popularity bias always a problem?

    No. Some items are popular because they are better, cheaper or more relevant to most users, and a recommender that ignores that signal serves people worse. The problem starts when the system shows popular items to more users than their own histories justify, or when tail items never receive enough impressions to show whether anyone wants them.

    That is why the metrics compare served lists with user profiles and with the previous model rather than with a uniform distribution.

  • Does popularity bias hurt sellers on a marketplace?

    It can, because impressions are how sellers reach buyers. When a few listings absorb most impressions, the sellers behind them absorb most of the demand the recommender creates, while small and new sellers get few chances to earn interactions.

    Sum impressions per seller, compute the same Gini and coverage metrics on that distribution and report them next to the item-level readings. Exploration slots are the usual first fix for new sellers, because their listings have no history for re-ranking to work with.

  • Which open-source libraries compute popularity bias metrics?

    RecBole includes AveragePopularity, GiniIndex, TailPercentage, ItemCoverage and ShannonEntropy among its evaluator metrics. Microsoft Recommenders provides catalog coverage, distributional coverage and novelty.

    Both are designed around offline evaluation, so for production monitoring the usual approach is to reuse their definitions while feeding them impression logs instead of test splits, and to set the tail cut-off yourself, because RecBole's default tail is only the bottom 10% of items.

  • Does reducing popularity bias cost accuracy?

    Abdollahpouri, Burke and Mobasher report that their personalized re-ranker raised the share of less popular items while keeping what they call acceptable recommendation accuracy. That result is offline, and offline accuracy is scored on interactions the old, popularity-heavy policy produced, so it says little about live behavior. The trade-off has to be measured in an A/B test, deciding on business metrics such as conversion or retention and treating exposure metrics as guardrails.

  • Which metric should a small team start with?

    As a matter of engineering practice, start with catalog coverage and exposure Gini computed over impression logs that include zero-impression items. Both need only item IDs and impression counts, with no user profiles or category taxonomy, and together they show whether the catalog is served at all and how unevenly.

    Add popularity lift once users can be grouped by how mainstream their histories are, and miscalibration once item categories are clean.

  • Does every recommender need a non-profiling option under the DSA?

    No. The Article 38 duty applies only to very large online platforms and very large online search engines, on top of Article 27. Every online platform that uses recommender systems, very large ones included, has the Article 27 transparency duties unless it qualifies as a micro or small enterprise and is not designated a very large online platform.

    Whether a given business is an online platform at all is a legal question worth settling early, since a retailer selling only its own catalog may sit outside the definition.

I work with startup founders who need a dedicated software development team but don’t want to gamble on hiring, random outsourcing, or opaque delivery.
Most founders face the same problem sooner or later.
Early technical and team decisions lock the product into tech debt, slow delivery, missed milestones and constant re-hiring. By the time this becomes visible, fixing it is already expensive.

As a CTO and software architect, I help founders design, build and run dedicated development teams that work as a true extension of the startup. Not as a black-box vendor.

My focus is on complex products where mistakes are costly:

  • Web3 and blockchain platforms
  • FinTech and regulated products
  • High-load startup systems
  • MVP → scale transitions

We don’t do body-shopping.
We don’t sell generic outsourcing.

Instead, we help founders:

  • build the right team structure from day one
  • keep technical ownership and transparency
  • scale delivery without losing control
  • avoid vendor lock-in and hidden risks

Teams are aligned with the product roadmap, business goals and long-term architecture. Not just short-term velocity.

Dmytro Nasyrov, Founder and CTO at Pharos Production
Dmytro Nasyrov Founder & CTO Let's work together!

Your business results matter

Achieve them with minimized risk through our bespoke innovation capabilities

Your contact details
Please enter your name
Please enter a valid email address
Please enter your message

We use your details only to reply to your request. Data Privacy and Legal Notice

We typically reply within 24 hours

What happens next?

  1. Contact us

    Contact us today to discuss your project. We're ready to review your request promptly and guide you on the best next steps for collaboration

    Same day
  2. NDA

    We're committed to keeping your information confidential, so we'll sign a Non-Disclosure Agreement

    1 day
  3. Plan the Goals

    After we chat about your goals and needs, we'll craft a comprehensive proposal detailing the project scope, team, timeline and budget

    3-5 days
  4. Finalize the Details

    Let's connect on Google Meet to go through the proposal and confirm all the details together!

    1-2 days
  5. Sign the Contract

    As soon as the contract is signed, our dedicated team will jump into action on your project!

    Same day