Content Relevance : The Invisible Product Series

Seven Essays on Ranking, Relevance, and the Future of Discovery

Series thesis

Every social or entertainment product makes a promise about what deserves a person’s attention. Ranking is where that promise becomes executable.

This series argues that relevance is not simply the probability of a click, view, or like. It is a product judgment about the right content, person, or conversation for a particular user, in a particular context, with an acceptable set of consequences—for the user, creators, the platform, and the culture that forms around it.

I came to this view after working across ranking and relevance problems in comments, Stories, Notes, sharing, advertising, and discovery. These surfaces may share models and infrastructure, but they do not share the same definition of value. A comment ranker should help someone understand or join a conversation. A Stories tray should help someone keep up with people they care about. A sharing ranker should understand relationships, not just content. An entertainment homepage should reduce the distance between opening the app and finding something worth starting.

The mistake is to call all of these “engagement optimization.” The real work is translating a surface’s distinct promise into candidates, objectives, feedback loops, safeguards, and product controls.

The series

  1. The Feed Is a Product Constitution — Why ranking begins with a product promise, not a model.
  2. The Candidate Set Is Destiny — How retrieval silently determines what a ranker can ever choose.
  3. The Objective Function Is the Product — Why metric design is a values decision in mathematical clothing.
  4. Your User Is Not a Vector — How intent, context, relationships, and changing tastes redefine personalization.
  5. Serendipity Is a Systems Problem — Why exploration is essential for users, new content, and creator mobility.
  6. Every Ranker Creates Its Own Reality — How recommendation systems shape their future data and govern an ecosystem.
  7. After the Ranker — How sequence models, foundation models, and generative AI are changing the unit of recommendation.

Part 1: The Feed Is a Product Constitution

Why ranking is the most important product users never see

TL;DR: Ranking has become the invisible operating system of social media and entertainment: it translates a platform’s values into millions of daily decisions about who and what deserves attention. The right place to begin is not with a model or engagement metric, but with a clear contract among user intent, surface purpose, contextual value, and the long-term consequences of repeatedly making the same choice.

Open almost any social or entertainment app and pull down to refresh.

In less time than it takes to blink, the product makes a remarkable set of decisions. It decides which friend you should hear from, which creator deserves a chance, which conversation is worth entering, which song fits the moment, which show you might commit an evening to, and which unfamiliar idea should interrupt everything the system already knows about you.

The interface makes these decisions feel natural. A feed simply appears. A Stories tray looks inevitable. The next video begins. The homepage seems as if it has always been arranged that way.

But nothing about the order is inevitable.

Every position is a choice. Every omission is also a choice. And across billions of requests, those choices become the product.

Over the years, I have worked on ranking and relevance across Instagram Comments, Stories, Notes, and sharing, as well as advertising, marketplace, and discovery experiences. The lasting lesson was not that one model architecture consistently wins. It was that ranking problems that look similar from a distance become profoundly different once you ask what the surface is supposed to do for a person.

A Stories tray is not a miniature Explore page. A comments section is not a feed made of text. A sharing suggestion is not a content recommendation with names substituted for posts. A streaming homepage and a short-video feed both recommend entertainment, but one asks for a meaningful commitment while the other asks for a few seconds of attention.

The systems may share infrastructure. They do not share the same definition of relevance.

This is the starting point for The Invisible Product, a seven-part series about ranking, relevance, and discovery. The series moves from the product philosophy of ranking into the machinery beneath it: retrieval, objectives, user and intent modeling, exploration, feedback loops, ecosystem governance, and the emerging shift toward generative recommendation.

The core argument is simple:

A ranking system is product strategy made executable.

To understand why, we first need to understand how ranking became the center of the product.

From organizing content to allocating attention

The early social web was largely organized around explicit choice. You followed people, subscribed to channels, joined communities, or searched for something. Products still sorted and filtered, but the user’s declared graph did much of the work.

That world changed as three forms of abundance arrived at once.

First, content supply became effectively infinite. Smartphones turned billions of people into potential creators. Short-form video lowered the cost of production and consumption. Global catalogs placed more music, shows, podcasts, posts, and communities within reach than any person could evaluate.

Second, platforms moved beyond the social graph. The most compelling item might come from someone a user has never met or followed. TikTok’s public description of the For You feed made this interest-driven model explicit: user interactions and video information help rank content for each person, while deliberate diversification introduces material outside already expressed interests. YouTube has said that recommendations drive more viewing than subscriptions or search.

Third, AI is making supply grow faster still. Generative tools reduce the cost of producing text, images, audio, and video. At the same time, foundation models are changing recommendation itself: Netflix’s 2026 GenRec work, for example, translates user history and context into language and uses an LLM-backed system to rank catalog items against long-term member value.

In a world of scarce content, the platform’s job is to help people find enough. In a world of abundant content, its job is to decide what deserves attention.

That is a much more consequential role.

Ranking is no longer an optimization layer sitting beneath the product. It determines the effective inventory, the perceived community, the distribution of opportunity, and the emotional texture of each session. It affects what users believe the platform is, because most of the platform they experience is the subset the system chooses to show.

“The algorithm” is the wrong mental model

Users understandably talk about “the algorithm” as if a single intelligence controls the experience. Inside a platform, the reality is more fragmented.

There may be separate systems for Feed, Stories, Reels, comments, notifications, search, people recommendations, sharing suggestions, and ads. Each can have multiple candidate generators, prediction models, value formulas, integrity classifiers, business rules, and reranking stages. Meta has published separate system cards for a range of Facebook and Instagram experiences, reflecting the fact that ranking is a family of systems rather than one universal formula.

This distinction matters because it changes how teams diagnose problems.

If a feed feels repetitive, the final ranker may be behaving exactly as designed while retrieval keeps supplying near-duplicates. If new creators receive little exposure, the problem may be candidate eligibility or exploration rather than a creator-quality prediction. If a comment section feels hostile, optimizing the probability of a reply will not fix the experience if the system cannot distinguish productive conversation from conflict.

Saying “improve the algorithm” is like saying “improve the economy.” The request names the outcome but not the mechanism.

The product leader’s job is to make the system legible: what decision is each layer making, what definition of value guides it, and where can the product promise break?

Every surface makes a different promise

When ranking conversations begin with signals or model architecture, they begin too late. The more important first question is:

What promise is this surface making to the user?

That question can feel philosophical. In practice, it is highly operational because the answer determines what should be retrieved, predicted, optimized, constrained, and measured.

Surface

The implicit user question

A plausible product promise

What “relevant” might mean

Home Feed

What matters to me now?

Keep me meaningfully connected and informed

The right mix of relationships, interests, freshness, and discovery

Stories

Whose update should I not miss?

Help me keep up with people I care about

Relationship value, recency, and likelihood of meaningful consumption

Comments

What is worth reading or responding to?

Help me understand and participate in the conversation

Context, quality, author interaction, social relevance, and conversational health

Sharing

Who would appreciate this from me?

Help me turn content into connection

Recipient relevance, relationship context, and sender-recipient meaning

Short video

Is the next moment worth my attention?

Deliver immediate value while keeping discovery fresh

Fast intent inference, consumption quality, novelty, and session balance

Entertainment Home

What am I willing to start?

Reduce the distance from opening to satisfying consumption

Current intent, time budget, continuation, household context, and long-term satisfaction

These are not copywriting variations. They produce different labels and failure modes.

On Stories, a skip can mean “not now,” not “I do not care about this person.” In comments, a reply can signal genuine conversation or escalating conflict. In sharing, a frequently messaged person is not necessarily the right recipient for a particular post. On a streaming homepage, a click that does not lead to meaningful viewing may represent selection failure rather than success.

The surface promise gives behavior its meaning.

Relevance is not similarity

Early in my work on ranking, I thought about relevance largely as a matching problem: understand the user, understand the item, and estimate the strength of the match.

That model is useful. It is also incomplete.

Similarity answers, Is this item close to things the user has liked before? Relevance asks a harder question: Is this the right item for this person, on this surface, in this moment, given what will happen next?

The same person can arrive at adjacent surfaces with entirely different jobs to be done. Someone who wants close-friend updates in Stories may want novelty in Explore. Someone reading comments may value context, wit, disagreement, or the author’s own response. Someone sharing a post may choose a recipient because of a private joke or a recent conversation that no content-only representation captures.

Even the same piece of content can move from highly relevant to irrelevant as context changes. A 40-minute podcast may be perfect during a commute and unusable between meetings. A breaking-news clip may be valuable once and repetitive the fifth time. A child’s favorite song may be correct for a family speaker and wrong for the parent’s personal discovery feed.

The true unit of relevance is therefore not the user-item pair. It is:

user × item × intent × context × consequence

I call this the relevance contract. It has four clauses.

1. Intent: What is the person trying to do now?

Intent is more specific than interest. A person can love documentaries but currently want a five-minute laugh. They can care deeply about a friend but open the app to follow a live event. They can regularly listen to music yet prefer a podcast on weekday mornings.

A good system uses historical preference as evidence without allowing history to overrule the present.

2. Value: What would make this recommendation worthwhile?

The answer depends on the surface. It might be meaningful consumption, a useful reply, a successful share, a new creator relationship, a satisfying session, or simply helping the user complete a task quickly.

Value cannot be defined only as “more.” More clicks, comments, or minutes may be useful leading indicators, but volume is not the same as benefit.

3. Context: What changes the meaning of the decision?

Context includes surface, time, device, session history, location, relationship, freshness, available attention, and the items surrounding the recommendation. It determines which part of a person’s history should matter and what form the recommendation should take.

Context also explains why one global ranker rarely serves every surface well. The question is not merely what the user likes, but which version of the user has arrived.

4. Consequence: What happens if the system repeats this choice?

This is the clause most often left implicit.

A single sensational video may satisfy a user’s immediate curiosity. A feed filled with increasingly sensational content is a different product. One recommendation from a familiar creator may be ideal. A system that gives familiar creators nearly all exposure may make discovery and creator mobility impossible.

Ranking happens one impression at a time, but its consequences accumulate at the level of sessions, habits, markets, and culture.

Behavior is evidence, not ground truth

Modern rankers often predict several outcomes—click, watch, like, share, comment, hide, report—and combine those estimates into a value score. Meta’s public description of Instagram Explore shows a multi-stage system in which predicted probabilities for positive and negative actions can be weighted into an expected-value function before final reranking.

This architecture is sensible. But the formula cannot decide what the weights should mean.

Should a send be worth more than a like? Often, because sharing can turn content into a relationship event. But a send can also communicate ridicule or outrage.

Should a long watch beat a short one? Not if the short item answers the question immediately while the long one withholds the answer.

Should a comment count as positive engagement? Not automatically. It could reflect belonging, confusion, correction, or conflict.

Should a skip be negative? Perhaps the user already saw the item elsewhere, recognized it instantly, or got what they needed from the opening frame.

Every behavioral label compresses several possible human meanings into one observable action.

The product team’s task is therefore not to discover a magical engagement metric. It is to form a theory of value and continually test whether observed behavior remains a good proxy for that value.

YouTube’s public account of its recommendation evolution is instructive. The system moved beyond clicks after the limits of click-through optimization became clear, then beyond raw watch time by incorporating surveys and predicting “valued watchtime.” That progression was not simply an advance in modeling. It reflected a more mature answer to the product question: What makes time on YouTube worthwhile?

When the metric evolves, the product evolves.

Every recommendation makes two bets

A ranker makes two bets every time it allocates an impression.

The first is a response bet:

What will this person do if we show this item now?

The second is a consequence bet:

What will repeated decisions like this do to the person, creators, the platform, and the future inventory?

The response bet is easier to train because feedback arrives quickly. Did the user watch, like, share, skip, or report? These events can be logged within seconds.

The consequence bet is harder. It may appear later as satisfaction, return behavior, creator retention, content quality, trust, habit formation, or regret. Its causal path is longer and noisier. Yet much of the durable value—and much of the risk—lives in this second bet.

This matters because social and entertainment platforms are not inert catalogs. They are adaptive ecosystems.

  • Users learn what the product tends to show and change how they behave.
  • Creators learn what receives distribution and change what they make.
  • Product teams change surfaces in response to the measured behavior.
  • Models train on the resulting data and become more confident in the environment earlier models helped create.

Ranking is therefore not merely a mirror of preference. It is an intervention in preference and production.

That is why relevance must include consequence.

Why I call the feed a constitution

A constitution does more than declare values. It establishes how decisions get made, which rights and constraints cannot be traded away, and how competing interests are resolved.

A ranking stack performs a similar function for attention.

It decides which content is eligible to compete, which sources receive candidate capacity, what predicted outcomes count as value, when integrity overrides engagement, how much space is reserved for exploration, and whether users can correct the system’s assumptions.

In practice, this constitution is distributed across several layers:

Layer

Constitutional question

Eligibility

What is allowed to compete for attention on this surface?

Retrieval

Which parts of the content universe deserve consideration?

Pre-ranking

Which candidates receive expensive evaluation—and which are dismissed early?

Ranking

What value does each candidate offer this user in this request?

Reranking

What changes when individually strong items form a repetitive or unhealthy slate?

Integrity

Which risks can constrain or override predicted engagement?

Exploration

How much opportunity goes to uncertain users, interests, items, and creators?

User control

How can a person correct, steer, explain, or reset the system?

Measurement

How will the platform distinguish immediate response from durable value?

This framing reveals why improving the final ranking model is not enough.

If promising new creators never enter retrieval, a late-stage fairness weight cannot rescue them. If eligibility admits too much low-quality material, the value model is forced to compensate. If pre-ranking systematically removes a content class, the expensive model never gets to disagree. If individually relevant items form a monotonous slate, item-level accuracy can coexist with a poor session.

The experience is created by the constitution as a whole.

The characteristic failures of simple objectives

Every proxy creates a recognizable product when optimized too aggressively.

  • Pure click optimization tends toward clickbait and exaggerated packaging.
  • Pure watch-time optimization can reward padding, repetition, or content that is difficult to leave.
  • Pure freshness creates noise and undervalues durable quality.
  • Pure popularity makes discovery circular: content is shown because it is popular and becomes more popular because it is shown.
  • Pure relationship strength can make a social product feel closed and static.
  • Pure novelty can make a feed surprising but exhausting.
  • Pure posting recency or frequency can encourage creators to produce more rather than better.

These failure modes do not mean the signals are bad. They mean no proxy contains the complete product goal.

A strong team asks not only, “What metric will this improve?” but also, “What experience would emerge if this metric won every trade-off for a year?”

That question converts metric review into product review.

Start with a ranking brief

Before debating architectures, I recommend writing a one-page ranking brief. It forces the product definition into the open and gives product, engineering, data science, design, research, integrity, and creator teams a shared contract.

The brief should make seven decisions.

1. Surface promise

In one sentence, what should a successful visit help the user do? Avoid vague language such as “show relevant content.” Define the user outcome.

2. Primary value

What observed or modeled outcome best represents that promise? Where is it strong, and where is it ambiguous?

3. Intent model

Which distinct user intents arrive at this surface? How should the experience change among them?

4. Time horizon

Which immediate, session-level, and long-term outcomes matter? What evidence connects the leading metrics to durable value?

5. Non-negotiables

Which safety, quality, privacy, fairness, or ecosystem conditions are constraints rather than tradable weights?

6. Exploration policy

How will new content, creators, and interests receive enough exposure to generate evidence? Where is exploration welcome, and where is it costly?

7. Failure signature

If the system over-optimizes its proxy, what will the product feel like? What leading indicators will reveal the drift before the north star moves?

A ranking brief does not eliminate disagreement. It makes the disagreement concrete enough to resolve.

Ranking changes how product teams should work

In a conventional feature, teams can often specify the desired behavior directly: the user presses a button and the product performs an action. In a ranked experience, teams specify an objective, a candidate universe, and a set of constraints—then observe an emergent distribution of behavior.

That demands a different operating model.

Product managers need enough modeling fluency to understand what the system can learn and where data is biased. Machine-learning engineers need the product context to know which improvements are meaningful. Data scientists need to connect offline metrics with online and long-term outcomes. Designers need to treat position, presentation, explanation, and feedback controls as parts of the decision system. Researchers need to test whether behavioral proxies match people’s reported experience. Integrity and policy teams need influence upstream, not only after launch. Creator teams need visibility into how ranking incentives alter supply.

Most importantly, the team must inspect the experience itself.

Dashboards can show that predicted relevance improved while the feed becomes repetitive. Aggregate engagement can rise while a small but important cohort loses value. A creator-exposure metric can look healthy while the system repeatedly tests new creators with the wrong audience.

Session reviews, user research, creator feedback, and cohort analysis are not soft supplements to model evaluation. They are how a team learns what its metrics cannot see.

Why this series, and why now

Ranking is entering another transition.

Retrieval is becoming learned and more expressive. User histories are being modeled as sequences rather than hand-built aggregates. Multimodal models understand more of what appears inside an image, video, or song. Foundation models can interpret natural-language intent and help construct entire slates. Generative AI is making content supply even more abundant.

These advances will improve personalization. They will also make ranking policies more capable—and therefore more consequential.

The central questions will not disappear:

  • What enters the candidate set?
  • Which behavior represents value?
  • How should immediate response trade off against long-term satisfaction?
  • How can a system learn interests a user has not yet expressed?
  • What opportunity should new creators receive?
  • How do prior ranking decisions distort future training data?
  • What must never be optimized away?
  • How can users correct the model’s view of them?

The remaining essays take those questions in order.

Part 2, “The Candidate Set Is Destiny,” examines retrieval—the hidden stage that determines what the final ranker can ever choose.

Part 3, “The Objective Function Is the Product,” explores why metric design is a values decision in mathematical clothing.

Part 4, “Your User Is Not a Vector,” looks at intent, context, relationships, changing tastes, and the importance of forgetting.

Part 5, “Serendipity Is a Systems Problem,” explains why disciplined exploration is necessary for discovery, cold start, and creator mobility.

Part 6, “Every Ranker Creates Its Own Reality,” examines feedback loops, exposure bias, ecosystem governance, and user agency.

Part 7, “After the Ranker,” considers how sequence models, foundation models, and generative systems are changing recommendation from selecting items to constructing experiences.

Together, they make the case that ranking should be treated not as a narrow machine-learning function but as a product discipline.

The product is the distribution of attention

In many products, strategy becomes visible through features, navigation, and design. In ranked products, strategy also appears in a less visible place: the distribution of impressions.

Which friend appears first? Which creator receives an initial audience? Which comment establishes the tone? Which show gets a second chance after a weak launch? Which topic is interrupted before it becomes repetitive? Which unfamiliar interest is given room to develop?

These are product decisions even when models make them millions of times per second.

This is why ranking should not be treated as a service that personalizes the experience after the product has been designed. For social media and entertainment platforms, ranking is part of the product design. It determines what the product feels like, what behavior it encourages, what creators make, and what ecosystem eventually forms.

The feed is a product constitution.

The model is how that constitution is executed at request time.

Leave a Comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.