FUURAA AI Knowledge Library · recommendation evaluation

How to evaluate AI recommendation and personalization capability claims

Use six evidence gates to turn “understands users better,” “delivers more relevant recommendations” or “personalisation lifts conversion” into an exact surface, catalogue, user and outcome, traceable exposure and labels, candidate-generation and ranking pipeline, leakage-free offline evaluation, complete online denominator, feedback loops, operating cost and a dated applicability boundary.

Published31 August 2026Evidence statusMethod synthesis grounded in primary industrial systems, public datasets and exposure-bias researchScopeContent, product, news, advertising and sequential recommendation research

Higher click-through rate is not complete user value; one offline metric is not transferable evidence

Recommendation capability must keep exposure, labels, candidates, ranking, policy, complete outcomes and feedback loops in one evidence chain.

A recommender determines what users can see, so logged data is not a neutral sample. Evaluation must distinguish unexposed from disliked, offline relevance from online outcomes, and model scores from the complete policy, while documenting how personalisation changes future data.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, platform, user or leaderboard. It does not guarantee engagement, sales, retention, safety, fairness, user wellbeing or creator and provider outcomes.

Six rejectable evidence gates

Each gate requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the surface, catalogue, user and outcome

Decision question
What exact list, slot, session, eligible catalogue, user population, action and decision outcome does the claim cover?
Minimum evidence
Verbatim claim and date; exact system, model, feature, candidate-generation, ranking, reranking and policy versions; surface, slots, eligibility, user and session definition; click, watch, purchase, retention, satisfaction or other declared objective; constraints and fallback.
Stop condition
Stop when “personalised for everyone” or “better recommendations” replaces a named surface, population, catalogue, objective and acceptance decision.
02

Version catalogue, exposure, interaction and labels

Decision question
What could each user actually see, under which logging policy, and what does a missing interaction mean?
Minimum evidence
Catalogue and user provenance, snapshots and availability times; impression, position and logging policy; clicked, skipped, watched, purchased, hidden and reported events; label windows; repeated exposure; missing and unexposed treatment; provider and catalogue change; declared permissions and exclusions.
Stop condition
Stop when unexposed items are labelled as dislikes, non-clicks are treated as stable preferences, or future interactions enter training or evaluation.
03

Rebuild candidate generation, ranking and policy

Decision question
Can every recommendation be traced through retrieval, ranking, rules, filters, caches and fallbacks?
Minimum evidence
Candidate-source coverage and recall; feature-time availability; model scores and calibration; ranking and reranking order; safety, moderation, freshness, diversity, deduplication and inventory rules; exploration; caches, unavailable-item handling, fallbacks and per-impression trace.
Stop condition
Stop when a model score is presented as the complete recommender while policy rules, filters, candidate omissions or fallback traffic remain invisible.
04

Test offline relevance without confusing it with user value

Decision question
Do leakage-free offline results beat appropriate baselines across relevant slices, and what outcome do the metrics not measure?
Minimum evidence
Chronological, user and item splits; popularity, recency and simple collaborative baselines; Recall, Precision, NDCG, MRR, AUC, log loss and calibration with declared definitions; catalogue coverage, novelty and diversity; position and exposure-bias analysis; cold-start, tail, language, group and provider slices.
Stop condition
Stop when one offline metric, benchmark rank or click proxy replaces satisfaction, downstream value, calibration, critical slices and declared trade-offs.
05

Count every impression, user, item, failure and online effect

Decision question
What happened across the complete eligible population, experiments, failures and downstream guardrails?
Minimum evidence
All eligible users, sessions, impressions and items; empty, duplicate, unavailable, expired and unsafe results; latency, retries, timeouts, cache misses and fallbacks; repeated-run variance; experiments with allocation, interference and guardrails; satisfaction, retention, complaints, provider outcomes, human overrides and total cost.
Stop condition
Stop when selected sessions or aggregate click lift hide empty lists, harmful concentration, critical cohorts, unstable experiments, manual rescue, slow tails or cost.
06

Transfer through feedback loops and expire the claim

Decision question
Does the evidence remain valid after users, inventory, policy, seasonality and the recommender itself change future data?
Minimum evidence
Shadow, staged or randomized target testing; drift in users, items, exposure and outcomes; popularity concentration, filter-bubble and provider-distribution monitoring; exploration and counterfactual checks; change triggers, owner, fallback, rollback and expiry date.
Stop condition
Stop when a static logged-data result transfers directly to a live ranking policy without target validation, feedback-loop monitoring and rollback.

Minimum recommendation claim failure matrix

Check these conditions deliberately before transferring selected sessions, click lift or offline rank into real personalisation capability.

  • 01
    Unexposed items are treated as negative preferences

    Record the affected surface, user, item, position, metric and outcome; preserve the narrowest conclusion that remains and specify whether to repair logging, resplit data, rerun, expand exploration, restrict use, fall back to a policy or withdraw the claim.

  • 02
    Position and logging-policy bias make yesterday's recommender look like ground truth

    Record the affected surface, user, item, position, metric and outcome; preserve the narrowest conclusion that remains and specify whether to repair logging, resplit data, rerun, expand exploration, restrict use, fall back to a policy or withdraw the claim.

  • 03
    Future interactions, item popularity or availability leak into evaluation

    Record the affected surface, user, item, position, metric and outcome; preserve the narrowest conclusion that remains and specify whether to repair logging, resplit data, rerun, expand exploration, restrict use, fall back to a policy or withdraw the claim.

  • 04
    Click optimisation reduces satisfaction, retention or user control

    Record the affected surface, user, item, position, metric and outcome; preserve the narrowest conclusion that remains and specify whether to repair logging, resplit data, rerun, expand exploration, restrict use, fall back to a policy or withdraw the claim.

  • 05
    Popular items dominate while cold-start, niche and long-tail items disappear

    Record the affected surface, user, item, position, metric and outcome; preserve the narrowest conclusion that remains and specify whether to repair logging, resplit data, rerun, expand exploration, restrict use, fall back to a policy or withdraw the claim.

  • 06
    Aggregate relevance hides weak language, user, item or provider slices

    Record the affected surface, user, item, position, metric and outcome; preserve the narrowest conclusion that remains and specify whether to repair logging, resplit data, rerun, expand exploration, restrict use, fall back to a policy or withdraw the claim.

  • 07
    Duplicate, expired, unavailable or unsafe items reach the final list

    Record the affected surface, user, item, position, metric and outcome; preserve the narrowest conclusion that remains and specify whether to repair logging, resplit data, rerun, expand exploration, restrict use, fall back to a policy or withdraw the claim.

  • 08
    Policy, seasonality or feedback loops invalidate historical evidence

    Record the affected surface, user, item, position, metric and outcome; preserve the narrowest conclusion that remains and specify whether to repair logging, resplit data, rerun, expand exploration, restrict use, fall back to a policy or withdraw the claim.

Minimum recommendation capability evaluation record

Let the next reviewer reconstruct the conclusion with the same exposure log, catalogue, policy, system and complete denominator.

  1. 01Claim, date, owner, surface, slots, eligible catalogue, users, sessions, objective and decision
  2. 02System, models, features, candidate generation, ranking, reranking, policy and deployment versions
  3. 03Catalogue and user provenance, snapshots, availability times, permissions and exclusions
  4. 04Exposure, position, logging policy, interactions, labels, missingness and negative rules
  5. 05Splits, feature-time controls, baselines, metric definitions, seeds and reference code
  6. 06Candidate coverage, relevance, calibration, diversity, novelty, freshness and critical slices
  7. 07All users, sessions, impressions, items, empty results, failures, repeats and fallbacks
  8. 08Experiment design, allocation, interference, guardrails, downstream outcomes and confidence
  9. 09Latency, cache, compute, moderation, human review, provider impact and total cost
  10. 10Target evidence, feedback-loop and drift monitoring, triggers, rollback, owner and expiry

Common evidence states

Bind conclusions to the exact pipeline, surface, exposure policy, complete denominator, target outcome and date.

Supported

The exact pipeline meets declared relevance, calibration, diversity, critical-slice, failure, latency, cost and target-outcome thresholds on representative users and inventory, with a current review date.

Conditional

Support holds only for named surfaces, users, catalogue states, positions, objectives, policies, seasons or operating controls.

Mixed

Offline relevance, online outcomes, calibration, diversity, critical cohorts, stability, latency or cost vary materially.

Insufficient

Exposure policy, labels, pipeline, leakage control, baseline, complete denominator, target experiment, feedback-loop boundary or expiry is missing.

FUURAA analysisThe minimum decision unit for an AI recommendation and personalisation capability claim is exact system, model, feature, candidate-generation, ranking, reranking and policy version × surface, user, session, product or content catalogue, position and target decision × exposure, interaction, labels, availability time, permissions, missingness and negative rules × data splits, baselines, training, offline evaluation, experiments and operating constraints × relevance, calibration, diversity, novelty, coverage, critical errors and user outcomes × all users, impressions, items, empty results, failures, variance, human handling, latency and cost × feedback loops, target-workflow boundary and cut-off date. Higher click-through rate or offline rank is an observation, not transferable proof of personalisation capability.

Primary sources and non-transfer boundaries

These sources constrain candidates and ranking, memorisation and generalisation, interest modelling, news recommendation, exposure bias and randomized sequential evaluation.

Sources rechecked 31 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published September 2016Deep Neural Networks for YouTube Recommendations

Describes a large-scale industrial architecture that separates candidate generation from ranking and reports online experiments alongside offline evaluation.

BoundaryA high-level account of one proprietary video service and selected metrics does not transfer to another catalogue, user population, policy, objective or long-term outcome.

Open primary source ↗
Published 24 June 2016Wide & Deep Learning for Recommender Systems

Combines memorisation and generalisation in a production recommender and evaluates app acquisitions in Google Play.

BoundaryAcquisition gains in one app store do not establish long-term user value, transfer, calibration, fairness or performance for another surface.

Open primary source ↗
Published 21 June 2017Deep Interest Network for Click-Through Rate Prediction

A Chinese-led study models candidate-conditioned user interests for click-through-rate prediction and reports deployment in Alibaba advertising.

BoundaryCTR gains on advertising data do not prove satisfaction, purchase, retention, user wellbeing or performance in another domain.

Open primary source ↗
Published July 2020MIND — A Large-scale Dataset for News Recommendation

A Chinese-led Microsoft Research Asia dataset provides English news, impression logs and clicked histories for large-scale news recommendation research.

BoundaryEnglish Microsoft News clicks, sampled users and logged impressions do not represent every language, current event, editorial goal or unobserved preference.

Open primary source ↗
Published 22 February 2022KuaiRec — A Fully-observed Dataset for Recommender Systems

A Chinese-led Kuaishou study builds a near-fully observed user-item matrix to expose how missing-not-at-random interactions and exposure bias alter evaluation.

BoundaryA small opt-in short-video population and near-complete exposure do not reproduce a live catalogue, ranking policy, creator ecosystem or long-term feedback.

Open primary source ↗
Published 18 August 2022KuaiRand — An Unbiased Sequential Recommendation Dataset

A Chinese-led Kuaishou dataset uses randomized exposure to support evaluation of sequential recommendation under reduced policy bias.

BoundaryRandomized short-video interactions do not reproduce another platform, language, objective, moderation system, creator economics or long-term intervention.

Open primary source ↗

Continue checking

Move from recommendation into data analysis, factuality, information extraction or the full library.

Enter the AI Knowledge LibraryEvaluate data analysisEvaluate factualityEvaluate information extractionEnter AI Evidence Atlas