Pith. sign in

REVIEW 11 cited by

RecSim: A Configurable Simulation Platform for Recommender Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.04847 v2 pith:K3ZEWERY submitted 2019-09-11 cs.LG cs.HCcs.IRstat.ML

RecSim: A Configurable Simulation Platform for Recommender Systems

classification cs.LG cs.HCcs.IRstat.ML
keywords recsimuserenvironmentsbehaviorconfigurableitemplatformrecommender
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We propose RecSim, a configurable platform for authoring simulation environments for recommender systems (RSs) that naturally supports sequential interaction with users. RecSim allows the creation of new environments that reflect particular aspects of user behavior and item structure at a level of abstraction well-suited to pushing the limits of current reinforcement learning (RL) and RS techniques in sequential interactive recommendation problems. Environments can be easily configured that vary assumptions about: user preferences and item familiarity; user latent state and its dynamics; and choice models and other user response behavior. We outline how RecSim offers value to RL and RS researchers and practitioners, and how it can serve as a vehicle for academic-industrial collaboration.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. D4RL: Datasets for Deep Data-Driven Reinforcement Learning

    cs.LG 2020-04 accept novelty 8.0

    D4RL supplies new offline RL benchmarks and datasets from expert and mixed sources to expose weaknesses in existing algorithms and standardize evaluation.

  2. RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents

    cs.IR 2026-05 unverdicted novelty 7.0

    RecoAtlas is a benchmark that evaluates LLM recommendation agents on behavior-grounded metrics for relevance, complementarity, and diversity in addition to semantic coherence.

  3. Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces

    cs.CL 2026-04 unverdicted novelty 7.0

    OmniBehavior benchmark demonstrates that LLMs simulating real human behavior converge on hyper-active positive average personas, losing long-tail individual differences.

  4. Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces

    cs.CL 2026-04 unverdicted novelty 7.0

    Introduces OmniBehavior benchmark from real-world data and shows LLMs exhibit hyper-activity, persona homogenization, and utopian bias in behavior simulation.

  5. Hierarchical Residual Policy Optimization for Generative Recommendations

    cs.IR 2026-08 conditional novelty 6.0

    HRPO decomposes item-level rewards into token-level 'residual credits' along semantic identifier hierarchies and optimizes the generator with a PPO-style objective, improving session utility in KuaiSim and production ...

  6. CoSimRec: Measuring Coordinated-Content Penetration in Recommender Feedback Loops

    cs.IR 2026-07 conditional novelty 6.0

    Under controlled simulation, coordinated content reaches non-bot recommendation slots when rankers reward popularity or feedback (APR-Lift up to 0.47 on LastFM), while random ranking shows none.

  7. How Does Empowering Users with Greater System Control Affect News Filter Bubbles?

    cs.IR 2026-06 conditional novelty 6.0

    Users who could adjust a news recommender's stance and topic sliders changed how extreme their feed became depending on their starting point, but did not consistently increase political diversity.

  8. Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation

    cs.HC 2025-08 unverdicted novelty 6.0

    A two-phase data construction framework generates explanatory rationales from user feedback and applies uncertainty-based distillation to fine-tune lightweight LLMs as preference-aligned user simulators for recommende...

  9. Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising

    cs.IR 2026-07 conditional novelty 5.5

    DASH folds cross-domain user histories, distills teacher thinking traces, and RL-tunes a small LLM with action plus rubric rewards to jointly predict ad actions and decision traces on Tencent data.

  10. Exploitation Over Exploration: Unmasking the Bias in Linear Bandit Recommender Offline Evaluation

    cs.LG 2025-07 unverdicted novelty 5.0

    Greedy linear models without exploration consistently achieve top-tier performance in over 90% of offline dataset evaluations for linear bandit recommenders, with hyperparameter tuning favoring minimal exploration and...

  11. Affective Music Recommendation: A Rollout-Based World Model for Offline Preference Optimization

    cs.LG 2026-05 unverdicted novelty 4.0

    AMRS deploys a rollout-based causal transformer world model for offline DPO-based affective music recommendation under cold-start conditions on health platforms.