A new benchmark pipeline shows that LLM recommenders produce gender- and nationality-dependent top-k lists in cold-start settings, and model size affects bias non-monotonically.
Faireval: A benchmark for evaluating user-level fairness in large language model-based recommender systems
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting
A new benchmark pipeline shows that LLM recommenders produce gender- and nationality-dependent top-k lists in cold-start settings, and model size affects bias non-monotonically.