Pith. sign in

hub Mixed citations

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders

Mixed citation behavior. Most common role is background (44%).

61 Pith papers citing it
4 external citations · Pith
Background 44% of classified citations
abstract

Feature engineering has long been central to recommender systems, yet effectively leveraging textual item features remains challenging. Recent advances in large language models (LLMs) have enabled their use as semantic encoders for recommendation, but their roles and behaviors in this setting are still not well understood. Prior studies often rely on general-purpose embedding benchmarks (e.g., MTEB) when selecting LLMs, overlooking the unique characteristics of recommendation tasks. To address this gap, we introduce BLaIR, a comprehensive benchmark for evaluating LLMs as semantic encoders in recommendation scenarios. We contribute (1) a new large-scale Amazon Reviews 2023 dataset with over 570 million reviews and 48 million items, (2) a unified benchmark covering sequential recommendation, collaborative filtering, and product search, and (3) a new complex-query product search task featuring both semi-synthetic and real-world evaluation datasets. Experiments with 11 leading LLMs show that their rankings on BLaIR show little correlation with MTEB, highlighting the unique challenges of semantic encoding in recommendation.

hub tools

citation-role summary

dataset 4 background 3 method 2

citation-polarity summary

years

2026 54 2025 7

representative citing papers

SemCEB: A Cardinality Estimation Benchmark for Semantic Operators

cs.DB · 2026-06-22 · unverdicted · novelty 7.0

SemCEB is the first benchmark for cardinality estimation over semantic operators, evaluating sampling methods and Semantic Histograms on accuracy, cost, latency, and memory using 102 queries on a real-world dataset.

Breaking the Information Silo: Semantic Personas for Cross-Domain Recommendation

cs.IR · 2026-06-01 · unverdicted · novelty 7.0

SPHERE uses LLM-generated semantic personas and a dual-tower architecture to transfer recommendation knowledge across disjoint domains without shared users or items, showing consistent gains over NCF, SVD++, and LightGCN on Amazon Books, Goodreads, and Steam while linking transfer success to target

fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery

cs.LG · 2026-05-10 · conditional · novelty 7.0

fmxcoders improve cross-layer feature recovery in transformers via factorized weights and layer masking, delivering 10-30 point probing F1 gains, 25-50% lower MSE, doubled functional coherence, and 3-13x more coherent latents than standard crosscoders on GPT2-Small, Pythia, and Gemma2 models.

HORIZON: A Benchmark for In-the-wild User Behaviour Modeling

cs.IR · 2026-04-19 · unverdicted · novelty 7.0

HORIZON creates a cross-domain, long-horizon user modeling benchmark from Amazon Reviews that tests generalization across time, domains, and unseen users, exposing gaps in sequential and LLM-based recommendation models.

Multimodal Graph Negative Learning

cs.LG · 2026-06-11 · unverdicted · novelty 6.0

GraphMNL applies negative learning as cross-branch guidance in multimodal graphs to mitigate semantic imbalance without propagating bias from dominant branches.

citing papers explorer

Showing 50 of 61 citing papers.