Pith. sign in

REVIEW 8 cited by

Tapping the Potential of Large Language Models as Recommender Systems: A Comprehensive Framework and Empirical Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.04997 v2 pith:YMXMWFUE submitted 2024-01-10 cs.IR

Tapping the Potential of Large Language Models as Recommender Systems: A Comprehensive Framework and Empirical Analysis

classification cs.IR
keywords llmsrecommendationanalysisanalyzeframeworkimpactlanguagerecommender
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recently, Large Language Models~(LLMs) such as ChatGPT have showcased remarkable abilities in solving general tasks, demonstrating the potential for applications in recommender systems. To assess how effectively LLMs can be used in recommendation tasks, our study primarily focuses on employing LLMs as recommender systems through prompting engineering. We propose a general framework for utilizing LLMs in recommendation tasks, focusing on the capabilities of LLMs as recommenders. To conduct our analysis, we formalize the input of LLMs for recommendation into natural language prompts with two key aspects, and explain how our framework can be generalized to various recommendation scenarios. As for the use of LLMs as recommenders, we analyze the impact of public availability, tuning strategies, model architecture, parameter scale, and context length on recommendation results based on the classification of LLMs. As for prompt engineering, we further analyze the impact of four important components of prompts, \ie task descriptions, user interest modeling, candidate items construction and prompting strategies. In each section, we first define and categorize concepts in line with the existing literature. Then, we propose inspiring research questions followed by detailed experiments on two public datasets, in order to systematically analyze the impact of different factors on performance. Based on our empirical analysis, we finally summarize promising directions to shed lights on future research.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees

    cs.LG 2026-05 unverdicted novelty 7.0

    InvEvolve evolves white-box inventory policies from LLMs with statistical safety guarantees and outperforms classical and deep learning methods on synthetic and real retail data.

  2. InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees

    cs.LG 2026-05 unverdicted novelty 6.0

    InvEvolve evolves inventory policies using LLMs with RL and provides statistical safety guarantees, outperforming classical and DL methods on synthetic and real data.

  3. InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees

    cs.LG 2026-05 unverdicted novelty 6.0

    InvEvolve uses LLMs and RL to generate certified inventory policies that outperform classical and deep learning methods on synthetic and real data while providing multi-period performance guarantees.

  4. GraphRAG-IRL: Personalized Recommendation with Graph-Grounded Inverse Reinforcement Learning and LLM Re-ranking

    cs.IR 2026-04 unverdicted novelty 6.0

    GraphRAG-IRL fuses graph-grounded MaxEnt IRL pre-ranking with persona-guided LLM re-ranking to deliver up to 16.8% NDCG@10 gains over IRL-only baselines on MovieLens and consistent 4-6% gains on KuaiRand.

  5. Tokenizing Numerical and Embedding Features for LLM RecSys

    cs.IR 2026-07 conditional novelty 5.0

    Interaction-based soft-token fusion of numerical and embedding features improves LLM two-tower retrieval over text-only and direct-concatenation baselines on three Amazon datasets.

  6. Tokenizing Numerical and Embedding Features for LLM RecSys

    cs.IR 2026-07 conditional novelty 5.0

    Interaction-based soft-token fusion of numerical and embedding features beats text-only and direct-concatenation LLM two-tower recommenders on Beauty, Sports and Outdoors, and Toys and Games.

  7. Tokenizing Numerical and Embedding Features for LLM RecSys

    cs.IR 2026-07 conditional novelty 5.0

    A soft-token fusion framework that turns price and dense-embedding features into LLM-readable tokens improves retrieval in an LLM-based two-tower recommender, and interaction-based fusion beats concatenation on three ...

  8. From Hidden Profiles to Governable Personalization: Recommender Systems in the Age of LLM Agents

    cs.IR 2026-04 unverdicted novelty 5.0

    LLM agents enable a shift in recommender systems from opaque hidden profiles to governable, inspectable, and portable user representations.