Pith. sign in

REVIEW 12 cited by

Tapping the Potential of Large Language Models as Recommender Systems: A Comprehensive Framework and Empirical Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.04997 v2 pith:YMXMWFUE submitted 2024-01-10 cs.IR

classification cs.IR
keywords llmsrecommendationanalysisanalyzeframeworkimpactlanguagerecommender
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, Large Language Models~(LLMs) such as ChatGPT have showcased remarkable abilities in solving general tasks, demonstrating the potential for applications in recommender systems. To assess how effectively LLMs can be used in recommendation tasks, our study primarily focuses on employing LLMs as recommender systems through prompting engineering. We propose a general framework for utilizing LLMs in recommendation tasks, focusing on the capabilities of LLMs as recommenders. To conduct our analysis, we formalize the input of LLMs for recommendation into natural language prompts with two key aspects, and explain how our framework can be generalized to various recommendation scenarios. As for the use of LLMs as recommenders, we analyze the impact of public availability, tuning strategies, model architecture, parameter scale, and context length on recommendation results based on the classification of LLMs. As for prompt engineering, we further analyze the impact of four important components of prompts, \ie task descriptions, user interest modeling, candidate items construction and prompting strategies. In each section, we first define and categorize concepts in line with the existing literature. Then, we propose inspiring research questions followed by detailed experiments on two public datasets, in order to systematically analyze the impact of different factors on performance. Based on our empirical analysis, we finally summarize promising directions to shed lights on future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Biases in LLM-Generated Musical Taste Profiles for Recommendation

    cs.IR 2025-07 conditional novelty 7.0 of 10

    Users identify more with LLM-generated music taste profiles for some genres and user groups than others, and these biases differ across models.

  2. PageLLM: A Multi-Grained Reward Framework for Whole-Page Optimization with Large Language Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A multi-grained reward framework fine-tunes an LLM with PPO to generate whole-page recommendations, showing that page-level and item-level reward heads are complementary.

  3. Tokenizing Numerical and Embedding Features for LLM RecSys

    cs.IR 2026-07 conditional novelty 5.0 of 10

    Interaction-based soft-token fusion of numerical and embedding features improves LLM two-tower retrieval over text-only and direct-concatenation baselines on three Amazon datasets.

  4. R4ec: A Reasoning, Reflection, and Refinement Framework for Recommendation Systems

    cs.IR 2025-07 conditional novelty 5.0 of 10

    R4ec trains a small reflection model to critique and refine LLM-generated user and item knowledge, which then improves downstream recommendation accuracy.

  5. Just Ask for Music (JAM): Multimodal and Personalized Natural Language Music Recommendation

    cs.IR 2025-07 conditional novelty 5.0 of 10

    JAM models natural-language music queries as vector translations from user to item embeddings, and with cross-attention fusion of audio, lyrics, and collaborative-filtering signals it outperforms the tested baselines ...

  6. Prompt Tuning for Item Cold-start Recommendation

    cs.IR 2024-12 reject novelty 5.0 of 10

    PROMO uses top positive-feedback users as item prompts with per-item prompt networks, reporting state-of-the-art cold-start recommendation, but the offline evaluation as written may leak the test label through the prompt.

  7. LIBER: Lifelong User Behavior Modeling Based on Large Language Models

    cs.IR 2024-11 conditional novelty 5.0 of 10

    LIBER partitions lifelong user behavior into fixed chunks, uses LLMs to summarize each chunk and detect interest shifts, and fuses these summaries to improve CTR prediction.

  8. Preserving Privacy and Utility in LLM-Based Product Recommendations

    cs.IR 2025-05 conditional novelty 4.0 of 10

    A hybrid system that filters sensitive purchases out of LLM-based recommendation prompts and generates those recommendations locally nearly matches full-data recommendation quality while keeping most sensitive data of...

  9. A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms

    cs.IR 2025-04 conditional novelty 4.0 of 10

    A survey that organizes foundation-model recommender systems into feature-based, generative, and agentic paradigms and reviews tasks, empirical results, and open challenges.

  10. Cold-Start Recommendation towards the Era of Large Language Models (LLMs): A Comprehensive Survey and Roadmap

    cs.IR 2025-01 conditional novelty 4.0 of 10

    A comprehensive survey and roadmap that groups cold-start recommendation methods into four knowledge scopes and defines nine cold-start problem types.

  11. Are Longer Prompts Always Better? Prompt Selection in Large Language Models for Recommendation Systems

    cs.IR 2024-12 conditional novelty 4.0 of 10

    No single LLM recommendation prompt is best across datasets; validation-based prompt selection with a ratio-based RPI indicator beats fixed prompts in most tested cases.

  12. Demystifying ChatGPT: How It Masters Genre Recognition

    cs.CL 2025-07 reject novelty 3.0 of 10

    A benchmark reports that ChatGPT outperforms other LLMs and traditional classifiers on multi-label movie genre prediction, but the result is undermined by likely pretraining contamination and weak baselines.

Pith tools