Pith. sign in

REVIEW 4 cited by

Are Longer Prompts Always Better? Prompt Selection in Large Language Models for Recommendation Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.14454 v1 pith:PFEH4MAD submitted 2024-12-19 cs.IR cs.CL

classification cs.IRcs.CL
keywords promptspromptrecommendationaccuracydatallm-rssselectionlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In large language models (LLM)-based recommendation systems (LLM-RSs), accurately predicting user preferences by leveraging the general knowledge of LLMs is possible without requiring extensive training data. By converting recommendation tasks into natural language inputs called prompts, LLM-RSs can efficiently solve issues that have been difficult to address due to data scarcity but are crucial in applications such as cold-start and cross-domain problems. However, when applying this in practice, selecting the prompt that matches tasks and data is essential. Although numerous prompts have been proposed in LLM-RSs and representing the target user in prompts significantly impacts recommendation accuracy, there are still no clear guidelines for selecting specific prompts. In this paper, we categorize and analyze prompts from previous research to establish practical prompt selection guidelines. Through 450 experiments with 90 prompts and five real-world datasets, we examined the relationship between prompts and dataset characteristics in recommendation accuracy. We found that no single prompt consistently outperforms others; thus, selecting prompts on the basis of dataset characteristics is crucial. Here, we propose a prompt selection method that achieves higher accuracy with minimal validation data. Because increasing the number of prompts to explore raises costs, we also introduce a cost-efficient strategy using high-performance and cost-efficient LLMs, significantly reducing exploration costs while maintaining high prediction accuracy. Our work offers valuable insights into the prompt selection, advancing accurate and efficient LLM-RSs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AutoData: A Multi-Agent System for Open Web Data Collection

    cs.IR 2025-05 conditional novelty 6.0 of 10

    AutoData, a multi-agent system with a hypergraph message cache, automates web dataset collection from a sentence instruction and outperforms general agent baselines on the new Instruct2DS benchmark.

  2. CTG-Insight: A Multi-Agent Interpretable LLM Framework for Cardiotocography Analysis and Classification

    cs.LG 2025-07 conditional novelty 5.0 of 10

    CTG-Insight reports 96.4% accuracy on a 50-sample subset of the NeuroFetalNet test set, but the comparison against deep learning baselines is uneven and the system code and exact prompts are not released.

  3. Unravelling the Probabilistic Forest: Arbitrage in Prediction Markets

    cs.CR 2025-08 unverdicted novelty 4.0 of 10

    Claims two forms of Polymarket arbitrage and $40 million extracted profit, but the body text provided is an unrelated plasma physics paper.

  4. Fine-tuning on simulated data outperforms prompting for agent tone of voice

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Fine-tuning a 1B-parameter LLM on as few as 100 synthetically generated, readability-filtered samples achieved conversational tone more reliably than a verbose system prompt.

Pith tools