Pith. sign in

REVIEW 10 cited by

LaMP: When Large Language Models Meet Personalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.11406 v4 pith:SYYPW6VF submitted 2023-04-22 cs.CL

classification cs.CL
keywords languagemodelslamptaskspersonalizationretrievalaugmentationbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper highlights the importance of personalization in large language models and introduces the LaMP benchmark -- a novel benchmark for training and evaluating language models for producing personalized outputs. LaMP offers a comprehensive evaluation framework with diverse language tasks and multiple entries for each user profile. It consists of seven personalized tasks, spanning three text classification and four text generation tasks. We additionally propose two retrieval augmentation approaches that retrieve personal items from each user profile for personalizing language model outputs. To this aim, we study various retrieval models, including term matching, semantic matching, and time-aware methods. Extensive experiments on LaMP for zero-shot and fine-tuned language models demonstrate the efficacy of the proposed retrieval augmentation approach and highlight the impact of personalization in various natural language tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

    cs.CL 2026-07 conditional novelty 6.5 of 10

    On 800 jointly constrained trip tasks with a deterministic scorer and achievable gold, the best of 15 LLM agents fully solves only 46.2% of feasible plans, with unstated persona needs as the universal bottleneck.

  2. Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

    cs.LG 2026-07 reject novelty 6.0 of 10

    IRIS learns and iteratively refines natural-language user personas from implicit interaction streams; on 100 Reddit AITA commenters it predicts held-out verdicts at 61% accuracy, within a statistically untested 56-61% band.

  3. Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory

    cs.AI 2026-07 accept novelty 6.0 of 10

    Consolidation beats retention for language-agent memory only when raw evidence exceeds the token budget; LongMemEval and LoCoMo show a budget-dependent crossover, with Abstract/Merge preferred under compression.

  4. From Profiles to Steering Vectors: Global Sparse Priors and Local Semantic Calibration for Personalized Text Generation

    cs.AI 2026-06 conditional novelty 6.0 of 10

    GLASS personalizes text generation by injecting global and local style vectors—built from sparse autoencoder features of a user's history—into different layers of a frozen LLM, improving ROUGE and judge ratings over p...

  5. Evaluating Style-Personalized Text Generation: Challenges and Directions

    cs.CL 2025-08 reject novelty 6.0 of 10

    A new style-discrimination benchmark for personalized text generation shows ensemble metrics give only a marginal, possibly test-fitted, edge over the best single judge.

  6. PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs

    cs.CL 2025-08 conditional novelty 6.0 of 10

    PREF is a reference-free, two-stage LLM judge that personalizes a quality rubric with a user profile and scores candidates against it, beating reminder-only baselines on the PrefEval implicit preference subset.

  7. Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs

    cs.CL 2026-08 conditional novelty 5.0 of 10

    On consumer GPUs, LoRA+ gives the best energy-focused fine-tuning score in 19 of 24 small-model task configurations, while QLoRA wins the memory-focused score when peak VRAM is the binding constraint.

  8. "Pragmatic Tools or Empowering Friends?" Discovering and Co-Designing Personality-Aligned AI Writing Companions

    cs.HC 2025-09 conditional novelty 5.0 of 10

    Writers grouped into four MBTI-based profiles showed divergent preferences for AI writing companion features, demonstrated by two contrasting prototypes in a small proof-of-concept study.

  9. Personas within Parameters: Fine-Tuning Small Language Models with Low-Rank Adapters to Mimic User Behaviors

    cs.IR 2025-08 conditional novelty 5.0 of 10

    Persona-level LoRA fine-tuning lets a 3.8B small language model simulate MovieLens users about as accurately as a much larger frozen LLM, at lower cost.

  10. AI Agent Behavioral Science

    q-bio.NC 2025-06 conditional novelty 4.0 of 10

    AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.

Pith tools