Pith. sign in

REVIEW 28 cited by

Personalization of Large Language Models: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.00027 v3 pith:LM3TOY6R submitted 2024-10-29 cs.CL

classification cs.CL
keywords llmspersonalizationpersonalizedapplicationsusagechallengesexistingfacets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Personalization of Large Language Models (LLMs) has recently become increasingly important with a wide range of applications. Despite the importance and recent progress, most existing works on personalized LLMs have focused either entirely on (a) personalized text generation or (b) leveraging LLMs for personalization-related downstream applications, such as recommendation systems. In this work, we bridge the gap between these two separate main directions for the first time by introducing a taxonomy for personalized LLM usage and summarizing the key differences and challenges. We provide a formalization of the foundations of personalized LLMs that consolidates and expands notions of personalization of LLMs, defining and discussing novel facets of personalization, usage, and desiderata of personalized LLMs. We then unify the literature across these diverse fields and usage scenarios by proposing systematic taxonomies for the granularity of personalization, personalization techniques, datasets, evaluation methods, and applications of personalized LLMs. Finally, we highlight challenges and important open problems that remain to be addressed. By unifying and surveying recent research using the proposed taxonomies, we aim to provide a clear guide to the existing literature and different facets of personalization in LLMs, empowering both researchers and practitioners.

Discussion (0). Sign in to comment.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. User as Engram: Internalizing Per-User Memory as Local Parametric Edits

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    User facts are internalized as surgical local edits to a hash-keyed Engram memory table with reasoning skill held in a shared adapter, claimed to match LoRA recall, improve indirect reasoning 5.6x on average, and comp...

  2. VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    VitaBench 2.0 introduces a benchmark for long-term personalized and proactive agent behavior, with results indicating substantial gaps in current frontier LLMs.

  3. OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    OLIVIA treats LLM agent action selection as a contextual linear bandit over frozen hidden states and applies UCB exploration to adapt online, yielding consistent gains over static ReAct and prompt-based baselines on f...

  4. Skill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    CMIB uses a conditional multimodal information bottleneck to create reusable agent skills that separate verbalizable text content from predictive perceptual residuals, improving execution stability.

  5. Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

    cs.CL 2026-03 conditional novelty 7.0 of 10

    Personalized deep research systems need evaluation with real users because LLM judges overlook nuanced errors that matter to researchers.

  6. Tailored untruths: How personalisation challenges LLM safeguards

    cs.CL 2025-10 conditional novelty 7.0 of 10

    A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.

  7. PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing

    cs.CL 2026-07 conditional novelty 6.5 of 10

    A training-free four-stage LLM agent that routes papers into personal folksonomy folders by inspecting member papers and metadata, lifting Recall@1 from 0.39 to 0.61 on real libraries.

  8. Auditing Alignment Controllability in LLMs via Political Axes

    cs.CY 2026-07 conditional novelty 6.0 of 10

    On a 63,700-response Political Compass stress test of seven frontier LLMs, system-prompt framing dominates model identity, and steerability needs dispersion, symmetry, saturation, and refusal-floor metrics.

  9. Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging

    cs.CV 2026-07 conditional novelty 6.0 of 10

    PRAC mines preference-rich images and merges LoRA adapters from aesthetically similar users to achieve state-of-the-art personalized aesthetic rating prediction.

  10. FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A tile-based memory layout for mobile GPUs that unifies forward and backward data access, eliminating most transpose/reshape overhead and speeding LLM fine-tuning 2.2–5.7× in the paper's measurements.

  11. Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

    cs.CL 2026-07 unverdicted novelty 6.0 of 10

    Persona-driven generations by LLMs in MCQA tasks exhibit instability that differs systematically by model family, size, domain, and prompt format.

  12. Retrieval-Augmented Personalization with Foundation Models for Wearable Stress Detection

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Retrieval from out-of-domain foundation models enables personalization of a lightweight transformer for stress detection, yielding +3.92% accuracy and +4.76% F1 gains on WESAD without user labels.

  13. Personal Visual Context Learning in Large Multimodal Models

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Introduces Personal VCL formalization and benchmark revealing LMM context gaps, plus an Agentic Context Bank baseline that boosts personalized visual reasoning.

  14. Skill-R1: Agent Skill Evolution via Reinforcement Learning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Skill-R1 applies bi-level group-relative policy optimization to evolve skills recurrently from verified outcomes, yielding gains over baselines on multi-step tasks.

  15. Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    LLMs show mixed results on authorship verification, post generation, and attribute inference from Twitter data, with new frameworks and user studies establishing benchmarks for these analytics tasks.

  16. PersonaVLM: Long-Term Personalized Multimodal LLMs

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    PersonaVLM adds memory extraction, multi-turn retrieval-based reasoning, and personality inference to multimodal LLMs, yielding 22.4% gains on a new long-term personalization benchmark and outperforming GPT-4o.

  17. Synthetic Interaction Data for Scalable Personalization in Large Language Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    PersonaGym simulates noisy multi-turn user–assistant interactions to build PersonaAtlas, and PPOpt learns to rewrite user prompts from interaction history, improving judged personalization on synthetic benchmarks.

  18. No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A broad measurement study shows that false refusals in LLMs depend more on model and task choice than on sociodemographic personas, with newer models refusing far less.

  19. PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs

    cs.CL 2025-08 conditional novelty 6.0 of 10

    PREF is a reference-free, two-stage LLM judge that personalizes a quality rubric with a user profile and scores candidates against it, beating reminder-only baselines on the PrefEval implicit preference subset.

  20. Invisible Impact of Empathy on Behavioral Change: Isolating the Effect of Empathy in Long-term Physical Activity Coaching Chatbot Interactions

    cs.HC 2026-06 unverdicted novelty 5.0 of 10

    Higher-empathy chatbots were associated with larger step count increases and faster intention improvements despite users struggling to distinguish empathy levels and sometimes preferring the non-empathetic version.

  21. MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    MemSlides introduces a three-part memory hierarchy (user profile, working, tool) with scoped local revision for multi-turn personalized slide generation.

  22. High-Stakes Personalization: Rethinking LLM Customization for Individual Investor Decision-Making

    cs.CL 2026-04 conditional novelty 5.0 of 10

    Individual investing exposes four structural limits of LLM personalization—evolving contradictory memory, long-horizon thesis drift, style-vs-signal conflict, and no ground-truth labels—requiring new architectures bey...

  23. TiMem: Temporal-Hierarchical Memory Consolidation for Long-Horizon Conversational Agents

    cs.CL 2026-01 unverdicted novelty 5.0 of 10

    TiMem introduces a Temporal Memory Tree that consolidates conversational history into hierarchical persona representations, reaching 75.30% on LoCoMo and 76.88% on LongMemEval-S while cutting recalled length by 52%.

  24. Effects of Personality- and Opinion-Alignment in Human-AI Interaction

    cs.HC 2025-11 conditional novelty 5.0 of 10

    People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.

  25. Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents

    cs.HC 2025-10 conditional novelty 5.0 of 10

    A 3x3 between-subjects experiment finds that risk-contingent autonomy in LLM agents attenuates personalization's negative effects on privacy concerns and trust via increased perceived control.

  26. PrefReward: Learning User Preference Matrix for Personalized Text Generation

    cs.CL 2026-07 conditional novelty 4.0 of 10

    PrefReward selects the most style-aligned LLM output via a KL-divergence reward against an explicit user preference matrix, beating retrieval baselines on LongLaMP.

  27. SenseJudge: Human-Centric Preference-Driven Judgment Framework

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    SenseJudge proposes a human-preference-driven LLM judgment framework and SenseBench benchmark that claims to outperform prior methods in personalized judging and produce human-aligned model rankings.

  28. A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

    cs.AI 2025-07 accept novelty 4.0 of 10

    The paper delivers the first systematic review of self-evolving agents, structured around what components evolve, when adaptation occurs, and how it is implemented.

Pith tools