REVIEW 18 cited by
LongLaMP: A Benchmark for Personalized Long-form Text Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Long-text generation is seemingly ubiquitous in real-world applications of large language models such as generating an email or writing a review. Despite the fundamental importance and prevalence of long-text generation in many practical applications, existing work on personalized generation has focused on the generation of very short text. To overcome these limitations, we study the problem of personalized long-text generation, that is, generating long-text that is personalized for a specific user while being practically useful for the vast majority of real-world applications that naturally require the generation of longer text. In this work, we demonstrate the importance of user-specific personalization for long-text generation tasks and develop the Long-text Language Model Personalization (LongLaMP) Benchmark. LongLaMP provides a comprehensive and diverse evaluation framework for personalized long-text generation. Extensive experiments on LongLaMP for zero-shot and fine-tuned language tasks demonstrate the effectiveness of the proposed benchmark and its utility for developing and evaluating techniques for personalized long-text generation across a wide variety of long-text generation tasks. The results highlight the importance of personalization across a wide variety of long-text generation tasks. Finally, we release the benchmark for others to use for this important problem.
Forward citations
Cited by 18 Pith papers
-
ClawRec: A Claw-Native Recommender System
ClawRec turns cross-platform behavior into a temporally managed user state and role-aware complementary slates, beating agentic baselines on a new synthetic life-event benchmark.
-
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
A new human-centric agentic AI paradigm, Combodied Agents, organizes perception, memory, prediction, and intervention around the evolving human state and agency over time.
-
Know It, Act on It: Investigating Memory Utilization in LLM Personalization
LLM agents often pass a direct recall question about a user's preference yet fail to act on the same preference in a realistic request — a Know–Act gap that persists even in the best systems and is widest, on average,...
-
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
Persona2Web is a new open-web benchmark where agents must infer a user's preferences from synthetic browsing history to solve intentionally ambiguous queries; current best agents score 13% success.
-
Evaluating Style-Personalized Text Generation: Challenges and Directions
A new style-discrimination benchmark for personalized text generation shows ensemble metrics give only a marginal, possibly test-fitted, edge over the best single judge.
-
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
PREF is a reference-free, two-stage LLM judge that personalizes a quality rubric with a user profile and scores candidates against it, beating reminder-only baselines on the PrefEval implicit preference subset.
-
From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation
ProxyReward trains long-form generation models by rewarding how well an AI judge can answer generated yes/no questions about the response, improving open-source models on ProxyQA.
-
ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation
ExPerT is a reference-based LLM evaluation metric that extracts and matches atomic aspects, scores content and style, and reports 0.74 human alignment on LongLaMP, a 7.2% relative gain over GEMBA and G-Eval.
-
Persona-SQ: A Personalized Suggested Question Generation Framework For Real-world Documents
Persona-based prompting plus quality filtering yields more diverse suggested questions, and synthetic data from the pipeline trains a 360M model to near-GPT-4o levels.
-
Accelerating Retrieval-Augmented Generation
Exact nearest neighbor search, accelerated by a near-memory CXL device called IKS, can make retrieval-augmented generation faster and more accurate end-to-end than approximate search.
-
On the Way to LLM Personalization: Learning to Remember User Conversations
Finetuning a LoRA adapter on self-generated question-answer pairs lets Llama 3 8B recall conversation topics with 81.5% accuracy, close to RAG at 83.5% but without retrieval.
-
CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs
CURP represents users as sparse combinations of discrete prototype codebook embeddings and uses them as frozen-LLM prefixes, outperforming personalization baselines on four text-generation tasks with about 20M trainab...
-
Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation
REST-PG trains LLMs to reason over user profiles and self-train on high-reward outputs, improving personalized long-form generation by 14.5% over SFT on LongLaMP.
-
PrefReward: Learning User Preference Matrix for Personalized Text Generation
PrefReward selects the most style-aligned LLM output via a KL-divergence reward against an explicit user preference matrix, beating retrieval baselines on LongLaMP.
-
Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs
PsPLUG, a soft-prompt plug-in trained with style-conditioned preference pairs, preserves user identity under explicit style instructions and lets users tune personalization strength via an α scalar.
-
Position: It's Time to Act on the Risk of Efficient Personalized Text Generation
Fine-tuned open LLMs can imitate individual writing styles from small samples, evade detection tools, and are not yet addressed by current safeguards or law.
-
Personalized Graph-Based Retrieval for Large Language Models
PGraphRAG adds neighbor-user reviews to LLM prompts and claims improved personalized generation, but its own ablations show the user's history contributes little beyond item context.
-
Personalized Multimodal Large Language Models: A Survey
The paper provides a survey and taxonomy of personalization techniques for multimodal LLMs across text generation, image generation, recommendation, and retrieval.
Discussion (0). Continue with ORCID to comment.