Introduces MPT benchmark and PRefine method that models user preferences as evolving hypotheses to improve personalized tool calling accuracy with 1.24% of full-history token cost.
arXiv preprint arXiv:2601.02702 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Human data reveals LLMs struggle at extracting attributes from real conversations, selecting relevant ones, and generating responses humans rate no better than generic ones, with reward models showing only modest human correlation.
citing papers explorer
-
Latent Preference Modeling for Cross-Session Personalized Tool Calling
Introduces MPT benchmark and PRefine method that models user preferences as evolving hypotheses to improve personalized tool calling accuracy with 1.24% of full-history token cost.
-
Re-Centering Humans in LLM Personalization
Human data reveals LLMs struggle at extracting attributes from real conversations, selecting relevant ones, and generating responses humans rate no better than generic ones, with reward models showing only modest human correlation.