Trip+ benchmark evaluates language model agents on generating and revising personalized minute-level travel itineraries under dynamic interactions, finding consistent gaps where models produce feasible but exhausting plans that ignore traveler profiles.
arXiv preprint arXiv:2602.16173 , year=
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5verdicts
UNVERDICTED 5representative citing papers
Proposes PARPO for reward decoupling with user anchors and PSGM for preference-aligned skill memory in personalized agentic RL, reporting outperformance on ETAPP benchmarks.
Hedwig is a coding agent that dynamically adjusts its autonomy by learning behavioral guidelines from developer decisions and feedback over time.
AgentClick is a localhost npm server and skill-based plugin that connects terminal AI agents to a structured web UI for human review of plans, code execution, memory, and errors.
Position paper advocating personalized preference learning in LLMs over aggregated approaches, grounded in social choice theory and demographic variation.
citing papers explorer
-
Trip+: Benchmarking Agents in Personalized Interactive Travel Planning
Trip+ benchmark evaluates language model agents on generating and revising personalized minute-level travel itineraries under dynamic interactions, finding consistent gaps where models produce feasible but exhausting plans that ignore traveler profiles.
-
From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning
Proposes PARPO for reward decoupling with user anchors and PSGM for preference-aligned skill memory in personalized agentic RL, reporting outperformance on ETAPP benchmarks.
-
Hedwig: Dynamic Autonomy for Coding Agents Under Local Oversight
Hedwig is a coding agent that dynamically adjusts its autonomy by learning behavioral guidelines from developer decisions and feedback over time.
-
AgentClick: A Skill-Based Human-in-the-Loop Review Layer for Terminal AI Agents
AgentClick is a localhost npm server and skill-based plugin that connects terminal AI agents to a structured web UI for human review of plans, code execution, memory, and errors.
-
Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences
Position paper advocating personalized preference learning in LLMs over aggregated approaches, grounded in social choice theory and demographic variation.