Pith. sign in

REVIEW 7 cited by

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.14106 v2 pith:VOVGTRYI submitted 2025-05-20 cs.CL cs.AI

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

classification cs.CL cs.AI
keywords personalizedbenchmarkconversationalllmspersonaconvbenchclassificationcontextconversations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We present PersonaConvBench, a large-scale benchmark for evaluating personalized reasoning and generation in multi-turn conversations with large language models (LLMs). Unlike existing work that focuses on either personalization or conversational structure in isolation, PersonaConvBench integrates both, offering three core tasks: sentence classification, impact regression, and user-centric text generation across ten diverse Reddit-based domains. This design enables systematic analysis of how personalized conversational context shapes LLM outputs in realistic multi-user scenarios. We benchmark several commercial and open-source LLMs under a unified prompting setup and observe that incorporating personalized history yields substantial performance improvements, including a 198 percent relative gain over the best non-conversational baseline in sentiment classification. By releasing PersonaConvBench with evaluations and code, we aim to support research on LLMs that adapt to individual styles, track long-term context, and produce contextually rich, engaging responses.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking

    cs.LG 2026-05 unverdicted novelty 7.0

    F-GRPO factorizes group-relative policy optimization into generation and ranking phases within one autoregressive sequence, using order-invariant coverage and position-aware utility rewards to improve top-ranked perfo...

  2. Agent Safety Is Action Alignment

    cs.AI 2026-06 unverdicted novelty 6.0

    Agent safety cannot be achieved via model refusal training and instead requires external least-privilege enforcement evaluated as action alignment.

  3. Benchmarking the Personalization Capabilities of Large Language Models

    cs.AI 2026-05 conditional novelty 6.0

    Across a new sales-outreach benchmark (SDR-Bench), frontier LLMs recover at most ~55% of the strategic pitch points from human-authored deal-winning messages, with no model statistically separating successful from uns...

  4. TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation

    cs.IR 2026-04 unverdicted novelty 6.0

    TimeMM proposes a time-as-operator spectral filtering framework with adaptive mixing and modality routing to model non-stationary multimodal user preferences in recommendation systems.

  5. A Survey on LLM-based Conversational User Simulation

    cs.CL 2026-04 unverdicted novelty 6.0

    A survey that introduces a taxonomy for LLM-based conversational user simulation, analyzes core techniques and evaluation methods, and identifies open challenges in the field.

  6. Cat-DPO: Category-Adaptive Safety Alignment

    cs.CL 2026-04 unverdicted novelty 6.0

    Cat-DPO applies per-category adaptive safety margins during direct preference optimization to reduce variance in safety across harm categories.

  7. Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

    cs.AI 2026-02 conditional novelty 5.0

    BAO, a behavior-enhanced SFT plus regularized RL pipeline, improves proactive agents' task performance while lowering user-involvement rate, beating UserRL baselines on three UserRL gym tasks.