REVIEW 2 cited by
User-Specific Dialogue Generation with User Profile-Aware Pre-Training Model and Parameter-Efficient Fine-Tuning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper addresses user-specific dialogs. In contrast to previous research on personalized dialogue focused on achieving virtual user dialogue as defined by persona descriptions, user-specific dialogue aims to reproduce real-user dialogue beyond persona-based dialogue. Fine-tuning using the target user's dialogue history is an efficient learning method for a user-specific model. However, it is prone to overfitting and model destruction due to the small amount of data. Therefore, we propose a learning method for user-specific models by combining parameter-efficient fine-tuning with a pre-trained dialogue model that includes user profiles. Parameter-efficient fine-tuning adds a small number of parameters to the entire model, so even small amounts of training data can be trained efficiently and are robust to model destruction. In addition, the pre-trained model, which is learned by adding simple prompts for automatically inferred user profiles, can generate speech with enhanced knowledge of the user's profile, even when there is little training data during fine-tuning. In experiments, we compared the proposed model with large-language-model utterance generation using prompts containing users' personal information. Experiments reproducing real users' utterances revealed that the proposed model can generate utterances with higher reproducibility than the compared methods, even with a small model.
Forward citations
Cited by 2 Pith papers
-
Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning
AlignXada uses verbal reinforcement learning to learn reusable text-rewriting policies that compress universal user preference profiles into task-specific ones, improving downstream personalization accuracy on most te...
-
Personalized LLM for Generating Customized Responses to the Same Query from Different Users
A dual-tower LLM with a low-rank querier-specific encoder and cluster-restricted contrastive learning generates responses tailored to the person asking, evaluated on a new 173-querier multi-source dialogue dataset.
Discussion (0). Continue with ORCID to comment.