Pith. sign in

REVIEW

You Don't Know My Favorite Color: Preventing Dialogue Representations from Revealing Speakers' Private Personas

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.10228 v1 pith:ZE7WW4KY submitted 2022-04-26 cs.CL cs.CRcs.LG

classification cs.CLcs.CRcs.LG
keywords chatbotslanguagemodelsobjectivesaccuracydefensehiddenlarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Social chatbots, also known as chit-chat chatbots, evolve rapidly with large pretrained language models. Despite the huge progress, privacy concerns have arisen recently: training data of large language models can be extracted via model inversion attacks. On the other hand, the datasets used for training chatbots contain many private conversations between two individuals. In this work, we further investigate the privacy leakage of the hidden states of chatbots trained by language modeling which has not been well studied yet. We show that speakers' personas can be inferred through a simple neural network with high accuracy. To this end, we propose effective defense objectives to protect persona leakage from hidden states. We conduct extensive experiments to demonstrate that our proposed defense objectives can greatly reduce the attack accuracy from 37.6% to 0.5%. Meanwhile, the proposed objectives preserve language models' powerful generation ability.

Discussion (0). Sign in to comment.

Pith tools