Pith. sign in

REVIEW

MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2602.24188 v2 pith:QU27U3ZN submitted 2026-02-27 cs.CL cs.LG

classification cs.CLcs.LG
keywords informationmodelscollaborativelanguagemulti-turnprivateagentcollaboration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a scalable and verifiable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games that require effective communication about private information. This enables an interactive scaling analysis, in which a fixed token budget is divided over a variable number of turns. We find that language models often fail to use interactive collaboration to improve over the non-interactive baseline in which one agent summarizes its information and the other agent immediately acts, despite substantial headroom. This suggests that state-of-the-art models still suffer from significant weaknesses in planning and executing multi-turn collaborative conversations. We analyze the linguistic features of these dialogues, assessing the roles of sycophancy, information density, and discourse coherence. While there is no single linguistic explanation for the collaborative weaknesses of contemporary language models, we note that humans achieve comparable task success at superior token efficiency by producing more coherent dialogues. The proactive management of private information is a defining feature of real-world communication, and this work is designed to drive further progress on this capability.

Discussion (0). Continue with ORCID to comment.

Pith tools