A 96-dyad benchmark with matched human ratings shows multimodal LLMs match human crowd accuracy on familiarity inference but rely on a stranger response bias and underuse visible behavior.
``Was That Your Mother on the Phone?'': Classifying Interpersonal Relationships between Dialog Participants with Lexical and Acoustic Properties , shorttitle =
1 Pith paper cite this work, alongside 4 external citations. Polarity classification is still indexing.
1
Pith paper citing it
4
external citations · OpenAlex
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
A 96-dyad benchmark with matched human ratings shows multimodal LLMs match human crowd accuracy on familiarity inference but rely on a stranger response bias and underuse visible behavior.