A 96-dyad benchmark with matched human ratings shows multimodal LLMs match human crowd accuracy on familiarity inference but rely on a stranger response bias and underuse visible behavior.
Psychological Methods , volume =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
A 96-dyad benchmark with matched human ratings shows multimodal LLMs match human crowd accuracy on familiarity inference but rely on a stranger response bias and underuse visible behavior.