Pith. sign in

REVIEW 1 cited by

BQA: Body Language Question Answering Dataset for Video Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13206 v3 pith:54Z3ZJEH submitted 2024-10-17 cs.CL

classification cs.CL
keywords languagebodydatasetvideollmslargevideoansweringanswers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A large part of human communication relies on nonverbal cues such as facial expressions, eye contact, and body language. Unlike language or sign language, such nonverbal communication lacks formal rules, requiring complex reasoning based on commonsense understanding. Enabling current Video Large Language Models (VideoLLMs) to accurately interpret body language is a crucial challenge, as human unconscious actions can easily cause the model to misinterpret their intent. To address this, we propose a dataset, BQA, a body language question answering dataset, to validate whether the model can correctly interpret emotions from short clips of body language comprising 26 emotion labels of videos of body language. We evaluated various VideoLLMs on BQA and revealed that understanding body language is challenging, and our analyses of the wrong answers by VideoLLMs show that certain VideoLLMs made significantly biased answers depending on the age group and ethnicity of the individuals in the video. The dataset is available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Using a new LLC Score, translation and geo-culture cloze experiments on three small multilingual LLMs show that latent-language consistency does not reliably predict task accuracy, contradicting the paper's initial hy...

Pith tools