Pith. sign in

REVIEW 2 cited by

StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.00304 v2 pith:76U4PEG4 submitted 2024-08-31 cs.CV

classification cs.CV
keywords reasoningstimuvarvideoaffectiveawarenessemotionalmllmsspatiotemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Predicting and reasoning how a video would make a human feel is crucial for developing socially intelligent systems. Although Multimodal Large Language Models (MLLMs) have shown impressive video understanding capabilities, they tend to focus more on the semantic content of videos, often overlooking emotional stimuli. Hence, most existing MLLMs fall short in estimating viewers' emotional reactions and providing plausible explanations. To address this issue, we propose StimuVAR, a spatiotemporal Stimuli-aware framework for Video Affective Reasoning (VAR) with MLLMs. StimuVAR incorporates a two-level stimuli-aware mechanism: frame-level awareness and token-level awareness. Frame-level awareness involves sampling video frames with events that are most likely to evoke viewers' emotions. Token-level awareness performs tube selection in the token space to make the MLLM concentrate on emotion-triggered spatiotemporal regions. Furthermore, we create VAR instruction data to perform affective training, steering MLLMs' reasoning strengths towards emotional focus and thereby enhancing their affective reasoning ability. To thoroughly assess the effectiveness of VAR, we provide a comprehensive evaluation protocol with extensive metrics. StimuVAR is the first MLLM-based method for viewer-centered VAR. Experiments demonstrate its superiority in understanding viewers' emotional responses to videos and providing coherent and insightful explanations. Our code is available at https://github.com/EthanG97/StimuVAR

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Multimodal LLMs underperform humans at directly rating charts' experiential impact, but they are substantially better at pairwise comparisons, especially when the human ratings differ clearly.

  2. Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Anomaly-OV, trained on the new Anomaly-Instruct-125k dataset, improves zero-shot detection of image anomalies and their textual explanations over generalist MLLMs.

Pith tools