X-Stream benchmark shows SOTA MLLMs score ~50% on concurrent multi-stream tasks and lack proactive ability, using a dual-verification pipeline to avoid single-stream bias.
Video action differencing
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 3verdicts
UNVERDICTED 3representative citing papers
New benchmark diagnoses directional, attributional, and temporal hallucinations in multimodal motion comparison models and demonstrates gains from explicit measurement verification.
A gaze-only student model distilled from a joint gaze-video teacher achieves high skill-assessment accuracy using 73x less power than prior methods.
citing papers explorer
-
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
X-Stream benchmark shows SOTA MLLMs score ~50% on concurrent multi-stream tasks and lack proactive ability, using a dual-verification pipeline to avoid single-stream bias.
-
MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning
New benchmark diagnoses directional, attributional, and temporal hallucinations in multimodal motion comparison models and demonstrates gains from explicit measurement verification.
-
SkillSight: Efficient First-Person Skill Assessment with Gaze
A gaze-only student model distilled from a joint gaze-video teacher achieves high skill-assessment accuracy using 73x less power than prior methods.