Pith. sign in

REVIEW 3 cited by

IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07577 v1 pith:DUBHSILN submitted 2024-07-10 cs.CV cs.AI

classification cs.CVcs.AI
keywords visuallvlmsacrosslargeunderstandingvision-languagecomplexcontent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid advancement of Large Vision-Language models (LVLMs) has demonstrated a spectrum of emergent capabilities. Nevertheless, current models only focus on the visual content of a single scenario, while their ability to associate instances across different scenes has not yet been explored, which is essential for understanding complex visual content, such as movies with multiple characters and intricate plots. Towards movie understanding, a critical initial step for LVLMs is to unleash the potential of character identities memory and recognition across multiple visual scenarios. To achieve the goal, we propose visual instruction tuning with ID reference and develop an ID-Aware Large Vision-Language Model, IDA-VLM. Furthermore, our research introduces a novel benchmark MM-ID, to examine LVLMs on instance IDs memory and recognition across four dimensions: matching, location, question-answering, and captioning. Our findings highlight the limitations of existing LVLMs in recognizing and associating instance identities with ID reference. This paper paves the way for future artificial intelligence systems to possess multi-identity visual inputs, thereby facilitating the comprehension of complex visual narratives like movies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MentalThink: Shaping Thoughts in Mental SVG World

    cs.AI 2026-07 conditional novelty 7.0 of 10

    MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.

  2. I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A new person-centric video reasoning benchmark and 7B model that link a reference image of a person to their appearances and actions in a video.

  3. Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Prompt-A-Video refines text prompts for video diffusion models via evolutionary search and DPO alignment, improving generated video quality on Open-Sora and CogVideoX.

Pith tools