Pith. sign in

REVIEW 6 cited by

Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08046 v3 pith:YH6OHZOM submitted 2023-11-14 cs.CV

classification cs.CV
keywords visualchat-univiimagesvideosrepresentationmodeltokensunified
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations. However, existing methods encounter challenges in effectively handling both image and video understanding, particularly with limited visual tokens. In this work, we introduce Chat-UniVi, a Unified Vision-language model capable of comprehending and engaging in conversations involving images and videos through a unified visual representation. Specifically, we employ a set of dynamic visual tokens to uniformly represent images and videos. This representation framework empowers the model to efficiently utilize a limited number of visual tokens to simultaneously capture the spatial details necessary for images and the comprehensive temporal relationship required for videos. Moreover, we leverage a multi-scale representation, enabling the model to perceive both high-level semantic concepts and low-level visual details. Notably, Chat-UniVi is trained on a mixed dataset containing both images and videos, allowing direct application to tasks involving both mediums without requiring any modifications. Extensive experimental results demonstrate that Chat-UniVi consistently outperforms even existing methods exclusively designed for either images or videos. Code is available at https://github.com/PKU-YuanGroup/Chat-UniVi.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

    cs.AI 2026-07 conditional novelty 6.0 of 10

    EmoAgent-R1 combines dynamic agent routing with a token-reweighted GRPO variant (P-GRPO) to reach 77.85% mean on MER-UniBench, exceeding AffectGPT-R1 by 1.90 points.

  2. EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    Separating temporal grounding from answer reasoning, plus low-confidence full-video re-reading, modestly improves long-video QA on Qwen3-VL across five benchmarks.

  3. GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Introduces GLIMPSE, a video-QA benchmark whose questions cannot be answered from single frames; best model GPT-o3 scores 66.43% vs 94.82% human accuracy.

  4. TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Weakly supervised vision-language model jointly generating open-ended video QA answers with temporal groundings, reporting SOTA on NExT-GQA, MSVD-QA, and ActivityNet-QA.

  5. Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning

    cs.CV 2026-03 conditional novelty 5.5 of 10

    Goal-driven selection of 1× multimodal instruction subsets reaches a 512k Uni-10x baseline after ~27–35k samples and improves accuracy by up to +3.08 pp under a fixed Qwen3-VL recipe.

  6. LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs

    cs.CV 2025-06 conditional novelty 4.0 of 10

    LeanPO improves Video-LLM alignment by using a reference-free average-likelihood reward, self-generated winning/losing pairs, and dynamic label smoothing, yielding gains on six video benchmarks.

Pith tools