Pith. sign in

hub Mixed citations

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Mixed citation behavior. Most common role is background (57%).

18 Pith papers citing it
Background 57% of classified citations

hub tools

citation-role summary

background 4 dataset 2 method 1

citation-polarity summary

representative citing papers

Cambrian-P: Pose-Grounded Video Understanding

cs.CV · 2026-05-21 · conditional · novelty 6.0

Adding per-frame camera-pose supervision to a video MLLM improves spatial and general video question answering by 2–6% and yields SOTA streaming pose estimates on ScanNet.

Building a Precise Video Language with Human-AI Oversight

cs.CV · 2026-04-22 · unverdicted · novelty 6.0

CHAI framework pairs AI pre-captions with expert human critiques to produce precise video descriptions, enabling open models to outperform closed ones like Gemini-3.1-Pro and improve fine-grained control in video generation models.

NVILA: Efficient Frontier Visual Language Models

cs.CV · 2024-12-05 · unverdicted · novelty 5.0

NVILA improves on VILA with a scale-then-compress visual token strategy and full-lifecycle efficiency optimizations, matching or exceeding leading VLMs on image and video benchmarks while reducing training cost 1.9-5.1x and latencies 1.2-2.8x.

LLaVA-OneVision: Easy Visual Task Transfer

cs.CV · 2024-08-06 · unverdicted · novelty 5.0

LLaVA-OneVision is the first single open LMM to simultaneously achieve strong performance in single-image, multi-image, and video scenarios with cross-scenario transfer capabilities.

citing papers explorer

Showing 18 of 18 citing papers.