Pith. sign in

hub

Otterhd: A high- resolution multi-modality model

13 Pith papers cite this work, alongside 6 external citations. Polarity classification is still indexing.

13 Pith papers citing it
6 external citations · external index

hub tools

citation-role summary

background 4

citation-polarity summary

fields

cs.CV 13

roles

background 4

polarities

background 3 unclear 1

representative citing papers

Otter: A Multi-Modal Model with In-Context Instruction Tuning

cs.CV · 2023-05-05 · unverdicted · novelty 6.0

Otter is a multi-modal model instruction-tuned on the MIMIC-IT dataset of over 3 million in-context instruction-response pairs to improve convergence and generalization on tasks with multiple images and videos.

Make Your LVLM KV Cache More Lightweight

cs.CV · 2026-05-01 · unverdicted · novelty 5.0

LightKV compresses vision-token KV cache in LVLMs to 55% size via prompt-guided cross-modality aggregation, halving cache memory, cutting compute 40%, and maintaining performance on benchmarks.

Qwen2.5-VL Technical Report

cs.CV · 2025-02-19 · unverdicted · novelty 5.0

Qwen2.5-VL reports a vision-language model family using native dynamic-resolution ViT and absolute time encoding that matches GPT-4o on document and diagram tasks while supporting hour-long videos with second-level localization.

citing papers explorer

Showing 13 of 13 citing papers.