Pith. sign in

REVIEW 10 cited by

OtterHD: A High-Resolution Multi-modality Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.04219 v1 pith:UQHG5HRZ submitted 2023-11-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelshigh-resolutionmodelotterhd-8bvisualabilityencodersinput
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we present OtterHD-8B, an innovative multimodal model evolved from Fuyu-8B, specifically engineered to interpret high-resolution visual inputs with granular precision. Unlike conventional models that are constrained by fixed-size vision encoders, OtterHD-8B boasts the ability to handle flexible input dimensions, ensuring its versatility across various inference requirements. Alongside this model, we introduce MagnifierBench, an evaluation framework designed to scrutinize models' ability to discern minute details and spatial relationships of small objects. Our comparative analysis reveals that while current leading models falter on this benchmark, OtterHD-8B, particularly when directly processing high-resolution inputs, outperforms its counterparts by a substantial margin. The findings illuminate the structural variances in visual information processing among different models and the influence that the vision encoders' pre-training resolution disparities have on model effectiveness within such benchmarks. Our study highlights the critical role of flexibility and high-resolution input capabilities in large multimodal models and also exemplifies the potential inherent in the Fuyu architecture's simplicity for handling complex visual data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Frontier VLMs perform poorly on VisualOverload, a new dense-scene VQA benchmark, with the best model scoring 69.5% overall and only 19.6% on questions that models collectively found hardest.

  2. LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LLaVA-SP adds six spatial tokens produced by multi-scale cropping or pooling and cross-attention to MLLMs, improving 10/11 benchmarks over LLaVA-1.5 with nearly unchanged latency.

  3. LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A hierarchical window transformer that injects image details into upsampled CLIP features and compresses them with cross-scale window attention improves MLLM fine-grained perception by 3.7% on average over LLaVA-UHD.

  4. Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Compressing visual tokens and suppressing attention to irrelevant image regions improves instruction-following in vision-language models with little loss on visual understanding benchmarks.

  5. FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FocusLLaVA compresses visual tokens to 39% using a vision-guided region sampler plus a text-guided attention sampler, beating its LLaVA-NeXT baseline on 9 of 10 benchmarks.

  6. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  7. Visual Instruction Tuning with Chain of Region-of-Interest

    cs.CV 2025-05 reject novelty 5.0 of 10

    CoRoI injects a chain of language-guided image regions into LLM hidden layers and reports improved MLLM benchmark scores at 7B-34B scale.

  8. Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization

    cs.CV 2025-04 conditional novelty 5.0 of 10

    KeCO updates the visual feature keys of a small coreset with all leftover support images, and its diversity-based update outperforms retrieval from the five-times-larger full support set for LVLM in-context image clas...

  9. SignEye: Traffic Sign Interpretation from Vehicle First-Person View

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A vision-language pipeline, SignEye, reads traffic signs from a vehicle's first-person view, assigns them to the current, left, or right lane and road, and generates driving-plan suggestions, supported by a new Chines...

  10. Multimodal Instruction Tuning with Hybrid State Space Models

    cs.CV 2024-11 reject novelty 4.0 of 10

    MMJamba applies a hybrid transformer-Mamba LLM backbone to multimodal instruction tuning, claiming 4x faster inference on long visual contexts and state-of-the-art accuracy on image and video benchmarks.

Pith tools