Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Inference Compute-Optimal Video Vision Language Models

cs.CV · 2025-05-24 · conditional · novelty 7.0

For video vision-language models under a fixed inference compute budget, the optimal setup scales language model size, frame count, and tokens per frame jointly, and shifts toward smaller LMs with more frames and tokens as finetuning data grows.

citing papers explorer

Showing 1 of 1 citing paper.

  • Inference Compute-Optimal Video Vision Language Models cs.CV · 2025-05-24 · conditional · none · ref 7

    For video vision-language models under a fixed inference compute budget, the optimal setup scales language model size, frame count, and tokens per frame jointly, and shifts toward smaller LMs with more frames and tokens as finetuning data grows.