LAVE lets video agents reuse pre-verbal visual hidden states from past tool calls via timestamp-aligned residual injection, improving Video-MME by 3.76 points over the strongest baseline without extra training or frames.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents
LAVE lets video agents reuse pre-verbal visual hidden states from past tool calls via timestamp-aligned residual injection, improving Video-MME by 3.76 points over the strongest baseline without extra training or frames.