VLM-TSI, which interleaves vision and text tokens along a shared timeline, outperforms the turn-based VideoLLM-Online on TGLG, a new benchmark for real-time temporally-grounded language generation.
Collecting highly parallel data for paraphrase evaluation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
VLM-TSI, which interleaves vision and text tokens along a shared timeline, outperforms the turn-based VideoLLM-Online on TGLG, a new benchmark for real-time temporally-grounded language generation.