REVIEW 1 cited by
Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Querying generative AI models, e.g., large language models (LLMs), has become a prevalent method for information acquisition. However, existing query-answer datasets primarily focus on textual responses, making it challenging to address complex user queries that require visual demonstrations or explanations for better understanding. To bridge this gap, we construct a benchmark, RealVideoQuest, designed to evaluate the abilities of text-to-video (T2V) models in answering real-world, visually grounded queries. It identifies 7.5K real user queries with video response intents from Chatbot-Arena and builds 4.5K high-quality query-video pairs through a multistage video retrieval and refinement process. We further develop a multi-angle evaluation system to assess the quality of generated video answers. Experiments indicate that current T2V models struggle with effectively addressing real user queries, pointing to key challenges and future research opportunities in multimodal AI.
Forward citations
Cited by 1 Pith paper
-
T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
A 1,200-prompt benchmark across six world-knowledge domains reports that ten state-of-the-art text-to-video models average below 0.70 on a 0 to 1 scale for producing videos consistent with real-world knowledge.
Discussion (0). Sign in to comment.