Pith. sign in

REVIEW 44 cited by

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.13503 v1 pith:3PRRUHGF submitted 2024-11-20 cs.CV

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

classification cs.CV
keywords videogenerationmodelsvbenchevaluationcomprehensivedimensionshuman
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal evaluation system should provide insights to inform future developments of video generation. To this end, we present VBench, a comprehensive benchmark suite that dissects "video generation quality" into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench has several appealing properties: 1) Comprehensive Dimensions: VBench comprises 16 dimensions in video generation (e.g., subject identity inconsistency, motion smoothness, temporal flickering, and spatial relationship, etc). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investigate the gaps between video and image generation models. 4) Versatile Benchmarking: VBench++ supports evaluating text-to-video and image-to-video. We introduce a high-quality Image Suite with an adaptive aspect ratio to enable fair evaluations across different image-to-video generation settings. Beyond assessing technical quality, VBench++ evaluates the trustworthiness of video generative models, providing a more holistic view of model performance. 5) Full Open-Sourcing: We fully open-source VBench++ and continually add new video generation models to our leaderboard to drive forward the field of video generation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 44 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Do generative video models understand physical principles?

    cs.CV 2025-01 unverdicted novelty 8.0

    Physics-IQ benchmark reveals that generative video models exhibit limited physical understanding unrelated to their visual quality.

  2. WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

    cs.CV 2026-06 unverdicted novelty 7.0

    WorldRoamBench is a new benchmark for interactive world models that evaluates four stability dimensions with custom metrics and finds no tested model performs reliably across all.

  3. WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

    cs.CV 2026-06 accept novelty 7.0

    A 600+ case open-world benchmark finds no interactive world model is simultaneously action-faithful, visually stable, physically plausible, and memory-consistent over long WASD interaction.

  4. OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation

    cs.CV 2026-06 unverdicted novelty 7.0

    DRIVE-CHOREO uses three LLM agents to create a unified position-aware token sequence co-compressed with multi-view video, achieving SOTA BEV mAP of 21.6 and +2.4 NDS improvement on nuScenes.

  5. WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

    cs.CV 2026-05 unverdicted novelty 7.0

    WorldJen is a new benchmark for generative video models that uses VLM-judged multi-dimensional Likert questionnaires validated against human preferences to achieve perfect tier agreement.

  6. WorldMark: A Unified Benchmark Suite for Interactive Video World Models

    cs.CV 2026-04 unverdicted novelty 7.0

    WorldMark is the first public benchmark that standardizes scenes, trajectories, and control interfaces across heterogeneous interactive image-to-video world models.

  7. HumanScore: Benchmarking Human Motions in Generated Videos

    cs.CV 2026-04 unverdicted novelty 7.0

    HumanScore defines six metrics for kinematic plausibility, temporal stability, and biomechanical consistency to benchmark human motions in videos from thirteen state-of-the-art generation models, revealing gaps betwee...

  8. CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

    cs.CV 2026-01 unverdicted novelty 7.0

    CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.

  9. VACE: All-in-One Video Creation and Editing

    cs.CV 2025-03 unverdicted novelty 7.0

    VACE unifies reference-to-video generation, video-to-video editing, and masked video-to-video editing in one Diffusion Transformer framework using a Video Condition Unit for inputs and a Context Adapter for task injection.

  10. Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

    cs.CV 2026-07 conditional novelty 6.0

    Explicit instance-level physical parameter conditioning with routing attention improves physical-law consistency in image-to-video generation, as measured on the authors' new simulator-based benchmark.

  11. Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation

    cs.CV 2026-07 conditional novelty 6.0

    Imperceptible image perturbations that suppress text-conditioned attention and boost visual-only pathways disrupt text-guided I2V generation more efficiently than prior I2V protection.

  12. Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation

    cs.CV 2026-07 conditional novelty 6.0

    Anti-Prompt protects images from text-guided image-to-video generation by suppressing text-conditioned attention during denoising, producing visible generation failures.

  13. WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

    cs.CV 2026-06 unverdicted novelty 6.0

    WorldOdysseyBench introduces four new evaluation dimensions and metrics for interactive world models and shows that none of 10+ tested models reliably pass all of them.

  14. A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks

    cs.MM 2026-06 unverdicted novelty 6.0

    EffectivePresentationScorer evaluates paper-to-video talks for instructional quality by checking clear explanation of ideas, prerequisite concepts, and links to contributions, finding that current systems cover topics...

  15. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    MAVEN introduces a multi-agent system for refining prompts in multicultural text-to-video generation and releases a benchmark of 243 prompts and 972 videos showing improved cultural relevance via parallel agent specia...

  16. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 conditional novelty 6.0

    Parallel, role-specialized prompt agents improve cultural relevance in text-to-video generation, with a new cross-cultural benchmark showing the largest gains for location cues.

  17. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    MAVEN is a multi-agent prompt refinement framework that improves cultural fidelity in text-to-video generation, demonstrated on a new benchmark of 243 prompts and 972 videos across Chinese, American, and Romanian cultures.

  18. Compositional Video Generation via Inference-Time Guidance

    cs.CV 2026-05 unverdicted novelty 6.0

    CVG improves compositional faithfulness in frozen text-to-video diffusion models by steering early denoising steps with gradients from a classifier trained on the model's own cross-attention features.

  19. PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation

    cs.CV 2026-05 conditional novelty 6.0

    PhyMotion scores generated human videos by grounding recovered 3D poses in a physics simulator across kinematic, contact, and dynamic axes, yielding stronger human correlation and larger RL post-training gains than pr...

  20. LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention

    cs.CV 2026-05 unverdicted novelty 6.0

    LIVEditor-14B applies a new sparse attention method (ISA) that prunes context and uses query-sharpness routing to cut attention latency ~60% with no loss in editing quality on standard benchmarks.

  21. WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

    cs.CV 2026-05 conditional novelty 6.0

    WorldJen is a multi-dimensional video generation benchmark using VLM-graded Likert questionnaires on joint prompts, validated to match human three-tier rankings.

  22. Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos

    cs.CV 2026-04 unverdicted novelty 6.0

    EgoIn uses a fine-tuned vision-language model to infer transition steps and a conditioning module plus auxiliary supervision to generate coherent egocentric video sequences of object state changes.

  23. InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

    cs.CV 2025-12 conditional novelty 6.0

    InsertAnywhere inserts a reference object into arbitrary videos by reconstructing 4D geometry to propagate a user-given placement across frames and fine-tuning video diffusion on ROSE++, a removal-to-insertion dataset...

  24. WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

    cs.CV 2025-12 conditional novelty 6.0

    A five-aspect, 24-metric benchmark, a 26K human-annotated dataset, and an AI evaluator show that today's driving world models cannot simultaneously look real, respect geometry, and behave safely.

  25. InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem

    cs.CV 2025-12 conditional novelty 6.0

    A training-free, near-zero-overhead inverse solver for novel-view video generation and inpainting that projects masks into continuous multi-channel latent masks and applies DDS with conjugate gradient in latent space.

  26. Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation

    cs.CV 2025-12 conditional novelty 6.0

    Reward Forcing combines EMA-Sink tokens and Rewarded Distribution Matching Distillation to deliver state-of-the-art streaming video generation at 23.1 FPS without copying initial frames.

  27. SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation

    cs.CV 2025-11 unverdicted novelty 6.0

    SteadyDancer is an I2V framework using condition reconciliation, synergistic pose modulation, and staged training to achieve robust first-frame preservation and coherent motion control in human image animation.

  28. RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling

    cs.CV 2025-10 unverdicted novelty 6.0

    RAPO++ is a three-stage prompt optimization framework combining retrieval-augmented refinement, closed-loop test-time scaling, and LLM fine-tuning to enhance text-to-video generation quality.

  29. LongLive: Real-time Interactive Long Video Generation

    cs.CV 2025-09 conditional novelty 6.0

    LongLive is a causal autoregressive video generator that produces up to 240-second interactive videos at 20.7 FPS on one H100 GPU after 32 GPU-days of fine-tuning from a 1.3B short-clip model.

  30. Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

    cs.RO 2025-08 unverdicted novelty 6.0

    Genie Envisioner unifies robotic policy learning, simulation, and evaluation inside one instruction-conditioned video diffusion framework using GE-Base, GE-Act, and GE-Sim.

  31. Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

    cs.RO 2025-07 unverdicted novelty 6.0

    RIGVid shows that filtered AI-generated videos can serve as effective supervision for complex robotic manipulation tasks without any real demonstrations.

  32. MAGI-1: Autoregressive Video Generation at Scale

    cs.CV 2025-05 unverdicted novelty 6.0

    MAGI-1 is a 24B-parameter autoregressive video world model that predicts denoised frame chunks sequentially with increasing noise to enable causal, scalable, streaming generation up to 4M token contexts.

  33. VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

    cs.CV 2025-03 accept novelty 6.0

    VBench-2.0 is a benchmark suite that automatically evaluates video generative models on five dimensions of intrinsic faithfulness: Human Fidelity, Controllability, Creativity, Physics, and Commonsense using VLMs, LLMs...

  34. Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

    cs.CV 2026-07 conditional novelty 5.0

    A chunk-wise streaming video model with bounded multi-scale memory and streaming 4K upscaling reports real-time interactive long-form generation and top Arena preference/stability scores.

  35. Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation

    cs.CV 2026-07 unverdicted novelty 5.0

    Anti-Prompt adds imperceptible perturbations to images that disrupt text-guided I2V generation by attenuating text-conditioned pathways, achieving protection on two model architectures with a new Video-LLM evaluation ...

  36. Bridging Video Understanding and Generation in a Unified Framework

    cs.CV 2026-06 unverdicted novelty 5.0

    Vega unifies video understanding and generation via shared vocabulary and hybrid autoregressive-diffusion architecture, reporting strong results on VBench and VideoMME.

  37. Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

    cs.CV 2026-06 unverdicted novelty 5.0

    A decoupled-control autoregressive video model using Fast-Slow Memory training, dynamic projection, and staged camera control to produce stable long-horizon outputs with human and viewpoint guidance.

  38. Physics-IQ Verified

    cs.CV 2026-06 unverdicted novelty 5.0

    Physics-IQ Verified refines 57.6% of samples and 34.8% of prompts from the original benchmark and produces moderate ranking shifts (Kendall's τ = 0.46) across six image-to-video models.

  39. LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention

    cs.CV 2026-05 unverdicted novelty 5.0

    ISA prunes low-saliency context tokens and routes queries by sharpness to either full or 0-th order Taylor sparse attention, enabling LIVEditor to cut attention latency ~60% while beating prior video editing methods o...

  40. Resilience Revisited: A Multidimensional Framework Derived from Realistic Attack Scenarios

    eess.SY 2026-04 unverdicted novelty 5.0

    A Multidimensional Resilience Index on IEEE 39-bus shows coordinated multivector attacks amplify resilience loss ~8× via interdimensional coupling, and ~15× when climatic and regulatory stressors are included.

  41. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    cs.RO 2026-07 conditional novelty 4.0

    Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.

  42. Resilience Revisited: A Multidimensional Framework Derived from Realistic Attack Scenarios

    eess.SY 2026-04 unverdicted novelty 4.0

    A new MDRI framework decomposes power system resilience into five dimensions with a calibrated multiplicative term, showing multi-vector attacks cause 5.6 times more degradation than linear sums plus 60.6% from coupling.

  43. Toward Native Multimodal Modeling: A Roadmap

    cs.CV 2026-05 unverdicted novelty 3.0

    A roadmap that defines architectural nativity for multimodal models and categorizes them into Multi-to-Text, Multi-to-Target, and Multi-to-Multi types while outlining an industrial pipeline toward unified transformer-...

  44. Evolution of Video Generative Foundations

    cs.CV 2026-04 unverdicted novelty 2.0

    This survey traces video generation technology from GANs to diffusion models and then to autoregressive and multimodal approaches while analyzing principles, strengths, and future trends.