Pith. sign in

REVIEW 9 cited by

Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09645 v3 pith:BZ6G66BM submitted 2024-12-10 cs.CV cs.AIcs.CL

Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models

classification cs.CV cs.AIcs.CL
keywords evaluationmodelsagentefficientframeworkgenerativevisualdiverse
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent advancements in visual generative models have enabled high-quality image and video generation, opening diverse applications. However, evaluating these models often demands sampling hundreds or thousands of images or videos, making the process computationally expensive, especially for diffusion-based models with inherently slow sampling. Moreover, existing evaluation methods rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. In contrast, humans can quickly form impressions of a model's capabilities by observing only a few samples. To mimic this, we propose the Evaluation Agent framework, which employs human-like strategies for efficient, dynamic, multi-round evaluations using only a few samples per round, while offering detailed, user-tailored analyses. It offers four key advantages: 1) efficiency, 2) promptable evaluation tailored to diverse user needs, 3) explainability beyond single numerical scores, and 4) scalability across various models and tools. Experiments show that Evaluation Agent reduces evaluation time to 10% of traditional methods while delivering comparable results. The Evaluation Agent framework is fully open-sourced to advance research in visual generative models and their efficient evaluation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MultiAnimate: Pose-Guided Image Animation Made Extensible

    cs.CV 2026-02 unverdicted novelty 7.0

    MultiAnimate adds Identifier Assigner and Identifier Adapter modules to diffusion video models so they can handle multiple characters without identity mix-ups, generalizing from two-character training data to more characters.

  2. TunerDiT: Training-free Progressive Steering of Diffusion Transformer for Multi-Event Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    TunerDiT adds event-partitioned masking and cross-event prompt fusion to diffusion transformers for training-free multi-event video generation, with gains scaling by event count on a new Meve benchmark.

  3. Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing

    cs.CV 2026-05 unverdicted novelty 6.0

    Occlusion-aware keyframe selection via structural, cycle-consistent tracking, and vision-language criteria improves diffusion video editing robustness without manual annotations.

  4. Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability

    cs.CV 2025-12 conditional novelty 6.0

    A video VAE whose latents are biased toward low frequencies and a few dominant channel modes improves text-to-video diffusion convergence and reward scores.

  5. Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation

    cs.CV 2025-11 unverdicted novelty 6.0

    Fine-tuning text-to-video models on sparse low-quality synthetic data for physical camera controls outperforms fine-tuning on photorealistic data.

  6. EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

    cs.CV 2025-09 unverdicted novelty 6.0

    EditVerse unifies image and video editing and generation in one transformer model via unified token sequences and in-context learning, trained jointly on curated video editing data plus image/video corpora and evaluat...

  7. VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

    cs.CV 2025-03 accept novelty 6.0

    VBench-2.0 is a benchmark suite that automatically evaluates video generative models on five dimensions of intrinsic faithfulness: Human Fidelity, Controllability, Creativity, Physics, and Commonsense using VLMs, LLMs...

  8. TRON: Tracing Rays to Orchestrate a Neural Renderer for 3D Gaussian Reconstructions

    cs.CV 2026-06 unverdicted novelty 5.0

    Hybrid system that uses ray-traced 3D Gaussians to supply radiometric guidance and material regularization to a neural renderer for editable, realistic output from captured scenes.

  9. Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing

    cs.CV 2026-05 unverdicted novelty 5.0

    A new keyframe selection framework combines structural, tracking, and semantic criteria to select reliable anchor frames for diffusion-based video editing under occlusion.