Pith. sign in

REVIEW 17 cited by

Scaling Image and Video Generation via Test-Time Evolutionary Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.17618 v1 pith:JIB7PST7 submitted 2025-05-23 cs.CV cs.AIcs.LG

Scaling Image and Video Generation via Test-Time Evolutionary Search

classification cs.CV cs.AIcs.LG
keywords scalingevosearchimagemodelstest-timevideoacrossdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

As the marginal cost of scaling computation (data and parameters) during model pre-training continues to increase substantially, test-time scaling (TTS) has emerged as a promising direction for improving generative model performance by allocating additional computation at inference time. While TTS has demonstrated significant success across multiple language tasks, there remains a notable gap in understanding the test-time scaling behaviors of image and video generative models (diffusion-based or flow-based models). Although recent works have initiated exploration into inference-time strategies for vision tasks, these approaches face critical limitations: being constrained to task-specific domains, exhibiting poor scalability, or falling into reward over-optimization that sacrifices sample diversity. In this paper, we propose \textbf{Evo}lutionary \textbf{Search} (EvoSearch), a novel, generalist, and efficient TTS method that effectively enhances the scalability of both image and video generation across diffusion and flow models, without requiring additional training or model expansion. EvoSearch reformulates test-time scaling for diffusion and flow models as an evolutionary search problem, leveraging principles from biological evolution to efficiently explore and refine the denoising trajectory. By incorporating carefully designed selection and mutation mechanisms tailored to the stochastic differential equation denoising process, EvoSearch iteratively generates higher-quality offspring while preserving population diversity. Through extensive evaluation across both diffusion and flow architectures for image and video generation tasks, we demonstrate that our method consistently outperforms existing approaches, achieves higher diversity, and shows strong generalizability to unseen evaluation metrics. Our project is available at the website https://tinnerhrhe.github.io/evosearch.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

    cs.CV 2026-06 unverdicted novelty 7.0

    PhyEditBench is a new benchmark for physics-aware image editing with real and synthetic instances plus a training-free PhyWorld baseline that uses test-time scaling to outperform SOTA models.

  2. PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

    cs.CV 2026-06 unverdicted novelty 7.0

    PhyEditBench is a new benchmark with real-world and synthetic instances that reveals limitations in current image editing models' physics reasoning and proposes a video-generation-based baseline called PhyWorld.

  3. SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation

    eess.AS 2026-06 unverdicted novelty 7.0

    SMC-ITA applies sequential Monte Carlo resampling with lookahead-based multi-dimensional cross-modal rewards to improve inference-time alignment in video-to-audio generation, reporting 55.67% DeSync reduction and gain...

  4. Inference-Time Scaling for Joint Audio-Video Generation

    cs.MM 2026-06 unverdicted novelty 7.0

    Presents multi-verifier framework and Adaptive Reward Weighting (ARW) for inference-time scaling in joint audio-video generation, reporting gains in alignment and synchronization on VGGSound and JavisBench-mini.

  5. VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

    cs.CV 2026-06 unverdicted novelty 7.0

    VLMs act as teachers by deriving differentiable rewards from task rules to adapt VGMs via test-time LoRA optimization, delivering 16.7-point average gains on symbolic and general video reasoning benchmarks.

  6. VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

    cs.CV 2026-06 unverdicted novelty 7.0

    VLMs formulate differentiable rewards from task-specific rules to enable test-time online LoRA optimization of VGMs, delivering 16.7-point gains on symbolic and general video reasoning benchmarks over VLM-as-solver an...

  7. Inference-Time Scaling in Diffusion Models through Iterative Partial Refinement

    cs.LG 2026-05 unverdicted novelty 7.0

    IPR improves valid solution rates on MNIST Sudoku from 55.8% to 75.0% by iteratively refining partial regions in sequential diffusion models without external verifiers or reward models.

  8. CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

    cs.CV 2026-05 unverdicted novelty 7.0

    CollabVR improves video reasoning performance by coupling vision-language models and video generation models in a closed-loop step-level collaboration that detects and repairs generation failures.

  9. CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

    cs.AI 2026-07 conditional novelty 6.0

    Using cached drafts to select the winner and regenerating only that winner captures 94.7% of best-of-8 search gain at 63% of the cost.

  10. Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0

    Under wall-clock budgets, cheap multi-knob drafts plus multi-stage verification outperform guided intermediate search for diffusion T2I inference-time scaling.

  11. SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    SpecLoR rectifies the amplitude spectrum of lookahead-estimated clean latents to natural-video priors during early ODE sampling steps, cutting physical artifacts with only four extra NFEs.

  12. Stream-T1: Test-Time Scaling for Streaming Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    Stream-T1 is a test-time scaling framework for streaming video generation using scaled noise propagation from history, reward pruning across short and long windows, and feedback-guided memory sinking to improve tempor...

  13. NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training

    cs.LG 2026-05 unverdicted novelty 6.0

    NoiseRater meta-learns instance-level importance scores for noise in diffusion training via bilevel optimization, then uses a two-stage pipeline to improve efficiency and generation quality on FFHQ and ImageNet.

  14. ELT: Elastic Looped Transformers for Visual Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    Elastic Looped Transformers share weights across recurrent blocks and apply intra-loop self-distillation to deliver 4x parameter reduction while matching competitive FID and FVD scores on ImageNet and UCF-101.

  15. Superbunched random fiber laser

    physics.optics 2026-03 unverdicted novelty 6.0

    A fiber-integrated random laser uses Rayleigh scattering, cascaded Brillouin scattering, and four-wave mixing to generate multi-wavelength superbunched light with g(2)(0) up to ~26 and improved temporal ghost imaging.

  16. LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion

    cs.CV 2026-03 accept novelty 6.0

    LatSearch improves video diffusion quality and efficiency by scoring intermediate latents with a trained reward model and performing reward-guided resampling plus final pruning.

  17. Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning

    cs.CV 2026-06 unverdicted novelty 5.0

    A survey of test-time scaling for multimodal foundation models that introduces a three-way taxonomy of sampling, feedback, and search approaches along with applications and benchmarks.