REVIEW 3 major objections 6 minor 45 references
Long video generation can stay consistent at real-time speed by treating memory and denoising compute as online allocation problems driven by the model's own 'surprise' signals.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 15:25 UTC pith:JYBUA53D
load-bearing objection A practical training-free method that improves long-video consistency and compute allocation; needs error bars and code, but the central claim holds. the 3 major comments →
Surprise Forcing: What to Remember, When to Skip in Long Video Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that selective memory and selective computation, driven by surprise signals within the inference trajectory, improve long-horizon consistency and visual quality without changing the trained generator. Concretely, a Surprise-Gated Memory Bank scores evicted frames by the deviation of their mean-pooled value-token descriptor from the bank, admits them through a feedback-controlled budget in normalized score space, evicts by a composite priority of surprise, usage, and age, and routes only top-k relevant entries to attention. In parallel, Surprise-Aware Denoising estimates chunk difficulty from the maximum adjacent-frame cosine distance after the first denoising pass and as
What carries the argument
The load-bearing mechanism is the surprise score: for memory, a mixture of global deviation and nearest-neighbor novelty computed on L2-normalized mean-pooled value tokens (Equation 6), gated by Budget-Norm Gating, an online controller that adjusts the admission threshold to hold a target write ratio; for denoising, the maximum adjacent-frame cosine distance in the first-step latent (Equation 13), ranked within a sliding window to decide step skipping. These signals are causal, require no auxiliary network, and convert 'what is surprising' into concrete resource-allocation decisions.
Load-bearing premise
The whole memory pipeline assumes that a single mean-pooled, L2-normalized value-token vector per frame is enough to tell which frames hold information worth keeping; if two frames have similar global content but different layout or pose, the bank will treat them as the same and may discard the frame that later matters.
What would settle it
Generate a video where a subject's average color and texture stay constant while its pose or layout changes drastically (e.g., a person walking across a uniformly colored room); if the Surprise-Gated Memory Bank fails to retain the earlier view and subject identity drifts on VBench-2.0 Human Identity, that confirms the descriptor collapses spatial distinctions. A cleaner test: replace the mean-pooled descriptor with spatially-partitioned patch descriptors and check whether consistency scores rise — if they do, the mean pooling is the limiting factor.
If this is right
- If the central claim holds, streaming video generators can improve long-horizon consistency without retraining, by reallocating memory and compute at inference time.
- The reported speedup of ~15.8% denoising work at r_skip=0.4 with minimal quality loss suggests that difficulty-adaptive step skipping is a viable real-time technique.
- The method's strong Multi-View Consistency gain implies that preserving early establishing views in an external bank can recover viewpoint consistency that rolling caches lose.
- Since the framework is training-free and complementary to distillation-based streaming generators, it can be layered on top of future base models.
Where Pith is reading between the lines
- The mean-pooled descriptor is the critical bottleneck: if a scene contains two states with identical global color/texture but different spatial layout, the bank may treat them as redundant and evict the discriminative frame, undermining the consistency claim; a multi-resolution or spatially-partitioned descriptor would be a natural test extension.
- The feedback-controlled admission ratio r and the priority weights are fixed hyperparameters; a higher-level controller that adapts these to prompt structure or scene dynamics could further improve the stability–freshness trade-off the paper identifies.
- Because the surprise signal is self-referential, it should be tested on longer-than-one-minute rollouts and interactive trajectories where error accumulation may change the statistics of surprise; the paper explicitly flags this as unestablished.
- The binary {0,3} versus full schedule could be generalized to intermediate step counts or to variable step allocations across chunks, which the paper lists as future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Surprise Forcing, a training-free inference-time framework for streaming autoregressive video diffusion. It adds two controllers on top of a pretrained streaming generator: (1) a Surprise-Gated Memory Bank that decides which evicted frames to store using a mean-pooled value-token descriptor, a dual-component surprise score, a feedback-controlled admission threshold, priority-based eviction, and query-dependent routing; and (2) a Surprise-Aware Denoising scheduler that estimates chunk difficulty from the maximum adjacent-frame cosine distance after the first denoising pass and skips intermediate steps for chunks ranked below a local percentile. The authors evaluate on VBench, VBench-Long, and selected VBench-2.0 consistency metrics, reporting improved scores over CausVid, Self Forcing, Rolling Forcing, and LongLive while retaining real-time throughput (17.18 FPS on 60-second 832x480 rollouts).
Significance. If the reported results hold, the paper makes a useful contribution to efficient long-video generation. The core idea—extracting control signals from the generator's own intermediate representations—is interesting and goes beyond common cache/eviction heuristics. The explicit formulation of budget-norm gating, the component ablations, and the use of public benchmarks are strengths. The paper also provides a falsifiable check in Figure 4, where the intra-chunk surprise score is correlated with residual denoising error (PLCC 0.695, SRCC 0.730). However, the central claims are currently supported only by single-point benchmark scores without variance or significance testing, and the memory descriptor that drives admission and routing is a mean-pooled summary whose discriminative limits are acknowledged by the authors but not experimentally probed. These issues leave the strength of the central consistency claim somewhat uncertain.
major comments (3)
- [§3.3, Eq. (5)] The frame descriptor d = mean(V_l)/||mean(V_l)|| is a permutation-invariant summary of spatial value tokens. Since admission (Eq. 6), priority-based eviction (Eq. 11), and dynamic routing (§3.3.2) all operate on this descriptor, the memory bank cannot distinguish frames that share similar global statistics but differ in spatial layout, object pose, or identity. This is directly relevant to the claimed improvements on VBench-2.0 Human Identity and Multi-View Consistency in Table 2. The authors acknowledge the issue in Section 5, but no experiment quantifies its impact. Please either compare against a spatially-aware descriptor (e.g., token-wise max-pooling or a small set of spatial-region descriptors) under the same pipeline, or provide a targeted failure analysis on prompts with repeated poses/viewpoints. Without such evidence, the central consistency claim is not fully established.
- [Tables 1-11] All quantitative results are single point estimates with no confidence intervals, standard deviations, or significance tests. Some of the headline improvements over LongLive are small on individual metrics (e.g., Table 1: VBench Subject Consistency 96.28 vs 96.19; VBench-Long Subject Consistency 98.51 vs 98.42). Because video generation is stochastic, these differences may be within run-to-run noise. Please report at least 3 seeds with standard deviations (or paired bootstrap/permutation tests on the fixed prompt sets) for the main benchmark comparisons and the key ablations, especially Tables 1, 2, 5, and 9. This is necessary to support the claim that the improvements are reliable.
- [§3.4, Table 5] The intra-chunk surprise predictor is validated on only 50 MovieGen prompts with moderate correlation (PLCC 0.695, SRCC 0.730). The scheduler then uses this predictor to skip denoising work, with the default r_skip=0.4 reported as reducing 'denoising work by 15.8%'. However, the reduced schedule {0,3} has half the transformer passes of {0,1,2,3}, so the relationship between the 'Acc. Ratio' and the fraction of actually skipped chunks is unclear and should be stated explicitly. Please (a) clarify what 'Acc. Ratio' measures, (b) report the scheduler-only effect on quality by disabling the memory bank or holding it constant, and (c) show, if possible, the distribution of residual errors for skipped versus non-skipped chunks to confirm that the skipped chunks are indeed the easy ones. This is load-bearing for the speed-quality trade-off claim.
minor comments (6)
- [Table 5] The row labeled 'rskip = 084.03' appears to be a formatting error; it should presumably read 'rskip = 0.0' followed by the score 84.03.
- [§3.3.2] The query descriptor d_q is used in cosine similarity for routing but is never explicitly defined. It should be stated whether d_q is the same mean-pooled value-token descriptor as Eq. (5) computed for the current generation query, and over which tokens.
- [Eq. (5)] Please specify which transformer layer(s) provide the value tokens V_l, the index l (e.g., spatial position, head, or token), and how L is defined. This is needed for reproducibility.
- [§3.3] The phrase 'neighboring-chunk throttling' (used to explain why final commit rate can be lower than r) is never defined. Please clarify whether it refers to a separate mechanism or to the warmup/eviction behavior.
- [Table 1 caption and §4.2] The caption claims 'the speedup becomes more pronounced when generating longer videos', but Table 1 shows Surprise Forcing at 16.45 FPS (5s) and 17.18 FPS (60s), both below LongLive's 18.01 FPS. This statement is not supported by the reported numbers. Please rephrase to 'throughput remains real-time and comparable' or provide the intended baseline for the claimed speedup.
- [§4.2] The phrase 'state-of-the-art performance across all metrics' is too broad for a comparison against five streaming baselines. It should be restricted to 'among the compared streaming alternatives'.
Circularity Check
No significant circularity: the central improvements are external-benchmark measurements, and the surprise signals are causal heuristics rather than definitions of the target metrics.
full rationale
The paper's central claims are empirical: Surprise Forcing is a training-free inference-time controller, and its reported gains on VBench, VBench-Long, and VBench-2.0 are measured against external benchmarks and compared with streaming baselines, not derived from the method's own definitions. The memory-bank admission rule (Eqs. 5-10) and the step-skipping rule (Eqs. 12-14) are engineering heuristics whose parameters are ablated; they are not fitted to the benchmark scores in a way that would make the evaluation a restatement of the fit. The only apparent self-citations (e.g., [28] in the intro list of text-to-video models) are non-load-bearing related-work citations. A genuine limitation is acknowledged in Section 5: the mean-pooled value-token descriptor 'may not distinguish spatial arrangements that share similar global content.' That is a correctness/failure-mode concern for long-horizon consistency, not a circularity: it means the descriptor may under-discriminate, which would weaken the claimed improvement rather than force it by construction. The intra-chunk difficulty predictor is validated by PLCC/SRCC against residual denoising error on separate MovieGen prompts, so it is not a fitted input renamed as a prediction. No step in the paper reduces an output to an input by definition or via a load-bearing self-citation.
Axiom & Free-Parameter Ledger
free parameters (10)
- alpha (surprise mixture) =
0.7
- EMA momentum m =
0.95
- controller step size eta =
0.1
- target admission ratio r =
0.3
- initial threshold tau_0 =
0.002
- priority weights w_s, w_u, w_a =
1.8, 1.0, 0.4
- routing top-k k =
3
- bank capacity C =
6 (long), 3 (short)
- skip ratio r_skip =
0.4
- percentile window size and warmup =
20 recent chunks; first 5 chunks full
axioms (5)
- domain assumption The causal factorization p(x_1:N|c) with a bounded rolling KV cache is a valid model of long-video generation.
- domain assumption Mean-pooled value tokens form a stable and sufficient content descriptor for comparing frames.
- domain assumption The maximum adjacent-frame cosine distance after the first denoising pass estimates residual denoising error.
- ad hoc to paper Percentile rank within a sliding window of recent chunks is a valid normalization for chunk difficulty.
- domain assumption The distribution-matching distillation objective (KL between student and teacher score distributions) is valid for the base model.
read the original abstract
Streaming autoregressive diffusion makes minute-scale video synthesis practical, but its bounded context and fixed denoising schedule allocate resources uniformly across a highly non-stationary sequence. A rolling key-value cache forgets distant visual evidence even when that evidence remains important, while every generated chunk receives the same number of denoising passes irrespective of its actual difficulty. We introduce Surprise Forcing, a training-free framework that treats both limitations as online resource-allocation problems. A Surprise-Gated Memory Bank summarizes evicted frames with value-token descriptors, evaluates them using complementary global-deviation and nearest-neighbor novelty signals, and regulates admission through a feedback-controlled budget in normalized score space. Priority-based replacement and relevance-aware routing then keep the external memory compact and useful. In parallel, Surprise-Aware Denoising estimates chunk difficulty from the maximum adjacent-frame cosine distance after the first denoising pass and uses a local percentile scheduler to skip intermediate steps for comparatively easy chunks. Experiments on VBench, VBench-Long, and VBench-2.0 show that the proposed allocation strategy improves long-horizon consistency and visual quality while retaining real-time streaming throughput.
Reference graph
Works this paper leans on
-
[1]
S. Cai, C. Yang, L. Zhang, Y. Guo, J. Xiao, Z. Yang, Y. Xu, Z. Yang, A. Yuille, L. Guibas, et al. Mixture of contexts for long video generation.arXiv preprint arXiv:2508.21058, 2025
arXiv 2025
-
[2]
B. Chen, D. Martí Monsó, Y. Du, M. Simchowitz, R. Tedrake, and V. Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion.Advances in Neural Information Processing Systems, 37:24081–24125, 2024
2024
-
[3]
G. Chen, D. Lin, J. Yang, C. Lin, J. Zhu, M. Fan, H. Zhang, S. Chen, Z. Chen, C. Ma, et al. Skyreels-v2: Infinite-length film generative model.arXiv preprint arXiv:2504.13074, 2025
Pith/arXiv arXiv 2025
-
[4]
J. Cui, J. Wu, M. Li, T. Yang, X. Li, R. Wang, A. Bai, Y. Ban, and C.-J. Hsieh. Self-forcing++: Towards minute-scale high-quality video generation.arXiv preprint arXiv:2510.02283, 2025
Pith/arXiv arXiv 2025
-
[5]
Henschel, L
R. Henschel, L. Khachatryan, H. Poghosyan, D. Hayrapetyan, V. Tadevosyan, Z. Wang, S. Navasardyan, and H. Shi. Streamingt2v: Consistent, dynamic, and extendable long video generation from text. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 2568–2577, 2025
2025
-
[6]
J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
2020
- [7]
-
[8]
X. Huang, Z. Li, G. He, M. Zhou, and E. Shechtman. Self forcing: Bridging the train-test gap in autoregressive video diffusion.arXiv preprint arXiv:2506.08009, 2025
Pith/arXiv arXiv 2025
-
[9]
Huang, Y
Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, et al. Vbench: Comprehensive benchmark suite for video generative models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21807–21818, 2024
2024
-
[10]
Huang, F
Z. Huang, F. Zhang, X. Xu, Y. He, J. Yu, Z. Dong, Q. Ma, N. Chanpaisit, C. Si, Y. Jiang, Y. Wang, X. Chen, Y.-C. Chen, L. Wang, D. Lin, Y. Qiao, and Z. Liu. VBench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
2025
-
[11]
S. Ji, X. Chen, S. Yang, X. Tao, P. Wan, and H. Zhao. Memflow: Flowing adaptive memory for consistent and efficient long video narratives.arXiv preprint arXiv:2512.14699, 2025
arXiv 2025
-
[12]
Jiang, W
J. Jiang, W. Li, J. Ren, Y. Qiu, R. Pei, F. Song, Y. Guo, X. Xu, H. Wu, and W. Zuo. Lovic: Efficient long video generation with context compression. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4022–4034, 2026
2026
-
[13]
Kahatapitiya, H
K. Kahatapitiya, H. Liu, S. He, D. Liu, M. Jia, C. Zhang, M. S. Ryoo, and T. Xie. Adaptive caching for faster video generation with diffusion transformers. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 15240–15252, 2025
2025
-
[14]
J. Kim, J. Kang, J. Choi, and B. Han. Fifo-diffusion: Generating infinite videos from text without training.Advances in Neural Information Processing Systems, 37:89834–89868, 2024
2024
-
[15]
Kodaira, T
A. Kodaira, T. Hou, J. Hou, M. Georgopoulos, F. Juefei-Xu, M. Tomizuka, and Y. Zhao. Streamdit: Real-time streaming text-to-video generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 29200–29210, 2026
2026
-
[16]
Lewis, E
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020
2020
-
[17]
R. Li, P. Torr, A. Vedaldi, and T. Jakab. Vmem: Consistent interactive video scene generation with surfel-indexed view memory. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 25690–25699, 2025. 13
2025
-
[18]
X. Li, M. Li, T. Cai, H. Xi, S. Yang, Y. Lin, L. Zhang, S. Yang, J. Hu, K. Peng, et al. Radial attention: O(nlogn)sparse attention with energy decay for long video generation.Advances in Neural Information Processing Systems, 38:16822–16852, 2026
2026
-
[19]
K. Liu, W. Hu, J. Xu, Y. Shan, and S. Lu. Rolling forcing: Autoregressive long video diffusion in real time.arXiv preprint arXiv:2509.25161, 2025
Pith/arXiv arXiv 2025
-
[20]
Y. Lu, Y. Liang, L. Zhu, and Y. Yang. Freelong: Training-free long video generation with spectralblend temporal attention.Advances in Neural Information Processing Systems, 37:131434–131455, 2024
2024
-
[21]
Y. Lu and Y. Yang. Freelong++: Training-free long video generation via multi-band spectralfusion. arXiv preprint arXiv:2507.00162, 2025
Pith/arXiv arXiv 2025
-
[22]
A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y. Ma, C.-Y. Chuang, et al. Movie gen: A cast of media foundation models.arXiv preprint arXiv:2410.13720, 2024
Pith/arXiv arXiv 2024
-
[23]
H. Qiu, S. Liu, Z. Zhou, Z. An, W. Ren, Z. Liu, J. Schult, S. He, S. Chen, Y. Cong, et al. Histream: Efficient high-resolution video generation via redundancy eliminated streaming. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4603–4613, 2026
2026
-
[24]
H. Qiu, M. Xia, Y. Zhang, Y. He, X. Wang, Y. Shan, and Z. Liu. Freenoise: Tuning-free longer video diffusion via noise rescheduling. InInternational Conference on Learning Representations, volume 2024, pages 5260–5274, 2024
2024
-
[25]
D. Ruhe, J. Heek, T. Salimans, and E. Hoogeboom. Rolling diffusion models.arXiv preprint arXiv:2402.09470, 2024
Pith/arXiv arXiv 2024
-
[26]
J. Song, C. Meng, and S. Ermon. Denoising diffusion implicit models.arXiv preprint arXiv:2010.02502, 2020
Pith/arXiv arXiv 2010
-
[27]
K. Song, B. Chen, M. Simchowitz, Y. Du, R. Tedrake, and V. Sitzmann. History-guided video diffusion. arXiv preprint arXiv:2502.06764, 2025
Pith/arXiv arXiv 2025
-
[28]
S. Tan, B. Gong, Y. Feng, K. Zheng, D. Zheng, S. Shi, Y. Shen, J. Chen, and M. Yang. Mimir: Improving video diffusion models for precise text understanding. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 23978–23988, 2025
2025
-
[29]
H. Teng, H. Jia, L. Sun, L. Li, M. Li, M. Tang, S. Han, T. Zhang, W. Zhang, W. Luo, et al. Magi-1: Autoregressive video generation at scale.arXiv preprint arXiv:2505.13211, 2025
Pith/arXiv arXiv 2025
-
[30]
T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Pith/arXiv arXiv 2025
-
[31]
B. Wu, C. Zou, C. Li, D. Huang, F. Yang, H. Tan, J. Peng, J. Wu, J. Xiong, J. Jiang, et al. Hunyuanvideo 1.5 technical report.arXiv preprint arXiv:2511.18870, 2025
Pith/arXiv arXiv 2025
-
[32]
X. Wu, G. Zhang, Z. Xu, Y. Zhou, Q. Lu, and X. He. Pack and force your memory: Long-form and consistent video generation.arXiv preprint arXiv:2510.01784, 2025
arXiv 2025
-
[33]
H. Xi, S. Yang, Y. Zhao, C. Xu, M. Li, X. Li, Y. Lin, H. Cai, J. Zhang, D. Li, et al. Sparse videogen: Ac- celerating video diffusion transformers with spatial-temporal sparsity.arXiv preprint arXiv:2502.01776, 2025
Pith/arXiv arXiv 2025
-
[34]
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis. Efficient streaming language models with attention sinks. InInternational Conference on Learning Representations, volume 2024, pages 21875–21895, 2024
2024
-
[35]
Z. Xiao, Y. Lan, Y. Zhou, W. Ouyang, S. Yang, Y. Zeng, and X. Pan. Worldmem: Long-term consistent world simulation with memory.Advances in Neural Information Processing Systems, 38:49632–49652, 2026
2026
-
[36]
S. Yang, W. Huang, R. Chu, Y. Xiao, Y. Zhao, X. Wang, M. Li, E. Xie, Y. Chen, Y. Lu, et al. Longlive: Real-time interactive long video generation.arXiv preprint arXiv:2509.22622, 2025. 14
Pith/arXiv arXiv 2025
-
[37]
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer.arXiv preprint arXiv:2408.06072, 2024
Pith/arXiv arXiv 2024
-
[38]
J. Yi, W. Jang, P. H. Cho, J. Nam, H. Yoon, and S. Kim. Deep forcing: Training-free long video generation with deep sink and participative compression.arXiv preprint arXiv:2512.05081, 2025
arXiv 2025
-
[39]
T. Yin, M. Gharbi, T. Park, R. Zhang, E. Shechtman, F. Durand, and B. Freeman. Improved distribu- tion matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455–47487, 2024
2024
-
[40]
T. Yin, M. Gharbi, R. Zhang, E. Shechtman, F. Durand, W. T. Freeman, and T. Park. One-step diffusion with distribution matching distillation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6613–6623, 2024
2024
-
[41]
T. Yin, Q. Zhang, R. Zhang, W. T. Freeman, F. Durand, E. Shechtman, and X. Huang. From slow bidirectional to fast autoregressive video diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22963–22974, 2025
2025
-
[42]
J. Yu, J. Bai, Y. Qin, Q. Liu, X. Wang, P. Wan, D. Zhang, and X. Liu. Context as memory: Scene- consistent interactive long video generation with memory retrieval. InProceedings of the SIGGRAPH Asia 2025 Conference Papers, pages 1–11, 2025
2025
-
[43]
Zhang, S
L. Zhang, S. Cai, M. Li, G. Wetzstein, and M. Agrawala. Frame context packing and drift prevention in next-frame-prediction video diffusion models. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025
2025
-
[44]
L. Zhang, S. Cai, M. Li, C. Zeng, B. Lu, A. Rao, S. Han, G. Wetzstein, and M. Agrawala. Pretraining frame preservation in autoregressive video memory compression.arXiv preprint arXiv:2512.23851, 2025
Pith/arXiv arXiv 2025
-
[45]
D. Zheng, Z. Huang, H. Liu, K. Zou, Y. He, F. Zhang, Y. Zhang, J. He, W.-S. Zheng, Y. Qiao, and Z. Liu. VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025. 15
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.