REVIEW 4 major objections 1 minor 102 references
From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms
T0 review · 4 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper claims that an explainable feature-based model can predict fidelity, fluency, and language use in consecutive interpreting well enough to support automated diagnostic feedback.
desk verdict Full text is the wrong paper; abstract-only review points to a plausible but unverified contribution for interpreting pedagogy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The framework is the combination of (1) feature engineering over construct-relevant, transparent signals—BLEURT and CometKiwi scores for semantic fidelity, pause-related temporal features for fluency, and phraseological diversity metrics for language use; (2) data augmentation to mitigate data scarcity and imbalance; and (3) Shapley Value (SHAP) analysis to attribute each prediction to these features. The SHAP step is what converts a predictive score into a diagnostic explanation.
What would settle it
A concrete test: take a new English-Chinese consecutive interpreting corpus with at least two independent human raters per recording; compute inter-rater reliability, train the same feature set, and check whether (a) the model's predictions correlate with held-out raters beyond chance, (b) BLEURT and CometKiwi remain the top SHAP features for fidelity, and (c) the SHAP rankings are stable across raters. If inter-rater agreement is low or feature rankings flip, the claim of reliable transparent prediction fails.
Extended reading notes
Core claim
The paper's central claim is that an interpretable feature-based model can predict the three standard dimensions of consecutive interpreting quality—fidelity, fluency, and language use—with enough accuracy to replace or augment human evaluation. Using only construct-relevant, transparent features and Shapley-Value-based explanations, the model identifies BLEURT and CometKiwi quality scores as the strongest predictors of fidelity, pause-related features as the strongest for fluency, and Chinese-specific phraseological diversity metrics as the strongest for language use. The authors position this as a move from 'black box' predictions to a scalable, reliable, transparent assessment that can pr
Load-bearing premise
The load-bearing premise is that the human quality scores used as training labels are valid and reliable, and that features trained on written translation quality (BLEURT, CometKiwi) and simple pause/diversity metrics transfer to spoken consecutive interpreting as raters judge it; if these do not hold, the reported predictive strength may just be the features re-encoding the target construct, especially given the paper's own concession of data scarcity and imbalance.
Editorial extensions
If this is right
- If the model works, automated interpreting assessment can move from a single opaque score to per-dimension diagnostic feedback (fidelity, fluency, language use) that learners can act on.
- SHAP-based feature attributions let learners see which concrete behaviors (e.g., long pauses, low lexical diversity) pulled their score down.
- The same explainable-feature pipeline could be adapted to other language pairs and to simultaneous interpreting if the feature definitions are re-specified.
- The finding that BLEURT and CometKiwi transfer to spoken interpreting would support using these metrics more widely in interpreting research rather than only in written translation.
Reading between the lines
- The supplied full text is a different manuscript (about physically plausible video generation), not the interpreting-assessment paper; the claims above rest on the abstract alone, and the body's empirical details (dataset size, augmentation scheme, model class) are not verifiable here.
- The reliance on BLEURT and CometKiwi as fidelity proxies for spoken interpreting is an untested transfer: written-translation metrics may encode source difficulty or text formality rather than interpreter error, so a good fidelity score could partly reflect the source text being easier.
- Pause features conflate strategic and disfluent pausing; a testable extension is to separate pause positions (clause boundaries vs. mid-clause) to sharpen fluency diagnostics.
- Data augmentation via rephrasing may create renditions that are fluent but inaccurate, so the augmentation step should be validated against human fidelity judgments before relying on it for training.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims to present a multi-dimensional, explainable machine-learning framework for automated interpreting quality assessment, using a novel English-Chinese consecutive interpreting dataset. According to the abstract, the model predicts three quality dimensions—fidelity, fluency, and language use—from transparent, construct-relevant features, with BLEURT and CometKiwi as strongest predictors of fidelity, pause-related features for fluency, and Chinese-specific phraseological diversity metrics for language use. The authors assert strong predictive performance and position the framework as a scalable, reliable, transparent alternative to human evaluation. However, the full text supplied for arXiv:2508.10860 is not this manuscript: it is arXiv:2508.10858v1, a video-generation paper titled 'Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation.' Consequently, none of the methods, experiments, feature sets, or results described in the abstract can be located or evaluated.
Significance. If the claimed results were properly supported, the work would address a real gap in automated interpreting assessment by combining feature-based prediction with SHAP-based explainability and by targeting three quality dimensions. However, the manuscript as supplied provides no evidence for these claims. The abstract itself contains no dataset size, no performance numbers, no comparison against human raters or black-box models, and no external validation. The paper's own mention of 'data scarcity and imbalance' further undercuts the generalization claim. The mismatch between the abstract and the supplied full text is a decisive, unverifiable defect: the central claims cannot be assessed at all.
major comments (4)
- [Full text (entire manuscript)] The body text provided is not the interpreting-assessment paper described in the abstract. It is the manuscript for arXiv:2508.10858v1, 'Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation' (PhysHPO), with a different title, different authors, different subject matter, and different results. Sections, equations, tables, and references all concern video generation, not consecutive interpreting, fidelity, fluency, or language-use assessment. The central claim of the abstract—strong predictive performance from BLEURT, CometKiwi, pause, and phraseological diversity features—therefore has no accompanying methodology or evidence in the submitted text. This is not a local omission but a complete absence of the claimed manuscript.
- [Abstract] Even taking the abstract as the sole evidence, 'strong predictive performance' is unsupported. No evaluation metrics (e.g., correlation with human raters, RMSE, accuracy, F1), no dataset size, no train/test split details, no baseline comparisons, and no confidence intervals are reported. The concluding claim that the framework is a 'scalable, reliable, and transparent alternative to traditional human evaluation' exceeds what an abstract without any quantitative result can support.
- [Abstract (feature validity)] The choice of BLEURT and CometKiwi as fidelity predictors raises a circularity concern that the manuscript does not address in the available text: both are learned metrics trained on human translation-quality judgments, so regressing human-rated fidelity on these scores may partly re-encode the target construct rather than independently measuring interpreting fidelity. The manuscript should demonstrate incremental validity over simpler baseline features and report how transfer from written translation metrics to spoken consecutive interpreting was validated. Since the full text is missing, this concern cannot be checked against any experimental design.
- [Abstract (data scarcity)] The abstract concedes 'data scarcity and imbalance' as a known modeling challenge, and the proposed solution is data augmentation followed by feature-based machine learning. For classroom-scale datasets, SHAP-based feature importance can be unstable, and augmentation can introduce label-preserving but construct-distorting samples. The submitted text contains no details on corpus size, augmentation technique, cross-validation scheme, or stability analysis, so the claim that the identified features are the strongest predictors is not established.
minor comments (1)
- [General] The manuscript should be re-submitted with the correct full text. In addition, the abstract should report concrete performance numbers and a comparison against a human-rater baseline, since those are essential for the 'reliable alternative to human evaluation' claim.
Circularity Check
No circularity demonstrated: the supplied full text is a different paper (PhysHPO, arXiv:2508.10858v1), so the interpreting paper's derivation chain cannot be inspected; the abstract alone shows no definitional reduction.
full rationale
The submitted full text is not the manuscript for arXiv:2508.10860; it is the PhysHPO video-generation paper. Because the equations, feature definitions, label construction, and model-fitting details of the interpreting paper are absent, I cannot exhibit the kind of reduction required by the circularity standard (e.g., Eq. X = Eq. Y by construction, or a fitted parameter renamed as a prediction). The abstract's statement that 'BLEURT and CometKiwi scores [are] the strongest predictive features for fidelity' does raise a construct-validity question, since those metrics are themselves trained on human quality judgments; however, the abstract does not define the fidelity labels as BLEURT/CometKiwi scores, so the claimed prediction is not equivalent to its inputs by definition. Similarly, the acknowledgment of 'data scarcity and imbalance' is a generalizability limitation, not a circular step. Under the hard rules, no circularity can be flagged without quoteable evidence of a specific reduction, and no such evidence is available from the material provided.
Assumptions & free parameters
free parameters (2)
- Construct-relevant feature set (BLEURT, CometKiwi, pause, phraseological diversity metrics)
- Data augmentation strategy
assumptions (3)
- domain assumption Human quality scores on the novel English-Chinese consecutive interpreting dataset are a valid and reliable ground truth.
- domain assumption BLEURT and CometKiwi, pretrained on written translation quality judgments, transfer to spoken consecutive interpreting as valid fidelity proxies.
- domain assumption SHAP values on the trained model faithfully reveal the true relationship between each feature and the rated quality construct.
Cite this review
Pith. "Pith review of From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms." pith.science (2026). https://pith.science/paper/AWZCY3FL
@misc{pith2026250810860,
author = {Pith},
title = {Pith review of: From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms},
year = {2026},
howpublished = {\url{https://pith.science/paper/AWZCY3FL}},
note = {Machine review of arXiv:2508.10860}
}
read the original abstract
Recent advancements in machine learning have spurred growing interests in automated interpreting quality assessment. Nevertheless, existing research suffers from insufficient examination of language use quality, unsatisfactory modeling effectiveness due to data scarcity and imbalance, and a lack of efforts to explain model predictions. To address these gaps, we propose a multi-dimensional modeling framework that integrates feature engineering, data augmentation, and explainable machine learning. This approach prioritizes explainability over ``black box'' predictions by utilizing only construct-relevant, transparent features and conducting Shapley Value (SHAP) analysis. Our results demonstrate strong predictive performance on a novel English-Chinese consecutive interpreting dataset, identifying BLEURT and CometKiwi scores to be the strongest predictive features for fidelity, pause-related features for fluency, and Chinese-specific phraseological diversity metrics for language use. Overall, by placing particular emphasis on explainability, we present a scalable, reliable, and transparent alternative to traditional human evaluation, facilitating the provision of detailed diagnostic feedback for learners and supporting self-regulated learning advantages not afforded by automated scores in isolation.
Reference graph
Works this paper leans on
-
[1]
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Morcos. Semd- edup: Data-efficient learning at web-scale through semantic deduplication.arXiv preprint arXiv:2303.09540, 2023
arXiv 2023
-
[2]
Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025
Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025
arXiv 2025
-
[3]
A survey on data selection for language models.arXiv preprint arXiv:2402.16827, 2024
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. A survey on data selection for language models.arXiv preprint arXiv:2402.16827, 2024
arXiv 2024
-
[4]
Qwen2.5-vl technical report.arXiv preprint arXiv:2502.13923, 2025
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report.ar...
arXiv 2025
-
[5]
Impossible videos.arXiv preprint arXiv:2503.14378, 2025
Zechen Bai, Hai Ci, and Mike Zheng Shou. Impossible videos.arXiv preprint arXiv:2503.14378, 2025
arXiv 2025
-
[6]
Videophy: Evaluating physical commonsense for video generation.arXiv preprint arXiv:2406.03520, 2024
Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chen- fanfu Jiang, Yizhou Sun, Kai-Wei Chang, and Aditya Grover. Videophy: Evaluating physical commonsense for video generation.arXiv preprint arXiv:2406.03520, 2024
arXiv 2024
-
[7]
Color-filter: Conditional loss reduction filtering for targeted language model pre- training.Advances in Neural Information Processing Systems, 37:97618–97649, 2024
David Brandfonbrener, Hanlin Zhang, Andreas Kirsch, Jonathan Richard Schwarz, and Sham Kakade. Color-filter: Conditional loss reduction filtering for targeted language model pre- training.Advances in Neural Information Processing Systems, 37:97618–97649, 2024
2024
-
[8]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. 2024. URL https://openai.com/research/ video-generation-models-as-world-simulators
2024
Show all 102 references
-
[9]
Dspo: Direct semantic preference optimization for real-world image super-resolution.arXiv preprint arXiv:2504.15176, 2025
Miaomiao Cai, Simiao Li, Wei Li, Xudong Huang, Hanting Chen, Jie Hu, and Yunhe Wang. Dspo: Direct semantic preference optimization for real-world image super-resolution.arXiv preprint arXiv:2504.15176, 2025
2025 arXiv
-
[10]
Skyreels-v2: Infinite-length film generative model.arXiv preprint arXiv:2504.13074, 2025
Guibin Chen, Dixuan Lin, Jiangping Yang, Chunze Lin, Juncheng Zhu, Mingyuan Fan, Hao Zhang, Sheng Chen, Zheng Chen, Chengchen Ma, et al. Skyreels-v2: Infinite-length film generative model.arXiv preprint arXiv:2504.13074, 2025
2025 arXiv
-
[12]
Beyond generation: Unlocking universal editing via self-supervised fine-tuning.arXiv preprint arXiv:2412.02114, 2024
Harold Haodong Chen, Harry Yang, and Ser-Nam Lim. Beyond generation: Unlocking universal editing via self-supervised fine-tuning.arXiv preprint arXiv:2412.02114, 2024
2024 arXiv
-
[13]
Temporal regularization makes your video generator stronger
Harold Haodong Chen, Haojian Huang, Xianfeng Wu, Yexin Liu, Yajing Bai, Wen-Jie Shu, Harry Yang, and Ser-Nam Lim. Temporal regularization makes your video generator stronger. arXiv preprint arXiv:2503.15417, 2025
2025 arXiv
-
[14]
Alpagasus: Training a better alpaca with fewer data.arXiv preprint arXiv:2307.08701, 2023
Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, et al. Alpagasus: Training a better alpaca with fewer data.arXiv preprint arXiv:2307.08701, 2023
2023 arXiv
-
[15]
Goku: Flow based video generative foundation models.arXiv preprint arXiv:2502.04896, 2025
Shoufa Chen, Chongjian Ge, Yuqi Zhang, Yida Zhang, Fengda Zhu, Hao Yang, Hongxiang Hao, Hui Wu, Zhichao Lai, Yifei Hu, Ting-Che Lin, Shilong Zhang, Fu Li, Chuan Li, Xing Wang, Yanghua Peng, Peize Sun, Ping Luo, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Xiaobing Liu. Goku: Flow ...
2025 arXiv
-
[16]
Discriminator-free direct preference optimization for video diffusion.arXiv preprint arXiv:2504.08542, 2025
Haoran Cheng, Qide Dong, Liang Peng, Zhizhou Sha, Weiguo Feng, Jinghui Xie, Zhao Song, Shilei Wen, Xiaofei He, and Boxi Wu. Discriminator-free direct preference optimization for video diffusion.arXiv preprint arXiv:2504.08542, 2025
2025 arXiv
-
[17]
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna
Wei-Lin Chiang, Zhuohan Li, Ziqing Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna. lmsys. org (accessed 14 April 2023)...
2023
-
[18]
Ultrafeedback: Boosting language models with scaled ai feedback.arXiv preprint arXiv:2310.01377, 2023
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan Ni, Guotong Xie, Ruobing Xie, Yankai Lin, et al. Ultrafeedback: Boosting language models with scaled ai feedback.arXiv preprint arXiv:2310.01377, 2023
2023 arXiv
-
[19]
One-minute video generation with test-time training.arXiv preprint arXiv:2504.05298, 2025
Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Chun Cheung, Jan Kautz, Carlos Guestrin, et al. One-minute video generation with test-time training.arXiv preprint arXiv:2504.05298, 2025
2025 arXiv
-
[20]
Enhancing chat language models by scaling high-quality instructional conversations.arXiv preprint arXiv:2305.14233, 2023
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. Enhancing chat language models by scaling high-quality instructional conversations.arXiv preprint arXiv:2305.14233, 2023
2023 arXiv
-
[21]
What’s in my big data?arXiv preprint arXiv:2310.20707, 2023
Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, et al. What’s in my big data?arXiv preprint arXiv:2310.20707, 2023
2023 arXiv
-
[22]
Wave: Warping ddim inversion features for zero-shot text-to-video editing
Yutang Feng, Sicheng Gao, Yuxiang Bao, Xiaodi Wang, Shumin Han, Juan Zhang, Baochang Zhang, and Angela Yao. Wave: Warping ddim inversion features for zero-shot text-to-video editing. InEuropean Conference on Computer Vision, pages 38–55. Springer, 2024
2024
-
[23]
CHip: Cross-modal hierarchical direct preference optimization for multimodal LLMs
Jinlan Fu, huangfushenzhen, Hao Fei, Xiaoyu Shen, Bryan Hooi, Xipeng Qiu, and See- Kiong Ng. CHip: Cross-modal hierarchical direct preference optimization for multimodal LLMs. InThe Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.n...
2025
-
[24]
Clustering and ranking: Diversity-preserved instruc- tion selection through expert-aligned quality estimation.arXiv preprint arXiv:2402.18191, 2024
Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Hongxia Ma, Li Zhang, Boxing Chen, Hao Yang, et al. Clustering and ranking: Diversity-preserved instruc- tion selection through expert-aligned quality estimation.arXiv preprint arXiv:2402.18191, 2024
2024 arXiv
-
[25]
Task-adaptive pretrained lan- guage models via clustered-importance sampling
David Grangier, Simin Fan, Skyler Seto, and Pierre Ablin. Task-adaptive pretrained lan- guage models via clustered-importance sampling. InThe Thirteenth International Confer- ence on Learning Representations, 2025. URL https://openreview.net/forum?id= p6ncr0eTKE
2025
-
[26]
A survey on llm-as-a-judge.arXiv preprint arXiv:2411.15594, 2024
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. A survey on llm-as-a-judge.arXiv preprint arXiv:2411.15594, 2024
2024 arXiv
-
[27]
Detecting and preventing hallucinations in large vision language models
Anisha Gunjal, Jihan Yin, and Erhan Bas. Detecting and preventing hallucinations in large vision language models. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 18135–18143, 2024
2024
-
[28]
Long context tuning for video generation.arXiv preprint arXiv:2503.10589, 2025
Yuwei Guo, Ceyuan Yang, Ziyan Yang, Zhibei Ma, Zhijie Lin, Zhenheng Yang, Dahua Lin, and Lu Jiang. Long context tuning for video generation.arXiv preprint arXiv:2503.10589, 2025
2025 arXiv
-
[29]
Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
2020
-
[30]
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Li Hu. Animate anyone: Consistent and controllable image-to-video synthesis for character animation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8153–8163, 2024. 11
2024
-
[31]
Vistadpo: Video hierarchical spatial-temporal direct preference optimization for large video models.arXiv preprint arXiv:2504.13122, 2025
Haojian Huang, Haodong Chen, Shengqiong Wu, Meng Luo, Jinlan Fu, Xinya Du, Han- wang Zhang, and Hao Fei. Vistadpo: Video hierarchical spatial-temporal direct preference optimization for large video models.arXiv preprint arXiv:2504.13122, 2025
2025 arXiv
-
[32]
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Information Systems, 43(2):1–55, 2025
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qiang- long Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Informati...
2025
-
[33]
Conceptmaster: Multi-concept video customization on diffusion transformer models without test-time tuning.arXiv preprint arXiv:2501.04698, 2025
Yuzhou Huang, Ziyang Yuan, Quande Liu, Qiulin Wang, Xintao Wang, Ruimao Zhang, Pengfei Wan, Di Zhang, and Kun Gai. Conceptmaster: Multi-concept video customization on diffusion transformer models without test-time tuning.arXiv preprint arXiv:2501.04698, 2025
2025 arXiv
-
[34]
VBench: Comprehensive benchmark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive benchmark suite for video generative models. I...
2024
-
[35]
Camels in a changing climate: Enhancing lm adaptation with tulu 2.arXiv preprint arXiv:2311.10702, 2023
Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A Smith, Iz Beltagy, et al. Camels in a changing climate: Enhancing lm adaptation with tulu 2.arXiv preprint arXiv:2311.10702, 2023
2023 arXiv
-
[36]
Huvidpo: Enhancing video generation through direct preference optimization for human-centric alignment.arXiv preprint arXiv:2502.01690, 2025
Lifan Jiang, Boxi Wu, Jiahui Zhang, Xiaotong Guan, and Shuang Chen. Huvidpo: Enhancing video generation through direct preference optimization for human-centric alignment.arXiv preprint arXiv:2502.01690, 2025
2025 arXiv
-
[37]
Miradata: A large-scale video dataset with long durations and structured captions.Advances in Neural Information Processing Systems, 37:48955–48970, 2024
Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xintao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long durations and structured captions.Advances in Neural Information Processing Systems, 37:48955–48970, 2024
2024
-
[38]
How far is video generation from world model: A physical law perspective.arXiv preprint arXiv:2411.02385, 2024
Bingyi Kang, Yang Yue, Rui Lu, Zhijie Lin, Yang Zhao, Kaixin Wang, Gao Huang, and Jiashi Feng. How far is video generation from world model: A physical law perspective.arXiv preprint arXiv:2411.02385, 2024
2024 arXiv
-
[39]
Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024
2024 arXiv
-
[40]
Differentiable physics simulation of dynamics- augmented neural objects.IEEE Robotics and Automation Letters, 8(5):2780–2787, 2023
Simon Le Cleac’h, Hong-Xing Yu, Michelle Guo, Taylor Howell, Ruohan Gao, Jiajun Wu, Zachary Manchester, and Mac Schwager. Differentiable physics simulation of dynamics- augmented neural objects.IEEE Robotics and Automation Letters, 8(5):2780–2787, 2023
2023
-
[41]
Pisa experiments: Exploring physics post-training for video diffusion models by watching stuff drop.arXiv preprint arXiv:2503.09595, 2025
Chenyu Li, Oscar Michel, Xichen Pan, Sainan Liu, Mike Roberts, and Saining Xie. Pisa experiments: Exploring physics post-training for video diffusion models by watching stuff drop.arXiv preprint arXiv:2503.09595, 2025
2025 arXiv
-
[42]
Worldmodelbench: Judging video generation models as world models.arXiv preprint arXiv:2502.20694, 2025
Dacheng Li, Yunhao Fang, Yukang Chen, Shuo Yang, Shiyi Cao, Justin Wong, Michael Luo, Xiaolong Wang, Hongxu Yin, Joseph E Gonzalez, et al. Worldmodelbench: Judging video generation models as world models.arXiv preprint arXiv:2502.20694, 2025
2025 arXiv
-
[43]
Magicid: Hybrid preference optimization for id-consistent and dynamic-preserved video customization.arXiv preprint arXiv:2503.12689, 2025
Hengjia Li, Lifan Jiang, Xi Xiao, Tianyang Wang, Hongwei Yi, Boxi Wu, and Deng Cai. Magicid: Hybrid preference optimization for id-consistent and dynamic-preserved video customization.arXiv preprint arXiv:2503.12689, 2025
2025 arXiv
-
[44]
Science-t2i: Addressing scientific illusions in image synthesis.arXiv preprint arXiv:2504.13129, 2025
Jialuo Li, Wenhao Chai, Xingyu Fu, Haiyang Xu, and Saining Xie. Science-t2i: Addressing scientific illusions in image synthesis.arXiv preprint arXiv:2504.13129, 2025
2025
-
[45]
From quantity to quality: Boosting llm performance with self-guided data selection for instruction tuning.arXiv preprint arXiv:2308.12032, 2023
Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, and Jing Xiao. From quantity to quality: Boosting llm performance with self-guided data selection for instruction tuning.arXiv preprint arXiv:2308.12032, 2023. 12
2023 arXiv
-
[46]
Selective reflection-tuning: Student-selected data recycling for llm instruction-tuning
Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Jiuxiang Gu, and Tianyi Zhou. Selective reflection-tuning: Student-selected data recycling for llm instruction-tuning. InFindings of the Association for Computational Linguistics ACL 2024, pages 16189–16211, 2024
2024
-
[47]
Superfiltering: Weak-to-strong data filtering for fast instruction-tuning.arXiv preprint arXiv:2402.00530, 2024
Ming Li, Yong Zhang, Shwai He, Zhitao Li, Hongyu Zhao, Jianzong Wang, Ning Cheng, and Tianyi Zhou. Superfiltering: Weak-to-strong data filtering for fast instruction-tuning.arXiv preprint arXiv:2402.00530, 2024
2024 arXiv
-
[48]
Open-sora plan: Open-source large video generation model.arXiv preprint arXiv:2412.00131, 2024
Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, et al. Open-sora plan: Open-source large video generation model.arXiv preprint arXiv:2412.00131, 2024
2024 arXiv
-
[49]
Yu, and Meng Cao
Aiwei Liu, Haoping Bai, Zhiyun Lu, Yanchao Sun, Xiang Kong, Xiaoming Simon Wang, Jiulong Shan, Albin Madappally Jose, Xiaojiang Liu, Lijie Wen, Philip S. Yu, and Meng Cao. TIS-DPO: Token-level importance sampling for direct preference optimization with estimated weights. InThe...
2025
-
[50]
Safetydpo: Scalable safety alignment for text-to-image generation.arXiv preprint arXiv:2412.10493, 2024
Runtao Liu, Chen I Chieh, Jindong Gu, Jipeng Zhang, Renjie Pi, Qifeng Chen, Philip Torr, Ashkan Khakzar, and Fabio Pizzati. Safetydpo: Scalable safety alignment for text-to-image generation.arXiv preprint arXiv:2412.10493, 2024
2024 arXiv
-
[51]
Videodpo: Omni-preference alignment for video diffusion generation.arXiv preprint arXiv:2412.14167, 2024
Runtao Liu, Haoyu Wu, Zheng Ziqiang, Chen Wei, Yingqing He, Renjie Pi, and Qifeng Chen. Videodpo: Omni-preference alignment for video diffusion generation.arXiv preprint arXiv:2412.14167, 2024
2024 arXiv
-
[52]
Physgen: Rigid-body physics-grounded image-to-video generation
Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, and Shenlong Wang. Physgen: Rigid-body physics-grounded image-to-video generation. InEuropean Conference on Computer Vision, pages 360–378. Springer, 2024
2024
-
[53]
What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning
Wei Liu, Weihao Zeng, Keqing He, Yong Jiang, and Junxian He. What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning. In The Twelfth International Conference on Learning Representations, 2024. URL https: //openreview.net/foru...
2024
-
[54]
Towards world simulator: Crafting physical commonsense- based benchmark for video generation.arXiv preprint arXiv:2410.05363, 2024
Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao, and Ping Luo. Towards world simulator: Crafting physical commonsense- based benchmark for video generation.arXiv preprint arXiv:2410.05363, 2024
2024 arXiv
-
[55]
Motioncraft: Physics-based zero-shot video generation.Advances in Neural Information Processing Systems, 37:123155–123181, 2024
Antonio Montanaro, Luca Savant Aira, Emanuele Aiello, Diego Valsesia, and Enrico Magli. Motioncraft: Physics-based zero-shot video generation.Advances in Neural Information Processing Systems, 37:123155–123181, 2024
2024
-
[56]
Do generative video models learn physical principles from watching videos?arXiv preprint arXiv:2501.09038, 2025
Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, and Robert Geirhos. Do generative video models learn physical principles from watching videos?arXiv preprint arXiv:2501.09038, 2025
2025 arXiv
-
[57]
Openvid-1m: A large-scale high-quality dataset for text-to-video generation.arXiv preprint arXiv:2407.02371, 2024
Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Zhijie Chen, Xiang Li, Jian Yang, and Ying Tai. Openvid-1m: A large-scale high-quality dataset for text-to-video generation.arXiv preprint arXiv:2407.02371, 2024
2024 arXiv
-
[58]
Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730...
2022
-
[59]
G- dig: Towards gradient-based diverse and high-quality instruction data selection for machine translation.arXiv preprint arXiv:2405.12915, 2024
Xingyuan Pan, Luyang Huang, Liyan Kang, Zhicheng Liu, Yu Lu, and Shanbo Cheng. G- dig: Towards gradient-based diverse and high-quality instruction data selection for machine translation.arXiv preprint arXiv:2405.12915, 2024
2024 arXiv
-
[60]
Freenoise: Tuning-free longer video diffusion via noise rescheduling.arXiv preprint arXiv:2310.15169, 2023
Haonan Qiu, Menghan Xia, Yong Zhang, Yingqing He, Xintao Wang, Ying Shan, and Ziwei Liu. Freenoise: Tuning-free longer video diffusion via noise rescheduling.arXiv preprint arXiv:2310.15169, 2023. 13
2023 arXiv
-
[61]
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36:53728–53741, 2023
2023
-
[62]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022
2022
-
[63]
Towards nsfw-free text-to-image generation via safety-constraint direct preference optimization.arXiv preprint arXiv:2504.14290, 2025
Shouwei Ruan, Zhenyu Wu, Yao Huang, Ruochen Zhang, Yitong Sun, Caixin Kang, and Xingx- ing Wei. Towards nsfw-free text-to-image generation via safety-constraint direct preference optimization.arXiv preprint arXiv:2504.14290, 2025
2025
-
[64]
Seaweed-7b: Cost-effective training of video generation foundation model.arXiv preprint arXiv:2504.08685, 2025
Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al. Seaweed-7b: Cost-effective training of video generation foundation model.arXiv preprint arXiv:2504.08685, 2025
2025 arXiv
-
[65]
Finephys: Fine-grained human action generation by explicitly incorporating physical laws for effective skeletal guidance
Dian Shao, Mingfei Shi, Shengda Xu, Haodong Chen, Yongle Huang, and Binglu Wang. Finephys: Fine-grained human action generation by explicitly incorporating physical laws for effective skeletal guidance. InProceedings of the Computer Vision and Pattern Recognition Conference, p...
1905
-
[66]
Deep unsuper- vised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsuper- vised learning using nonequilibrium thermodynamics. InInternational conference on machine learning, pages 2256–2265. pmlr, 2015
2015
-
[67]
Conifer: Improving complex constrained instruction-following ability of large language models
Haoran Sun, Lixin Liu, Junjie Li, Fengyu Wang, Baohua Dong, Ran Lin, and Ruohui Huang. Conifer: Improving complex constrained instruction-following ability of large language models. arXiv preprint arXiv:2404.02823, 2024
2024 arXiv
-
[68]
Dsv: Exploiting dynamic sparsity to accelerate large-scale video dit training
Xin Tan, Yuetao Chen, Yimin Jiang, Xing Chen, Kun Yan, Nan Duan, Yibo Zhu, Daxin Jiang, and Hong Xu. Dsv: Exploiting dynamic sparsity to accelerate large-scale video dit training. arXiv preprint arXiv:2502.07590, 2025
2025
-
[69]
Stanford alpaca: An instruction-following llama model, 2023
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. Stanford alpaca: An instruction-following llama model, 2023
2023
-
[70]
D4: Improving llm pretraining via document de-duplication and diversification.Advances in Neural Information Processing Systems, 36:53983–53995, 2023
Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari Morcos. D4: Improving llm pretraining via document de-duplication and diversification.Advances in Neural Information Processing Systems, 36:53983–53995, 2023
2023
-
[71]
Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timo- thée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[72]
Diffusion model alignment using direct preference optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model alignment using direct preference optimization. InProceedings of the IEEE/CVF Conference on Computer Vision and ...
2024
-
[73]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
2025 arXiv
-
[74]
A survey on data selection for llm instruction tuning.arXiv preprint arXiv:2402.05123, 2024
Jiahao Wang, Bolin Zhang, Qianlong Du, Jiajun Zhang, and Dianhui Chu. A survey on data selection for llm instruction tuning.arXiv preprint arXiv:2402.05123, 2024
2024 arXiv
-
[75]
Wisa: World simulator assistant for physics-aware text-to- video generation.arXiv preprint arXiv:2503.08153, 2025
Jing Wang, Ao Ma, Ke Cao, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng, et al. Wisa: World simulator assistant for physics-aware text-to- video generation.arXiv preprint arXiv:2503.08153, 2025
2025 arXiv
-
[76]
Internvid: A large-scale video-text dataset for multimodal understanding and generation.arXiv preprint arXiv:2307.06942, 2023
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation.arXiv preprint arXiv:2307.06942, 2023. 14
2023 arXiv
-
[77]
Self-instruct: Aligning language models with self-generated instruc- tions.arXiv preprint arXiv:2212.10560, 2022
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language models with self-generated instruc- tions.arXiv preprint arXiv:2212.10560, 2022
2022 arXiv
-
[78]
Lightgen: Efficient image generation through knowledge distillation and direct preference optimization.arXiv preprint arXiv:2503.08619, 2025
Xianfeng Wu, Yajing Bai, Haoze Zheng, Harold Haodong Chen, Yexin Liu, Zihao Wang, Xuran Ma, Wen-Jie Shu, Xianzu Wu, Harry Yang, et al. Lightgen: Efficient image generation through knowledge distillation and direct preference optimization.arXiv preprint arXiv:2503.08619, 2025
2025 arXiv
-
[79]
Deepseek-vl2: Mixture-of-experts vision- language models for advanced multimodal understanding.arXiv preprint arXiv:2412.10302, 2024
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, et al. Deepseek-vl2: Mixture-of-experts vision- language models for advanced multimodal understanding.arXiv preprint arXiv:2412.10302, 2024
2024 arXiv
-
[80]
LESS: Selecting influential data for targeted instruction tuning
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. LESS: Selecting influential data for targeted instruction tuning. InInternational Conference on Machine Learning (ICML), 2024
2024
-
[81]
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang. Data selection for language models via importance resampling. InThirty-seventh Conference on Neural Information Processing Systems, 2023. URLhttps://openreview.net/forum?id=uPSQv0leAu
2023
-
[82]
Tooncrafter: Generative cartoon interpolation.ACM Transactions on Graphics (TOG), 43(6):1–11, 2024
Jinbo Xing, Hanyuan Liu, Menghan Xia, Yong Zhang, Xintao Wang, Ying Shan, and Tien-Tsin Wong. Tooncrafter: Generative cartoon interpolation.ACM Transactions on Graphics (TOG), 43(6):1–11, 2024
2024
-
[83]
Make-your-video: Customized video generation using textual and structural guidance.IEEE Transactions on Visualization and Computer Graphics, 2024
Jinbo Xing, Menghan Xia, Yuxin Liu, Yuechen Zhang, Yong Zhang, Yingqing He, Hanyuan Liu, Haoxin Chen, Xiaodong Cun, Xintao Wang, et al. Make-your-video: Customized video generation using textual and structural guidance.IEEE Transactions on Visualization and Computer Graphics, 2024
2024
-
[84]
Wizardlm: Empowering large language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244, 2023
2023 arXiv
-
[85]
Phyt2v: Llm-guided iterative self- refinement for physics-grounded text-to-video generation.arXiv preprint arXiv:2412.00596, 2024
Qiyao Xue, Xiangyu Yin, Boyuan Yang, and Wei Gao. Phyt2v: Llm-guided iterative self- refinement for physics-grounded text-to-video generation.arXiv preprint arXiv:2412.00596, 2024
2024 arXiv
-
[86]
Qwen2.5 technical report.arXiv preprint arXiv:2412.15115, 2024
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Me...
2024 arXiv
-
[87]
Rethinking video tokenization: A conditioned diffusion-based approach.arXiv preprint arXiv:2503.03708, 2025
Nianzu Yang, Pandeng Li, Liming Zhao, Yang Li, Chen-Wei Xie, Yehui Tang, Xudong Lu, Zhihang Liu, Yun Zheng, Yu Liu, et al. Rethinking video tokenization: A conditioned diffusion-based approach.arXiv preprint arXiv:2503.03708, 2025
2025 arXiv
-
[88]
Vlipp: Towards physically plausible video generation with vision and language informed physical prior.arXiv e-prints, pages arXiv–2503, 2025
Xindi Yang, Baolu Li, Yiming Zhang, Zhenfei Yin, Lei Bai, Liqian Ma, Zhiyong Wang, Jianfei Cai, Tien-Tsin Wong, Huchuan Lu, et al. Vlipp: Towards physically plausible video generation with vision and language informed physical prior.arXiv e-prints, pages arXiv–2503, 2025
2025
-
[89]
Decoding data quality via synthetic corruptions: Embedding-guided pruning of code data.arXiv preprint arXiv:2312.02418, 2023
Yu Yang, Aaditya K Singh, Mostafa Elhoushi, Anas Mahmoud, Kushal Tirumala, Fabian Gloeckle, Baptiste Rozière, Carole-Jean Wu, Ari S Morcos, and Newsha Ardalani. Decoding data quality via synthetic corruptions: Embedding-guided pruning of code data.arXiv preprint arXiv:2312.02418, 2023
2023 arXiv
-
[90]
Cogvideox: Text-to-video diffusion models with an expert transformer.arXiv preprint arXiv:2408.06072, 2024
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer.arXiv preprint arXiv:2408.06072, 2024. 15
2024 arXiv
-
[91]
Limo: Less is more for reasoning.arXiv preprint arXiv:2502.03387, 2025
Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, and Pengfei Liu. Limo: Less is more for reasoning.arXiv preprint arXiv:2502.03387, 2025
2025 arXiv
-
[92]
Gamefactory: Creating new games with generative interactive videos.arXiv preprint arXiv:2501.08325, 2025
Jiwen Yu, Yiran Qin, Xintao Wang, Pengfei Wan, Di Zhang, and Xihui Liu. Gamefactory: Creating new games with generative interactive videos.arXiv preprint arXiv:2501.08325, 2025
2025
-
[93]
Magictime: Time-lapse video generation models as metamorphic simulators.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
Shenghai Yuan, Jinfa Huang, Yujun Shi, Yongqi Xu, Ruijie Zhu, Bin Lin, Xinhua Cheng, Li Yuan, and Jiebo Luo. Magictime: Time-lapse video generation models as metamorphic simulators.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
2025
-
[94]
Onlinevpo: Align video diffusion model with online video-centric preference optimization
Jiacheng Zhang, Jie Wu, Weifeng Chen, Yatai Ji, Xuefeng Xiao, Weilin Huang, and Kai Han. Onlinevpo: Align video diffusion model with online video-centric preference optimization. arXiv preprint arXiv:2412.15159, 2024
2024
-
[95]
Tagcos: Task-agnostic gradient clustered coreset selection for instruction tuning data.arXiv preprint arXiv:2407.15235, 2024
Jipeng Zhang, Yaxuan Qin, Renjie Pi, Weizhong Zhang, Rui Pan, and Tong Zhang. Tagcos: Task-agnostic gradient clustered coreset selection for instruction tuning data.arXiv preprint arXiv:2407.15235, 2024
2024 arXiv
-
[96]
Packing input frame contexts in next-frame prediction models for video generation.arXiv preprint arXiv:2504.12626, 2025
Lvmin Zhang and Maneesh Agrawala. Packing input frame contexts in next-frame prediction models for video generation.arXiv preprint arXiv:2504.12626, 2025
2025
-
[97]
Fast video generation with sliding tile attention, 2025
Peiyuan Zhang, Yongqi Chen, Runlong Su, Hangliang Ding, Ion Stoica, Zhenghong Liu, and Hao Zhang. Fast video generation with sliding tile attention, 2025. URL https://arxiv. org/abs/2502.04507
2025 arXiv
-
[98]
Synthetic video enhances physical fidelity in video synthesis.arXiv preprint arXiv:2503.20822, 2025
Qi Zhao, Xingyu Ni, Ziyu Wang, Feng Cheng, Ziyan Yang, Lu Jiang, and Bohan Wang. Synthetic video enhances physical fidelity in video synthesis.arXiv preprint arXiv:2503.20822, 2025
2025 arXiv
-
[99]
Open-sora: Democratizing efficient video production for all
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all. arXiv preprint arXiv:2412.20404, 2024
2024 arXiv
-
[100]
Deco: Decoupled human-centered diffusion video editing with motion consistency
Xiaojing Zhong, Xinyi Huang, Xiaofeng Yang, Guosheng Lin, and Qingyao Wu. Deco: Decoupled human-centered diffusion video editing with motion consistency. InEuropean Conference on Computer Vision, pages 352–370. Springer, 2024
2024
-
[101]
Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023
2023
-
[102]
Aligning anime video generation with human feedback.arXiv preprint arXiv:2504.10044, 2025
Bingwen Zhu, Yudong Jiang, Baohan Xu, Siqian Yang, Mingyu Yin, Yidi Wu, Huyang Sun, and Zuxuan Wu. Aligning anime video generation with human feedback.arXiv preprint arXiv:2504.10044, 2025
2025 arXiv
-
[103]
Champ: Controllable and consistent human image animation with 3d parametric guidance
Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Zilong Dong, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu. Champ: Controllable and consistent human image animation with 3d parametric guidance. InEuropean Conference on Computer Vision, pages 145–162. Springer, 2024. 16 A Mor...
2024
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.