Pith. sign in

REVIEW 4 major objections 6 minor 32 references

Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform

T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Tiger200K is a manually curated 170,000-clip bilingual video dataset built to give open text-to-video models a higher-quality fine-tuning resource.

desk verdict A genuinely new 170k-clip bilingual video dataset with a transparent pipeline, but the 'high visual quality' claim is self-validated and needs external evidence. read the letter →

arxiv 2504.15182 v1 pith:2STWHH45 submitted 2025-04-21 cs.CV

classification cs.CV
keywords videodatasettext-to-videogenerationuser-generatedcontentdatacurationbilingualcaptioningshotboundarydetectionsafezonefine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Tiger200K is built on a simple premise: for fine-tuning a video generation model, a small set of videos that humans have judged to look good beats a huge algorithmically filtered set. The paper claims that existing open datasets such as Koala-36M, though large, do not meet the visual quality bar for post-training or quality-tuning, and that careful human selection of creators and clips from UGC platforms can close that gap. The result offered is an open dataset of 170,000 fixed-length, temporally consistent clips from 4,151 source videos, each with Chinese and English captions plus safe-zone crop information. If the claim holds, the dataset gives the open-source community a directly usable resource for the supervised fine-tuning stage of text-to-video models.

What carries the argument

The load-bearing object is the manually selected source collection combined with the safe-zone computation. A safe zone is the region of a frame left after subtracting watermarks, subtitles, logos, and black borders, found by running PaddleOCR over every frame and scanning for persistent black regions, and clips whose safe zone falls below half the frame are discarded. Around this, the pipeline uses TransNetV2 for shot boundary detection with scenes shorter than 121 frames dropped, optical-flow-based motion filtering to remove static clips, and Qwen2.5-VL to produce dense bilingual captions. The paper's argument is that these mechanisms preserve temporal consistency and clean frames, while the manual curation at the front end supplies the aesthetic quality that algorithmic filtering misses.

What would settle it

A controlled fine-tuning experiment would settle the claim: fine-tune identical copies of one open video generation model on Tiger200K and on a matched random subset of Koala-36M with comparable clip counts and captions, then compare outputs on a fixed prompt suite through a human preference study or automated quality metric. If the Tiger200K-tuned model does not show measurably better visual quality or prompt adherence, the paper's quality claim is not supported. A second check is to compute an objective aesthetic score distribution on random clips from both datasets; if the distributions overlap heavily, the claimed quality gap disappears.

Watch

Extended reading notes

Core claim

The central claim is that human expertise in data curation, applied at the input stage, is what separates data good enough for fine-tuning from data merely good enough for pretraining. The paper argues that UGC platforms now contain professionally made content, and that selecting top creators, searching by camera model and production keywords, and relying on recommendation systems yields videos whose visual and aesthetic quality exceeds what threshold-based algorithmic filtering of older web-scraped video can guarantee. The dataset then applies a pipeline of TransNetV2 shot detection, OCR and border-based safe-zone computation, optical-flow motion filtering, and Qwen2.5-VL bilingual captioning to turn those source videos into 85,000 scenes and 170,000 fixed 121-frame cuts. The intended contribution is a high-visual-quality, temporally consistent, bilingual video-text corpus for post-training and quality-tuning of video generation models.

Load-bearing premise

The load-bearing premise is that the author's own judgment of 'visual and aesthetic quality' during manual selection and review is a reliable measure of the quality that matters for fine-tuning video generation models, since the paper never defines or measures that quality against an external benchmark.

Editorial extensions

If this is right

  • Open-source text-to-video models gain a fine-tuning set of 170,000 clips whose frames are cropped to overlay-free safe zones and whose captions support both Chinese and English prompt following.
  • Fine-tuning on Tiger200K should shift generation quality toward the polished, high-resolution look of professionally produced UGC rather than the average quality of large web-scraped corpora.
  • The per-clip bilingual captions make the dataset usable as a benchmark for caption quality and video-text alignment, not only as training material.
  • Because half the source videos are 4K or above, the dataset also supplies high-resolution material for video super-resolution and high-resolution generation research.
  • The pipeline itself is reusable: shot detection, safe-zone cropping, motion filtering, and VLM captioning can be applied to any new UGC source for continued dataset expansion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper does not run is a controlled fine-tuning comparison: train the same video model on Tiger200K and on an equally sized random sample of Koala-36M, then measure generation quality and prompt adherence on a held-out prompt set.
  • The subjective quality claim could be made measurable by collecting pairwise human preferences between clips from Tiger200K and clips from existing open datasets; the paper leaves that quantification to future work.
  • If the manual-curation approach transfers, the same platform-focused strategy could be applied to other regional UGC platforms to produce culturally diverse high-quality corpora rather than one platform's aesthetic.
  • One implicit consequence is that dataset curation effort may shift from building smarter automatic filters to building better creator-discovery and review workflows, since the paper locates the quality gain in human judgment at the input stage.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces Tiger200K, a manually curated video dataset sourced from the UGC platform Bilibili, intended for post-training and quality-tuning of text-to-video generation models. The construction pipeline has five stages: manual curation of creators/videos based on aesthetic criteria, TransNetV2-based shot boundary detection with fixed-length cut generation, safe-zone computation using OCR and black-border detection, motion filtering and manual quality review, and bilingual (Chinese-English) dense captioning with Qwen2.5-VL. The paper reports 4,151 source videos yielding 85k scenes and 170k clips, with statistics on resolution, caption length, and safe-zone retention area. The central claim is that the dataset exhibits high visual quality and temporal consistency, with the quality advantage attributed to human curation at the input and review stages.

Significance. If the quality claim can be substantiated, Tiger200K would be a useful open resource for the video-generation community, potentially filling a gap left by algorithmically filtered datasets such as Koala-36M for the fine-tuning stage. The manuscript is commendably transparent about its pipeline, provides algorithmic details for safe-zone detection, and reports aggregate statistics over a large number of clips. The use of bilingual captions and the focus on 4K/UGC sources are also practically relevant. However, the significance is currently conditional: the paper's core value proposition is high visual quality, but that proposition is not tested against any external benchmark, human preference study, or downstream generation experiment, and the dataset is not yet accessible for independent verification.

major comments (4)
  1. [Sections 1, 2.1, 2.4, and 4] The central claim that Tiger200K has 'strong competitiveness in visual quality' is supported only by the author's own manual curation and self-review via random sampling. No inter-annotator agreement, no quantitative aesthetic metric, no comparison against Koala-36M by independent human raters, and no downstream fine-tuning experiment is reported. Because the dataset's raison d'être is quality, this absence is load-bearing. Please add at least one of the following: a human preference study comparing Tiger200K clips against Koala-36M clips, a downstream video-generation fine-tuning experiment with quantitative metrics, or a reproducible quality-rating protocol with reported agreement statistics.
  2. [Section 2.2 and Figure 5] The claim that TransNetV2 'achieves consistent segmentation performance across both synthetic test videos and real-world UGC content' is based on visual timeline comparisons only. Since temporal consistency is part of the dataset's stated value, the shot-boundary detection selection should be supported by quantitative metrics such as precision, recall, and F1 on the synthetic and real test sets, including cross-dissolve transitions. The construction of the ground-truth test set should also be described.
  3. [Section 3] The descriptive statistics presented do not establish 'high quality.' In particular, the statement that a safe-zone retention area above 85% 'demonstrates the high quality of the processed data' conflates the filter's output distribution with an independent measure of visual quality. Resolution, caption length, and safe-zone area are pipeline statistics, not quality evaluations. Please separate these descriptive statistics from any quality-validity evidence, or add appropriate quantitative quality measures.
  4. [Abstract and Conclusion] The dataset is announced as 'will be released,' but no release URL, sample download, or reviewer-access mechanism is provided. For a dataset paper, the contribution cannot be independently verified without access to at least a substantial sample with metadata and captions. Please provide an anonymous review link or release the data (or a representative subset) at revision time.
minor comments (6)
  1. [Section 2.3, Algorithm 1] Variables X1, Y1, X2, Y2 are used in the return statement but are never defined in the pseudocode; please define them explicitly as the final safe-zone coordinates. Also, line 31 compares a normalized area against the threshold 0.5, so the units should be stated.
  2. [Section 2.4] The motion filter description is underspecified: 'filter out frames below a predefined threshold' does not state whether clips are dropped based on a per-frame threshold, a fraction of low-motion frames, or a clip-level aggregate. Please clarify the exact criterion and report the threshold value.
  3. [Section 2.3] The phrase 'black broader' should be corrected to 'black border' in the text and pseudocode comments.
  4. [Figure 3] The statement that 'the quantities in the figure are relative' makes the figure hard to interpret; please either provide actual dataset sizes or remove the misleading quantitative appearance.
  5. [Section 2.5] Please specify the exact version of Qwen2.5-VL used and include the captioning prompt or template, as caption style substantially affects reproducibility.
  6. [Section 2.2] The text says scenes shorter than 121 frames are discarded, then 'center-based segmentation' produces 121-frame cuts; please state explicitly whether cuts may overlap and whether every valid scene yields floor(length/121) or another number of cuts, to make the relation between 85k scenes and 170k clips clearer.

Circularity Check

1 steps flagged · score 2.0 of 10

No formal derivation is circular, but the headline quality claim is self-referential: 'visual quality' is both the manual selection criterion and the asserted outcome, with no external quality benchmark, inter-annotator agreement, or downstream validation.

  1. self definitional [Section 2.1, Section 2.4, and Section 4 (Conclusion)]
    "The primary criteria for our video selection emphasize visual and aesthetic quality. ... Additionally, we conduct manual random sampling of videos to verify the accuracy of the aforementioned procedures and ensure overall video quality. ... our dataset demonstrates strong competitiveness in visual quality, a direct result of rigorous human quality control at the input stage."

    The claimed property, high visual quality, is the same subjective judgment used to select the content: videos are admitted because they satisfy the author's visual and aesthetic criteria, and then the dataset is asserted to have strong visual quality on the basis of that same manual review. No independent quality metric, inter-annotator agreement, or downstream generation benchmark is used, so the conclusion does not test the premise; it restates the curation criterion as a result. This is a self-definitional labeling issue rather than a fitted prediction or a derived equation, and it does not affect the technical pipeline description.

full rationale

The paper's processing stages (TransNetV2 shot detection, PaddleOCR and border-based safe-zone computation, optical-flow motion filtering, and Qwen2.5-VL captioning) are standard tools and are not fitted to the outcome; the reported statistics (85k scenes, 170k cuts, 4K resolution share, caption lengths) are descriptive, not predictions derived from the claimed quality. The only self-referential element is the subjective quality label: manual selection for aesthetic quality is reused as evidence of aesthetic quality. This is a validation limitation rather than a circular derivation in the mathematical sense, so the score is low. Self-citations involving the author, references [26] and [32], appear only as downstream-task context and are not load-bearing for the dataset construction claim.

Assumptions & free parameters 7 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new theoretical entities. It relies on hand-chosen pipeline thresholds and several untested domain assumptions about the accuracy of off-the-shelf tools. The most consequential assumption is that the author's manual curation is a sufficient guarantee of 'high visual quality,' as no external metric or downstream validation is provided.

free parameters (7)
  • Minimum scene length = 121 frames
    Scenes shorter than 121 frames are discarded; chosen without a reported sensitivity analysis.
  • Safe zone minimum area ratio = 0.5
    Video is dropped if the safe zone area is less than 50% of the original frame, per Algorithm 1 line 31.
  • OCR confidence threshold = tau (unspecified)
    PaddleOCR confidence threshold in Algorithm 1; the numeric value is not reported.
  • Text area ratio threshold = epsilon (unspecified)
    Videos with excessive text coverage are dropped; the threshold epsilon is not reported.
  • Subtitle region fraction = alpha (unspecified)
    Defines the top and bottom 20% regions in Algorithm 1; prose says 20% but the parameter value is not explicitly listed.
  • Border detection threshold = T (unspecified)
    Algorithm 1 takes border threshold T to scan for black borders; the value is not reported.
  • Motion filter threshold = undefined
    Farneback optical flow threshold is described as 'predefined' in Section 2.4 but never specified.
assumptions (5)
  • domain assumption TransNetV2 shot boundary detection is accurate on UGC content and cross dissolves.
    The paper tests TransNetV2 on one synthetic and one real video (Figure 5) and extrapolates to the entire corpus; this assumption drives scene segmentation.
  • domain assumption PaddleOCR detects all relevant subtitles and watermarks.
    The safe zone computation depends on OCR recall; missed text would leave unwanted overlays in the cropped area.
  • domain assumption Qwen2.5-VL generates accurate bilingual captions.
    Captions are generated by a VLM with no reported caption quality check or human verification.
  • ad hoc to paper Manual curation by the author ensures visual quality.
    The core quality guarantee is the author's own aesthetic judgment, which is not measured or inter-annotator-validated.
  • domain assumption Platform-declared resolution and quality labels are reliable.
    Resolution statistics in Figure 6 assume BiliBili's 4K and HDR labels are accurate, with no independent verification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform." pith.science (2026). https://pith.science/paper/2STWHH45

@misc{pith2026250415182,
  author       = {Pith},
  title        = {Pith review of: Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2STWHH45}},
  note         = {Machine review of arXiv:2504.15182}
}
read the original abstract

The recent surge in open-source text-to-video generation models has significantly energized the research community, yet their dependence on proprietary training datasets remains a key constraint. While existing open datasets like Koala-36M employ algorithmic filtering of web-scraped videos from early platforms, they still lack the quality required for fine-tuning advanced video generation models. We present Tiger200K, a manually curated high visual quality video dataset sourced from User-Generated Content (UGC) platforms. By prioritizing visual fidelity and aesthetic quality, Tiger200K underscores the critical role of human expertise in data curation, and providing high-quality, temporally consistent video-text pairs for fine-tuning and optimizing video generation architectures through a simple but effective pipeline including shot boundary detection, OCR, border detecting, motion filter and fine bilingual caption. The dataset will undergo ongoing expansion and be released as an open-source initiative to advance research and applications in video generative models. Project page: https://tinytigerpan.github.io/tiger200k/

Figures

Figures reproduced from arXiv: 2504.15182 by the authors.

Figure 1
Figure 1. Visualization of randomly sampled clips in Tiger200k dataset. These clips demonstrate [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The pipeline of data construction. First, the selected video will segment by scene and [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The inheritance relationship among Koala-36m and other datasets. Most are collected [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: The content differences between the video frames of dissolve transition are relatively [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Comparison of the shot detecting results of PySceneDetect with different parameters and [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 6
Figure 6. Figure 6: Statistical results of video level. Videos with 4K and 1080P resolutions each account for [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Statistical results of annotation level. The length of bilingual annotations is mostly con [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

32 extracted references · 15 canonical work pages

  1. [1]

    https://openai.com/sora/

    Sora | OpenAI. https://openai.com/sora/. 2

  2. [2]

    Youtube-8m: A large-scale video classifica- tion benchmark

    Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large-scale video classifica- tion benchmark. arXiv preprint arXiv:1609.08675, 2016. 3

  3. [3]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923, 2025. 6

  4. [4]

    Improving image generation with better captions

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Jun- tang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023. 6

  5. [5]

    Activitynet: A large-scale video benchmark for human activity understanding

    Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015. 2

  6. [6]

    PySceneDetect

    Brandon Castellano. PySceneDetect. URL https://github.com/Breakthrough/ PySceneDetect. 4

  7. [7]

    Panda-70m: Captioning 70m videos with multiple cross-modality teachers

    Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, et al. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13320–13331, 2024. 2, 3

  8. [8]

    From lifestyle vlogs to everyday interactions

    David F Fouhey, Wei-cheng Kuo, Alexei A Efros, and Jitendra Malik. From lifestyle vlogs to everyday interactions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4991–5000, 2018. 3

Show all 32 references
  1. [9]

    Ltx-video: Realtime video latent diffusion

    Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, Poriya Panet, Sapir Weissbuch, Victor Kulikov, Yaki Bitterman, Zeev Melumian, and Ofir Bibi. Ltx-video: Realtime video latent diffusion. a...

  2. [10]

    Svd: A large-scale short video dataset for near-duplicate video retrieval

    Qing-Yuan Jiang, Yi He, Gen Li, Jian Lin, Lei Li, and Wu-Jun Li. Svd: A large-scale short video dataset for near-duplicate video retrieval. InProceedings of the IEEE/CVF international conference on computer vision, pages 5281–5289, 2019. 2

  3. [11]

    Deepstory: Video story qa by deep embedded memory networks

    Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang. Deepstory: Video story qa by deep embedded memory networks. arXiv preprint arXiv:1707.00836, 2017. 2

  4. [12]

    Hunyuanvideo: A systematic framework for large video generative models

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. 2

  5. [13]

    Pp-ocrv3: More attempts for the improvement of ultra lightweight ocr system

    Chenxia Li, Weiwei Liu, Ruoyu Guo, Xiaoting Yin, Kaitao Jiang, Yongkun Du, Yuning Du, Lingfeng Zhu, Baohua Lai, Xiaoguang Hu, et al. Pp-ocrv3: More attempts for the improvement of ultra lightweight ocr system. arXiv preprint arXiv:2206.03001, 2022. 6

  6. [14]

    Realcam-i2v: Real-world image-to-video generation with interactive complex camera control.arXiv preprint arXiv:2502.10059, 2025

    Teng Li, Guangcong Zheng, Rui Jiang, Tao Wu, Yehao Lu, Yining Lin, Xi Li, et al. Realcam-i2v: Real-world image-to-video generation with interactive complex camera control.arXiv preprint arXiv:2502.10059, 2025. 2

  7. [15]

    Stylecrafter: Enhancing stylized text-to-video generation with style adapter

    Gongye Liu, Menghan Xia, Yong Zhang, Haoxin Chen, Jinbo Xing, Yibo Wang, Xintao Wang, Yujiu Yang, and Ying Shan. Stylecrafter: Enhancing stylized text-to-video generation with style adapter. arXiv preprint arXiv:2312.00330, 2023. 2 8

  8. [16]

    Follow-your-click: Open-domain regional image animation via short prompts

    Yue Ma, Yingqing He, Hongfa Wang, Andong Wang, Chenyang Qi, Chengfei Cai, Xiu Li, Zhifeng Li, Heung-Yeung Shum, Wei Liu, et al. Follow-your-click: Open-domain regional image animation via short prompts. arXiv preprint arXiv:2403.08268, 2024. 2

  9. [17]

    Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation

    Yue Ma, Hongyu Liu, Hongfa Wang, Heng Pan, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Heung-Yeung Shum, Wei Liu, et al. Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation. In SIGGRAPH Asia 2024 Conference Papers, pages 1–12,

  10. [18]

    Howto100m: Learning a text-video embedding by watching hundred million nar- rated video clips

    Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million nar- rated video clips. In Proceedings of the IEEE/CVF international conference on computer vision , pag...

  11. [19]

    A dataset for movie description

    Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele. A dataset for movie description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015. 2

  12. [20]

    How2: a large-scale dataset for multimodal language understanding

    Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Lo ¨ıc Barrault, Lucia Spe- cia, and Florian Metze. How2: a large-scale dataset for multimodal language understanding. arXiv preprint arXiv:1811.00347, 2018. 2

  13. [21]

    Transnet v2: An effective deep network architecture for fast shot transition detection

    Tom ´as Soucek and Jakub Lokoc. Transnet v2: An effective deep network architecture for fast shot transition detection. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 11218–11221, 2024. 2, 5

  14. [22]

    Movieqa: Understanding stories in movies through question-answering

    Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. Movieqa: Understanding stories in movies through question-answering. In Pro- ceedings of the IEEE conference on computer vision and pattern recognition, pages 4631–4640, 2016. 2

  15. [23]

    Wan: Open and advanced large-scale video generative models

    Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025. 2

  16. [24]

    Koala-36m: A large-scale video dataset improving con- sistency between fine-grained conditions and video content

    Qiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen, Ke Lin, Jiahao Wang, Boyuan Jiang, Haotian Yang, Mingwu Zheng, Xin Tao, et al. Koala-36m: A large-scale video dataset improving con- sistency between fine-grained conditions and video content. arXiv preprint arXiv:2410.08260,

  17. [25]

    Videomaker: Zero-shot customized video generation with the inherent force of video diffusion models

    Tao Wu, Yong Zhang, Xiaodong Cun, Zhongang Qi, Junfu Pu, Huanzhang Dou, Guangcong Zheng, Ying Shan, and Xi Li. Videomaker: Zero-shot customized video generation with the inherent force of video diffusion models. arXiv preprint arXiv:2412.19645, 2024. 2

  18. [26]

    Customcrafter: Customized video generation with preserving motion and concept composition abilities

    Tao Wu, Yong Zhang, Xintao Wang, Xianpan Zhou, Guangcong Zheng, Zhongang Qi, Ying Shan, and Xi Li. Customcrafter: Customized video generation with preserving motion and concept composition abilities. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 39,...

  19. [27]

    Msr-vtt: A large video description dataset for bridging video and language

    Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016. 2

  20. [28]

    Advancing high-resolution video-language representation with large-scale video transcriptions

    Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo. Advancing high-resolution video-language representation with large-scale video transcriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...

  21. [29]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 2

  22. [30]

    Merlot: Multimodal neural script knowledge models

    Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. Merlot: Multimodal neural script knowledge models. Advances in neural information processing systems, 34:23634–23651, 2021. 3

  23. [31]

    Cami2v: Camera- controlled image-to-video diffusion model

    Guangcong Zheng, Teng Li, Rui Jiang, Yehao Lu, Tao Wu, and Xi Li. Cami2v: Camera- controlled image-to-video diffusion model. arXiv preprint arXiv:2410.15957, 2024. 2

  24. [32]

    Realcam-vid: High-resolution video dataset with dynamic scenes and metric-scale camera movements

    Guangcong Zheng, Teng Li, Xianpan Zhou, and Xi Li. Realcam-vid: High-resolution video dataset with dynamic scenes and metric-scale camera movements. arXiv preprint arXiv:2504.08212, 2025. 2 10

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.