REVIEW 4 major objections 6 minor 32 references
Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Tiger200K is a manually curated 170,000-clip bilingual video dataset built to give open text-to-video models a higher-quality fine-tuning resource.
desk verdict A genuinely new 170k-clip bilingual video dataset with a transparent pipeline, but the 'high visual quality' claim is self-validated and needs external evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the manually selected source collection combined with the safe-zone computation. A safe zone is the region of a frame left after subtracting watermarks, subtitles, logos, and black borders, found by running PaddleOCR over every frame and scanning for persistent black regions, and clips whose safe zone falls below half the frame are discarded. Around this, the pipeline uses TransNetV2 for shot boundary detection with scenes shorter than 121 frames dropped, optical-flow-based motion filtering to remove static clips, and Qwen2.5-VL to produce dense bilingual captions. The paper's argument is that these mechanisms preserve temporal consistency and clean frames, while the manual curation at the front end supplies the aesthetic quality that algorithmic filtering misses.
What would settle it
A controlled fine-tuning experiment would settle the claim: fine-tune identical copies of one open video generation model on Tiger200K and on a matched random subset of Koala-36M with comparable clip counts and captions, then compare outputs on a fixed prompt suite through a human preference study or automated quality metric. If the Tiger200K-tuned model does not show measurably better visual quality or prompt adherence, the paper's quality claim is not supported. A second check is to compute an objective aesthetic score distribution on random clips from both datasets; if the distributions overlap heavily, the claimed quality gap disappears.
Extended reading notes
Core claim
The central claim is that human expertise in data curation, applied at the input stage, is what separates data good enough for fine-tuning from data merely good enough for pretraining. The paper argues that UGC platforms now contain professionally made content, and that selecting top creators, searching by camera model and production keywords, and relying on recommendation systems yields videos whose visual and aesthetic quality exceeds what threshold-based algorithmic filtering of older web-scraped video can guarantee. The dataset then applies a pipeline of TransNetV2 shot detection, OCR and border-based safe-zone computation, optical-flow motion filtering, and Qwen2.5-VL bilingual captioning to turn those source videos into 85,000 scenes and 170,000 fixed 121-frame cuts. The intended contribution is a high-visual-quality, temporally consistent, bilingual video-text corpus for post-training and quality-tuning of video generation models.
Load-bearing premise
The load-bearing premise is that the author's own judgment of 'visual and aesthetic quality' during manual selection and review is a reliable measure of the quality that matters for fine-tuning video generation models, since the paper never defines or measures that quality against an external benchmark.
Editorial extensions
If this is right
- Open-source text-to-video models gain a fine-tuning set of 170,000 clips whose frames are cropped to overlay-free safe zones and whose captions support both Chinese and English prompt following.
- Fine-tuning on Tiger200K should shift generation quality toward the polished, high-resolution look of professionally produced UGC rather than the average quality of large web-scraped corpora.
- The per-clip bilingual captions make the dataset usable as a benchmark for caption quality and video-text alignment, not only as training material.
- Because half the source videos are 4K or above, the dataset also supplies high-resolution material for video super-resolution and high-resolution generation research.
- The pipeline itself is reusable: shot detection, safe-zone cropping, motion filtering, and VLM captioning can be applied to any new UGC source for continued dataset expansion.
Reading between the lines
- A natural extension the paper does not run is a controlled fine-tuning comparison: train the same video model on Tiger200K and on an equally sized random sample of Koala-36M, then measure generation quality and prompt adherence on a held-out prompt set.
- The subjective quality claim could be made measurable by collecting pairwise human preferences between clips from Tiger200K and clips from existing open datasets; the paper leaves that quantification to future work.
- If the manual-curation approach transfers, the same platform-focused strategy could be applied to other regional UGC platforms to produce culturally diverse high-quality corpora rather than one platform's aesthetic.
- One implicit consequence is that dataset curation effort may shift from building smarter automatic filters to building better creator-discovery and review workflows, since the paper locates the quality gain in human judgment at the input stage.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Tiger200K, a manually curated video dataset sourced from the UGC platform Bilibili, intended for post-training and quality-tuning of text-to-video generation models. The construction pipeline has five stages: manual curation of creators/videos based on aesthetic criteria, TransNetV2-based shot boundary detection with fixed-length cut generation, safe-zone computation using OCR and black-border detection, motion filtering and manual quality review, and bilingual (Chinese-English) dense captioning with Qwen2.5-VL. The paper reports 4,151 source videos yielding 85k scenes and 170k clips, with statistics on resolution, caption length, and safe-zone retention area. The central claim is that the dataset exhibits high visual quality and temporal consistency, with the quality advantage attributed to human curation at the input and review stages.
Significance. If the quality claim can be substantiated, Tiger200K would be a useful open resource for the video-generation community, potentially filling a gap left by algorithmically filtered datasets such as Koala-36M for the fine-tuning stage. The manuscript is commendably transparent about its pipeline, provides algorithmic details for safe-zone detection, and reports aggregate statistics over a large number of clips. The use of bilingual captions and the focus on 4K/UGC sources are also practically relevant. However, the significance is currently conditional: the paper's core value proposition is high visual quality, but that proposition is not tested against any external benchmark, human preference study, or downstream generation experiment, and the dataset is not yet accessible for independent verification.
major comments (4)
- [Sections 1, 2.1, 2.4, and 4] The central claim that Tiger200K has 'strong competitiveness in visual quality' is supported only by the author's own manual curation and self-review via random sampling. No inter-annotator agreement, no quantitative aesthetic metric, no comparison against Koala-36M by independent human raters, and no downstream fine-tuning experiment is reported. Because the dataset's raison d'être is quality, this absence is load-bearing. Please add at least one of the following: a human preference study comparing Tiger200K clips against Koala-36M clips, a downstream video-generation fine-tuning experiment with quantitative metrics, or a reproducible quality-rating protocol with reported agreement statistics.
- [Section 2.2 and Figure 5] The claim that TransNetV2 'achieves consistent segmentation performance across both synthetic test videos and real-world UGC content' is based on visual timeline comparisons only. Since temporal consistency is part of the dataset's stated value, the shot-boundary detection selection should be supported by quantitative metrics such as precision, recall, and F1 on the synthetic and real test sets, including cross-dissolve transitions. The construction of the ground-truth test set should also be described.
- [Section 3] The descriptive statistics presented do not establish 'high quality.' In particular, the statement that a safe-zone retention area above 85% 'demonstrates the high quality of the processed data' conflates the filter's output distribution with an independent measure of visual quality. Resolution, caption length, and safe-zone area are pipeline statistics, not quality evaluations. Please separate these descriptive statistics from any quality-validity evidence, or add appropriate quantitative quality measures.
- [Abstract and Conclusion] The dataset is announced as 'will be released,' but no release URL, sample download, or reviewer-access mechanism is provided. For a dataset paper, the contribution cannot be independently verified without access to at least a substantial sample with metadata and captions. Please provide an anonymous review link or release the data (or a representative subset) at revision time.
minor comments (6)
- [Section 2.3, Algorithm 1] Variables X1, Y1, X2, Y2 are used in the return statement but are never defined in the pseudocode; please define them explicitly as the final safe-zone coordinates. Also, line 31 compares a normalized area against the threshold 0.5, so the units should be stated.
- [Section 2.4] The motion filter description is underspecified: 'filter out frames below a predefined threshold' does not state whether clips are dropped based on a per-frame threshold, a fraction of low-motion frames, or a clip-level aggregate. Please clarify the exact criterion and report the threshold value.
- [Section 2.3] The phrase 'black broader' should be corrected to 'black border' in the text and pseudocode comments.
- [Figure 3] The statement that 'the quantities in the figure are relative' makes the figure hard to interpret; please either provide actual dataset sizes or remove the misleading quantitative appearance.
- [Section 2.5] Please specify the exact version of Qwen2.5-VL used and include the captioning prompt or template, as caption style substantially affects reproducibility.
- [Section 2.2] The text says scenes shorter than 121 frames are discarded, then 'center-based segmentation' produces 121-frame cuts; please state explicitly whether cuts may overlap and whether every valid scene yields floor(length/121) or another number of cuts, to make the relation between 85k scenes and 170k clips clearer.
Circularity Check
No formal derivation is circular, but the headline quality claim is self-referential: 'visual quality' is both the manual selection criterion and the asserted outcome, with no external quality benchmark, inter-annotator agreement, or downstream validation.
-
self definitional
[Section 2.1, Section 2.4, and Section 4 (Conclusion)]
"The primary criteria for our video selection emphasize visual and aesthetic quality. ... Additionally, we conduct manual random sampling of videos to verify the accuracy of the aforementioned procedures and ensure overall video quality. ... our dataset demonstrates strong competitiveness in visual quality, a direct result of rigorous human quality control at the input stage."
The claimed property, high visual quality, is the same subjective judgment used to select the content: videos are admitted because they satisfy the author's visual and aesthetic criteria, and then the dataset is asserted to have strong visual quality on the basis of that same manual review. No independent quality metric, inter-annotator agreement, or downstream generation benchmark is used, so the conclusion does not test the premise; it restates the curation criterion as a result. This is a self-definitional labeling issue rather than a fitted prediction or a derived equation, and it does not affect the technical pipeline description.
full rationale
The paper's processing stages (TransNetV2 shot detection, PaddleOCR and border-based safe-zone computation, optical-flow motion filtering, and Qwen2.5-VL captioning) are standard tools and are not fitted to the outcome; the reported statistics (85k scenes, 170k cuts, 4K resolution share, caption lengths) are descriptive, not predictions derived from the claimed quality. The only self-referential element is the subjective quality label: manual selection for aesthetic quality is reused as evidence of aesthetic quality. This is a validation limitation rather than a circular derivation in the mathematical sense, so the score is low. Self-citations involving the author, references [26] and [32], appear only as downstream-task context and are not load-bearing for the dataset construction claim.
Assumptions & free parameters
free parameters (7)
- Minimum scene length =
121 frames
- Safe zone minimum area ratio =
0.5
- OCR confidence threshold =
tau (unspecified)
- Text area ratio threshold =
epsilon (unspecified)
- Subtitle region fraction =
alpha (unspecified)
- Border detection threshold =
T (unspecified)
- Motion filter threshold =
undefined
assumptions (5)
- domain assumption TransNetV2 shot boundary detection is accurate on UGC content and cross dissolves.
- domain assumption PaddleOCR detects all relevant subtitles and watermarks.
- domain assumption Qwen2.5-VL generates accurate bilingual captions.
- ad hoc to paper Manual curation by the author ensures visual quality.
- domain assumption Platform-declared resolution and quality labels are reliable.
Cite this review
Pith. "Pith review of Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform." pith.science (2026). https://pith.science/paper/2STWHH45
@misc{pith2026250415182,
author = {Pith},
title = {Pith review of: Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform},
year = {2026},
howpublished = {\url{https://pith.science/paper/2STWHH45}},
note = {Machine review of arXiv:2504.15182}
}
read the original abstract
The recent surge in open-source text-to-video generation models has significantly energized the research community, yet their dependence on proprietary training datasets remains a key constraint. While existing open datasets like Koala-36M employ algorithmic filtering of web-scraped videos from early platforms, they still lack the quality required for fine-tuning advanced video generation models. We present Tiger200K, a manually curated high visual quality video dataset sourced from User-Generated Content (UGC) platforms. By prioritizing visual fidelity and aesthetic quality, Tiger200K underscores the critical role of human expertise in data curation, and providing high-quality, temporally consistent video-text pairs for fine-tuning and optimizing video generation architectures through a simple but effective pipeline including shot boundary detection, OCR, border detecting, motion filter and fine bilingual caption. The dataset will undergo ongoing expansion and be released as an open-source initiative to advance research and applications in video generative models. Project page: https://tinytigerpan.github.io/tiger200k/
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
Youtube-8m: A large-scale video classifica- tion benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large-scale video classifica- tion benchmark. arXiv preprint arXiv:1609.08675, 2016. 3
arXiv 2016
-
[3]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923, 2025. 6
arXiv 2025
-
[4]
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Jun- tang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023. 6
work page 2023
-
[5]
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015. 2
work page 2015
-
[6]
Brandon Castellano. PySceneDetect. URL https://github.com/Breakthrough/ PySceneDetect. 4
-
[7]
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, et al. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13320–13331, 2024. 2, 3
work page 2024
-
[8]
From lifestyle vlogs to everyday interactions
David F Fouhey, Wei-cheng Kuo, Alexei A Efros, and Jitendra Malik. From lifestyle vlogs to everyday interactions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4991–5000, 2018. 3
work page 2018
Show all 32 references
-
[9]
Ltx-video: Realtime video latent diffusion
Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, Poriya Panet, Sapir Weissbuch, Victor Kulikov, Yaki Bitterman, Zeev Melumian, and Ofir Bibi. Ltx-video: Realtime video latent diffusion. a...
2024 arXiv
-
[10]
Svd: A large-scale short video dataset for near-duplicate video retrieval
Qing-Yuan Jiang, Yi He, Gen Li, Jian Lin, Lei Li, and Wu-Jun Li. Svd: A large-scale short video dataset for near-duplicate video retrieval. InProceedings of the IEEE/CVF international conference on computer vision, pages 5281–5289, 2019. 2
2019
-
[11]
Deepstory: Video story qa by deep embedded memory networks
Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang. Deepstory: Video story qa by deep embedded memory networks. arXiv preprint arXiv:1707.00836, 2017. 2
2017 arXiv
-
[12]
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. 2
2024 arXiv
-
[13]
Pp-ocrv3: More attempts for the improvement of ultra lightweight ocr system
Chenxia Li, Weiwei Liu, Ruoyu Guo, Xiaoting Yin, Kaitao Jiang, Yongkun Du, Yuning Du, Lingfeng Zhu, Baohua Lai, Xiaoguang Hu, et al. Pp-ocrv3: More attempts for the improvement of ultra lightweight ocr system. arXiv preprint arXiv:2206.03001, 2022. 6
2022 arXiv
-
[14]
Realcam-i2v: Real-world image-to-video generation with interactive complex camera control.arXiv preprint arXiv:2502.10059, 2025
Teng Li, Guangcong Zheng, Rui Jiang, Tao Wu, Yehao Lu, Yining Lin, Xi Li, et al. Realcam-i2v: Real-world image-to-video generation with interactive complex camera control.arXiv preprint arXiv:2502.10059, 2025. 2
2025 arXiv
-
[15]
Stylecrafter: Enhancing stylized text-to-video generation with style adapter
Gongye Liu, Menghan Xia, Yong Zhang, Haoxin Chen, Jinbo Xing, Yibo Wang, Xintao Wang, Yujiu Yang, and Ying Shan. Stylecrafter: Enhancing stylized text-to-video generation with style adapter. arXiv preprint arXiv:2312.00330, 2023. 2 8
2023 arXiv
-
[16]
Follow-your-click: Open-domain regional image animation via short prompts
Yue Ma, Yingqing He, Hongfa Wang, Andong Wang, Chenyang Qi, Chengfei Cai, Xiu Li, Zhifeng Li, Heung-Yeung Shum, Wei Liu, et al. Follow-your-click: Open-domain regional image animation via short prompts. arXiv preprint arXiv:2403.08268, 2024. 2
2024 arXiv
-
[17]
Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation
Yue Ma, Hongyu Liu, Hongfa Wang, Heng Pan, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Heung-Yeung Shum, Wei Liu, et al. Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation. In SIGGRAPH Asia 2024 Conference Papers, pages 1–12,
2024
-
[18]
Howto100m: Learning a text-video embedding by watching hundred million nar- rated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million nar- rated video clips. In Proceedings of the IEEE/CVF international conference on computer vision , pag...
2019
-
[19]
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele. A dataset for movie description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015. 2
2015
-
[20]
How2: a large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Lo ¨ıc Barrault, Lucia Spe- cia, and Florian Metze. How2: a large-scale dataset for multimodal language understanding. arXiv preprint arXiv:1811.00347, 2018. 2
2018 arXiv
-
[21]
Transnet v2: An effective deep network architecture for fast shot transition detection
Tom ´as Soucek and Jakub Lokoc. Transnet v2: An effective deep network architecture for fast shot transition detection. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 11218–11221, 2024. 2, 5
2024
-
[22]
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. Movieqa: Understanding stories in movies through question-answering. In Pro- ceedings of the IEEE conference on computer vision and pattern recognition, pages 4631–4640, 2016. 2
2016
-
[23]
Wan: Open and advanced large-scale video generative models
Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025. 2
2025 arXiv
-
[24]
Koala-36m: A large-scale video dataset improving con- sistency between fine-grained conditions and video content
Qiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen, Ke Lin, Jiahao Wang, Boyuan Jiang, Haotian Yang, Mingwu Zheng, Xin Tao, et al. Koala-36m: A large-scale video dataset improving con- sistency between fine-grained conditions and video content. arXiv preprint arXiv:2410.08260,
-
[25]
Videomaker: Zero-shot customized video generation with the inherent force of video diffusion models
Tao Wu, Yong Zhang, Xiaodong Cun, Zhongang Qi, Junfu Pu, Huanzhang Dou, Guangcong Zheng, Ying Shan, and Xi Li. Videomaker: Zero-shot customized video generation with the inherent force of video diffusion models. arXiv preprint arXiv:2412.19645, 2024. 2
2024 arXiv
-
[26]
Customcrafter: Customized video generation with preserving motion and concept composition abilities
Tao Wu, Yong Zhang, Xintao Wang, Xianpan Zhou, Guangcong Zheng, Zhongang Qi, Ying Shan, and Xi Li. Customcrafter: Customized video generation with preserving motion and concept composition abilities. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 39,...
2025
-
[27]
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016. 2
2016
-
[28]
Advancing high-resolution video-language representation with large-scale video transcriptions
Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo. Advancing high-resolution video-language representation with large-scale video transcriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...
2022
-
[29]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 2
2024 arXiv
-
[30]
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. Merlot: Multimodal neural script knowledge models. Advances in neural information processing systems, 34:23634–23651, 2021. 3
2021
-
[31]
Cami2v: Camera- controlled image-to-video diffusion model
Guangcong Zheng, Teng Li, Rui Jiang, Yehao Lu, Tao Wu, and Xi Li. Cami2v: Camera- controlled image-to-video diffusion model. arXiv preprint arXiv:2410.15957, 2024. 2
2024 arXiv
-
[32]
Realcam-vid: High-resolution video dataset with dynamic scenes and metric-scale camera movements
Guangcong Zheng, Teng Li, Xianpan Zhou, and Xi Li. Realcam-vid: High-resolution video dataset with dynamic scenes and metric-scale camera movements. arXiv preprint arXiv:2504.08212, 2025. 2 10
2025 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.