REVIEW 3 major objections 5 minor 112 references
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read BounTCHA claims a CAPTCHA can rest on people's ability to spot where a real video ends and an AI-generated extension begins, while current multimodal models largely miss it.
desk verdict A genuinely new CAPTCHA idea with a solid human-performance study, but the security analysis ignores the cheapest attack class—classical cut detection—so the resilience claim is not yet established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the guided AI-extended composite video: a raw real-camera clip concatenated with a prompted AI-generated continuation, with the true seam at frame $F^*_n$ serving as ground truth. The decision mechanism is the human time-bias acceptance window: a submitted boundary time is accepted when it falls in $[\beta_1,\beta_2]$, where $\beta_i = \mu \pm \sigma \cdot \mathrm{ppf}(1-\alpha/2)$ and $\mu, \sigma$ come from the human study; the significance level $\alpha$ and the video length $L$ are the tunable knobs that control how narrow the window is. The generation pipeline carries the argument by making the seam salient to human perception but not to the tested multimodal models. The paper also uses shift cutting to relocate the boundary inside a single generated video, turning one generation into multiple CAPTCHA variants.
What would settle it
Run a standard cut-detection pass, such as frame-by-frame histogram difference or FFmpeg's scene-detection filter, on the 25 released composite videos and check whether the detected seam falls inside the human acceptance window at $\alpha = 0.25$. If a cheap detector locates the boundary with far higher accuracy than the reported MLLM success rates, BounTCHA's security claim is refuted.
Extended reading notes
Core claim
The central claim is that the boundary between a raw video and a guided AI-generated continuation is itself a usable human test: people can locate it reliably, leading multimodal models cannot, and the resulting distribution of human errors supports a simple pass/fail rule. The paper calls the boundary the 'real boundary frame' $F^*_n$ and treats it as ground truth. It reports a human time-bias distribution with mean $0.332$ s and standard deviation $0.406$ s, and defines the acceptance window from that distribution as $\mu \pm \sigma \cdot \mathrm{ppf}(1-\alpha/2)$. Leave-one-out cross-validation over the 25 videos gives human success rates that stay at or above 82% for $\alpha \le 0.25$; under the same window the four tested MLLMs succeed in at most 17.33% of trials and take over 20 seconds per attempt, against 14.2 seconds average for humans. The paper's conclusion is that a boundary-identification task of this kind can currently separate humans from state-of-the-art multimodal AI.
Load-bearing premise
The security case assumes attackers will only guess randomly, search a database of known videos, or ask a multimodal AI model; it never tests simple, cheap algorithms that look for an abrupt visual cut at the stitching point, and such an algorithm would break the CAPTCHA if it works.
Editorial extensions
If this is right
- If BounTCHA's numbers hold, generative AI itself becomes a source of CAPTCHA material, and the same models that erode text- and image-based CAPTCHAs can be used to build new challenges.
- A deployment can set the acceptance window from the reported human distribution and tune $\alpha$ and video length to trade human success against bot success.
- Current multimodal LLMs would need to reach well above 17.33% boundary-location accuracy before an MLLM-based solver becomes a practical threat to BounTCHA.
- Because human time-bias statistics are roughly normal and age-group differences are small, one window can serve users across the tested age range.
- Shift cutting provides a cheap way to generate fresh variants from each produced video, which the database-attack analysis treats as protection against replay.
Reading between the lines
- The paper never tests cheap, non-semantic seam detectors such as frame-histogram differences or optical-flow discontinuities; my inference is that those, not MLLMs, are the threat most likely to break BounTCHA, because the composite video is made by simple concatenation.
- A testable extension would be to soften the seam with a crossfade or motion continuity and then re-measure both human and classical-detector performance; if humans stay accurate while detectors degrade, the mechanism becomes much stronger.
- The same 'boundary between real and AI-generated' design could generalize to other perceptual discontinuities, such as audio-visual mismatches, physics-violating motion, or story-consistency breaks, wherever humans' world knowledge currently outpaces MLLMs.
- If video-language models keep improving, the acceptance window should be recalibrated over time; the paper's normal-distribution model gives a direct procedure for doing that, but the paper itself does not propose a retraining or update schedule.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes BounTCHA, a CAPTCHA in which a user watches a short video formed by concatenating a raw clip with a generative-AI extension and drags a progress bar to the reported boundary. The generation pipeline uses Tarsier for video description, GPT-4o for prompt synthesis, and Kling for video extension (Section 3). A user study with 186 participants on 25 videos measures the time bias between the true boundary and the user-reported boundary; the acceptance interval is defined in Eq. (4) from the fitted mean and standard deviation of that bias. The paper claims human success rates of at least 82% at significance level alpha <= 0.25, and evaluates uniform and truncated-normal random attacks, a database attack, and MLLM attacks (Tarsier, the model referred to as MiniVPM-V 2.6, GPT-4V, and Claude 3.5 Sonnet), with claimed attack success rates no higher than 17.33%.
Significance. If the security claims held, BounTCHA would be a timely and genuinely novel CAPTCHA: it exploits a perceptual capability (sensitivity to abrupt video transitions) that current MLLMs do not reliably exhibit, and the authors ship an open-sourced prototype, the video set, and a real user study. The arithmetic of the random-attack analysis is straightforward, and the database-attack model is clearly derived. However, the headline numbers are weakened by two load-bearing gaps: the human acceptance threshold is derived from the same data used to report human success, and the security analysis omits cheap classical video-segmentation attacks that are directly aimed at the kind of concatenation boundary the scheme creates. The MLLM evaluation is also too small and under-specified to support a robust upper-bound claim.
major comments (3)
- [Section 5.2 and Eq. (4)] The reported human success rates are partly true by construction. The acceptance interval [beta1, beta2] is defined from the fitted mean mu and standard deviation sigma of the same 186-participant time-bias data that are then classified as successes, so under the normal model the aggregate success rate at confidence 1-alpha is approximately 1-alpha; the green curve in Fig. 9 therefore inherits its value from the Gaussian fit rather than from an independent behavioral measurement. The leave-one-out validation divides the 25 videos into five groups but reuses the same participants in every fold, so it does not establish generalization to new users. Please report held-out participant success (participant-level cross-validation or a fresh deployment), fix the threshold before evaluation, and show the distribution of per-participant success instead of only the aggregate curve.
- [Section 6 and Section 3.3 (Eq. 3)] The security analysis omits classical video-boundary detection, which is the most obvious low-cost attack. Each challenge video is the direct concatenation of V_in and V_ext, with F*_n as the transition frame; per-frame color-histogram distance, optical-flow discontinuity scoring, and FFmpeg scene detection are designed to locate exactly such abrupt changes and can run in milliseconds. The paper should measure at least one such baseline on the 25 published videos and report how often the detected boundary falls within [beta1, beta2] at alpha = 0.25. Without this measurement, the claim that BounTCHA is resilient against various types of attacks is unsupported for the cheapest attack class.
- [Section 6.3 and Table 2] The MLLM attack evaluation is too small and under-specified to support the stated upper bound. With 25 videos and three rounds per model, a 17.33% success rate corresponds to a 95% confidence interval of roughly 9-28%, so the claim that the models succeed in no more than 17.33% of trials is not a robust statistical statement. The paper does not report the prompt template, decoding parameters, or frame-sampling strategy for the four models, which limits reproducibility. Please provide this information, report per-video and per-round breakdowns, and compare against human performance using a non-circular acceptance threshold, ideally at fixed human false-positive rates.
minor comments (5)
- [Section 6.3 and Figure 15] The sentence claiming that the oldest group's worst-case time is notably higher than the worst-case time for other MLLMs contradicts Table 2, where the human worst-case time is 26.6s and all MLLM worst-case times are between 37.1s and 43.2s; please correct the text or the figure labels.
- [Table 2 and Section 6.3] The model name 'MiniVPM-V 2.6' appears to be a typo for MiniCPM-V 2.6 (reference [100]); please standardize the name throughout.
- [Eq. (4)] Writing beta_i = mu ± sigma * ppf(1 - alpha/2) with i in {1,2} is ambiguous; please define beta_1 = mu - sigma * ppf(1 - alpha/2) and beta_2 = mu + sigma * ppf(1 - alpha/2) explicitly.
- [Section 3.2.2] The generation prompt says the extension should 'differ significantly' from the original video while also avoiding 'drastic changes in the visuals'; please clarify how these constraints are reconciled, since both directly affect task difficulty.
- [Throughout] Minor typographical issues include 'CATPCHA' in Figure 1, 'gameified' in Section 2.1.3, and 'not only happened' in Section 1; a light copyedit would improve readability.
Circularity Check
The human success-rate claim restates the fitted confidence interval; MLLM attack results are external, so circularity is partial.
-
fitted input called prediction
[Section 5.2 (Figure 9), Section 6 Eq. (4), Table 2]
"All the data of time bias falls between -1.6s and 2.6s, and it follows a normal distribution. Therefore, we can increase the difficulty of BounTCHA by adjusting the significance level to narrow the time range. ... β_i = μ ± σ · ppf(1−α/2), i∈{1,2} (4) ... Human ≥ 82% (α≤0.25), The relevant data is from Figure 9."
The acceptance interval [β1, β2] in Eq. (4) is defined as the central (1−α) interval of the normal distribution fitted to the same 186-participant time-bias data (μ = 0.332, σ = 0.406). A human drawn from that same distribution therefore falls inside the interval with probability approximately 1−α by construction. Table 2's headline 'Human ≥ 82% (α≤0.25)' is thus a restatement of the fitted interval's coverage, not an independent measurement of human ability. The leave-one-out cross-validation only verifies that the fitted distribution generalizes to held-out videos from the same dataset; it does not break the definitional link between the fitted confidence level and the reported success rate. The MLLM attack results are genuinely external, which keeps the circularity partial.
full rationale
The main derivation chain is not globally circular: the video-generation pipeline, the 186-participant user study, and the MLLM attack benchmarks are independent empirical contributions. The one genuinely circular element is the presentation of the human success rate. Because Eq. (4) defines the acceptance window as a central (1−α) interval of the fitted human time-bias distribution, the success rates plotted in Figure 9 and quoted in Table 2 are close to 1−α by construction; the cross-validation gives limited independence but does not remove the definitional dependence. The self-citation [90] is motivational only and is backed by this paper's own new user study, so it is not load-bearing. The security analysis in Section 6 omits classical shot/cut-detection baselines (e.g., histogram-based or optical-flow boundary detection); that is a security-evaluation gap rather than a circularity and does not affect the circularity score. Overall score 4 reflects one partial by-construction result while the core MLLM-vs-human comparison remains externally measured.
Assumptions & free parameters
free parameters (3)
- mu =
0.332 s
- sigma =
0.406 s
- alpha =
0.25 (recommended)
assumptions (4)
- domain assumption Human time bias follows a normal distribution
- domain assumption Variants within a video group have uniformly distributed boundaries
- ad hoc to paper The concatenation point is not detectable by low-level video segmentation
- standard math Trial independence for the binomial attack model
Cite this review
Pith. "Pith review of BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos." pith.science (2026). https://pith.science/paper/MW3TXQXA
@misc{pith2026250118565,
author = {Pith},
title = {Pith review of: BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/MW3TXQXA}},
note = {Machine review of arXiv:2501.18565}
}
read the original abstract
In recent years, the rapid development of artificial intelligence (AI) especially multi-modal Large Language Models (MLLMs), has enabled it to understand text, images, videos, and other multimedia data, allowing AI systems to execute various tasks based on human-provided prompts. However, AI-powered bots have increasingly been able to bypass most existing CAPTCHA systems, posing significant security threats to web applications. This makes the design of new CAPTCHA mechanisms an urgent priority. We observe that humans are highly sensitive to shifts and abrupt changes in videos, while current AI systems still struggle to comprehend and respond to such situations effectively. Based on this observation, we design and implement BounTCHA, a CAPTCHA mechanism that leverages human perception of boundaries in video transitions and disruptions. By utilizing generative AI's capability to extend original videos with prompts, we introduce unexpected twists and changes to create a pipeline for generating guided short videos for CAPTCHA purposes. We develop a prototype and conduct experiments to collect data on humans' time biases in boundary identification. This data serves as a basis for distinguishing between human users and bots. Additionally, we perform a detailed security analysis of BounTCHA, demonstrating its resilience against various types of attacks. We hope that BounTCHA will act as a robust defense, safeguarding millions of web applications in the AI-driven era.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Suhas Aggarwal. 2013. Animated CAPTCHAs and games for advertising. In Proceedings of the 22nd International Conference on World Wide Web . 1167–1174
2013
-
[2]
YeonChan Ahn, Namsoo Kim, and Yoo-Sung Kim. 2013. A user-friendly image-text fusion CAPTCHA for secure web services. In Proceedings of International Conference on Information Integration and Web-based Applications & Services . 550–554
2013
-
[3]
Omar Alonso, Catherine Marshall, and Marc Najork. 2013. A human-centered framework for ensuring reliability on crowdsourced labeling tasks. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol. 1. 2–3
2013
-
[4]
Omar Alonso, Catherine Marshall, and Marc Najork. 2014. Crowdsourcing a subjective labeling task: a human-centered framework to ensure reliable results. Microsoft Res., Redmond (2014)
2014
-
[5]
Fatmah H Alqahtani and Fawaz A Alsulaiman. 2020. Is image-based CAPTCHA secure against attacks based on machine learning? An experimental study. Computers & Security 88 (2020), 101635
2020
-
[6]
Babak Amin Azad, Oleksii Starov, Pierre Laperdrix, and Nick Nikiforakis. 2020. Web runner 2049: Evaluating third-party anti-bot services. In Detection of Intrusions and Malware, and Vulnerability Assessment: 17th International Conference, DIMV A 2020, Lisbon, Portugal, June 24–26, 2020, Proceedings 17. Springer, 135–159
2020
-
[7]
K Anjitha and IK Rijin. 2015. Captcha as graphical passwords-enhanced with video-based captcha for secure services. In 2015 International Conference on Applied and Theoretical Computing and Communication Technology (iCATccT) . IEEE, 213–217
2015
-
[8]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. (2023)
2023
Show all 112 references
-
[9]
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. 2023. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127 (2023)
2023 arXiv
-
[10]
Neha Pradyumna Bora and Dinesh Chandra Jain. 2023. A web authentication biometric 3D animated CAPTCHA system using artificial intelligence and machine learning approach. Journal of Artificial Intelligence and Technology 3, 3 (2023), 126–133
2023
-
[11]
Jose Brustoloni. 2002. Protecting electronic commerce from distributed denial-of-service attacks. In Proceedings of the 11th international conference on World Wide Web. 553–561. BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos 19
2002
-
[12]
Elie Bursztein, Steven Bethard, Celine Fabry, John C Mitchell, and Dan Jurafsky. 2010. How good are humans at solving CAPTCHAs? A large scale evaluation. In 2010 IEEE symposium on security and privacy . IEEE, 399–413
2010
-
[13]
Wenhao Chai, Xun Guo, Gaoang Wang, and Yan Lu. 2023. Stablevideo: Text-driven consistency-aware diffusion video editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 23040–23050
2023
-
[14]
Arindam Chaudhuri, Krupa Mandaviya, Pratixa Badelia, Soumya K Ghosh, Arindam Chaudhuri, Krupa Mandaviya, Pratixa Badelia, and Soumya K Ghosh. 2017. Optical character recognition systems . Springer
2017
-
[15]
Chaoran Chen, Leyang Li, Luke Cao, Yanfang Ye, Tianshi Li, Yaxing Yao, and Toby Jia-jun Li. 2024. Why am I seeing this: Democratizing End User Auditing for Online Content Recommendations. arXiv preprint arXiv:2410.04917 (2024)
2024 arXiv
-
[16]
Jun Chen, Xiangyang Luo, Yanqing Guo, Yi Zhang, and Daofu Gong. 2017. A Survey on Breaking Technique of Text-Based CAPTCHA. Security and communication networks 2017, 1 (2017), 6898617
2017
-
[17]
Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. 2018. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer vision (ECCV) . 801–818
2018
-
[18]
Monica Chew and J Doug Tygar. 2004. Image recognition captchas. In International Conference on Information Security . Springer, 268–279
2004
-
[19]
Joseph Cho, Fachrina Dewi Puspitasari, Sheng Zheng, Jingyao Zheng, Lik-Hang Lee, Tae-Ho Kim, Choong Seon Hong, and Chaoning Zhang. 2024. Sora as an agi world model? a complete survey on text-to-video generation. arXiv preprint arXiv:2403.05131 (2024)
2024
-
[20]
Paul Couairon, Clément Rambour, Jean-Emmanuel Haugeard, and Nicolas Thome. 2023. Videdit: Zero-shot and spatially aware text-driven video editing. Transactions on Machine Learning Research (2023)
2023
-
[21]
Yifan Cui, Xinyi Shan, and Jeanhun Chung. 2024. A Feasibility Study on RUNWAY GEN-2 for Generating Realistic Style Images. International Journal of Internet, Broadcasting and Communication 16, 1 (2024), 99–105
2024
-
[22]
Alexey Dosovitskiy. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)
2020 arXiv
-
[23]
John R Douceur. 2002. The sybil attack. In International workshop on peer-to-peer systems . Springer, 251–260
2002
-
[24]
Zhengjie Du, Yuekang Li, Yaowen Zheng, Xiaohan Zhang, Cen Zhang, Yi Liu, Sheikh Mahbub Habib, Xinghua Li, Linzhang Wang, Yang Liu, et al
-
[25]
Peter Eckersley. 2010. How unique is your web browser?. In Privacy Enhancing Technologies: 10th International Symposium, PETS 2010, Berlin, Germany, July 21-23, 2010. Proceedings 10 . Springer, 1–18
2010
-
[26]
Tuğrulcan Elmas. 2023. Analyzing activity and suspension patterns of twitter bots attacking turkish twitter trends by a longitudinal dataset. In Companion Proceedings of the ACM Web Conference 2023 . 1404–1412
2023
-
[27]
Christoph Fritsch, Michael Netter, Andreas Reisser, and Günther Pernul. 2010. Attacking Image Recognition Captcha s: A Naive but Effective Approach. In Trust, Privacy and Security in Digital Business: 7th International Conference, TrustBus 2010, Bilbao, Spain, August 30-31, 20...
2010
-
[28]
Haichang Gao, Dan Yao, Honggang Liu, Xiyang Liu, and Liming Wang. 2010. A novel image based CAPTCHA using jigsaw puzzle. In 2010 13th IEEE international conference on computational science and engineering . IEEE, 351–356
2010
-
[29]
Paolo Gasti, Gene Tsudik, Ersin Uzun, and Lixia Zhang. 2013. DoS and DDoS in named data networking. In 2013 22nd International Conference on Computer Communication and Networks (ICCCN) . IEEE, 1–7
2013
-
[30]
Nethanel Gelernter and Amir Herzberg. 2016. Tell me about yourself: The malicious captcha attack. In Proceedings of the 25th International Conference on World Wide Web. 999–1008
2016
-
[31]
Rich Gossweiler, Maryam Kamvar, and Shumeet Baluja. 2009. What’s up CAPTCHA? A CAPTCHA based on image orientation. In Proceedings of the 18th international conference on World wide web . 841–850
2009
-
[32]
Gilad Gressel, Rahul Pankajakshan, and Yisroel Mirsky. 2024. Discussion Paper: Exploiting LLMs for Scam Automation: A Looming Threat. In Proceedings of the 3rd ACM Workshop on the Security Implications of Deepfakes and Cheapfakes . 20–24
2024
-
[33]
Meriem Guerar, Luca Verderame, Mauro Migliardi, Francesco Palmieri, and Alessio Merlo. 2021. Gotta CAPTCHA’Em all: a survey of 20 Years of the human-or-computer Dilemma. ACM Computing Surveys (CSUR) 54, 9 (2021), 1–33
2021
-
[34]
Chris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk, and Philipp Zimmer. 2023. Simplistic collection and labeling practices limit the utility of benchmark datasets for Twitter bot detection. In Proceedings of the ACM web conference 2023 . 3660–3669
2023
-
[35]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
2016
-
[36]
Olivier J Hénaff, Robbe LT Goris, and Eero P Simoncelli. 2019. Perceptual straightening of natural videos.Nature neuroscience 22, 6 (2019), 984–991
2019
-
[37]
Carlos Javier Hernandez-Castro and Arturo Ribagorda. 2010. Pitfalls in CAPTCHA design and implementation: The Math CAPTCHA, a case study. computers & security 29, 1 (2010), 141–157
2010
-
[38]
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. 2022. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303 (2022)
2022 arXiv
-
[39]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022. Video diffusion models. Advances in Neural Information Processing Systems 35 (2022), 8633–8646
2022
-
[40]
Yaosi Hu, Chong Luo, and Zhenzhong Chen. 2022. Make it move: controllable image-to-video generation with text descriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18219–18228. 20 Lehao Lin, Ke Wang, Maha Abdallah, and Wei Cai
2022
-
[41]
Guodong Huang, Chuan Ma, Ming Ding, Yuwen Qian, Chunpeng Ge, Liming Fang, and Zhe Liu. 2023. Efficient and low overhead website fingerprinting attacks and defenses based on TCP/IP traffic. In Proceedings of the ACM Web Conference 2023 . 1991–1999
2023
-
[42]
Montree Imsamai and Suphakant Phimoltares. 2010. 3D CAPTCHA: A next generation of the CAPTCHA. In 2010 International Conference on Information Science and Applications. IEEE, 1–8
2010
-
[43]
Iat Long Iong, Xiao Liu, Yuxuan Chen, Hanyu Lai, Shuntian Yao, Pengbo Shen, Hao Yu, Yuxiao Dong, and Jie Tang. 2024. OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Ling...
2024
-
[44]
Rutvij H Jhaveri, Sankita J Patel, and Devesh C Jinwala. 2012. DoS attacks in mobile ad hoc networks: A survey. In 2012 second international conference on advanced computing & communication technologies . IEEE, 535–541
2012
-
[45]
Junfeng Jing, Shenjuan Liu, Gang Wang, Weichuan Zhang, and Changming Sun. 2022. Recent advances on image edge detection: A comprehensive review. Neurocomputing 503 (2022), 259–271
2022
-
[46]
Hongwen Kang, Kuansan Wang, David Soukal, Fritz Behr, and Zijian Zheng. 2010. Large-scale bot detection for search engines. In Proceedings of the 19th international conference on World wide web . 501–510
2010
-
[47]
Mohammad Karami, Youngsam Park, and Damon McCoy. 2016. Stress testing the booters: Understanding and undermining the business of DDoS services. In Proceedings of the 25th International Conference on World Wide Web . 1033–1043
2016
-
[48]
Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi
-
[49]
Suzi Kim and Sunghee Choi. 2019. Dotcha: A 3d text-based scatter-type captcha. In Web Engineering: 19th International Conference, ICWE 2019, Daejeon, South Korea, June 11–14, 2019, Proceedings 19 . Springer, 238–252
2019
-
[50]
Kurt Alfred Kluever and Richard Zanibbi. 2009. Balancing usability and security in a video CAPTCHA. In Proceedings of the 5th Symposium on Usable Privacy and Security . 1–11
2009
-
[51]
Karel Kubicek, Jakob Merane, Ahmed Bouhoula, and David Basin. 2024. Automating Website Registration for Studying GDPR Compliance. In Proceedings of the ACM on Web Conference 2024 . 1295–1306
2024
-
[52]
Pierre Laperdrix, Gildas Avoine, Benoit Baudry, and Nick Nikiforakis. 2019. Morellian analysis for browsers: Making web authentication stronger with canvas fingerprinting. In Detection of Intrusions and Malware, and Vulnerability Assessment: 16th International Conference, DIMV...
2019
-
[53]
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. 2021. Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 7331–7341
2021
-
[54]
Jiangtao Li, Ninghui Li, XiaoFeng Wang, and Ting Yu. 2009. Denial of service attacks and defenses in decentralized trust management.International Journal of Information Security 8 (2009), 89–101
2009
-
[55]
Linjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Ce Liu, and Lijuan Wang. 2023. Lavender: Unifying video-language understanding as masked language modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 23119–23129
2023
-
[56]
Qiujie Li. 2015. A computer vision attack on the ARTiFACIAL CAPTCHA. Multimedia Tools and Applications 74 (2015), 4583–4597
2015
-
[57]
Xigao Li, Babak Amin Azad, Amir Rahmati, and Nick Nikiforakis. 2023. Scan Me If You Can: Understanding and Detecting Unwanted Vulnerability Scanning. In Proceedings of the ACM Web Conference 2023 . 2284–2294
2023
-
[58]
Xigao Li, Babak Amin Azad, Amir Rahmati, and Nick Nikiforakis. 2021. Good bot, bad bot: Characterizing automated browsing activity. In 2021 IEEE symposium on security and privacy (sp) . IEEE, 1589–1605
2021
-
[59]
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, et al. 2024. Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv:2402.17177 (2024)
2024 arXiv
-
[60]
Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3431–3440
2015
-
[61]
Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. 2023. Videofusion: Decomposed diffusion models for high-quality video generation. arXiv preprint arXiv:2303.08320 (2023)
2023 arXiv
-
[62]
Xinbei Ma, Zhuosheng Zhang, and Hai Zhao. 2024. CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation. In Findings of the Association for Computational Linguistics ACL 2024 . 9097–9110
2024
-
[63]
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2023. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424 (2023)
2023 arXiv
-
[64]
Raman Maini and Himanshu Aggarwal. 2009. Study and comparison of various image edge detection techniques. International journal of image processing (IJIP) 3, 1 (2009), 1–11
2009
-
[65]
Peter Matthews, Andrew Mantel, and Cliff C Zou. 2010. Scene tagging: image-based CAPTCHA using image composition and object relationships. In Proceedings of the 5th ACM Symposium on Information, Computer and Communications Security . 345–350
2010
-
[66]
Manar Mohamed, Niharika Sachdeva, Michael Georgescu, Song Gao, Nitesh Saxena, Chengcui Zhang, Ponnurangam Kumaraguru, Paul C Van Oorschot, and Wei-Bang Chen. 2014. A three-way investigation of a game-captcha: automated attacks, relay attacks and usability. In Proceedings of th...
2014
-
[67]
Vu Duc Nguyen, Yang-Wai Chow, and Willy Susilo. 2012. Breaking a 3D-based CAPTCHA scheme. In Information Security and Cryptology-ICISC 2011: 14th International Conference, Seoul, Korea, November 30-December 2, 2011. Revised Selected Papers 14 . Springer, 391–405
2012
-
[68]
Vu Duc Nguyen, Yang-Wai Chow, and Willy Susilo. 2014. On the security of text-based 3D CAPTCHAs. Computers & security 45 (2014), 84–99
2014
-
[69]
Nick Nikiforakis, Alexandros Kapravelos, Wouter Joosen, Christopher Kruegel, Frank Piessens, and Giovanni Vigna. 2013. Cookieless monster: Exploring the ecosystem of web-based device fingerprinting. In 2013 IEEE Symposium on Security and Privacy . IEEE, 541–555
2013
-
[70]
Behzad Ousat, Esteban Schafir, Duc C Hoang, Mohammad Ali Tofighi, Cuong V Nguyen, Sajjad Arshad, Selcuk Uluagac, and Amin Kharraz
-
[71]
Nitisha Payal, Nidhi Chaudhary, and Parma Nand Astya. 2012. JigCAPTCHA: An Advanced Image-Based CAPTCHA Integrated with Jigsaw Piece Puzzle using AJAX. International Journal of Soft Computing and Engineering (IJSCE) 2, 5 (2012), 2231–2307
2012
-
[72]
M Kameswara Rao, MSVK Maniraj, and B Sneha Ganga. 2014. Improved video captcha. Journal of Emerging Technologies in Web Intelligence 6, 4 (2014), 416–416
2014
-
[73]
In Proceedings of the ACM on Web Conference 2024
The Matter of Captchas: An Analysis of a Brittle Security Feature on the Modern Web. In Proceedings of the ACM on Web Conference 2024 . 1835–1846
2024
-
[74]
Steven A Ross, J Alex Halderman, and Adam Finkelstein. 2010. Sketcha: a captcha based on line drawings of 3d models. In Proceedings of the 19th international conference on World wide web . 821–830
2010
-
[75]
Philip Sedgwick. 2012. Pearson’s correlation coefficient. Bmj 345 (2012)
2012
-
[76]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence 39, 6 (2016), 1137–1149
2016
-
[77]
Dongyu She and Kun Xu. 2022. An image-to-video model for real-time video enhancement. In Proceedings of the 30th ACM International Conference on Multimedia. 1837–1846
2022
-
[78]
Suphannee Sivakorn, Iasonas Polakis, and Angelos D Keromytis. 2016. I am robot:(deep) learning to break semantic image captchas. In 2016 IEEE European Symposium on Security and Privacy (EuroS&P) . IEEE, 388–403
2016
-
[79]
Asuman Senol, Alisha Ukani, Dylan Cutler, and Igor Bilogrevic. 2024. The Double Edged Sword: Identifying Authentication Pages and their Fingerprinting Behavior. In Proceedings of the ACM on Web Conference 2024 . 1690–1701
2024
-
[80]
Oleg Starostenko, Claudia Cruz-Perez, Fernando Uceda-Ponga, and Vicente Alarcon-Aquino. 2015. Breaking text-based CAPTCHAs with variable word and character orientation. Pattern Recognition 48, 4 (2015), 1101–1112
2015
-
[81]
Wenhao Sun, Rong-Cheng Tu, Jingyi Liao, and Dacheng Tao. 2024. Diffusion Model-Based Video Editing: A Survey.arXiv preprint arXiv:2407.07111 (2024)
2024 arXiv
-
[82]
Ray Smith. 2007. An overview of the Tesseract OCR engine. In Ninth international conference on document analysis and recognition (ICDAR 2007) , Vol. 2. IEEE, 629–633
2007
-
[83]
Mengyun Tang, Haichang Gao, Yang Zhang, Yi Liu, Ping Zhang, and Ping Wang. 2018. Research on deep learning techniques in breaking text-based captchas and designing image-based captcha. IEEE Transactions on Information Forensics and Security 13, 10 (2018), 2522–2537
2018
-
[84]
Isakwisa Gaddy Tende, Kentaro Aburada, Hisaaki Yamaba, Tetsuro Katayama, and Naonobu Okazaki. 2021. Development and Evaluation of Swahili Text Based CAPTCHA. In 2021 IEEE 3rd Global Conference on Life Sciences and Technologies (LifeTech) . IEEE, 293–297
2021
-
[85]
Mayumi Takaya, Yusuke Tsuruta, and Akihiro Yamamura. 2013. Reverse Turing Test using Touchscreens and CAPTCHA. J. Wirel. Mob. Networks Ubiquitous Comput. Dependable Appl. 4, 3 (2013), 41–57
2013
-
[86]
Upthrust. 2024. Runway, Luma, Kling, Pika, and Haiper: AI Video Generators Review Roundup. (August 2024). https://upthrust.co/2024/08/runway- luma-kling-pika-and-haiper-ai-video-generators-review-roundup Accessed: 2024-10-02
2024
-
[87]
Luis Von Ahn, Benjamin Maurer, Colin McMillen, David Abraham, and Manuel Blum. 2008. recaptcha: Human-based character recognition via web security measures. Science 321, 5895 (2008), 1465–1468
2008
-
[88]
Sheng Tian and Tao Xiong. 2020. A generic solver combining unsupervised learning and representation learning for breaking text-based captchas. In Proceedings of The Web Conference 2020 . 860–871
2020
-
[89]
Jiawei Wang, Liping Yuan, and Yuchen Zhang. 2024. Tarsier: Recipes for Training and Evaluating Large Video Description Models. arXiv preprint arXiv:2407.00634 (2024)
2024 arXiv
-
[90]
Ke Wang, Lehao Lin, Maha Abdallah, and Wei Cai. 2025. Where is the Boundary? Understanding How People Recognize and Evaluate Generative AI-extended Videos. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems
2025
-
[91]
Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Kevin Qinghong Lin, Satoshi Tsutsui, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, et al. 2023. All in one: Exploring unified video-language pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...
2023
-
[92]
Tingting Wang and Jørgen Bøegh. 2014. Multi-layer CAPTCHA based on Chinese character deformation. In Trustworthy Computing and Services: International Conference, ISCTCS 2013, Beijing, China, November 2013, Revised Selected Papers . Springer, 205–211
2014
-
[93]
Simon S Woo, Jingul Kim, Duoduo Yu, and Beomjun Kim. 2017. Exploration of 3D texture and projection for new CAPTCHA design. InInformation Security Applications: 17th International Workshop, WISA 2016, Jeju Island, Korea, August 25-27, 2016, Revised Selected Papers 17 . Springe...
2017
-
[94]
Ping Wang, Haichang Gao, Xiaoyan Guo, Chenxuan Xiao, Fuqi Qi, and Zheng Yan. 2023. An experimental investigation of text-based captcha attacks and their robustness. Comput. Surveys 55, 9 (2023), 1–38
2023
-
[95]
Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, and Yu-Gang Jiang. 2023. A survey on video diffusion models. arXiv preprint arXiv:2310.10647 (2023)
2023 arXiv
-
[96]
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. Videoclip: Contrastive pre-training for zero-shot video-text understanding. arXiv preprint arXiv:2109.14084 (2021)
2021 arXiv
-
[97]
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua. 2024. Can i trust your answer? visually grounded video question answering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13204–13214. 22 Lehao Lin, Ke Wang, Maha Abdallah, and Wei Cai
2024
-
[98]
Xin Xu, Lei Liu, and Bo Li. 2020. A survey of CAPTCHA technologies to distinguish between human and computer. Neurocomputing 408 (2020), 292–307
2020
-
[99]
Takumi Yamamoto, J Doug Tygar, and Masakatsu Nishigaki. 2010. CAPTCHA using strangeness in machine translation. In 2010 24th IEEE International Conference on Advanced Information Networking and Applications . IEEE, 430–437
2010
-
[100]
Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. 2024. PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning. arXiv:2404.16994 [cs.CV] https://arxiv.org/abs/2404.16994
2024 arXiv
-
[101]
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023. A survey on multimodal large language models. arXiv preprint arXiv:2306.13549 (2023)
2023 arXiv
-
[102]
Junnan Yu, Xuna Ma, and Ting Han. 2017. Usability investigation on the localization of text captchas: take chinese characters as a case study. In Transdisciplinary Engineering: A Paradigm Shift . IOS Press, 233–242
2017
-
[103]
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. 2024. MiniCPM-V: A GPT-4V Level MLLM on Your Phone. arXiv preprint arXiv:2408.01800 (2024)
2024 arXiv
-
[104]
Jerrold H Zar. 2005. Spearman rank correlation. Encyclopedia of biostatistics 7 (2005)
2005
-
[105]
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021. Merlot: Multimodal neural script knowledge models. Advances in neural information processing systems 34 (2021), 23634–23651
2021
-
[106]
Zhengqing Yuan, Yixin Liu, Yihan Cao, Weixiang Sun, Haolong Jia, Ruoxi Chen, Zhaoxu Li, Bin Lin, Li Yuan, Lifang He, et al. 2024. Mora: Enabling generalist video generation via a multi-agent framework. arXiv preprint arXiv:2403.13248 (2024)
2024 arXiv
-
[107]
Shijie Zhang and Jong-Hyouk Lee. 2019. Double-spending with a sybil attack in the bitcoin decentralized network. IEEE transactions on Industrial Informatics 15, 10 (2019), 5715–5722
2019
-
[108]
Ziyi Zhang, Shuofei Zhu, Jaron Mink, Aiping Xiong, Linhai Song, and Gang Wang. 2022. Beyond bot detection: combating fraudulent online survey takers. In Proceedings of the ACM Web Conference 2022 . 699–709
2022
-
[109]
Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[112]
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022. Learning to prompt for vision-language models. International Journal of Computer Vision 130, 9 (2022), 2337–2348
2022
-
[2023]
In Proceedings of the IEEE/CVF International Conference on Computer Vision
Text2video-zero: Text-to-image diffusion models are zero-shot video generators. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 15954–15964
-
[2024]
In Proceedings of the ACM on Web Conference 2024
Medusa: Unveil Memory Exhaustion DoS Vulnerabilities in Protocol Implementations. In Proceedings of the ACM on Web Conference 2024 . 1668–1679
2024
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.