REVIEW 4 major objections 6 minor 1 cited by
PopSweeper: Automatically Detecting and Resolving App-Blocking Pop-Ups to Assist Automated Mobile GUI Testing
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A screenshot-only pipeline can detect app-blocking pop-ups during mobile GUI tests and return close-button coordinates, resolving blockages in 87.1% of apps in end-to-end evaluation.
desk verdict Useful dataset and task definition, but the marquee end-to-end number is a replay-based upper bound, not a demonstrated live resolution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a three-stage real-time pipeline. First, a 100 ms sampler plus RGB histogram similarity (threshold 0.8) sends only visually changed frames onward, cutting the frame volume by roughly an order of magnitude. Second, a two-stage classifier stacks ResNet50 as primary detector and MobileNetV2 as a verification stage, with custom binary-classification heads, so that a frame is treated as a pop-up only when both models agree; this stacking is what lifts precision from about 0.87 (ResNet50 alone) to 0.917 while keeping recall near 0.935. Third, YOLO-World, fine-tuned on 832 manually labeled pop-up screenshots, outputs a 640x640 bounding box (x1,y1,x2,y2) for the close button. The pipeline's output is not a label but an actionable coordinate that a test script can click, which is what turns detection into resolution.
What would settle it
Run PopSweeper inside a real automated GUI test runner on a fresh set of 50 popular apps, letting the test script click the returned coordinates as soon as they arrive. If the app-level blockage-resolution rate falls substantially below the claimed 87.1%, or if clicks frequently land at the wrong moment because a pop-up is still animating or has already auto-closed, the central claim is not supported. A cleaner direct check is to count how often a click at the returned coordinates actually dismisses the pop-up (verified by screenshot before and after) in live runs versus the replay-based estimate.
Extended reading notes
Core claim
An app-blocking pop-up can be treated as a visual anomaly during GUI testing rather than a special case that the test script must predict. PopSweeper establishes that a lightweight screenshot-only pipeline can detect these pop-ups and localize their close buttons with enough accuracy and speed to be inserted into an existing test workflow: sample frames every 100 ms, drop near-duplicates by histogram similarity, classify remaining frames with a two-stage ResNet50/MobileNetV2 stack, and, when a pop-up is found, run YOLO-World to return a clickable bounding box. The central empirical claim is the combination of classification precision 91.7%, recall 93.5%; close-button detection BoxAP 93.9%, recall 89.2%; and an end-to-end app-level resolution rate of 87.1% across 155 apps, at an average processing cost of roughly 60 ms per selected frame. The authors claim this is the first screenshot-based system that both detects app-blocking pop-ups and provides the resolution action, and they position it as a complement to existing exploration and LLM-driven GUI testing agents rather than a replacement.
Load-bearing premise
The load-bearing premise is that the 60-second recorded replays used for end-to-end evaluation behave like a live test: if pop-up timing, animations, dynamically loaded content, or close-button hit behavior differs under a real test runner, the reported 87.1% resolution rate may not transfer, and the metrics also depend on the undocumented accuracy of the manual annotations.
Editorial extensions
If this is right
- Automated GUI test scripts can be extended with a few lines: send each 100 ms frame to PopSweeper, and if coordinates come back, tap there and continue; no per-app pop-up modeling is needed.
- Large-scale comparative and regression testing, where one team runs many apps across devices, would no longer stall at ads and system alerts that appear unpredictably.
- The reported error analysis says that misclassifying a pop-up as app content behaves like not using the tool, while the rare bad-coordinate clicks within a pop-up can still redirect, so the tool degrades gracefully rather than crashing the run.
- Because the system works on screenshots alone, it applies to any GUI test setup that can produce frames, without instrumenting the app or needing layout or XML access.
Reading between the lines
- Beyond the paper, the same pipeline could be dropped in front of LLM-driven GUI agents, which the paper notes currently analyze screenshots without pop-up awareness, letting agents receive a cleaned frame before choosing their next action.
- Beyond the paper, the histogram gate may miss pop-ups that appear and vanish within one 100 ms sampling window or that barely change the global RGB distribution, so a change-detection baseline that also tracks local regions would be a natural testable extension.
- Beyond the paper, the lower performance on fullscreen text-based and Chinese close buttons suggests a concrete next experiment: retrain the detector with a balanced multilingual close-button corpus and measure whether the 37.5% recall on fullscreen top-app pop-ups recovers.
- Beyond the paper, the public dataset of 832 labeled pop-up screenshots could support a benchmark for open-vocabulary pop-up detection, since YOLO-World is an open-vocabulary detector and the paper only uses it in a fixed fine-tuned mode.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. PopSweeper is a computer-vision pipeline for detecting and resolving app-blocking pop-ups during automated mobile GUI testing. It combines histogram-based frame differencing, a two-stage ResNet50+MobileNetV2 classifier, and a YOLO-World close-button detector, and returns close-button coordinates to a test script. The paper contributes a manually labeled dataset of 832 app-blocking pop-ups from the RICO dataset and 87 top-ranked apps, an empirical prevalence study, and an evaluation reporting 91.7% precision and 93.5% recall for classification, 93.9% BoxAP and 89.2% recall for close-button detection, and an app-level resolution rate of 87.1%. The two RQ3 experiments are both offline replays: recorded frames are passed through the pipeline and the returned close-button coordinates are compared with manual ground-truth labels; no tap is dispatched to a live app and no post-tap state is verified.
Significance. The dataset and code release (Section 8), the breadth of the manual labeling effort, and the comparison against several reasonable baselines (ResNet50, MobileNetV2, VGG19, custom CNN, CLIP, Faster R-CNN, YOLOv8) are genuine strengths. The second RQ3 experiment on 51 additional apps provides some out-of-distribution evidence. If the central claim were demonstrated in a live test runner, PopSweeper would be a useful complement to existing automated GUI testing workflows, with an attractively low per-frame overhead. The main weakness is that the headline 'resolved blockages in 87.1% of apps' claim is currently supported only by coordinate agreement on replayed recordings, not by actual resolution of pop-ups in a running test; this gap is acknowledged in Section 5.2 and is the key obstacle to accepting the paper's strongest claim.
major comments (4)
- [Section 4.4, Table 4] The central 'end-to-end resolution' claim is not demonstrated by the experiments as described. In both RQ3 experiments, recorded frames are replayed and the returned close-button coordinates are compared with manual ground-truth labels; no tap is dispatched, no UI hierarchy or screenshot is checked after a tap to verify that the pop-up disappeared, and no test script continues past the pop-up. The reported 87.1% app-level rate and the Table 4 precision/recall values therefore measure offline coordinate agreement, not successful resolution of live app-blocking pop-ups. Pop-up entrance animations, coordinate scaling under a real runner, overlapping views, and returned coordinates on non-clickable regions can all break the transfer. The paper's own Section 5.2 concedes that the RQ3 scenarios 'may not fully capture the real-world behavior of pop-ups.' The stress-test concern lands: either a live test-runner integration must be added, or the abstract and conclusions must be reworded to state that the result is close-button localization agreement rather than demonstrated resolution.
- [Section 4.4, app-level results] There is a numeric inconsistency in the key result. The first RQ3 experiment states that the test set contains pop-up screenshots from 154 unique apps, but the app-level result is reported as '135 out of 155 apps (87.1%).' If the denominator is 154, the rate is 87.7%; if it is 155, the earlier sentence should say 155. The manuscript must correct this and state the exact denominator for the headline 87.1% figure, since the contradiction undermines reproducibility of the abstract's main quantitative claim.
- [Section 4.4, Evaluation Metrics] The end-to-end metrics in Equations (10) and (11) are under-specified. The text says the authors 'compare the returned coordinates with the ground truth,' but it does not define a correctness criterion: no IoU threshold, no pixel-distance tolerance, and no requirement that the predicted point lie inside the ground-truth close-button box are given. Without this operational definition, the precision/recall values in Table 4 and the app-level resolution rate cannot be independently reproduced or compared with future detection-based approaches.
- [Section 5.2 (construct validity)] Manual annotation quality is not quantified. The paper relies on manually labeled pop-up regions and close-button boxes both for training and for ground-truth evaluation, and Section 5.2 acknowledges that 'the reliance on manual labeling ... may introduce subjectivity,' but no inter-annotator agreement (e.g., Cohen's kappa), no independent second-pass verification, and no detailed annotation protocol are reported. Because all reported metrics inherit the label quality, the paper should report at least a sampled agreement study or otherwise characterize label reliability.
minor comments (6)
- [Abstract vs. Conclusion] The abstract reports 91.7% precision while the conclusion reports 91.8%, and Table 2 shows 0.917; these numbers should be unified.
- [Section 5.2 and Section 5.3] Section 5.2 says PopSweeper achieved 'a high recall of 92.4%' in identifying pop-ups, which does not match the 93.5% recall reported in Table 2 and the abstract; Section 5.3 repeats the 92.4% figure. This inconsistency should be corrected.
- [Conclusion, Section 7] The conclusion says 'we manually reviewed over 7K screenshots from the RICO dataset,' but the dataset described throughout the paper is 72,218 screenshots; '7K' appears to be a typo.
- [References [14], [15], [1]] References [14] and [15] list 'John Doe and Jane Smith' and are not verifiable, and [1] is 'Anonymous' before the abstract. Placeholder or unverifiable citations are not acceptable in a journal submission; all bibliography entries should be real, complete, and traceable.
- [Section 3.2, sentence-level clarity] Section 3.2 contains 'We leverages a two-stage classification pipeline,' which is a subject-verb agreement error; a proofread pass would improve readability.
- [Section 4.1 vs. Section 4.4] Section 4.1 says 'we ultimately selected approximately 1,000 unique apps from the Rico dataset,' but the RQ3 test set is described as containing 154 unique apps; the relationship between the full curated set and the RQ3 test subset should be stated explicitly.
Circularity Check
No circular derivation: PopSweeper is an empirical ML pipeline evaluated on held-out splits; the replay-based end-to-end claim is a construct-validity threat, not circularity.
full rationale
PopSweeper is an empirical machine-learning pipeline; there is no equation-to-equation derivation in which a predicted quantity is equal to an input by construction. The pop-up classifier and close-button detector are trained on manually labeled RICO/top-app screenshots and evaluated on held-out test splits with no overlapping apps (Sections 4.2 and 4.3), and the second end-to-end experiment uses 4th/5th ranked apps not in the training set (Section 4.4). No fitted parameter is renamed as a prediction; the reported 'app-level resolution' metric is an overstatement because RQ3 replays recordings and compares returned coordinates with manual labels rather than dispatching clicks in live apps, and Section 5.2 admits 'the end-to-end app testing scenarios in RQ3 may not fully capture the real-world behavior of pop-ups.' That is a construct-validity threat, not a circularity. The only self-reference is the anonymous Zenodo data/code repository [1], which is non-load-bearing for the scientific claims; the models and comparisons rest on external baselines such as ResNet50, MobileNetV2, YOLO-World, Faster R-CNN, and CLIP. Therefore no circular step is present; the low score reflects only the minor non-load-bearing self-citation and the mild self-evaluation inherent to building one's own labeled dataset, not a derivation that reduces to its inputs.
Assumptions & free parameters
free parameters (3)
- Histogram similarity threshold =
0.8
- Frame sampling interval =
100 ms
- Oversampling duplication counts =
+369 pop-up instances, +360 content instances
assumptions (5)
- domain assumption The manual labeling of 832 pop-up screenshots is correct.
- domain assumption Replayed recordings faithfully reproduce live automated GUI testing.
- domain assumption A close-button tap is the correct way to resolve every blocking pop-up.
- domain assumption Histogram similarity at 0.8 catches all UI changes that matter.
- domain assumption Pre-trained ImageNet features transfer to GUI screenshots.
Cite this review
Pith. "Pith review of PopSweeper: Automatically Detecting and Resolving App-Blocking Pop-Ups to Assist Automated Mobile GUI Testing." pith.science (2026). https://pith.science/paper/I3LSAFNR
@misc{pith2026241202933,
author = {Pith},
title = {Pith review of: PopSweeper: Automatically Detecting and Resolving App-Blocking Pop-Ups to Assist Automated Mobile GUI Testing},
year = {2026},
howpublished = {\url{https://pith.science/paper/I3LSAFNR}},
note = {Machine review of arXiv:2412.02933}
}
read the original abstract
Graphical User Interfaces (GUIs) are the primary means by which users interact with mobile applications, making them crucial to both app functionality and user experience. However, a major challenge in automated testing is the frequent appearance of app-blocking pop-ups, such as ads or system alerts, which obscure critical UI elements and disrupt test execution, often requiring manual intervention. These interruptions lead to inaccurate test results, increased testing time, and reduced reliability, particularly for stakeholders conducting large-scale app testing. To address this issue, we introduce PopSweeper, a novel tool designed to detect and resolve app-blocking pop-ups in real-time during automated GUI testing. PopSweeper combines deep learning-based computer vision techniques for pop-up detection and close button localization, allowing it to autonomously identify pop-ups and ensure uninterrupted testing. We evaluated PopSweeper on over 72K app screenshots from the RICO dataset and 87 top-ranked mobile apps collected from app stores, manually identifying 832 app-blocking pop-ups. PopSweeper achieved 91.7% precision and 93.5% recall in pop-up classification and 93.9% BoxAP with 89.2% recall in close button detection. Furthermore, end-to-end evaluations demonstrated that PopSweeper successfully resolved blockages in 87.1% of apps with minimal overhead, achieving classification and close button detection within 60 milliseconds per frame. These results highlight PopSweeper's capability to enhance the accuracy and efficiency of automated GUI testing by mitigating pop-up interruptions.
Figures
Forward citations
Cited by 1 Pith paper
-
Screencast-Based Analysis of User-Perceived GUI Responsiveness
MobileGUIPerf detects on-screen touch indicators and frame-level visual changes in screencasts to measure GUI response and finish times, reaching 0.96 precision, 0.93 recall, and 50 ms and 100 ms timing accuracy on 2,...
Reference graph
Works this paper leans on
-
[14]
John Doe and Jane Smith. 2020. Transfer learning by fine-tuning pre-trained convolutional neural networks. Journal of Advanced Machine Learning Research (2020), 45–58
work page 2020
-
[15]
John Doe and Jane Smith. 2023. Customized ResNet for binary classification in medical image analysis. Research Square (2023). https://www.researchsquare.com/article/rs-2863523/v1.pdf
work page 2023
-
[1]
Anonymous. 2024. PopSweeper: Automatically Detecting and Resolving App-Blocking Pop-Ups to Assist Automated Mobile GUI Testing. https://doi.org/10.5281/zenodo.13754620
-
[2]
Carlos Bernal-Cárdenas, Nathan Cooper, Kevin Moran, Oscar Chaparro, Andrian Marcus, and Denys Poshyvanyk
-
[3]
Theodore Book, Adam Pridgen, and Dan S. Wallach. 2013. Longitudinal Analysis of Android Ad Library Permissions. arXiv:1303.0857 [cs.CR] https://arxiv.org/abs/1303.0857
work page Pith review arXiv 2013
-
[4]
Theodore Book and Dan S. Wallach. 2015. An Empirical Study of Mobile Ad Targeting. arXiv:1502.06577 [cs.CR]
work page Pith review arXiv 2015
-
[5]
Shaoheng Cao, Minxue Pan, Yu Pei, Wenhua Yang, Tian Zhang, Linzhang Wang, and Xuandong Li. 2024. Comprehensive Semantic Repair of Obsolete GUI Test Scripts for Mobile Applications. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering (Lisbon, Portugal) (ICSE ’24). Association for Computing Machinery, New York, NY, USA, Artic...
arXiv 2024
-
[6]
Gong Chen, Wei Meng, and John Copeland. 2019. Revisiting Mobile Advertising Threats with MAdLife. In The World Wide Web Conference (San Francisco, CA, USA) (WWW ’19). Association for Computing Machinery, New York, NY, USA, 207–217. https://doi.org/10.1145/3308558.3313549
arXiv 2019
Show all 71 references
-
[8]
Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, and Ying Shan. 2024. YOLO-World: Real-Time Open- Vocabulary Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (Vancouver, Canada) (CVPR ’24). IEEE, 1205–1...
2024
-
[9]
Eliane Collins, Arilo Neto, Auri Vincenzi, and José Maldonado. 2021. Deep Reinforcement Learning based Android Application GUI Testing. In Proceedings of the XXXV Brazilian Symposium on Software Engineering (Joinville, Brazil) (SBES ’21). Association for Computing Machinery, N...
2021
-
[10]
Riccardo Coppola, Maurizio Morisio, Marco Torchiano, and Luca Ardito. 2019. Scripted GUI testing of Android open-source apps: evolution of test code and fragility causes. Empirical Softw. Engg. 24, 5 (oct 2019), 3205–3248
2019
-
[11]
Luis Cruz, Rui Abreu, and David Lo. 2019. To the attention of mobile software developers: guess what, test your app! Empirical Softw. Engg. 24, 4 (aug 2019), 2438–2468. https://doi.org/10.1007/s10664-019-09701-0
2019 doi
-
[12]
Biplab Deka, Zhe Huang, Carl Franzen, John Hibschman, Michael Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. 2017. Rico: A mobile app dataset for building data-driven design applications. InProceedings of the 30th Annual ACM Symposium on User Interface Software and Tec...
2017
-
[13]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2009), 248–255
2009
-
[16]
Bissyandé, Tianming Liu, Guoai Xu, and Jacques Klein
Feng Dong, Haoyu Wang, Li Li, Yao Guo, Tegawendé F. Bissyandé, Tianming Liu, Guoai Xu, and Jacques Klein. 2018. FraudDroid: automated ad fraud detection for Android apps. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposiu...
2018
-
[17]
Sidong Feng, Chunyang Chen, and Zhenchang Xing. 2023. Video2Action: Reducing Human Interactions in Action Annotation of App Tutorial Videos. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23). Associati...
2023
-
[19]
R. F. Fouladi, C. E. Kayatas, and E. Anarim. 2016. Frequency based DDoS attack detection approach using naive Bayes classification. In International Conference on Telecommunications and Signal Processing (TSP) . 104–107
2016
-
[20]
Cuiyun Gao, Jichuan Zeng, David Lo, Xin Xia, Irwin King, and Michael R. Lyu. 2022. Understanding in-app advertising issues based on large scale app review analysis. Information and Software Technology 142 (2022), 106741
2022
-
[21]
Cuiyun Gao, Jichuan Zeng, Federica Sarro, David Lo, Irwin King, and Michael R. Lyu. 2021. Do users care about ad’s performance costs? Exploring the effects of the performance costs of in-app ads on user experience. Information and Software Technology 132 (2021), 106471. https:...
2021
-
[22]
Emanuel Giger, Martin Pinzger, and Harald Gall. 2010. Predicting the delay of issues with due dates in software projects. Empirical Software Engineering (2010), 52–56
2010
-
[23]
Google. 2023. UI/Application Exerciser Monkey. https://developer.android.com/studio/test/other-testing-tools/monkey
2023
-
[24]
Grace, Wu Zhou, Xuxian Jiang, and Ahmad-Reza Sadeghi
Michael C. Grace, Wu Zhou, Xuxian Jiang, and Ahmad-Reza Sadeghi. 2012. Unsafe exposure analysis of mobile in-app advertisements. In Proceedings of the Fifth ACM Conference on Security and Privacy in Wireless and Mobile Networks (Tucson, Arizona, USA) (WISEC ’12). Association f...
2012
-
[25]
Tianxiao Gu, Chengnian Sun, Xiaoxing Ma, Chun Cao, Chang Xu, Yuan Yao, Qirun Zhang, Jian Lu, and Zhendong Su
-
[26]
Jiaping Gui, Stuart Mcilroy, Meiyappan Nagappan, and William G. J. Halfond. 2015. Truth in Advertising: The Hidden Cost of Mobile Ads for Software Developers. In 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering, Vol. 1. 100–110. https://doi.org/10.1109/...
2015 doi
-
[27]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016), 770–778
2016
-
[28]
Yongxiang Hu, Jiazhen Gu, Shuqing Hu, Yu Zhang, Wenjie Tian, Shiyu Guo, Chaoyi Chen, and Yangfan Zhou. 2023. Appaction: Automatic GUI Interaction for Mobile Apps via Holistic Widget Perception. In Proceedings of the 31st ACM Joint European Software Engineering Conference and S...
2023
-
[29]
A Jaiswal, N Gianchandani, D Singh, V Kumar, and M Kaur. 2022. Automatic Cauliflower Disease Detection Using Fine-Tuning Transfer Learning Approach. SN Computer Science (2022)
2022
-
[30]
Anureet Kaur and Kulwant Kaur. 2022. Systematic literature review of mobile application development and testing. Journal of King Saud University – Computer and Information Sciences 34 (2022), 1–15
2022
-
[31]
KDnuggets. 2022. Evaluating Object Detection Models Using Mean Average Precision. KDnuggets (2022)
2022
-
[32]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems , F. Pereira, C.J. Burges, L. Bottou, and K.Q. Weinberger (Eds.), Vol. 25. Curran Associates, Inc
2012
-
[33]
Martin Kropp and Pamela Morales. 2010. Automated GUI Testing on the Android Platform. IMVS Fokus Report (01 2010), 67–72
2010
-
[34]
Yuanhong Lan, Yifei Lu, Zhong Li, Minxue Pan, Wenhua Yang, Tian Zhang, and Xuandong Li. 2024. Deeply Reinforcing Android GUI Testing with Deep Reinforcement Learning. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (Lisbon, Portugal) (ICSE ...
2024
-
[35]
Yuanchun Li, Ziyue Yang, Yao Guo, and Xiangqun Chen. 2017. DroidBot: a lightweight UI-guided test input generator for Android. In Proceedings of the 39th International Conference on Software Engineering Companion (Buenos Aires, Argentina) (ICSE-C ’17). IEEE Press, 23–26. https...
2017 doi
-
[36]
Tianming Liu, Haoyu Wang, Li Li, Xiapu Luo, Feng Dong, Yao Guo, Liu Wang, Tegawendé Bissyandé, and Jacques Klein. 2020. MadDroid: Characterizing and Detecting Devious Ad Contents for Android Apps. In Proceedings of The Web Conference 2020 (Taipei, Taiwan) (WWW ’20). Associatio...
2020
-
[37]
Zhe Liu, Chunyang Chen, Junjie Wang, Xing Che, Yuekai Huang, Jun Hu, and Qing Wang. 2023. Fill in the Blank: Context-Aware Automated Text Input Generation for Mobile GUI Testing. In Proceedings of the 45th International Conference on Software Engineering (Melbourne, Victoria, ...
2023
-
[38]
Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, and Qing Wang. 2024. Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware Decisions. In Proceedings of the IEEE/ACM 46th International Confer...
2024
-
[39]
Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Zhilin Tian, Yuekai Huang, Jun Hu, and Qing Wang. 2024. Testing the Limits: Unusual Text Inputs Generation for Mobile App Crash Detection with Large Language Model. In Proceedings of the IEEE/ACM 46th International C...
2024
-
[40]
Zhicheng Liu and Jeffrey Heer. 2014. The effects of interactive latency on exploratory visual analysis.IEEE transactions on visualization and computer graphics 20, 12 (2014), 2122–2131
2014
-
[41]
Aravind Machiry, Rohan Tahiliani, and Mayur Naik. 2013. Dynodroid: an input generation system for Android apps. In Proceedings of the 2013 9th Joint Meeting on Foundations of Software Engineering (Saint Petersburg, Russia) (ESEC/FSE 2013). Association for Computing Machinery, ...
2013
-
[42]
Keiron O’Shea and Ryan Nash. 2015. An Introduction to Convolutional Neural Networks. arXiv:1511.08458 [cs.NE]
2015 arXiv
-
[43]
Minxue Pan, An Huang, Guoxin Wang, Tian Zhang, and Xuandong Li. 2020. Reinforcement learning based curiosity- driven testing of Android applications. In Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis (Virtual Event, USA) (ISSTA 202...
2020
-
[44]
Minxue Pan, Tongtong Xu, Yu Pei, Zhong Li, Tian Zhang, and Xuandong Li. 2022. GUI-Guided Test Script Repair for Mobile Apps. IEEE Transactions on Software Engineering 48, 3 (2022), 910–929. https://doi.org/10.1109/TSE.2020.3007664
2022
-
[45]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.000...
2021 arXiv
-
[46]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. arXiv:1506.01497 [cs.CV] https://arxiv.org/abs/1506.01497
2016 arXiv
-
[47]
Andrea Romdhana, Alessio Merlo, Mariano Ceccato, and Paolo Tonella. 2022. Deep Reinforcement Learning for Black-box Testing of Android Apps. ACM Trans. Softw. Eng. Methodol. 31, 4, Article 65 (jul 2022), 29 pages
2022
-
[48]
Israel Ruiz, Meiyappan Nagappan, Bram Adams, Theodore Berger, Steffen Dienst, and Ahmed E. Hassan. 2014. Impact of Ad Libraries on Ratings of Android Mobile Apps.Software, IEEE 31 (11 2014), 86–92. https://doi.org/10.1109/MS.2014.79
2014 doi
-
[49]
Sahu and B
S. Sahu and B. M. Mehtre. 2015. Network intrusion detection system using J48 Decision Tree. InInternational Conference on Advances in Computing, Communications and Informatics (ICACCI) . 2023–2026
2015
-
[50]
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018. MobileNetV2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018), 4510–4520
2018
-
[51]
Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv:1409.1556 [cs.CV] https://arxiv.org/abs/1409.1556
2015 arXiv
-
[52]
Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations
2015
-
[53]
Ting Su, Guozhu Meng, Yuting Chen, Ke Wu, Weiming Yang, Yao Yao, Geguang Pu, Yang Liu, and Zhendong Su
-
[54]
A Subeesh and CR Mehta. 2022. Soil Image Classification Using Transfer Learning Approach: MobileNetV2 with CNN. SN Computer Science 3, 6 (2022), 1–13
2022
-
[55]
I. S. Thaseen and C. A. Kumar. 2016. Intrusion detection model using fusion of chi-square feature selection and multi class SVM. Journal of King Saud University-Computer and Information Sciences 29, 4 (2016), 462–472
2016
-
[56]
Rejin Varghese and Sambath M. 2024. YOLOv8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness. In 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS). 1–6. https://doi.org/10.1109/ADICS58448.2024.10533619
2024
-
[57]
Jue Wang, Yanyan Jiang, Chang Xu, Chun Cao, Xiaoxing Ma, and Jian Lu. 2020. ComboDroid: generating high-quality test inputs for Android apps via use case combinations. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering (Seoul, South Korea) (IC...
2020
-
[58]
Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. 2024. Mobile- Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception. arXiv preprint arXiv:2401.16158 (2024)
2024 arXiv
-
[59]
Li Wang, Wei Zhang, et al. 2024. Applying Optimized YOLOv8 for Heritage Conservation: Enhanced Object Detection in Jiangnan Traditional Private Gardens. Heritage Science 12 (2024), 45–60
2024
-
[60]
Licorish, Daniel Alencar da Costa, and Stephen G
Chathrie Wimalasooriya, Sherlock A. Licorish, Daniel Alencar da Costa, and Stephen G. MacDonell. 2024. Just-in-Time crash prediction for mobile apps. Empirical Softw. Engg. 29, 3 (may 2024), 62 pages
2024
-
[61]
Yutian Yan, Yunhui Zheng, Xinyue Liu, Nenad Medvidovic, and Weihang Wang. 2023. AdHere: Automated Detection and Repair of Intrusive Ads. In Proceedings of the 45th International Conference on Software Engineering (Melbourne, Victoria, Australia) (ICSE ’23). IEEE Press, 486–498...
2023
-
[62]
Suorong Yang, Weikang Xiao, Mengchen Zhang, Suhan Guo, Jian Zhao, and Furao Shen. 2023. Image Data Augmentation for Deep Learning: A Survey. arXiv:2204.08610 [cs.CV] https://arxiv.org/abs/2204.08610 PopSweeper: Automatically Detecting and Resolving App-Blocking Pop-Ups to Assi...
2023 arXiv
-
[63]
Faraz YazdaniBanafsheDaragh and Sam Malek. 2022. Deep GUI: black-box GUI input generation with deep learning. In Proceedings of the 36th IEEE/ACM International Conference on Automated Software Engineering (Melbourne, Australia) (ASE ’21). IEEE Press, 905–916. https://doi.org/1...
2022
-
[64]
Jiaming Ye, Ke Chen, Xiaofei Xie, Lei Ma, Ruochen Huang, Yingfeng Chen, Yinxing Xue, and Jianjun Zhao. 2021. An empirical study of GUI widget detection for industrial mobile games. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Sym...
2021
-
[65]
Shengcheng Yu, Chunrong Fang, Xin Li, Yuchen Ling, Zhenyu Chen, and Zhendong Su. 2024. Effective, Platform- Independent GUI Testing via Image Embedding and Reinforcement Learning. ACM Trans. Softw. Eng. Methodol. (jun 2024). https://doi.org/10.1145/3674728 Just Accepted
2024 doi
-
[66]
Shengcheng Yu, Chunrong Fang, Yexiao Yun, and Yang Feng. 2021. Layout and Image Recognition Driving Cross- Platform Automated Mobile Testing. InProceedings of the 43rd International Conference on Software Engineering (Madrid, Spain) (ICSE ’21). IEEE Press, 1561–1571. https://d...
2021
-
[67]
Shengcheng Yu, Chunrong Fang, Quanjun Zhang, Zhihao Cao, Yexiao Yun, Zhenfei Cao, Kai Mei, and Zhenyu Chen
-
[68]
Hevner Yu Xia, Hailiang Chen
Alan R. Hevner Yu Xia, Hailiang Chen. 2020. Third-Party SDKs and Mobile App Performance. AIS Electronic Library (2020)
2020
-
[69]
Chi Zhang, Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2023. AppAgent: Multimodal Agents as Smartphone Users. arXiv:2312.13771 [cs.CV]
2023 arXiv
-
[70]
Nengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng, Gang Wang, Yong Wu, Fang Zhou, Zhen Feng, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, and Dan Pei. 2020. Real-time incident prediction for online service systems. In Proceedings of the 28th ACM Joint Meeting on European Software Engi...
2020
-
[2017]
In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering (Paderborn, Germany) (ESEC/FSE 2017)
Guided, stochastic model-based GUI testing of Android apps. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering (Paderborn, Germany) (ESEC/FSE 2017). Association for Computing Machinery, New York, NY, USA, 245–256. https://doi.org/10.1145/31062...
2017
-
[2019]
In Proceedings of the 41st International Conference on Software Engineering (Montreal, Quebec, Canada) (ICSE ’19)
Practical GUI testing of Android applications via model abstraction and refinement. In Proceedings of the 41st International Conference on Software Engineering (Montreal, Quebec, Canada) (ICSE ’19). IEEE Press, 269–280
-
[2023]
IEEE Trans
Mobile App Crowdsourced Test Report Consistency Detection via Deep Image-and-Text Fusion Understanding. IEEE Trans. Softw. Eng. 49, 8 (aug 2023), 4115–4134. https://doi.org/10.1109/TSE.2023.3285787
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.