REVIEW 2 major objections 5 minor 59 references
The first systematic study of multi-modal agent bugs yields a three-level taxonomy and a runtime checker that finds dozens of real failures.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 10:40 UTC pith:MHTQAS33
load-bearing objection First solid empirical map of multi-modal agent bugs, with a working detector that actually finds new issues on held-out systems. the 2 major comments →
A Comprehensive Study of Implementation Bugs in Multi-modal Agents
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A top-down taxonomy of 158 multi-modal-agent bugs—organized by six global symptoms, sixteen component-level symptoms, and seven root causes—together with a runtime inter-component analyzer called MATester, covers 61.4 percent of known open issues and discovers 31 new bugs on twelve held-out agents, showing that the taxonomy is both descriptive of real failures and useful for automated detection.
What carries the argument
The three-level taxonomy (global symptoms → Perceptor/Planner/Executor symptoms → root causes) plus MATester’s comparison of successive environment, snapshot, plan and action traces; the taxonomy supplies the categories, the runtime traces supply the evidence that lets those categories be recognized automatically.
Load-bearing premise
That the 34 open-source agents and the 158 bugs mined from their public issue trackers form a representative sample of how multi-modal agents actually fail; if important industrial or closed designs are missing, both the frequency counts and the coverage numbers shrink.
What would settle it
Apply MATester (or an equivalent inter-component monitor) to a fresh corpus of multi-modal agents drawn from sources deliberately outside the original 34; if coverage of known open issues falls well below 60 percent and almost no new bugs are found, the claimed generality of the taxonomy collapses.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims to present the first systematic empirical study of M-agent-specific implementation bugs. From multi-source collection (GitHub, surveys, top venues) it filters 34 representative open-source M-agents, extracts 158 agent-specific bugs from 1,268 issue reports via keyword + dual-author manual review (Cohen’s Kappa rising to 0.84), and induces a three-level top-down taxonomy of 6 global symptoms, 16 functionality-component symptoms (Perceptor/Planner/Executor), and 7 root causes. It further implements MATester, a runtime analyzer of instrumented inter-component outputs (snapshot/plan/action/reflect), which on 12 held-out agents covers 61.4 % of known open issues and surfaces 31 previously unreported bugs, thereby demonstrating both descriptive and practical utility of the taxonomy.
Significance. If the results hold, the work supplies a timely, reusable reference for the rapidly growing M-agent literature: an open bug corpus, a multi-level taxonomy that distinguishes multi-modal interaction failures from ordinary agent bugs, concrete prevention/fix guidelines for developers, and a proof-of-concept detector (MATester) whose held-out performance already yields new bugs. The multi-stage QGS + snowballing + dual-author filtering pipeline, explicit 34/12 train/eval split, and public artifacts strengthen reproducibility. These contributions are of clear value to software-engineering and AI-systems venues concerned with reliability of agents deployed in safety-critical multi-modal settings.
major comments (2)
- [Section IV-C / Table VII] Section IV-C and Table VII: the headline claim that MATester “covers 61.4 % of known open issues and discovers 31 additional bugs” rests on a single judging LLM (GPT-4o, temperature 0.2, majority-of-3) and fixed resource bounds (5 min / 20 rounds). No ablation on judge model, temperature, or the time/round limits is reported, nor is inter-judge agreement with human raters beyond the single 97.6 % accuracy figure. Because these free parameters directly affect which symptoms are declared, a short sensitivity study is needed to underwrite the quantitative usefulness claim.
- [Section III-B] Section III-B and Threats to Validity: the 34-agent corpus is obtained by successive manual README and code inspections whose inclusion criteria (complete Perceptor–Planner–Executor modules, runnable artifacts) are only partially operationalized. While dual-author agreement is reported, the paper does not quantify how many candidates were discarded at each fine-grained filter step nor whether the retained set systematically under-samples particular architectures (e.g., multi-agent orchestration or closed-loop robotics). A brief characterization of the discarded set would strengthen the representativeness argument that underpins both the frequency tables and the taxonomy’s claimed generality.
minor comments (5)
- [Tables IV–VI] Tables IV–VI report both absolute counts and “percentage of M-agents”; with N=34 the percentages are coarse (e.g., 2.9 % = 1 agent). Adding the raw agent counts beside each percentage would improve readability.
- [Figure 3] Figure 3’s factor labels are rendered with Unicode control characters that become unreadable in some PDF viewers; replace with plain text or vector labels.
- [Section IV-B] Findings 3-1/3-2/3-3 and 5-1/5-2/5-3 restate the same co-occurrence numbers already shown in Figures 4–5; a single consolidated paragraph would reduce redundancy.
- [Abstract / Introduction] The abstract and Introduction use the non-ASCII ligature “traffic”; replace with ordinary ASCII “traffic” for indexing and accessibility.
- [Section II / IV-C] Section II’s notation (snapshot_i, environment_i, …) is clear, yet the same symbols later appear without subscripts in the MATester description; keep notation consistent.
Circularity Check
No significant circularity: taxonomy is bottom-up from independently collected GitHub issues; MATester applies those categories to held-out agents without fitted parameters or self-referential definitions.
full rationale
The paper is a standard empirical software-engineering study. Bugs are extracted from 1,268 public issue reports on 34 filtered open-source M-agents via dual-author manual labeling (Cohen’s Kappa 0.84). The three-level taxonomy (global symptoms, component-level symptoms, root causes) is induced directly from those labeled reports; frequencies and co-occurrence maps are descriptive counts, not predictions. MATester’s detection rules are then derived from the same taxonomy and executed on a disjoint set of 12 agents by inspecting instrumented inter-component outputs (snapshot/plan/action). Coverage of open issues (61.4 %) and discovery of 31 new bugs constitute external validation, not a closed loop. No equations equate a fitted quantity to a claimed prediction, no uniqueness theorem is imported from the authors’ prior work, and instrumentation is ordinary dynamic-analysis scaffolding rather than a definitional circularity. The only mild self-reference is the authors’ own labeling and instrumentation, which is transparent and does not force the results by construction. Hence the derivation chain is self-contained against the collected corpus and held-out evaluation.
Axiom & Free-Parameter Ledger
free parameters (2)
- time/round limits for MATester (5 min / 20 rounds) =
5 min / 20 rounds
- judging-LLM temperature and majority-of-3 =
0.2 / majority-of-3
axioms (3)
- domain assumption An M-agent must satisfy R1–R3 (LLM-based, multi-modal perception, goal-directed action).
- domain assumption Closed issues + merged PRs (plus open issues for low-activity repos) within a 15-month window adequately sample real bugs.
- ad hoc to paper Inter-component outputs (snapshot, plan, action, reflect) are sufficient to diagnose both global and component-level symptoms.
invented entities (2)
-
Three-level M-agent bug taxonomy (6 global + 16 component symptoms + 7 root causes)
no independent evidence
-
MATester runtime analyzer
no independent evidence
read the original abstract
Multi-Modal Agents (M-agents), empowered by Large Language Models (LLMs), excel in various complex, open-world scenarios such as autonomous driving and robotics. However, their unique requirements to interact with dynamic and diverse multi-modal environments introduce novel implementation challenges beyond those faced by traditional agents. Outdated perception, untrustworthy planning and inapplicable execution could cause traffic accident and financial loss. Despite growing study on agent issues, there has not been a systematic study focusing on M-agent-specific implementation bugs. To address this gap, we conducted the first systematic study of implementation bugs in M-agents. We collected 34 representative M-agents from diverse sources and, through meticulous filtering,identified 158 M-agent-specific bugs from 1,268 issue reports. Using a top-down strategy, we developed a comprehensive taxonomy that classifies bugs by global symptoms, functionality component-level symptoms, and root causes. We then implemented MATester, an automatic proof-of-concept bug identifier by analyzing runtime inter-component outputs. When applied to 12 extra M-agents, MATester successfully covered 61.4% of known open issues and discovered 31 additional bugs, demonstrating the practical usefulness of our study. Our work provides a comprehensive reference and guideline for classification, prevention and fix of M-agent bugs.
Figures
Reference graph
Works this paper leans on
-
[1]
MARS: toward more efficient multi-agent collaboration for LLM reasoning,
X. Wang, J. Wang, Y. Wang, P. Dang, S. Cao, and C. Zhang, “MARS: toward more efficient multi-agent collaboration for LLM reasoning,” CoRR, vol. abs/2509.20502, 2025
arXiv 2025
-
[2]
Agentic reasoning: A streamlined framework for enhancing LLM reasoning with agentic tools,
J. Wu, J. Zhu, Y. Liu, M. Xu, and Y. Jin, “Agentic reasoning: A streamlined framework for enhancing LLM reasoning with agentic tools,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, E...
2025
-
[3]
Agentless: Demystifying llm-based software engineering agents,
C. S. Xia, Y. Deng, S. Dunn, and L. Zhang, “Agentless: Demystifying llm-based software engineering agents,” CoRR, vol. abs/2407.01489, 2024
Pith/arXiv arXiv 2024
-
[4]
Au- tocoderover: Autonomous program improvement,
Y. Zhang, H. Ruan, Z. Fan, and A. Roychoudhury, “Au- tocoderover: Autonomous program improvement,” in Proceed- ings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2024, Vienna, Austria, September 16-20, 2024, M. Christakis and M. Pradel, Eds. ACM, 2024, pp. 1592–1604
2024
-
[5]
Openhands: An open platform for AI software developers as generalist agents,
X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y. Shao, N. Muennighoff, Y. Zhang, B. Hui, J. Lin, and et al., “Openhands: An open platform for AI software developers as generalist agents,” in The Thirteenth International Conference on Learning Representations, ICLR 2025,...
2025
-
[6]
H. Qiu and Z. Lan, “Interactive agents: Simulating counselor- client psychological counseling via role-playing llm-to-llm inter- actions,” CoRR, vol. abs/2408.15787, 2024
Pith/arXiv arXiv 2024
-
[7]
Gpt-driver: Learning to drive with GPT,
J. Mao, Y. Qian, H. Zhao, and Y. Wang, “Gpt-driver: Learning to drive with GPT,” CoRR, vol. abs/2310.01415, 2023
Pith/arXiv arXiv 2023
-
[8]
Drive like a human: Rethinking autonomous driving with large language models,
D. Fu, X. Li, L. Wen, M. Dou, P. Cai, B. Shi, and Y. Qiao, “Drive like a human: Rethinking autonomous driving with large language models,” in IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, W ACVW 2024 - Workshops, Waikoloa, HI, USA, January 1-6, 2024. IEEE, 2024, pp. 910–919
2024
-
[9]
On the road with gpt-4v(ision): Early explorations of visual-language model on autonomous driving,
L. Wen, X. Yang, D. Fu, X. Wang, P. Cai, X. Li, T. Ma, Y. Li, L. Xu, D. Shang, Z. Zhu, S. Sun, Y. Bai, X. Cai, M. Dou, S. Hu, B. Shi, and Y. Qiao, “On the road with gpt-4v(ision): Early explorations of visual-language model on autonomous driving,” CoRR, vol. abs/2311.05332, 2023
Pith/arXiv arXiv 2023
-
[10]
MP5: A multi-modal open-ended embodied system in minecraft via active perception,
Y. Qin, E. Zhou, Q. Liu, Z. Yin, L. Sheng, R. Zhang, Y. Qiao, and J. Shao, “MP5: A multi-modal open-ended embodied system in minecraft via active perception,” in IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024. IEEE, 2024, pp. 16 307–16 316
2024
-
[11]
JAR VIS-1: open- world multi-task agents with memory-augmented multimodal language models,
Z. Wang, S. Cai, A. Liu, Y. Jin, J. Hou, B. Zhang, H. Lin, Z. He, Z. Zheng, Y. Yang, X. Ma, and Y. Liang, “JAR VIS-1: open- world multi-task agents with memory-augmented multimodal language models,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 3, pp. 1894–1907, 2025
1907
-
[12]
Embodied multi-modal agent trained by an LLM from a parallel textworld,
Y. Yang, T. Zhou, K. Li, D. Tao, L. Li, L. Shen, X. He, J. Jiang, and Y. Shi, “Embodied multi-modal agent trained by an LLM from a parallel textworld,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024. IEEE, 2024, pp. 26 265– 26 275
2024
-
[13]
Z. Wang, S. Cai, A. Liu, X. Ma, and Y. Liang, “Describe, explain, plan and select: Interactive planning with large lan- guage models enables open-world multi-task agents,” CoRR, vol. abs/2302.01560, 2023
Pith/arXiv arXiv 2023
-
[14]
Webwise: Unlocking web interface control for llms via sequential exploration,
H. Tao, S. TV, M. Shlapentokh-Rothman, T. Gupta, H. Ji, and D. Hoiem, “Webwise: Unlocking web interface control for llms via sequential exploration,” in Findings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard, Eds. Association for Computational Linguistics, 2024,...
2024
-
[15]
Empowering LLM to use smartphone for intelligent task automation,
H. Wen, Y. Li, G. Liu, S. Zhao, T. Yu, T. J. Li, S. Jiang, Y. Liu, Y. Zhang, and Y. Liu, “Empowering LLM to use smartphone for intelligent task automation,” CoRR, vol. abs/2308.15272, 2023
Pith/arXiv arXiv 2023
-
[16]
GPT-4V in wonderland: Large multimodal models for zero- shot smartphone GUI navigation,
A. Yan, Z. Yang, W. Zhu, K. Lin, L. Li, J. Wang, J. Yang, Y. Zhong, J. J. McAuley, J. Gao, Z. Liu, and L. Wang, “GPT-4V in wonderland: Large multimodal models for zero- shot smartphone GUI navigation,” CoRR, vol. abs/2311.07562, 2023
Pith/arXiv arXiv 2023
-
[17]
You only look at screens: Multimodal chain-of-action agents,
Z. Zhang and A. Zhang, “You only look at screens: Multimodal chain-of-action agents,” in Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar, Eds. Association for Computational Linguistics, 2024, pp. 3132–3149
2024
-
[18]
Collision between vehicle controlled by developmental automated driving system and pedestrian,
U. N. T. S. Board, “Collision between vehicle controlled by developmental automated driving system and pedestrian,” ww w.ntsb.gov/investigations/AccidentReports/Reports/HAR190 3.pdf, 2019
2019
-
[19]
Microsoft’s ai agent spends 100% of its testing funds on online fraud—lessons for msft and ai-secured transactions
B. News, “Microsoft’s ai agent spends 100% of its testing funds on online fraud—lessons for msft and ai-secured transactions. ” https://blockchain.news/zh/flashnews/microsoft-ai-agents-spe nt-100-of-test-funds-on-online-scams-trading-takeaways-for-m sft-and-ai-security-plays-zh , 2025
2025
-
[20]
Chatgpt agent exposes vulnerability, potentially al- lowing “shadow leak
Netmag, “Chatgpt agent exposes vulnerability, potentially al- lowing “shadow leak”’ to steal sensitive gmail data. ” https:// netmag.tw/2025/09/24/chatgpt-deep-research-gmail-security , 2025
2025
-
[21]
Security debt in llm agent applications: A measurement study of vulnerabilities andmitigation trade-offs,
Z. Shen, J. Dai, Y. Zhang, and M. Yang, “Security debt in llm agent applications: A measurement study of vulnerabilities andmitigation trade-offs,” in 40th IEEE/ACM International Conference on Automated Software Engineering, ASE 2025, Seoul, South Korea, November 16-20, 2025. ACM/IEEE, 2025
2025
-
[22]
Eu- agent-bench: Measuring illegal behavior of LLM agents under EU law,
I. Lichkovski, A. Müller, M. Ibrahim, and T. Mhundwa, “Eu- agent-bench: Measuring illegal behavior of LLM agents under EU law,” CoRR, vol. abs/2510.21524, 2025
arXiv 2025
-
[23]
A. W. Rahardja, J. Liu, W. Chen, Z. Chen, and Y. Lou, “Can agents fix agent issues?” CoRR, vol. abs/2505.20749, 2025
arXiv 2025
-
[24]
Where LLM agents fail and how they can learn from failures,
K. Zhu, Z. Liu, B. Li, M. Tian, Y. Yang, J. Zhang, P. Han, Q. Xie, F. Cui, W. Zhang, X. Ma, X. Yu, G. Ramesh, J. Wu, Z. Liu, P. Lu, J. Zou, and J. You, “Where LLM agents fail and how they can learn from failures,” CoRR, vol. abs/2509.25370, 2025
arXiv 2025
-
[25]
Understanding software engineering agents: A study of thought-action-result trajectories,
I. Bouzenia and M. Pradel, “Understanding software engineering agents: A study of thought-action-result trajectories,” CoRR, vol. abs/2506.18824, 2025
arXiv 2025
-
[26]
Understanding software engineering agents through the lens of traceability: An empirical study,
I. Ceka, S. Pujar, S. Ramji, L. Buratti, G. E. Kaiser, and B. Ray, “Understanding software engineering agents through the lens of traceability: An empirical study,” CoRR, vol. abs/2506.08311, 2025
Pith/arXiv arXiv 2025
-
[27]
Exploring autonomous agents: A closer look at why they fail when completing tasks,
R. Lu, Y. Li, and Y. Huo, “Exploring autonomous agents: A closer look at why they fail when completing tasks,” in 40th IEEE/ACM International Conference on Automated Software Engineering, ASE 2025, Seoul, Korea, Republic of, November 16-20, 2025, 2025, pp. 3856–3860
2025
-
[28]
Safesearch: Automated red-teaming for the safety of llm-based search agents,
J. Dong, S. Guo, H. Wang, Z. Liu, T. Zhang, K. Xu, M. Huang, and H. Qiu, “Safesearch: Automated red-teaming for the safety of llm-based search agents,” CoRR, vol. abs/2509.23694, 2025
Pith/arXiv arXiv 2025
-
[29]
Why do multi-agent LLM systems fail?
M. Cemri, M. Z. Pan, S. Yang, L. A. Agrawal, B. Chopra, R. Tiwari, K. Keutzer, A. G. Parameswaran, D. Klein, K. Ram- chandran, M. Zaharia, J. E. Gonzalez, and I. Stoica, “Why do multi-agent LLM systems fail?” CoRR, vol. abs/2503.13657, 2025
Pith/arXiv arXiv 2025
-
[30]
X. Ma, X. Xie, Y. Wang, J. Wang, B. Wu, M. Li, and Q. Wang, “Diagnosing failure root causes in platform-orchestrated agentic systems: Dataset, taxonomy, and benchmark,” CoRR, vol. abs/2509.23735, 2025
arXiv 2025
-
[31]
Large multi- modal agents: A survey,
J. Xie, Z. Chen, R. Zhang, X. Wan, and G. Li, “Large multi- modal agents: A survey,” CoRR, vol. abs/2402.15116, 2024
Pith/arXiv arXiv 2024
-
[32]
Large language model for verilog code generation: Literature review and the road ahead,
G. Yang, W. Zheng, X. Chen, D. Liang, P. Hu, Y. Yang, S. Peng, Z. Li, J. Feng, X. Wei, K. Sun, D. Ma, H. Cheng, Y. Shen, X. Hu, T. Y. Zhuo, and D. Lo, “Large language model for verilog code generation: Literature review and the road ahead,” CoRR, vol. abs/2512.00020, 2025
arXiv 2025
-
[33]
Identifying relevant studies in software engineering,
H. Zhang, M. A. Babar, and P. Tell, “Identifying relevant studies in software engineering,” Inf. Softw. Technol., vol. 53, no. 6, pp. 625–637, 2011
2011
-
[34]
Guidelines for snowballing in systematic literature studies and a replication in software engineering,
C. Wohlin, “Guidelines for snowballing in systematic literature studies and a replication in software engineering,” in 18th Inter- national Conference on Evaluation and Assessment in Software Engineering, EASE ’14, London, England, United Kingdom, May 13-14, 2014, M. J. Shepperd, T. Hall, and I. Myrtveit, Eds. ACM, 2014, pp. 38:1–38:10
2014
-
[35]
A comprehensive study of autonomous vehicle bugs,
J. Garcia, Y. Feng, J. Shen, S. Almanee, Y. Xia, and Q. A. Chen, “A comprehensive study of autonomous vehicle bugs,” in ICSE ’20: 42nd International Conference on Software Engineering, Seoul, South Korea, 27 June - 19 July, 2020, G. Rothermel and D. Bae, Eds. ACM, 2020, pp. 385–396
2020
-
[36]
Faults in deep reinforcement learning programs: a taxonomy and a detection approach,
A. Nikanjam, M. M. Morovati, F. Khomh, and H. B. Braiek, “Faults in deep reinforcement learning programs: a taxonomy and a detection approach,” Autom. Softw. Eng., vol. 29, no. 1, p. 8, 2022
2022
-
[37]
A comprehensive study of deep learning compiler bugs,
Q. Shen, H. Ma, J. Chen, Y. Tian, S. Cheung, and X. Chen, “A comprehensive study of deep learning compiler bugs,” in ESEC/FSE ’21: 29th ACM Joint European Software Engineer- ing Conference and Symposium on the Foundations of Software Engineering, Athens, Greece, August 23-28, 2021, D. Spinellis, G. Gousios, M. Chechik, and M. D. Penta, Eds. ACM, 2021, pp. 968–980
2021
-
[38]
Catiss: An intelligent tool for categorizing issues re- ports using transformers,
M. Izadi, “Catiss: An intelligent tool for categorizing issues re- ports using transformers,” in 2022 IEEE/ACM 1st International Workshop on Natural Language-Based Software Engineering (NLBSE 2022), Co-located with ICSE 2022, Pittsburgh, PA, USA, May 8, 2022, A. D. Sorbo and S. Panichella, Eds. ACM/IEEE, 2022, pp. 44–47
2022
-
[39]
Analysis and detection of information types of open source software issue discussions,
D. M. Arya, W. Wang, J. L. C. Guo, and J. Cheng, “Analysis and detection of information types of open source software issue discussions,” in Proceedings of the 41st International Conference on Software Engineering, ICSE 2019, Montreal, QC, Canada, May 25-31, 2019, J. M. Atlee, T. Bultan, and J. Whittle, Eds. IEEE / ACM, 2019, pp. 454–464
2019
-
[40]
Bug characteristics in open source software,
L. Tan, C. Liu, Z. Li, X. Wang, Y. Zhou, and C. Zhai, “Bug characteristics in open source software,” Empir. Softw. Eng., vol. 19, no. 6, pp. 1665–1705, 2014
2014
-
[41]
Cohen’s kappa coefficient as a performance measure for feature selection,
S. M. Vieira, U. Kaymak, and J. M. C. Sousa, “Cohen’s kappa coefficient as a performance measure for feature selection,” in FUZZ-IEEE 2010, IEEE International Conference on Fuzzy Sys- tems, Barcelona, Spain, 18-23 July, 2010, Proceedings. IEEE, 2010, pp. 1–8
2010
-
[42]
CCFQA: A benchmark for cross-lingual and cross- modal speech and text factuality evaluation,
Y. Du, K. Liu, Y. Pan, Z. Chu, B. Yang, X. Feng, M. Liu, and Y. Xiang, “CCFQA: A benchmark for cross-lingual and cross- modal speech and text factuality evaluation,” in Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artifici...
2026
-
[43]
Matester
MATester, “Matester. ” https://github.com/MATester0/MAT ester, 2026
2026
-
[44]
Appagent: Multimodal agents as smartphone users,
C. Zhang, Z. Yang, J. Liu, Y. Li, Y. Han, X. Chen, Z. Huang, B. Fu, and G. Yu, “Appagent: Multimodal agents as smartphone users,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, YokohamaJapan, 26 April 2025- 1 May 2025, N. Yamashita, V. Evers, K. Yatani, S. X. Ding, B. Lee, M. Chetty, and P. O. T. Dugas, Eds. ACM...
2025
-
[45]
Droidbot-gpt: Gpt-powered UI automation for android,
H. Wen, H. Wang, J. Liu, and Y. Li, “Droidbot-gpt: Gpt-powered UI automation for android,” CoRR, vol. abs/2304.07061, 2023
Pith/arXiv arXiv 2023
-
[46]
Mobilegpt: Augmenting LLM with human-like app memory for mobile task automation,
S. Lee, J. Choi, J. Lee, M. H. Wasi, H. Choi, S. Y. Ko, S. Oh, and I. Shin, “Mobilegpt: Augmenting LLM with human-like app memory for mobile task automation,” in Proceedings of the 30th Annual International Conference on Mobile Computing and Networking, ACM MobiCom 2024, Washington D.C., DC, USA, November 18-22, 2024, W. Shi, D. Ganesan, and N. D. Lane, E...
2024
-
[47]
ASSISTGUI: task-oriented desktop graphical user interface automation,
D. Gao, L. Ji, Z. Bai, M. Ouyang, P. Li, D. Mao, Q. Wu, W. Zhang, P. Wang, X. Guo, H. Wang, L. Zhou, and M. Z. Shou, “ASSISTGUI: task-oriented desktop graphical user interface automation,” CoRR, vol. abs/2312.13108, 2023
Pith/arXiv arXiv 2023
-
[48]
Browsing like human: A multimodal web agent with experien- tial fast-and-slow thinking,
H. Luo, J. Kuang, W. Liu, Y. Shen, J. Luan, and Y. Deng, “Browsing like human: A multimodal web agent with experien- tial fast-and-slow thinking,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. ...
2025
-
[49]
Octopus: Embodied vision-language programmer from environmental feedback,
J. Yang, Y. Dong, S. Liu, B. Li, Z. Wang, H. Tan, C. Jiang, J. Kang, Y. Zhang, K. Zhou, and Z. Liu, “Octopus: Embodied vision-language programmer from environmental feedback,” in Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part I, ser. Lecture Notes in Computer Science, A. Leonardis, E. ...
2024
-
[50]
See and think: Embodied agent in virtual environment,
Z. Zhao, W. Chai, X. Wang, L. Boyi, S. Hao, S. Cao, T. Ye, and G. Wang, “See and think: Embodied agent in virtual environment,” in Computer Vision - ECCV 2024 - 18th Euro- pean Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part VIII, ser. Lecture Notes in Computer Science, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler,...
2024
-
[51]
Open-world planning via lifted regression with llm-inferred affordances for embodied agents,
X. Liu, A. Pesaranghader, H. Li, P. Sukcharoenchaikul, J. Kim, T. Sadhu, H. Jeon, and S. Sanner, “Open-world planning via lifted regression with llm-inferred affordances for embodied agents,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 20...
2025
-
[52]
Citynavagent: Aerial vision-and- language navigation with hierarchical semantic planning and global memory,
W. Zhang, C. Gao, S. Yu, R. Peng, B. Zhao, Q. Zhang, J. Cui, X. Chen, and Y. Li, “Citynavagent: Aerial vision-and- language navigation with hierarchical semantic planning and global memory,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 202...
2025
-
[53]
RILA: reflective and imaginative language agent for zero-shot semantic audio-visual navigation,
Z. Yang, J. Lin, P. Chen, A. Cherian, T. K. Marks, J. L. Roux, and C. Gan, “RILA: reflective and imaginative language agent for zero-shot semantic audio-visual navigation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024. IEEE, 2024, pp. 16 251–16 261
2024
-
[54]
Llava- interactive: An all-in-one demo for image chat, segmentation, generation and editing,
W. Chen, I. Spiridonova, J. Yang, J. Gao, and C. Li, “Llava- interactive: An all-in-one demo for image chat, segmentation, generation and editing,” CoRR, vol. abs/2311.00571, 2023
Pith/arXiv arXiv 2023
-
[55]
Visual chatgpt: Talking, drawing and editing with visual foundation models,
C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan, “Visual chatgpt: Talking, drawing and editing with visual foundation models,” CoRR, vol. abs/2303.04671, 2023
Pith/arXiv arXiv 2023
-
[56]
Genartist: Multimodal LLM as an agent for unified image generation and editing,
Z. Wang, A. Li, Z. Li, and X. Liu, “Genartist: Multimodal LLM as an agent for unified image generation and editing,” in Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet,...
2024
-
[57]
Loop copilot: Conducting AI ensembles for music generation and iterative editing,
Y. Zhang, A. Maezawa, G. Xia, K. Yamamoto, and S. Dixon, “Loop copilot: Conducting AI ensembles for music generation and iterative editing,” CoRR, vol. abs/2310.12404, 2023
Pith/arXiv arXiv 2023
-
[58]
Musicagent: An AI agent for music understanding and generation with large language models,
D. Yu, K. Song, P. Lu, T. He, X. Tan, W. Ye, S. Zhang, and J. Bian, “Musicagent: An AI agent for music understanding and generation with large language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023 - System Demonstrations, Singapore, December 6-10, 2023, Y. Feng and E. Lefever, Eds. Associat...
2023
-
[59]
Wavjourney: Compositional audio creation with large language models,
X. Liu, Z. Zhu, H. Liu, Y. Yuan, M. Cui, Q. Huang, J. Liang, Y. Cao, Q. Kong, M. D. Plumbley, and W. Wang, “Wavjourney: Compositional audio creation with large language models,” CoRR, vol. abs/2307.14335, 2023
Pith/arXiv arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.