Pith. sign in

REVIEW 2 major objections 5 minor 59 references

The first systematic study of multi-modal agent bugs yields a three-level taxonomy and a runtime checker that finds dozens of real failures.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 10:40 UTC pith:MHTQAS33

load-bearing objection First solid empirical map of multi-modal agent bugs, with a working detector that actually finds new issues on held-out systems. the 2 major comments →

arxiv 2607.04974 v1 pith:MHTQAS33 submitted 2026-07-06 cs.SE

A Comprehensive Study of Implementation Bugs in Multi-modal Agents

classification cs.SE
keywords multi-modal agentsLLM agentsimplementation bugsempirical studybug taxonomyruntime testingperception-planning-execution
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Multi-modal agents that see, plan, and act in the real world fail in ways single-modal chatbots never do: outdated snapshots of a changing scene, plans that ignore what the environment actually allows, and actions that crash when objects move or permissions change. This paper is the first large-scale look at those implementation bugs. From 34 open agents and more than a thousand GitHub issues the authors extract 158 agent-specific bugs, then arrange them into a top-down taxonomy that runs from what an end user sees (crash, misbehavior, silent hang) down through the three internal components (Perceptor, Planner, Executor) to seven root-cause families. The same taxonomy drives a lightweight runtime analyzer, MATester, that watches the messages passed between components. On twelve held-out agents the analyzer recovers most already-reported open issues and surfaces thirty-one previously unknown bugs. The practical claim is that once the distinctive failure modes of multi-modal interaction are named and instrumented, many of them become detectable without deep source-code analysis.

Core claim

A top-down taxonomy of 158 multi-modal-agent bugs—organized by six global symptoms, sixteen component-level symptoms, and seven root causes—together with a runtime inter-component analyzer called MATester, covers 61.4 percent of known open issues and discovers 31 new bugs on twelve held-out agents, showing that the taxonomy is both descriptive of real failures and useful for automated detection.

What carries the argument

The three-level taxonomy (global symptoms → Perceptor/Planner/Executor symptoms → root causes) plus MATester’s comparison of successive environment, snapshot, plan and action traces; the taxonomy supplies the categories, the runtime traces supply the evidence that lets those categories be recognized automatically.

Load-bearing premise

That the 34 open-source agents and the 158 bugs mined from their public issue trackers form a representative sample of how multi-modal agents actually fail; if important industrial or closed designs are missing, both the frequency counts and the coverage numbers shrink.

What would settle it

Apply MATester (or an equivalent inter-component monitor) to a fresh corpus of multi-modal agents drawn from sources deliberately outside the original 34; if coverage of known open issues falls well below 60 percent and almost no new bugs are found, the claimed generality of the taxonomy collapses.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper claims to present the first systematic empirical study of M-agent-specific implementation bugs. From multi-source collection (GitHub, surveys, top venues) it filters 34 representative open-source M-agents, extracts 158 agent-specific bugs from 1,268 issue reports via keyword + dual-author manual review (Cohen’s Kappa rising to 0.84), and induces a three-level top-down taxonomy of 6 global symptoms, 16 functionality-component symptoms (Perceptor/Planner/Executor), and 7 root causes. It further implements MATester, a runtime analyzer of instrumented inter-component outputs (snapshot/plan/action/reflect), which on 12 held-out agents covers 61.4 % of known open issues and surfaces 31 previously unreported bugs, thereby demonstrating both descriptive and practical utility of the taxonomy.

Significance. If the results hold, the work supplies a timely, reusable reference for the rapidly growing M-agent literature: an open bug corpus, a multi-level taxonomy that distinguishes multi-modal interaction failures from ordinary agent bugs, concrete prevention/fix guidelines for developers, and a proof-of-concept detector (MATester) whose held-out performance already yields new bugs. The multi-stage QGS + snowballing + dual-author filtering pipeline, explicit 34/12 train/eval split, and public artifacts strengthen reproducibility. These contributions are of clear value to software-engineering and AI-systems venues concerned with reliability of agents deployed in safety-critical multi-modal settings.

major comments (2)
  1. [Section IV-C / Table VII] Section IV-C and Table VII: the headline claim that MATester “covers 61.4 % of known open issues and discovers 31 additional bugs” rests on a single judging LLM (GPT-4o, temperature 0.2, majority-of-3) and fixed resource bounds (5 min / 20 rounds). No ablation on judge model, temperature, or the time/round limits is reported, nor is inter-judge agreement with human raters beyond the single 97.6 % accuracy figure. Because these free parameters directly affect which symptoms are declared, a short sensitivity study is needed to underwrite the quantitative usefulness claim.
  2. [Section III-B] Section III-B and Threats to Validity: the 34-agent corpus is obtained by successive manual README and code inspections whose inclusion criteria (complete Perceptor–Planner–Executor modules, runnable artifacts) are only partially operationalized. While dual-author agreement is reported, the paper does not quantify how many candidates were discarded at each fine-grained filter step nor whether the retained set systematically under-samples particular architectures (e.g., multi-agent orchestration or closed-loop robotics). A brief characterization of the discarded set would strengthen the representativeness argument that underpins both the frequency tables and the taxonomy’s claimed generality.
minor comments (5)
  1. [Tables IV–VI] Tables IV–VI report both absolute counts and “percentage of M-agents”; with N=34 the percentages are coarse (e.g., 2.9 % = 1 agent). Adding the raw agent counts beside each percentage would improve readability.
  2. [Figure 3] Figure 3’s factor labels are rendered with Unicode control characters that become unreadable in some PDF viewers; replace with plain text or vector labels.
  3. [Section IV-B] Findings 3-1/3-2/3-3 and 5-1/5-2/5-3 restate the same co-occurrence numbers already shown in Figures 4–5; a single consolidated paragraph would reduce redundancy.
  4. [Abstract / Introduction] The abstract and Introduction use the non-ASCII ligature “traffic”; replace with ordinary ASCII “traffic” for indexing and accessibility.
  5. [Section II / IV-C] Section II’s notation (snapshot_i, environment_i, …) is clear, yet the same symbols later appear without subscripts in the MATester description; keep notation consistent.

Circularity Check

0 steps flagged

No significant circularity: taxonomy is bottom-up from independently collected GitHub issues; MATester applies those categories to held-out agents without fitted parameters or self-referential definitions.

full rationale

The paper is a standard empirical software-engineering study. Bugs are extracted from 1,268 public issue reports on 34 filtered open-source M-agents via dual-author manual labeling (Cohen’s Kappa 0.84). The three-level taxonomy (global symptoms, component-level symptoms, root causes) is induced directly from those labeled reports; frequencies and co-occurrence maps are descriptive counts, not predictions. MATester’s detection rules are then derived from the same taxonomy and executed on a disjoint set of 12 agents by inspecting instrumented inter-component outputs (snapshot/plan/action). Coverage of open issues (61.4 %) and discovery of 31 new bugs constitute external validation, not a closed loop. No equations equate a fitted quantity to a claimed prediction, no uniqueness theorem is imported from the authors’ prior work, and instrumentation is ordinary dynamic-analysis scaffolding rather than a definitional circularity. The only mild self-reference is the authors’ own labeling and instrumentation, which is transparent and does not force the results by construction. Hence the derivation chain is self-contained against the collected corpus and held-out evaluation.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 2 invented entities

Empirical software-engineering study; almost no free parameters. Core axioms are standard qualitative-research practices plus domain definitions of what counts as an M-agent. The taxonomy categories and MATester itself are invented constructs whose only evidence is the paper’s own coding and evaluation.

free parameters (2)
  • time/round limits for MATester (5 min / 20 rounds) = 5 min / 20 rounds
    Chosen by reference to human completion times; different cut-offs would change Unrespond and Crash counts.
  • judging-LLM temperature and majority-of-3 = 0.2 / majority-of-3
    Temperature 0.2 and triple query used to stabilize GPT-4o verdicts; alternative settings could alter Misbehave/Mis-signal labels.
axioms (3)
  • domain assumption An M-agent must satisfy R1–R3 (LLM-based, multi-modal perception, goal-directed action).
    Table II; used as the inclusion filter that reduced 86 candidates to 46 runnable agents.
  • domain assumption Closed issues + merged PRs (plus open issues for low-activity repos) within a 15-month window adequately sample real bugs.
    Section III-B2; standard SE mining practice but still an unproved sampling assumption.
  • ad hoc to paper Inter-component outputs (snapshot, plan, action, reflect) are sufficient to diagnose both global and component-level symptoms.
    Core design premise of MATester (Section IV-C); intra-component bugs are acknowledged as out of scope.
invented entities (2)
  • Three-level M-agent bug taxonomy (6 global + 16 component symptoms + 7 root causes) no independent evidence
    purpose: Organise and diagnose the 158 collected bugs and guide MATester rules.
    Categories induced from the authors’ own coding; no external validation corpus exists yet.
  • MATester runtime analyzer no independent evidence
    purpose: Automatically flag global and component symptoms from instrumented inter-component traces.
    Proof-of-concept tool built solely to demonstrate taxonomy utility; independent re-implementation would be needed for external confirmation.

pith-pipeline@v1.1.0-grok45 · 31904 in / 2727 out tokens · 79818 ms · 2026-07-11T10:40:47.689441+00:00 · methodology

0 comments
read the original abstract

Multi-Modal Agents (M-agents), empowered by Large Language Models (LLMs), excel in various complex, open-world scenarios such as autonomous driving and robotics. However, their unique requirements to interact with dynamic and diverse multi-modal environments introduce novel implementation challenges beyond those faced by traditional agents. Outdated perception, untrustworthy planning and inapplicable execution could cause traffic accident and financial loss. Despite growing study on agent issues, there has not been a systematic study focusing on M-agent-specific implementation bugs. To address this gap, we conducted the first systematic study of implementation bugs in M-agents. We collected 34 representative M-agents from diverse sources and, through meticulous filtering,identified 158 M-agent-specific bugs from 1,268 issue reports. Using a top-down strategy, we developed a comprehensive taxonomy that classifies bugs by global symptoms, functionality component-level symptoms, and root causes. We then implemented MATester, an automatic proof-of-concept bug identifier by analyzing runtime inter-component outputs. When applied to 12 extra M-agents, MATester successfully covered 61.4% of known open issues and discovered 31 additional bugs, demonstrating the practical usefulness of our study. Our work provides a comprehensive reference and guideline for classification, prevention and fix of M-agent bugs.

Figures

Figures reproduced from arXiv: 2607.04974 by Chang Yue, Fuman Xie, Guangdong Bai, Kai Chen, Lei Bu, Shangqing Liu, Suwan Li, Yile Wang.

Figure 1
Figure 1. Figure 1: The structure and workflow of M-agents. unknown bugs, which demonstrate the applicability of our taxonomy. II. Multi-modal LLM Agents M-agents are built upon LLMs, leverage external tools to accomplish tasks, and interact with multi-modal en￾vironments. Owing to the advanced reasoning, cross￾modal processing, and tool-use capabilities, they have been deployed in complex real-world scenarios, including auto… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the analysis. TABLE I: Keywords used for searching. Type Large Language Model Multi-Modal Agent Keywords LLM, Large Language Model, Language Model, ChatGPT, AI Multi-modal, Multi modal, Visual, Vision, Audio, Speech Agent, Embodied, Embody, Robot B. Data Collection and Processing To the best of our knowledge, there is currently no comprehensive and up-to-date list of M-agents. Prior to studying… view at source ↗
Figure 4
Figure 4. Figure 4: Relationship between the global level symptom and [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 3
Figure 3. Figure 3: Factors related to Corner-cases and Error-cases distributed by components. TABLE VI: Distribution of various root causes and per￾centage of M-agents with specific root causes. Root. Concu. Corner-. Error-. No perst. mem. LLM’s limit. Bad param. Incor. semat. Num. 1 40 38 2 33 17 27 Perc. 2.9% 26.5% 17.6% 5.9% 29.4% 11.8% 35.9% (35.9%) are also major root causes. Concurrency, which represents the uncertain … view at source ↗
Figure 5
Figure 5. Figure 5: Relationship between the functionality component [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: The workflow of MATester. FINDING 6-3: Missing-A is related to Corner-case gaps due to mishandling of plans with diverse formats and tools with incomplete descriptions. Wrong-A is likely to be caused by Incorrect semantics owing to incorrect parsing from plans to actions. Finally, Inapplicable-A is highly connected to Error￾case gaps and Incorrect semantics, manifesting as ignoring un￾available object acce… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

59 extracted references · 17 linked inside Pith

  1. [1]

    MARS: toward more efficient multi-agent collaboration for LLM reasoning,

    X. Wang, J. Wang, Y. Wang, P. Dang, S. Cao, and C. Zhang, “MARS: toward more efficient multi-agent collaboration for LLM reasoning,” CoRR, vol. abs/2509.20502, 2025

  2. [2]

    Agentic reasoning: A streamlined framework for enhancing LLM reasoning with agentic tools,

    J. Wu, J. Zhu, Y. Liu, M. Xu, and Y. Jin, “Agentic reasoning: A streamlined framework for enhancing LLM reasoning with agentic tools,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, E...

  3. [3]

    Agentless: Demystifying llm-based software engineering agents,

    C. S. Xia, Y. Deng, S. Dunn, and L. Zhang, “Agentless: Demystifying llm-based software engineering agents,” CoRR, vol. abs/2407.01489, 2024

  4. [4]

    Au- tocoderover: Autonomous program improvement,

    Y. Zhang, H. Ruan, Z. Fan, and A. Roychoudhury, “Au- tocoderover: Autonomous program improvement,” in Proceed- ings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2024, Vienna, Austria, September 16-20, 2024, M. Christakis and M. Pradel, Eds. ACM, 2024, pp. 1592–1604

  5. [5]

    Openhands: An open platform for AI software developers as generalist agents,

    X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y. Shao, N. Muennighoff, Y. Zhang, B. Hui, J. Lin, and et al., “Openhands: An open platform for AI software developers as generalist agents,” in The Thirteenth International Conference on Learning Representations, ICLR 2025,...

  6. [6]

    Interactive agents: Simulating counselor- client psychological counseling via role-playing llm-to-llm inter- actions,

    H. Qiu and Z. Lan, “Interactive agents: Simulating counselor- client psychological counseling via role-playing llm-to-llm inter- actions,” CoRR, vol. abs/2408.15787, 2024

  7. [7]

    Gpt-driver: Learning to drive with GPT,

    J. Mao, Y. Qian, H. Zhao, and Y. Wang, “Gpt-driver: Learning to drive with GPT,” CoRR, vol. abs/2310.01415, 2023

  8. [8]

    Drive like a human: Rethinking autonomous driving with large language models,

    D. Fu, X. Li, L. Wen, M. Dou, P. Cai, B. Shi, and Y. Qiao, “Drive like a human: Rethinking autonomous driving with large language models,” in IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, W ACVW 2024 - Workshops, Waikoloa, HI, USA, January 1-6, 2024. IEEE, 2024, pp. 910–919

  9. [9]

    On the road with gpt-4v(ision): Early explorations of visual-language model on autonomous driving,

    L. Wen, X. Yang, D. Fu, X. Wang, P. Cai, X. Li, T. Ma, Y. Li, L. Xu, D. Shang, Z. Zhu, S. Sun, Y. Bai, X. Cai, M. Dou, S. Hu, B. Shi, and Y. Qiao, “On the road with gpt-4v(ision): Early explorations of visual-language model on autonomous driving,” CoRR, vol. abs/2311.05332, 2023

  10. [10]

    MP5: A multi-modal open-ended embodied system in minecraft via active perception,

    Y. Qin, E. Zhou, Q. Liu, Z. Yin, L. Sheng, R. Zhang, Y. Qiao, and J. Shao, “MP5: A multi-modal open-ended embodied system in minecraft via active perception,” in IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024. IEEE, 2024, pp. 16 307–16 316

  11. [11]

    JAR VIS-1: open- world multi-task agents with memory-augmented multimodal language models,

    Z. Wang, S. Cai, A. Liu, Y. Jin, J. Hou, B. Zhang, H. Lin, Z. He, Z. Zheng, Y. Yang, X. Ma, and Y. Liang, “JAR VIS-1: open- world multi-task agents with memory-augmented multimodal language models,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 3, pp. 1894–1907, 2025

  12. [12]

    Embodied multi-modal agent trained by an LLM from a parallel textworld,

    Y. Yang, T. Zhou, K. Li, D. Tao, L. Li, L. Shen, X. He, J. Jiang, and Y. Shi, “Embodied multi-modal agent trained by an LLM from a parallel textworld,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024. IEEE, 2024, pp. 26 265– 26 275

  13. [13]

    Describe, explain, plan and select: Interactive planning with large lan- guage models enables open-world multi-task agents,

    Z. Wang, S. Cai, A. Liu, X. Ma, and Y. Liang, “Describe, explain, plan and select: Interactive planning with large lan- guage models enables open-world multi-task agents,” CoRR, vol. abs/2302.01560, 2023

  14. [14]

    Webwise: Unlocking web interface control for llms via sequential exploration,

    H. Tao, S. TV, M. Shlapentokh-Rothman, T. Gupta, H. Ji, and D. Hoiem, “Webwise: Unlocking web interface control for llms via sequential exploration,” in Findings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard, Eds. Association for Computational Linguistics, 2024,...

  15. [15]

    Empowering LLM to use smartphone for intelligent task automation,

    H. Wen, Y. Li, G. Liu, S. Zhao, T. Yu, T. J. Li, S. Jiang, Y. Liu, Y. Zhang, and Y. Liu, “Empowering LLM to use smartphone for intelligent task automation,” CoRR, vol. abs/2308.15272, 2023

  16. [16]

    GPT-4V in wonderland: Large multimodal models for zero- shot smartphone GUI navigation,

    A. Yan, Z. Yang, W. Zhu, K. Lin, L. Li, J. Wang, J. Yang, Y. Zhong, J. J. McAuley, J. Gao, Z. Liu, and L. Wang, “GPT-4V in wonderland: Large multimodal models for zero- shot smartphone GUI navigation,” CoRR, vol. abs/2311.07562, 2023

  17. [17]

    You only look at screens: Multimodal chain-of-action agents,

    Z. Zhang and A. Zhang, “You only look at screens: Multimodal chain-of-action agents,” in Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar, Eds. Association for Computational Linguistics, 2024, pp. 3132–3149

  18. [18]

    Collision between vehicle controlled by developmental automated driving system and pedestrian,

    U. N. T. S. Board, “Collision between vehicle controlled by developmental automated driving system and pedestrian,” ww w.ntsb.gov/investigations/AccidentReports/Reports/HAR190 3.pdf, 2019

  19. [19]

    Microsoft’s ai agent spends 100% of its testing funds on online fraud—lessons for msft and ai-secured transactions

    B. News, “Microsoft’s ai agent spends 100% of its testing funds on online fraud—lessons for msft and ai-secured transactions. ” https://blockchain.news/zh/flashnews/microsoft-ai-agents-spe nt-100-of-test-funds-on-online-scams-trading-takeaways-for-m sft-and-ai-security-plays-zh , 2025

  20. [20]

    Chatgpt agent exposes vulnerability, potentially al- lowing “shadow leak

    Netmag, “Chatgpt agent exposes vulnerability, potentially al- lowing “shadow leak”’ to steal sensitive gmail data. ” https:// netmag.tw/2025/09/24/chatgpt-deep-research-gmail-security , 2025

  21. [21]

    Security debt in llm agent applications: A measurement study of vulnerabilities andmitigation trade-offs,

    Z. Shen, J. Dai, Y. Zhang, and M. Yang, “Security debt in llm agent applications: A measurement study of vulnerabilities andmitigation trade-offs,” in 40th IEEE/ACM International Conference on Automated Software Engineering, ASE 2025, Seoul, South Korea, November 16-20, 2025. ACM/IEEE, 2025

  22. [22]

    Eu- agent-bench: Measuring illegal behavior of LLM agents under EU law,

    I. Lichkovski, A. Müller, M. Ibrahim, and T. Mhundwa, “Eu- agent-bench: Measuring illegal behavior of LLM agents under EU law,” CoRR, vol. abs/2510.21524, 2025

  23. [23]

    Can agents fix agent issues?

    A. W. Rahardja, J. Liu, W. Chen, Z. Chen, and Y. Lou, “Can agents fix agent issues?” CoRR, vol. abs/2505.20749, 2025

  24. [24]

    Where LLM agents fail and how they can learn from failures,

    K. Zhu, Z. Liu, B. Li, M. Tian, Y. Yang, J. Zhang, P. Han, Q. Xie, F. Cui, W. Zhang, X. Ma, X. Yu, G. Ramesh, J. Wu, Z. Liu, P. Lu, J. Zou, and J. You, “Where LLM agents fail and how they can learn from failures,” CoRR, vol. abs/2509.25370, 2025

  25. [25]

    Understanding software engineering agents: A study of thought-action-result trajectories,

    I. Bouzenia and M. Pradel, “Understanding software engineering agents: A study of thought-action-result trajectories,” CoRR, vol. abs/2506.18824, 2025

  26. [26]

    Understanding software engineering agents through the lens of traceability: An empirical study,

    I. Ceka, S. Pujar, S. Ramji, L. Buratti, G. E. Kaiser, and B. Ray, “Understanding software engineering agents through the lens of traceability: An empirical study,” CoRR, vol. abs/2506.08311, 2025

  27. [27]

    Exploring autonomous agents: A closer look at why they fail when completing tasks,

    R. Lu, Y. Li, and Y. Huo, “Exploring autonomous agents: A closer look at why they fail when completing tasks,” in 40th IEEE/ACM International Conference on Automated Software Engineering, ASE 2025, Seoul, Korea, Republic of, November 16-20, 2025, 2025, pp. 3856–3860

  28. [28]

    Safesearch: Automated red-teaming for the safety of llm-based search agents,

    J. Dong, S. Guo, H. Wang, Z. Liu, T. Zhang, K. Xu, M. Huang, and H. Qiu, “Safesearch: Automated red-teaming for the safety of llm-based search agents,” CoRR, vol. abs/2509.23694, 2025

  29. [29]

    Why do multi-agent LLM systems fail?

    M. Cemri, M. Z. Pan, S. Yang, L. A. Agrawal, B. Chopra, R. Tiwari, K. Keutzer, A. G. Parameswaran, D. Klein, K. Ram- chandran, M. Zaharia, J. E. Gonzalez, and I. Stoica, “Why do multi-agent LLM systems fail?” CoRR, vol. abs/2503.13657, 2025

  30. [30]

    Diagnosing failure root causes in platform-orchestrated agentic systems: Dataset, taxonomy, and benchmark,

    X. Ma, X. Xie, Y. Wang, J. Wang, B. Wu, M. Li, and Q. Wang, “Diagnosing failure root causes in platform-orchestrated agentic systems: Dataset, taxonomy, and benchmark,” CoRR, vol. abs/2509.23735, 2025

  31. [31]

    Large multi- modal agents: A survey,

    J. Xie, Z. Chen, R. Zhang, X. Wan, and G. Li, “Large multi- modal agents: A survey,” CoRR, vol. abs/2402.15116, 2024

  32. [32]

    Large language model for verilog code generation: Literature review and the road ahead,

    G. Yang, W. Zheng, X. Chen, D. Liang, P. Hu, Y. Yang, S. Peng, Z. Li, J. Feng, X. Wei, K. Sun, D. Ma, H. Cheng, Y. Shen, X. Hu, T. Y. Zhuo, and D. Lo, “Large language model for verilog code generation: Literature review and the road ahead,” CoRR, vol. abs/2512.00020, 2025

  33. [33]

    Identifying relevant studies in software engineering,

    H. Zhang, M. A. Babar, and P. Tell, “Identifying relevant studies in software engineering,” Inf. Softw. Technol., vol. 53, no. 6, pp. 625–637, 2011

  34. [34]

    Guidelines for snowballing in systematic literature studies and a replication in software engineering,

    C. Wohlin, “Guidelines for snowballing in systematic literature studies and a replication in software engineering,” in 18th Inter- national Conference on Evaluation and Assessment in Software Engineering, EASE ’14, London, England, United Kingdom, May 13-14, 2014, M. J. Shepperd, T. Hall, and I. Myrtveit, Eds. ACM, 2014, pp. 38:1–38:10

  35. [35]

    A comprehensive study of autonomous vehicle bugs,

    J. Garcia, Y. Feng, J. Shen, S. Almanee, Y. Xia, and Q. A. Chen, “A comprehensive study of autonomous vehicle bugs,” in ICSE ’20: 42nd International Conference on Software Engineering, Seoul, South Korea, 27 June - 19 July, 2020, G. Rothermel and D. Bae, Eds. ACM, 2020, pp. 385–396

  36. [36]

    Faults in deep reinforcement learning programs: a taxonomy and a detection approach,

    A. Nikanjam, M. M. Morovati, F. Khomh, and H. B. Braiek, “Faults in deep reinforcement learning programs: a taxonomy and a detection approach,” Autom. Softw. Eng., vol. 29, no. 1, p. 8, 2022

  37. [37]

    A comprehensive study of deep learning compiler bugs,

    Q. Shen, H. Ma, J. Chen, Y. Tian, S. Cheung, and X. Chen, “A comprehensive study of deep learning compiler bugs,” in ESEC/FSE ’21: 29th ACM Joint European Software Engineer- ing Conference and Symposium on the Foundations of Software Engineering, Athens, Greece, August 23-28, 2021, D. Spinellis, G. Gousios, M. Chechik, and M. D. Penta, Eds. ACM, 2021, pp. 968–980

  38. [38]

    Catiss: An intelligent tool for categorizing issues re- ports using transformers,

    M. Izadi, “Catiss: An intelligent tool for categorizing issues re- ports using transformers,” in 2022 IEEE/ACM 1st International Workshop on Natural Language-Based Software Engineering (NLBSE 2022), Co-located with ICSE 2022, Pittsburgh, PA, USA, May 8, 2022, A. D. Sorbo and S. Panichella, Eds. ACM/IEEE, 2022, pp. 44–47

  39. [39]

    Analysis and detection of information types of open source software issue discussions,

    D. M. Arya, W. Wang, J. L. C. Guo, and J. Cheng, “Analysis and detection of information types of open source software issue discussions,” in Proceedings of the 41st International Conference on Software Engineering, ICSE 2019, Montreal, QC, Canada, May 25-31, 2019, J. M. Atlee, T. Bultan, and J. Whittle, Eds. IEEE / ACM, 2019, pp. 454–464

  40. [40]

    Bug characteristics in open source software,

    L. Tan, C. Liu, Z. Li, X. Wang, Y. Zhou, and C. Zhai, “Bug characteristics in open source software,” Empir. Softw. Eng., vol. 19, no. 6, pp. 1665–1705, 2014

  41. [41]

    Cohen’s kappa coefficient as a performance measure for feature selection,

    S. M. Vieira, U. Kaymak, and J. M. C. Sousa, “Cohen’s kappa coefficient as a performance measure for feature selection,” in FUZZ-IEEE 2010, IEEE International Conference on Fuzzy Sys- tems, Barcelona, Spain, 18-23 July, 2010, Proceedings. IEEE, 2010, pp. 1–8

  42. [42]

    CCFQA: A benchmark for cross-lingual and cross- modal speech and text factuality evaluation,

    Y. Du, K. Liu, Y. Pan, Z. Chu, B. Yang, X. Feng, M. Liu, and Y. Xiang, “CCFQA: A benchmark for cross-lingual and cross- modal speech and text factuality evaluation,” in Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artifici...

  43. [43]

    Matester

    MATester, “Matester. ” https://github.com/MATester0/MAT ester, 2026

  44. [44]

    Appagent: Multimodal agents as smartphone users,

    C. Zhang, Z. Yang, J. Liu, Y. Li, Y. Han, X. Chen, Z. Huang, B. Fu, and G. Yu, “Appagent: Multimodal agents as smartphone users,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, YokohamaJapan, 26 April 2025- 1 May 2025, N. Yamashita, V. Evers, K. Yatani, S. X. Ding, B. Lee, M. Chetty, and P. O. T. Dugas, Eds. ACM...

  45. [45]

    Droidbot-gpt: Gpt-powered UI automation for android,

    H. Wen, H. Wang, J. Liu, and Y. Li, “Droidbot-gpt: Gpt-powered UI automation for android,” CoRR, vol. abs/2304.07061, 2023

  46. [46]

    Mobilegpt: Augmenting LLM with human-like app memory for mobile task automation,

    S. Lee, J. Choi, J. Lee, M. H. Wasi, H. Choi, S. Y. Ko, S. Oh, and I. Shin, “Mobilegpt: Augmenting LLM with human-like app memory for mobile task automation,” in Proceedings of the 30th Annual International Conference on Mobile Computing and Networking, ACM MobiCom 2024, Washington D.C., DC, USA, November 18-22, 2024, W. Shi, D. Ganesan, and N. D. Lane, E...

  47. [47]

    ASSISTGUI: task-oriented desktop graphical user interface automation,

    D. Gao, L. Ji, Z. Bai, M. Ouyang, P. Li, D. Mao, Q. Wu, W. Zhang, P. Wang, X. Guo, H. Wang, L. Zhou, and M. Z. Shou, “ASSISTGUI: task-oriented desktop graphical user interface automation,” CoRR, vol. abs/2312.13108, 2023

  48. [48]

    Browsing like human: A multimodal web agent with experien- tial fast-and-slow thinking,

    H. Luo, J. Kuang, W. Liu, Y. Shen, J. Luan, and Y. Deng, “Browsing like human: A multimodal web agent with experien- tial fast-and-slow thinking,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. ...

  49. [49]

    Octopus: Embodied vision-language programmer from environmental feedback,

    J. Yang, Y. Dong, S. Liu, B. Li, Z. Wang, H. Tan, C. Jiang, J. Kang, Y. Zhang, K. Zhou, and Z. Liu, “Octopus: Embodied vision-language programmer from environmental feedback,” in Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part I, ser. Lecture Notes in Computer Science, A. Leonardis, E. ...

  50. [50]

    See and think: Embodied agent in virtual environment,

    Z. Zhao, W. Chai, X. Wang, L. Boyi, S. Hao, S. Cao, T. Ye, and G. Wang, “See and think: Embodied agent in virtual environment,” in Computer Vision - ECCV 2024 - 18th Euro- pean Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part VIII, ser. Lecture Notes in Computer Science, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler,...

  51. [51]

    Open-world planning via lifted regression with llm-inferred affordances for embodied agents,

    X. Liu, A. Pesaranghader, H. Li, P. Sukcharoenchaikul, J. Kim, T. Sadhu, H. Jeon, and S. Sanner, “Open-world planning via lifted regression with llm-inferred affordances for embodied agents,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 20...

  52. [52]

    Citynavagent: Aerial vision-and- language navigation with hierarchical semantic planning and global memory,

    W. Zhang, C. Gao, S. Yu, R. Peng, B. Zhao, Q. Zhang, J. Cui, X. Chen, and Y. Li, “Citynavagent: Aerial vision-and- language navigation with hierarchical semantic planning and global memory,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 202...

  53. [53]

    RILA: reflective and imaginative language agent for zero-shot semantic audio-visual navigation,

    Z. Yang, J. Lin, P. Chen, A. Cherian, T. K. Marks, J. L. Roux, and C. Gan, “RILA: reflective and imaginative language agent for zero-shot semantic audio-visual navigation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024. IEEE, 2024, pp. 16 251–16 261

  54. [54]

    Llava- interactive: An all-in-one demo for image chat, segmentation, generation and editing,

    W. Chen, I. Spiridonova, J. Yang, J. Gao, and C. Li, “Llava- interactive: An all-in-one demo for image chat, segmentation, generation and editing,” CoRR, vol. abs/2311.00571, 2023

  55. [55]

    Visual chatgpt: Talking, drawing and editing with visual foundation models,

    C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan, “Visual chatgpt: Talking, drawing and editing with visual foundation models,” CoRR, vol. abs/2303.04671, 2023

  56. [56]

    Genartist: Multimodal LLM as an agent for unified image generation and editing,

    Z. Wang, A. Li, Z. Li, and X. Liu, “Genartist: Multimodal LLM as an agent for unified image generation and editing,” in Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet,...

  57. [57]

    Loop copilot: Conducting AI ensembles for music generation and iterative editing,

    Y. Zhang, A. Maezawa, G. Xia, K. Yamamoto, and S. Dixon, “Loop copilot: Conducting AI ensembles for music generation and iterative editing,” CoRR, vol. abs/2310.12404, 2023

  58. [58]

    Musicagent: An AI agent for music understanding and generation with large language models,

    D. Yu, K. Song, P. Lu, T. He, X. Tan, W. Ye, S. Zhang, and J. Bian, “Musicagent: An AI agent for music understanding and generation with large language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023 - System Demonstrations, Singapore, December 6-10, 2023, Y. Feng and E. Lefever, Eds. Associat...

  59. [59]

    Wavjourney: Compositional audio creation with large language models,

    X. Liu, Z. Zhu, H. Liu, Y. Yuan, M. Cui, Q. Huang, J. Liang, Y. Cao, Q. Kong, M. D. Plumbley, and W. Wang, “Wavjourney: Compositional audio creation with large language models,” CoRR, vol. abs/2307.14335, 2023