Pith. sign in

REVIEW 4 major objections 5 minor 46 references

SAGE claims that gray-box game regression testing can be automated end-to-end by chaining LLM-guided exploration, Pareto test selection, and changelog-based prioritization.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 19:24 UTC pith:E4FNCQM6

load-bearing objection SAGE is a genuinely integrated LLM+RL gray-box game regression pipeline with real promise, but its headline step-reduction claim is contradicted by its own tables. the 4 major comments →

arxiv 2512.00560 v2 pith:E4FNCQM6 submitted 2025-11-29 cs.SE

SAGE: Semantic-Aware Gray-Box Game Regression Testing with Large Language Models

classification cs.SE
keywords game regression testinggray-box testinglarge language modelsreinforcement learningtest case selectionmulti-objective optimizationtest case prioritizationstate-action transition graph
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

SAGE is trying to establish that regression testing of live-service games need not wait for white-box access or manual test writing. The paper's central claim is that in a gray-box setting — where testers have runtime logs and changelogs but no source code — an LLM can act as a semantic orchestrator: it seeds goal-directed exploration, guides an RL agent to build a state–action graph, prunes that graph by balancing cost, coverage, and behavioral rarity, and finally reads natural-language update logs to decide which tests to run first. SAGE reports that this loop detects close to the same number of distinct bugs as human-recorded tests in two very different game genres, roughly 1.6x more distinct bugs than automated RL baselines, while completing regression cycles in about 10–40% of the interaction steps and 75–90% less wall-clock time. If true, it would make automated regression testing practical in the common industrial situation where testing teams are separate from development teams and source code is closed.

Core claim

The paper argues that the three obstacles of gray-box regression testing — building an initial test suite, keeping it compact, and choosing what to rerun after an update — are one problem solvable with semantics as the bridge. SAGE uses an LLM to generate seed trajectories for a task, behavior-clones them to initialize an RL agent that explores with a novelty bonus, and accumulates all observed transitions into a state–action graph. Candidate tests (paths from start to goal) are then encoded as multi-dimensional vectors and reduced to the Pareto front over cost, coverage, and a newly defined n-gram rarity that favors short unusual action sequences likely to trigger edge-case bugs. Finally, a

What carries the argument

The central object is the semantically annotated state–action transition graph G=(S,A,E). Every edge is a transition s→s′ under action a, carrying metadata from runtime logs: estimated execution time, objects touched, scenes, and UI components. Test cases are paths from an initial state to a goal state. Three mechanisms hang off this graph: (1) LLM-generated seed trajectories, cloned to give the RL policy a goal-directed prior; (2) multi-objective Pareto selection over cost (length, time), coverage (states, actions, objects, scenes, UI), and n-gram rarity — the inverse frequency of short action subsequences, which is the paper's main new metric; and (3) an LLM that turns update-log sentences

Load-bearing premise

SAGE's whole pipeline depends on the gray-box environment exposing rich, structured, pre-defined debugging logs — actions, object states, scenes, UI components — from which all cost, coverage, rarity, and semantic metrics can be computed; the paper says this in §3.2.3 and concedes in §4.6 that in more restricted gray-box environments such detailed metrics are rarely available.

What would settle it

Run SAGE on a released game whose only interface is raw pixels and controller input (no structured logs), keeping the LLM and optimization stages unchanged; if unique-bug diversity falls to baseline levels, the log-dependency is confirmed as the load-bearing premise. A lighter test: shuffle the update-log tags used by the prioritization stage and check whether the early-episode bug-discovery advantage disappears.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • With only runtime logs and changelogs as input, SAGE's generated suite achieves 37.5 and 57.7 unique bugs in the two Overcooked Plus regression phases, against 40.3 and 65.7 for human-recorded tests — near-human diversity without code access.
  • In Minecraft, the top-50% RTS configuration yields 39 and 42 unique bugs versus 41.3 and 45.4 human, and about 1.6x more than the automated RL baselines.
  • Execution cost drops to 21–69K steps and 6–16 seconds in the cooking game, versus 50–185K and 14–69 seconds for baselines; in the sandbox game, to 66–189K steps and roughly 134 seconds versus 120–240K and 269–778 seconds.
  • Taking only the top 10–30% of prioritized paths captures most of the bug diversity at a fraction of the cost, so the framework supports flexible trade-offs between testing window and coverage.
  • Ablations show both modules matter: removing optimization leads to rapid saturation of unique-bug coverage, while removing update-aware prioritization slows early bug discovery; combining them gives the best efficiency for a fixed budget.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the n-gram rarity objective is not game-specific; any black-box system where bugs hide in short unusual action sequences — web automation, robotics controllers, simulators — could reuse the same Pareto-selection idea.
  • Editorial extension: the framework's reliance on rich debugging logs means it is best suited to games instrumented during development. If applied to an off-the-shelf game, the log-dependency would have to be replaced by screen-differencing state inference, which the paper mentions as future work but does not test.
  • Editorial extension: a severity-weighted variant of the Pareto front would likely change the selected suite; the paper measures bug counts and diversity only, and notes that commercial teams may prioritize critical bugs. This is a natural next experiment rather than a refutation.
  • Editorial extension: the paper's own ablation suggests the optimizer can prune long-tail behaviors; an adaptive pruning strategy that preserves rare actions while still removing redundancy would be a falsifiable improvement.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes SAGE, a gray-box regression testing framework for games that combines LLM-guided RL exploration to generate test trajectories, a multi-objective Pareto optimization over cost/coverage/rarity to select a compact test suite, and an LLM-based update-log interpreter to prioritize test cases for a specific version update. The framework is evaluated on two custom environments (Overcooked Plus and a Minecraft mod) against Random, PPO, diff-Qlearning, and human-recorded tests, and is reported to achieve near-human unique-bug diversity while substantially reducing execution steps and wall-clock duration. Ablation studies are used to attribute gains to the optimization and prioritization modules.

Significance. If the reported results are accurate, SAGE would be a useful contribution to game regression testing in gray-box settings, an area where the paper correctly identifies a gap between white-box RTS methods and purely black-box exploration. The paper deserves credit for including human-recorded test cases as a baseline, for running multiple seeds, for open-sourcing the Overcooked Plus replication suite, and for candidly acknowledging limitations in §4.6 (restricted gray-box logging) and §4.7 (custom testbeds, severity-agnostic metrics). However, the empirical claims currently contain internal inconsistencies and unspecified configuration parameters that must be resolved before the central claims can be accepted.

major comments (4)
  1. [§4.4 (Summary of RQ2) and Table 2] The headline efficiency claim is not supported at the default configuration. The RQ2 summary states SAGE "us[es] only about 10–40% of the interaction steps required by baseline methods," and the abstract says "significantly lower execution cost." In Minecraft V2→V3, SAGE(top 50%) executes 188,592±14,826 steps versus 240,000 for PPO/diff-Qlearning — 78.6% of baseline, i.e., a 21.4% reduction. At top 90%, SAGE executes 250,517 steps, more than the baseline. The 60–90% step-reduction range is only attained at small RTS proportions in Overcooked; it is not a general property. The abstract and RQ2 summary must be qualified by environment and RTS proportion, or the step metric should be replaced by wall-clock duration, which is the metric where SAGE shows a robust advantage.
  2. [§3.4.2, Eq. (4); §3.3.2 vs. §4.1.3] Two load-bearing configuration parameters are never disclosed. (1) The balance weight λ in Eq. (4) is defined as λ∈[0,1] but no value or grid is reported; the prioritization results in RQ3 depend on this choice. (2) The n-gram length for the rarity objective is specified as "3-grams" in §3.3.2 but as "2-gram subsequences" in §4.1.3. Rarity is one of the three Pareto objectives, so changing the gram length changes the selected suite. The paper also never quantifies the path-search depth and maximum path count mentioned in §4.6. Without these values, the experiments are not reproducible and the Pareto front cannot be reconstructed.
  3. [§4.1.3 (Comparison Methods)] The baseline comparison is potentially confounded by reward design. SAGE is described as using "the same PPO configuration as above but with the task-specific reward design introduced below," while the reward function used to train the PPO baseline is never explicitly specified. If PPO is trained with a different reward (e.g., without the +1000/+10000 goal bonuses in Overcooked Plus), the 1.9× unique-bug advantage claimed for SAGE could be due to reward shaping rather than to the LLM-guided semantic mechanisms. The paper should state the reward function for every learning baseline and, ideally, run SAGE's PPO with the baseline rewards as an additional ablation.
  4. [Table 2 (Minecraft duration column)] The duration values for SAGE are identical between V1→V2 and V2→V3 at every RTS proportion (53.8, 94.2, 134.6, 161.5, 182.7 seconds), despite step counts that differ by roughly 3× at the same proportions (e.g., 66,094 vs. 188,592 steps at top 50%). The same wall-clock duration across threefold different interaction counts is implausible and suggests a data-entry error. Because §4.4 uses these durations to claim a 75–90% duration reduction, the table must be corrected and the statistics recomputed.
minor comments (5)
  1. [§3.3.2 / §4.1.3] Use a single n-gram length for the rarity metric, or explicitly state that both values were tested and report sensitivity.
  2. [Eq. (4)] Report the λ value or a sensitivity analysis over λ∈[0,1]; otherwise the prioritization is unfalsifiable.
  3. [Table 2] Verify the duration column; duplicated values across version transitions should be corrected.
  4. [§4.2 / Tables 1–2] State which RTS proportion is used in each row or figure; the paper sometimes refers to "SAGE" without the proportion, and the default is defined only in §4.1.3.
  5. [§4.7.1] The paper says results are reported with "standard deviations or confidence intervals" but Tables 1–2 use standard deviations; clarify the convention.

Circularity Check

0 steps flagged

No significant circularity: SAGE's empirical results are not definitionally entailed by its objectives; the only author-overlapping citation is the self-built Overcooked Plus testbed, which is not load-bearing for the method's derivation.

full rationale

SAGE is an empirical pipeline, not a formal derivation. Test cases are produced by LLM-seeded behavior cloning plus RL exploration; they are then filtered by Pareto optimization over cost, coverage, and n-gram rarity, and finally prioritized by LLM-extracted update tags with cosine similarity and a semantic complexity score. None of these quantities is fitted to, or defined in terms of, the reported bug counts or unique-bug diversity. The embedded bugs are independent, pre-defined triggers in the environments, so 'detecting' them is an external outcome rather than an algebraic consequence of the framework's own equations. The main self-citation is the Overcooked Plus testbed [41], used as an evaluation environment; it supports the experimental setting but does not force the method's outcomes. The paper itself flags the relevant limitations in Section 4.6 (gray-box debugging logs can be unavailable in more restricted environments) and Section 4.7.1 (evaluation relies on bug triggers recorded during environment development and on custom testbeds), which are external-validity concerns, not circularity. The RQ2 summary's '10–40% of interaction steps' statement is not consistently supported by Tables 1 and 2 (e.g., Minecraft V2→V3 at top-50% uses 188,592 vs. 240,000 baseline steps, about 79%), but that is an internal numerical inconsistency and scope-of-claim issue, not a circular reduction of a predicted quantity to its inputs.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 0 invented entities

The framework's practical validity rests on the availability of rich runtime logs and on unvalidated semantic-correlation assumptions; the evaluation additionally introduces several hand-set parameters and does not disclose λ.

free parameters (6)
  • λ (balance weight)
    Balances semantic similarity and complexity in Score(t_j) (Eq. in §3.4.2); value never reported in experiments.
  • n-gram length for rarity = 2 (evaluation) / 3 (method section)
    Inconsistency between §3.3.2 (3-grams) and §4.1.3 (2-grams); the metric's behavior depends on this choice.
  • reward shaping weights = Overcooked: +1000/+10000/-0.1; Minecraft: -0.5/+10/+100
    Hand-designed rewards guide RL exploration; not learned from data.
  • path search constraints
    Depth/breadth limits for graph path extraction are imposed but not specified; mentioned as mitigation in §4.6.
  • RTS proportion (default) = 50%
    Default configuration for main discussion; results vary across 10-90% proportions.
  • seed trajectory count = 20 per task
    Number of LLM-generated seeds for behavior cloning; not swept for sensitivity.
axioms (5)
  • domain assumption Gray-box environments expose rich, structured, pre-defined debugging runtime logs (actions, objects, scenes, UI components, states).
    Invoked in §3.2.3 and §3.3.1 to define objectives and semantic metadata; without it, the method cannot be applied.
  • domain assumption Semantic similarity between update-log tags and test-case metadata is predictive of regression-relevant test cases.
    Central to §3.4.2 prioritization; not validated in the paper.
  • domain assumption LLM-generated seed trajectories are behaviorally plausible and diverse enough to serve as a behavior-cloning prior for RL.
    Used in §3.2.1/3.2.2; no analysis of seed quality or diversity is provided.
  • domain assumption The planted bug set and reproducible triggers reflect a realistic distribution of game regressions.
    Acknowledged in §4.7.1; the RQ1 claim depends on the representativeness of these self-constructed bugs.
  • standard math Pareto-optimal path selection yields an effective regression suite.
    Adopted in §3.3.3; assumes the objectives capture test value and that dominance is a sound selection criterion.

pith-pipeline@v1.3.0-alltime-deepseek · 22482 in / 16930 out tokens · 155739 ms · 2026-08-03T19:24:16.645431+00:00 · methodology

0 comments
read the original abstract

The rapid iteration cycles of modern live-service games make regression testing indispensable for maintaining quality and stability. However, existing regression testing approaches face critical limitations, especially in common gray-box settings where full source code access is unavailable: they heavily rely on manual effort for test case construction, struggle to maintain growing suites plagued by redundancy, and lack efficient mechanisms for prioritizing relevant tests. These challenges result in excessive testing costs, limited automation, and insufficient bug detection. To address these issues, we propose SAGE, a semanticaware regression testing framework for gray-box game environments. SAGE systematically addresses the core challenges of test generation, maintenance, and selection. It employs LLM-guided reinforcement learning for efficient, goal-oriented exploration to automatically generate a diverse foundational test suite. Subsequently, it applies a semantic-based multi-objective optimization to refine this suite into a compact, high-value subset by balancing cost, coverage, and rarity. Finally, it leverages LLM-based semantic analysis of update logs to prioritize test cases most relevant to version changes, enabling efficient adaptation across iterations. We evaluate SAGE on two representative environments, Overcooked Plus and Minecraft, comparing against both automated baselines and human-recorded test cases. Across all environments, SAGE achieves superior bug detection with significantly lower execution cost, while demonstrating strong adaptability to version updates.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

46 extracted references · 2 canonical work pages

  1. [1]

    Accessed: 2024-10-29 (2024)

    Newzoo: Newzoo’s Global Games Market Report 2024 - Free Version. Accessed: 2024-10-29 (2024). https://newzoo.com/resources/trend-reports/ newzoos-global-games-market-report-2024-free-version

  2. [2]

    Wu, Y., Chen, Y., Xie, X., Yu, B., Fan, C., Ma, L.: Regression testing of massively multiplayeronlinerole-playinggames.In:2020IEEEInternationalConferenceonSoft- wareMaintenanceandEvolution(ICSME),pp.692–696(2020).https://doi.org/10.1109/ ICSME46990.2020.00074 26

  3. [3]

    In: 1993 Conference on Software Maintenance, pp

    Agrawal,H.,Horgan,J.R.,Krauser,E.W.,London,S.A.:Incrementalregressiontesting. In: 1993 Conference on Software Maintenance, pp. 348–357 (1993). https://doi.org/10. 1109/ICSM.1993.366927

  4. [4]

    In: Proceedings of the 2015 International Symposium on Software TestingandAnalysis.ISSTA2015,pp.211–222.AssociationforComputingMachinery, New York, NY, USA (2015)

    Gligoric, M., Eloussi, L., Marinov, D.: Practical regression test selection with dynamic file dependencies. In: Proceedings of the 2015 International Symposium on Software TestingandAnalysis.ISSTA2015,pp.211–222.AssociationforComputingMachinery, New York, NY, USA (2015). https://doi.org/10.1145/2771783.2771784

  5. [5]

    In: 2015 USENIX Annual Technical Conference (USENIX ATC 15), pp

    Netravali, R., Sivaraman, A., Das, S., Goyal, A., Winstein, K., Mickens, J., Balakr- ishnan, H.: Mahimahi: Accurate Record-and-Replay for HTTP. In: 2015 USENIX Annual Technical Conference (USENIX ATC 15), pp. 417–429. USENIX Association, Santa Clara, CA (2015). https://www.usenix.org/conference/atc15/technical-session/ presentation/netravali

  6. [6]

    GSTF Journal on Computing (JoC)3(2), 10 (2013) https://doi.org/10.7603/ s40601-013-0010-4

    Ostrowski, M., Aroudj, S.: Automated regression testing within video game devel- opment. GSTF Journal on Computing (JoC)3(2), 10 (2013) https://doi.org/10.7603/ s40601-013-0010-4

  7. [7]

    In: 2021 IEEE Conference on Games (CoG), pp

    Gordillo, C., Bergdahl, J., Tollmar, K., Gisslén, L.: Improving playtesting coverage via curiosity driven reinforcement learning agents. In: 2021 IEEE Conference on Games (CoG), pp. 1–8 (2021). https://doi.org/10.1109/CoG52621.2021.9619048

  8. [9]

    Duarte,Y.,Canella,H.,Durelli,V.,Nardi,P.,Endo,A.:Exploratorytestingforplatform videogames:strategiesandlessonslearned.JournalonInteractiveSystems15,657–669 (2024) https://doi.org/10.5753/jis.2024.4156

  9. [10]

    2023 10th International Conference on DependableSystemsandTheirApplications(DSA),734–742(2023)https://doi.org/10

    Mingyue, Z., Xiao-Yi, Z., Paolo, A., Fuyuki, I.: An investigation of the behaviours of machine learning agents used in the game of go. 2023 10th International Conference on DependableSystemsandTheirApplications(DSA),734–742(2023)https://doi.org/10. 1109/dsa59317.2023.00105

  10. [11]

    IEEE/ACM International Conference on Automation of Software Test (AST), 90–99 (2021) https://doi.org/10.1109/AST52587.2021.00018

    Politowski, C., Petrillo, F., Guéhéneuc, Y.-G.: A survey of video game testing. IEEE/ACM International Conference on Automation of Software Test (AST), 90–99 (2021) https://doi.org/10.1109/AST52587.2021.00018

  11. [12]

    Machine Learning63(3), 217–248 (2006) https://doi.org/10.1007/ s10994-006-6205-6 27

    Spronck, P., Ponsen, M., Sprinkhuizen-Kuyper, I., Postma, E.: Adaptive game ai with dynamic scripting. Machine Learning63(3), 217–248 (2006) https://doi.org/10.1007/ s10994-006-6205-6 27

  12. [13]

    1–8 (2025)

    Mioto,V.,Petrillo,F.:Amappingofrecording-basedgametestautomationtools.In:2025 IEEE/ACM 9th International Workshop on Games and Software Engineering (GAS), pp. 1–8 (2025). https://doi.org/10.1109/GAS66647.2025.00006

  13. [14]

    In: Proceedings of the Annual Symposium on Computer-Human Interaction in Play

    Stahlke, S., Nova, A., Mirza-Babaei, P.: Artificial players in the design process: Devel- oping an automated testing tool for game level and world design. In: Proceedings of the Annual Symposium on Computer-Human Interaction in Play. CHI PLAY ’20, pp. 267–280. Association for Computing Machinery, New York, NY, USA (2020). https://doi.org/10.1145/3410404.3...

  14. [15]

    In: 2015 ACM/IEEE 18th International Conference on Model Driven Engineering Languages and Systems (MODELS), pp

    Iftikhar,S.,Iqbal,M.Z.,Khan,M.U.,Mahmood,W.:Anautomatedmodelbasedtesting approach for platform games. In: 2015 ACM/IEEE 18th International Conference on Model Driven Engineering Languages and Systems (MODELS), pp. 426–435 (2015). https://doi.org/10.1109/MODELS.2015.7338274

  15. [16]

    In: Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems

    Stahlke, S.., Nova, A., Mirza-Babaei, P.: Artificial playfulness: A tool for automated agent-based playtesting. In: Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems. CHI EA ’19, pp. 1–6. Association for Computing Machinery, New York, NY, USA (2019). https://doi.org/10.1145/3290607.3313039

  16. [17]

    In: 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE), pp

    Zheng, Y., Xie, X., Su, T., Ma, L., Hao, J., Meng, Z., Liu, Y., Shen, R., Chen, Y., Fan, C.: Wuji: Automatic online combat game testing using evolutionary deep reinforcement learning. In: 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE), pp. 772–784 (2019). https://doi.org/10.1109/ASE.2019.00077

  17. [18]

    In: 2018 IEEE Conference on Computa- tionalIntelligenceandGames(CIG),pp.1–8(2018).https://doi.org/10.1109/CIG.2018

    Guerrero-Romero, C., Lucas, S.M., Perez-Liebana, D.: Using a team of general ai algorithms to assist game design and testing. In: 2018 IEEE Conference on Computa- tionalIntelligenceandGames(CIG),pp.1–8(2018).https://doi.org/10.1109/CIG.2018. 8490417

  18. [19]

    IEEE Transactions on Software Engineering50(4), 911–936 (2024) https://doi.org/10.1109/TSE.2024.3368208

    Wang, J., Huang, Y., Chen, C., Liu, Z., Wang, S., Wang, Q.: Software testing with large language models: Survey, landscape, and vision. IEEE Transactions on Software Engineering50(4), 911–936 (2024) https://doi.org/10.1109/TSE.2024.3368208

  19. [20]

    arXiv preprint arXiv:2009.05617 (2020)

    Tufano, M., et al.: Unit test case generation with transformers and focal context. arXiv preprint arXiv:2009.05617 (2020)

  20. [21]

    Information and Software Technology176, 107565 (2024) https: //doi.org/10.1016/j.infsof.2024.107565

    Alagarsamy,S.,Tantithamthavorn,C.,Aleti,A.:A3test:Assertion-augmentedautomated test case generation. Information and Software Technology176, 107565 (2024) https: //doi.org/10.1016/j.infsof.2024.107565

  21. [22]

    arXiv preprint arXiv:2305.04764 (2023)

    Xie, Y., et al.: Chatunitest: A chatgpt-based automated unit test generation tool. arXiv preprint arXiv:2305.04764 (2023)

  22. [23]

    Dakhel,A.M.,Nikanjam,A.,Majdinasab,V.,Khomh,F.,Desmarais,M.C.:Effectivetest 28 generationusingpre-trainedlargelanguagemodelsandmutationtesting.Informationand SoftwareTechnology171,107468(2024)https://doi.org/10.1016/j.infsof.2024.107468

  23. [24]

    Mastropaolo,A.,Cooper,N.,Palacio,D.N.,Scalabrino,S.,Poshyvanyk,D.,Oliveto,R., Bavota,G.:Usingtransferlearningforcode-relatedtasks.IEEETransactionsonSoftware Engineering49(4), 1580–1598 (2022) https://doi.org/10.1109/TSE.2022.3183297

  24. [25]

    Tufano, M., Drain, D., Svyatkovskiy, A., Sundaresan, N.: Generating accurate assert statements for unit test cases using pretrained transformers. In: Proceedings of the 3rd ACM/IEEEInternationalConferenceonAutomationofSoftwareTest.AST’22,pp.54– 64.AssociationforComputingMachinery,NewYork,NY,USA(2022).https://doi.org/ 10.1145/3524481.3527220 . https://doi....

  25. [27]

    In: 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), pp

    Liu, Z., Chen, C., Wang, J., Che, X., Huang, Y., Hu, J., Wang, Q.: Fill in the blank: Context-aware automated text input generation for mobile gui testing. In: 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), pp. 1355–1367 (2023). https://doi.org/10.1109/ICSE48619.2023.00119

  26. [28]

    In: Proceedings of the IEEE/ACM 46th International Conference on SoftwareEngineering.ICSE’24.AssociationforComputingMachinery,NewYork,NY, USA (2024)

    Liu, Z., Chen, C., Wang, J., Chen, M., Wu, B., Che, X., Wang, D., Wang, Q.: Make llm atestingexpert:Bringinghuman-likeinteractiontomobileguitestingviafunctionality- aware decisions. In: Proceedings of the IEEE/ACM 46th International Conference on SoftwareEngineering.ICSE’24.AssociationforComputingMachinery,NewYork,NY, USA (2024). https://doi.org/10.1145/3...

  27. [29]

    arXiv preprint arXiv:2304.02014 (2023)

    Deng, Z., et al.: Large language models are edge-case fuzzers: Testing deep learning libraries via fuzzgpt. arXiv preprint arXiv:2304.02014 (2023)

  28. [30]

    In: Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA), pp

    Deng,Y.,Xia,C.S.,Peng,H.,Yang,C.,Zhang,L.:Largelanguagemodelsarezero-shot fuzzers: Fuzzing deep-learning libraries via large language models. In: Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA), pp. 423–435 (2023). https://doi.org/10.1145/3597926.3598067

  29. [31]

    1637–1641 (2022)

    Zhang,T.,Irsan,I.C.,Thung,F.,Han,D.,Lo,D.,Jiang,L.:itiger:Anautomaticissuetitle generationtool.In:Proceedingsofthe30thACMJointEuropeanSoftwareEngineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE), pp. 1637–1641 (2022). https://doi.org/10.1145/3540250.3558934

  30. [32]

    findings-emnlp.57 29

    Bui, N.D.Q., Wang, Y., Hoi, S.C.H.: Detect–localize–repair: A unified framework for learning to debug with codet5, 812–823 (2022) https://doi.org/10.18653/v1/2022. findings-emnlp.57 29

  31. [33]

    arXiv preprint arXiv:2304.05128 (2023)

    Chen, S., et al.: Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128 (2023)

  32. [34]

    1–13 (2024)

    Feng, S., Chen, C.: Prompting is all you need: Automated android bug replay with largelanguagemodels.In:Proceedingsofthe46thIEEE/ACMInternationalConference on Software Engineering (ICSE), pp. 1–13 (2024). https://doi.org/10.1145/3597503. 3608137

  33. [35]

    https://doi.org/10.1145/3524459.3527350

    Lajkó, M., Csuvik, V., Vidács, L.: Towards javascript program repair with generative pre-trainedtransformer(gpt-2).In:ProceedingsoftheThirdInternationalWorkshopon AutomatedProgramRepair.APR’22,pp.61–68.AssociationforComputingMachinery, New York, NY, USA (2022). https://doi.org/10.1145/3524459.3527350

  34. [36]

    In: Proceedings of the 31st ACM Joint EuropeanSoftwareEngineeringConferenceandSymposiumontheFoundationsofSoft- ware Engineering (ESEC/FSE), pp

    Wang, W., Wang, Y., Joty, S., Hoi, S.C.H.: Rap-gen: Retrieval-augmented patch gener- ation with codet5 for automatic program repair. In: Proceedings of the 31st ACM Joint EuropeanSoftwareEngineeringConferenceandSymposiumontheFoundationsofSoft- ware Engineering (ESEC/FSE), pp. 146–158 (2023). https://doi.org/10.1145/3611643. 3616256

  35. [37]

    Accessed: 2025-10-14 (2024)

    iXie Gaming: A Comprehensive Review of Game Test Automation Tools. Accessed: 2025-10-14 (2024). https://www.ixiegaming.com/blog/ comprehensive-review-game-test-automation-tools/

  36. [38]

    In: 2014 Brazilian Conference on Intelligent Systems (BRACIS), pp

    De Souza, L.S., Prudêncio, R.B.C., Barros, F.d.A.: A hybrid binary multi-objective particle swarm optimization with local search for test case selection. In: 2014 Brazilian Conference on Intelligent Systems (BRACIS), pp. 414–419 (2014). https://doi.org/10. 1109/BRACIS.2014.80

  37. [39]

    In: 2015 IEEE 8th International Con- ference on Software Testing, Verification and Validation (ICST), pp

    Mondal, D., Hemmati, H., Durocher, S.: Exploring test suite diversification and code coverage in multi-objective test case selection. In: 2015 IEEE 8th International Con- ference on Software Testing, Verification and Validation (ICST), pp. 1–10 (2015). https://doi.org/10.1109/ICST.2015.7102588

  38. [40]

    In: 2011 IEEE 23rd International Conference on Tools withArtificialIntelligence,pp.245–252(2011).https://doi.org/10.1109/ICTAI.2011.45

    Souza, L.S.d., Miranda, P.B.C.d., Prudencio, R.B.C., Barros, F.d.A.: A multi-objective particle swarm optimization for test case selection based on functional requirements coverage and execution effort. In: 2011 IEEE 23rd International Conference on Tools withArtificialIntelligence,pp.245–252(2011).https://doi.org/10.1109/ICTAI.2011.45

  39. [41]

    In: 2024 IEEE International Conference on Autonomic Com- puting and Self-Organizing Systems Companion (ACSOS-C), pp

    Cai, J., Li, J., Li, N., Zhang, M., Yang, R., Tei, K.: Overcooked plus: A com- prehensive cooking scenario testbed for enhancing the evaluation of autonomous planning algorithms. In: 2024 IEEE International Conference on Autonomic Com- puting and Self-Organizing Systems Companion (ACSOS-C), pp. 146–151 (2024). https://doi.org/10.1109/ACSOS-C63493.2024.00046

  40. [42]

    https://www

    Ghost Town Games Ltd.: Overcooked! Accessed: 2024-08-26 (2024). https://www. 30 team17.com/games/overcooked/ Accessed 2024-08-26

  41. [43]

    https://www.minecraft.net

    Mojang Studios: Minecraft. https://www.minecraft.net. Sandbox video game developed by Mojang Studios (2009)

  42. [44]

    In: 2024 IEEE 35th International Symposium on Software Reliability Engineering Workshops (ISSREW), pp

    Hu, J., Zhang, M., Liu, B., Wu, Y., Chen, Y.: A language-guided acceleration method for smoke testing of game quests. In: 2024 IEEE 35th International Symposium on Software Reliability Engineering Workshops (ISSREW), pp. 7–12 (2024). https://doi. org/10.1109/ISSREW63542.2024.00039

  43. [45]

    Empirical Software Engineering27(6), 140 (2022) https://doi.org/10

    Li,Z.,Wu,Y.,Ma,L.,Xie,X.,Chen,Y.,Fan,C.:Gbgallery:Abenchmarkandframework for game testing. Empirical Software Engineering27(6), 140 (2022) https://doi.org/10. 1007/s10664-022-10158-x

  44. [46]

    https://arxiv.org/abs/1707.06347

    Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O.: Proximal Policy Optimization Algorithms (2017). https://arxiv.org/abs/1707.06347

  45. [47]

    Journal of machine learning research22(268), 1–8 (2021)

    Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., Dormann, N.: Stable- baselines3: Reliable reinforcement learning implementations. Journal of machine learning research22(268), 1–8 (2021)

  46. [48]

    https://store.steampowered.com/ 31

    ValveCorporation:Steam—TheUltimateOnlineGamePlatform.AccessedonOctober 30, 2025 (2025). https://store.steampowered.com/ 31