{"id":"754ece48-8dfe-4b30-9f83-1d50ecd2528e","arxiv_id":"2607.14548","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A 3B-scale GUI agent reportedly scores 82.6% on AndroidWorld and 42% on real-device tasks, but the evidence is not independently verified and may overlap with its RL training.","lead":"HyMobileAgent is a phone-operating AI that claims to beat much larger assistants on real app tasks while running at a ~3 billion parameter size. The report argues that clever training data and simulated phone environments matter more than parameter count.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AndroidWorld is both a training environment and the headline test: §5.3/Table 2 puts ~500 AndroidWorld devices in the online-RL pool, so the 82.6% score is not an independent measure of transfer.","rationale":"The reader's stated weakest_assumption is that PhoneWorld mock-app supervision transfers to real apps. That is a genuine concern, but it is secondary. The more load-bearing problem is explicit in the paper: AndroidWorld appears in the online-RL environment mixture (Table 2), and the same benchmark is the headline evaluation (Table 3). No independent held-out evaluation is provided. Even if PhoneWorld-to-real transfer works perfectly, the AndroidWorld number cannot support the central claim because the evaluation environment is also a training environment. The reader's rationale does mention this contamination possibility, so there is partial agreement, but the formal weakest_assumption focuses elsewhere. My concern moves the rejection rationale to a more direct and decisive issue. The verdict should remain REJECT: the central claim is not independently supported. If the authors provide a contamination-free evaluation and artifacts, conditional acceptance would become appropriate.","tokens_in":23287,"tokens_out":3146,"duration_ms":37538,"concrete_test":"Re-run the online-RL phase with AndroidWorld excluded from the environment pool (keeping real-app and PhoneWorld rollouts), then evaluate on the same AndroidWorld tasks. If success drops more than ~5 points from 82.6%, the headline gain is explained by training on the evaluation environment. Also ask authors to report the fraction of evaluation tasks appearing in RL rollouts and to release rollout logs for independent verification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that HyMobileAgent, at A3B scale, matches or beats much larger models on AndroidWorld because of data/environment co-scaling. But §5.3 and Table 2 explicitly state that the online-RL rollout pool includes approximately 500 AndroidWorld devices. AndroidWorld is then used in §6.1 as the primary public action-execution benchmark. This means the policy was optimized against reward signals drawn from the same environment and task distribution on which it is later evaluated. The comparison to Gemini 3.1, Seed 2.0 Pro, GPT-5.4-Pro, and Claude-4.7 Opus is therefore not a like-for-like test: those systems are reported as general-purpose models, not as models trained on the evaluation benchmark itself. The paper does not report any withheld AndroidWorld split, any per-task overlap analysis, or any ablation excluding AndroidWorld from the RL pool. The in-house HyMobileWorld benchmark is author-constructed, human-graded, and not released, so it cannot independently validate transfer either. The load-bearing condition for the headline claim—that the evaluation measures generalization beyond the training distribution—is not established. This is a correctness risk, not a disagreement with any field consensus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"HyMobileAgent is an A3B-scale mobile GUI agent built on Hy3.0-VL-A3B, trained with a joint data- and environment-scaling framework: a GUI perception flywheel, tutorial-based knowledge extraction, a large-scale action trajectory pipeline over sandbox and real devices, the PhoneWorld mock-app environment, a structured Planning-and-Reflection mechanism, and a three-stage training recipe (mid-training, SFT, offline/online RL). The paper reports an 82.6% success rate on AndroidWorld and 42% on an in-house HyMobileWorld benchmark, claiming to match or surpass substantially larger proprietary agents at a much smaller deployment footprint, and concludes that data/environment scaling can rival additional model parameters.","tokens_in":23691,"tokens_out":5752,"duration_ms":64903,"significance":"If independently verified, the central claim is significant: a ~3B-scale mobile agent outperforming or matching much larger general-purpose models on AndroidWorld would support the thesis that environment-grounded data scaling has high marginal return. The paper also contains a concrete decision structure, a deterministic dead-loop reflection mechanism, and a large-scale data pipeline design that are of interest to the GUI-agent community. However, the evidence as presented is not yet load-bearing: AndroidWorld appears in the online-RL training pool, the in-house benchmarks are unreleased, and no ablation isolates the contribution of PhoneWorld or mock-app data. These are not presentation issues; they directly affect whether the headline comparison measures transfer or overfitting to the evaluation environment.","major_comments":[{"comment":"The headline AndroidWorld result comes from a policy whose online-RL rollout pool includes approximately 500 AndroidWorld devices (Table 2), and AndroidWorld is then used in §6.1 as the primary public action-execution benchmark. No withheld split, per-task overlap analysis, or ablation excluding AndroidWorld from the RL pool is reported. The comparison to Gemini 3.1, Seed 2.0 Pro, GPT-5.4-Pro, and Claude-4.7 Opus is therefore not a like-for-like test of generalization: those systems were not trained on AndroidWorld rewards. This is the central load-bearing issue for the abstract's claim. A revision should either train a variant without AndroidWorld rollouts and report its AndroidWorld score, or provide a strict task-disjoint holdout analysis.","section":"§5.3, Table 2; §6.1, Table 3"},{"comment":"PhoneWorld is used for SFT-data generation, as the RL task pool (34,242 single-app and 500 cross-app tasks), and as part of the online-RL environment mixture (34 mock apps, ~1,000 VM instances). Yet PhoneWorld is also a self-cited prior paper (Tang et al., 2026) with overlapping authors. There is no controlled ablation that removes PhoneWorld or mock-app data from training, and no direct measurement of whether PhoneWorld tasks transfer to real Android apps. Without such an ablation, the causal attribution of the reported gains to 'environment co-scaling' is not established; the gains could be dominated by other data sources or by the real-app rollout pools.","section":"§4.4, §5.3"},{"comment":"All action-execution numbers are reported as single point estimates. AndroidWorld is a dynamic environment with stochastic task states, and HyMobileWorld involves human grading, yet no error bars, number of independent runs, per-task standard deviations, inter-annotator agreement, or statistical tests are provided. Differences such as 82.6 vs. 80.2 (AndroidWorld) and 42.0 vs. 44.7 (HyMobileWorld) are within plausible noise for such evaluations. The claims of 'matching or surpassing' are therefore not statistically supported as reported.","section":"§6.2, Table 3"},{"comment":"HyMobileWorld, HyMobileGrounding, and HyMobileQA are in-house benchmarks that are not released. For HyMobileWorld, the evaluation description is limited to 150 tasks ('50 native-app, 50 mini-program, 50 cross-app') with trained annotators performing blinded assessments; no task list, annotation protocol, release plan, or independent replication is provided. The 42% figure on HyMobileWorld and the in-house grounding/QA numbers therefore cannot be externally verified, which is particularly problematic because these in-house suites are used to support the central transfer claim.","section":"§6.1"}],"minor_comments":[{"comment":"The abstract refers to 'HyMobileOnline' but the benchmark is consistently called 'HyMobileWorld' elsewhere in the paper; please unify the name.","section":"Abstract"},{"comment":"The action-reward description states that 'the exact weighting and the per-primitive formulations are deferred to the released training configuration,' but no release URL, repository, or configuration artifact is provided. Please include the configuration or a pointer to a public release.","section":"§5.3"},{"comment":"Figure 1 and Table 3 report scores without any uncertainty or run count. Even if error bars are not possible for all proprietary baselines, the authors' own runs should include variability estimates, and the figure should distinguish published numbers from self-measured ones.","section":"§6.2/Figure 1"},{"comment":"The term 'A3B' is used throughout without a definition. Please specify whether it means 3 billion active parameters in a larger model, 3 billion total parameters, or something else, and clarify how it relates to the 'A3B-scale deployment footprint.'","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The decisive issue is the overlap between the online-RL training environment and the headline evaluation benchmark. I would not accept the paper in its current form because the 82.6% AndroidWorld claim is not an independent measure of generalization. The good news is that this is fixable within the paper's scope: the authors can retrain or ablate without AndroidWorld rollouts, report a strict holdout analysis, and add ablations that isolate PhoneWorld/mock-app contributions. Given the scale of the system, I would also weigh whether an unreleased in-house benchmark is acceptable as primary evidence for the journal's readership; if not, the authors should be asked to release at least HyMobileWorld's task list and grading protocol."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things: the engineering is serious, and the headline number is not trustworthy as reported. The system is a 3B-scale mobile GUI agent with a 32K context, a closed perception data flywheel, tutorial-video distillation, a resettable mock-app environment, and a two-phase RL recipe. If the results were independently confirmed, this would be a practical milestone for on-device agents. But they are not: Table 2 lists roughly 500 AndroidWorld devices in the online-RL sampling pool, and Section 6.1 uses AndroidWorld as the primary public action benchmark. That means the evaluation set overlaps the training distribution, and the comparison to general-purpose models like Gemini 3.1 or GPT-5.4-Pro is not like-for-like. There is no withheld split, no overlap analysis, and no ablation removing AndroidWorld from the RL pool. So the 82.6% figure does not establish generalization.\n\nWhat is new is the integration. Each component has prior art that the paper honestly cites. The novel bits are the scale of the combined pipeline, the deterministic dead-loop reflection, and the idea of co-scaling data and environment diversity. I credit the paper for explicitly disclosing the environment mixture; many would have hidden it. But disclosure does not fix the inferential gap.\n\nThe soft spots are predictable. All numbers are point estimates with no error bars. The in-house benchmarks are authored by the same team, human-graded, and unreleased, so they cannot validate transfer. PhoneWorld is a self-cited prior paper, supplies SFT data and RL tasks, and there is no controlled ablation separating mock from real apps. The appendix trajectories are all on PhoneWorld, not on third-party apps.\n\nI would still send this to peer review. The system is substantial, the writing is clear, and the contamination concern is addressable in revision. The referees should ask for a held-out external benchmark, an ablation excluding AndroidWorld from training, and release of the evaluation harness or at least the in-house benchmark. Without those, the central claim is a strong hypothesis, not a demonstrated result.","headline":"Serious engineering, plausible numbers, but the main AndroidWorld result is not independent because the policy trains in that same environment.","tokens_in":795,"tokens_out":1845,"would_cite":false,"duration_ms":42763,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A roughly 3-billion-parameter mobile agent matches or beats much larger rivals by co-scaling data and environments instead of parameters.","keywords":["mobile GUI agent","joint data-environment scaling","PhoneWorld","AndroidWorld","GUI grounding","reinforcement learning","long-horizon decision making","mock application environment"],"falsifier":"Train two otherwise identical agents, one with PhoneWorld mock-app data and the other with only real-device trajectories, then compare AndroidWorld and HyMobileWorld success rates. If the mock-only agent loses most of the gap, the mock-to-real transfer assumption is false and the co-scaling claim is weakened.","tokens_in":23241,"feed_emoji":"📱","tokens_out":5593,"duration_ms":58971,"temperature":0.7,"pith_summary":"HyMobileAgent is built on a ~3B-parameter vision-language model and claims that scaling training data and interaction environments, not model size, is what makes a mobile GUI agent succeed. The paper reports 82.6% strict success on AndroidWorld and 42% on its own real-device benchmark HyMobileWorld, matching or exceeding far larger general-purpose agents. It attributes the gains to a closed-loop data system (mock-interface synthesis, tutorial-to-data distillation, million-scale action trajectories), a resettable mock-app environment called PhoneWorld, and a Planning-and-Reflection decision structure with dead-loop detection, trained via mid-training, SFT, and two-phase RL. The broader claim is that, at a fixed deployment budget, environment-grounded reinforcement learning can deliver returns that rival or exceed adding parameters.","feed_headline":"82.6% on AndroidWorld with a 3B agent — data, not size, did it","feed_subtitle":"A 3B-parameter model matches much larger rivals when training scales data and interaction environments together.","key_machinery":"The load-bearing mechanism is the joint data- and environment-scaling loop. On the data side, a perception flywheel combines mock-interface synthesis, reject-sampling difficulty selection, and icon-specific augmentation; a knowledge pipeline turns tutorial videos into single-image planning data and multi-image state-transition data; and an action pipeline collects million-scale trajectories with automated failure attribution. On the environment side, PhoneWorld provides 34 resettable mock apps and over 34,000 verifiable tasks, making trajectory-level reinforcement learning possible without login, payment, or anti-bot barriers. The decision structure that carries long-horizon behavior is the","core_discovery":"HyMobileAgent, built on the Hy3.0-VL-A3B vision-language model, achieves an 82.6% strict success rate on AndroidWorld and 42.0% on the in-house HyMobileWorld real-device benchmark, matching or beating substantially larger general-purpose agents while keeping an A3B-scale deployment footprint. The paper's central claim is that the binding constraint for small mobile agents is not foundation-model capacity but the quality of surrounding data, interaction environments, and decision structure. It presents a joint data- and environment-centric scaling framework — a GUI perception flywheel, knowledge extraction from tutorial videos, a million-scale action-trajectory pipeline over 2,000+ sandbox an","pith_inferences":["The paper compares against other models but does not ablate the co-scaling components individually; a controlled study removing only the mock-app/PhoneWorld portion would test whether environment scaling, rather than the larger SFT corpus, drives the gains.","If mock-to-real transfer holds, mobile agents could be trained almost entirely in synthetic resettable environments and fine-tuned on a thin slice of real data, sharply reducing the cost of agent training.","Making recovery conditions deterministic and trainable, as in the dead-loop reflection design, could transfer beyond mobile GUIs to web and desktop agents.","The in-house benchmarks are described in the report but are not part of the public benchmark set, so independent re-evaluation awaits their release."],"forward_implications":["At a fixed deployment budget, investing in environment-grounded RL and self-correcting decision structures can close much of the gap to much larger models.","Verifiable, resettable mock-app environments make RL training feasible for tasks that real devices block (logins, payments, anti-bot), so environment design becomes a first-class scaling axis.","Tutorial videos and books can be distilled into structured planning and state-transition supervision, adding procedural knowledge without new human annotation.","The three-repeat dead-loop trigger gives error recovery a learnable, deterministic target, which should reduce compounding-error failures in long trajectories.","The large gap between AndroidWorld (82.6%) and HyMobileWorld (42%) suggests real-device, mini-program, and cross-app scenarios remain the hardest frontier for all agents."],"fun_headline_variants":["3B agent matches bigger rivals via data-env co-scaling, not model size","Data and environment co-scaling lifts 3B GUI agent to 82.6% on AndroidWorld","Small model, big gains: HyMobileAgent excels with co-scaled data and env","How a 3B agent outperforms larger LLMs: co-scaled data and environments","82.6% AndroidWorld with 3B parameters — the secret is co-scaled data"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The weakest load-bearing premise is that PhoneWorld's 34 mock apps and sandbox/real-device pipelines yield supervision that transfers to real mobile applications; the paper provides no controlled ablation measuring mock-to-real transfer, so if that transfer fails, the environment-scaling explanation for the gains collapses.","fun_headline_variants_meta":{"raw":{"variants":["3B agent matches bigger rivals via data-env co-scaling, not model size","Data and environment co-scaling lifts 3B GUI agent to 82.6% on AndroidWorld","Small model, big gains: HyMobileAgent excels with co-scaled data and env","How a 3B agent outperforms larger LLMs: co-scaled data and environments","82.6% AndroidWorld with 3B parameters — the secret is co-scaled data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000335,"raw_usage":{"total_tokens":1719,"prompt_tokens":792,"completion_tokens":927,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":806}},"tokens_in":536,"tokens_out":927,"duration_ms":9145,"temperature":1.0,"reasoning_tokens":806,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T01:46:13.979940+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train two otherwise identical agents, one with PhoneWorld mock-app data and the other with only real-device trajectories, then compare AndroidWorld and HyMobileWorld success rates. If the mock-only agent loses most of the gap, the mock-to-real transfer assumption is false and the co-scaling claim is weakened.","supporting_citations":[],"review_version":1}