{"id":"6e0efd5f-15a8-42b9-b547-fbdaa83995f5","arxiv_id":"2607.28074","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Deep, capability-targeted, co-evolving synthetic environments raise a 9B computer-use agent from 36.5% to 67.1% and enable RL gains the live web cannot supply.","lead":"Echoverse builds deep, resettable synthetic apps so computer-use agents can train on login-gated workflows graded against real database state. A 9B model nearly doubles accuracy, and the same worlds support reinforcement learning beyond imitation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Co-evolution lift is confounded with task-set expansion; the paper's own non-ablation leaves the central causal claim under-identified.","rationale":"The reader's weakest_assumption correctly names the soft spot: Sec. 6.7 openly declines environment/verifier-only ablations, so the co-evolution lever is a whole-loop effect. That is the single most load-bearing concern for the central claim that interior quality (especially co-evolution) dominates bulk count, because the other two levers have cleaner supports (shallow hurts live transfer; capability worlds transfer to held-out widgets and Online-Mind2Web). I do not escalate to REJECT: database-grounded grading, the negative shallow result, capability transfer, and RL on unmodified worlds are independent positive evidence, and the four-world release raises reproducibility above typical agent-systems papers. The concrete frozen-panel test would settle whether the ECHOSTAY doubling is capability or measurement drift without requiring a full redesign. Verdict stays CONDITIONAL; confidence remains moderate pending that disambiguation.","tokens_in":22566,"tokens_out":621,"duration_ms":10226,"concrete_test":"Freeze a panel of ECHOSTAY tasks that were already feasible and exported under v1 (same goals, same reference minting rules), grade both the v1-trained and v2-trained models on that identical panel with the v1 verifier (or a jointly agreed frozen DB check). If the 16.2%→38.5% lift largely disappears on the frozen panel while appearing only on v2-new tasks, the co-evolution headline is mostly task-set expansion; if it remains, the capability claim holds.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The thesis that gains come from interior properties (depth, targeting, co-evolution) rather than count rests heavily on Sec. 6.7's ECHOSTAY result (model 16.2%→38.5% after world v1→v2). The paper states that environment-only or verifier-only ablations are 'not well defined under this design' because repairing a control 'brings tasks into existence that could not previously be posed' and references are re-minted from the repaired DB. That means the measured lift conflates (a) cleaner supervision on a fixed task distribution with (b) a larger/harder/easier task set and re-grounded graders. If most of the jump is (b), co-evolution is mainly curriculum/task-factory improvement, not evidence that repairing world fidelity sharpens agent capability on a stable measure. The deep-vs-shallow live transfer (Sec. 6.2) is directionally supportive but small-N (two domains, ~dozens of tasks) and does not isolate co-evolution. Without a fixed held-out task panel scored under a frozen verifier across world versions, the load-bearing causal claim for co-evolution remains entangled.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"Echoverse argues that once synthetic computer-use environments are plentiful, returns come from three interior properties—behavioural depth (completeness w.r.t. a target workflow set), capability targeting (mass-varying failed controls), and co-evolution of environment, tasks, and verifier with the model—rather than from environment count. A factory compiles seeds into FastAPI/React/SQLite apps with database-grounded graders; a loop reads each graded rollout as both world repair and training signal. On twelve worlds, a 9B model rises from 36.5% to 67.1% across fourteen splits; shallow clones hurt live transfer while deep ones help; capability worlds transfer to held-out widgets and the open web; repairing ECHOSTAY lifts a model from 16.2% to 38.5%; and RL with a grounded trajectory reward plus dense per-step judge raises held-out score from 58.8% to 68.0%. Four worlds are released as a benchmark.","tokens_in":22864,"tokens_out":1377,"duration_ms":34172,"significance":"If the results hold, the paper usefully shifts the field from scaling environment count to interior quality and closed-loop repair, with an operational depth definition discharged by machine-checkable claims, database-grounded verification that is harder to game than screenshot judges, and unmodified worlds that meet RL reset/throughput/reward needs. The main scorecard, capability ablations with held-out widget families, dual scaling axes, live-web transfer, and the RL run are coherent contributions. Releasing runnable apps, seed data, and grounded graders is a concrete community asset. The work is complementary to bulk environment generators and to live-web benchmarks that cannot supply login-gated write workflows or exact reset.","major_comments":[{"comment":"Sec. 6.7 presents the ECHOSTAY v1→v2 model lift (16.2%→38.5%) as evidence for co-evolution, one of the three central levers. The paper states that environment-only or verifier-only ablations are “not well defined” because repairs bring new tasks into existence and re-mint references from the repaired DB. The measured jump therefore confounds cleaner supervision on a fixed distribution with task-set expansion and grader re-grounding. Without a frozen held-out task panel scored under a fixed verifier across world versions (or an explicit decomposition of solve-rate vs. corpus-composition effects), the causal claim that repairing world fidelity sharpens agent capability on a stable measure remains under-identified. Either add that panel or reframe co-evolution strictly as a joint loop effect and stop treating 16.2→38.5 as isolated evidence for depth/repair quality.","section":"Sec. 6.7"},{"comment":"Sec. 6.2 is the primary support for the depth thesis on live transfer, but uses only two WebVoyager domains (Allrecipes, Hugging Face) with a few dozen tasks each, and shallow vs. deep corpora that are “comparable” but not trajectory-matched. The paper itself reads this as directional evidence that shallow can be worse than no training, not as an estimate of depth’s value. That is appropriately cautious in the text, yet the abstract and thesis still lean on 80→75 vs. 80→85 / 48→65 as a main result. Strengthen with more domains, matched trajectory budgets, or demote the quantitative claim to a qualitative negative result with explicit N limits in the abstract.","section":"Sec. 6.2, Figure 5"},{"comment":"Phase-2 feasibility is white-box (Playwright plus source and DB access) and deliberately excludes the student policy, which is good, but the supervised corpus is still teacher-filtered (GPT-5.4 trajectories that pass the grounded verifier). Evaluation splits are therefore shaped by what the factory and teacher can complete. Table 4’s comparison of πSFT to GPT-5.4 is informative as distillation progress, but the paper should state more clearly which gaps are student-capacity vs. residual world/task hardness the teacher also fails, and whether any evaluation tasks were ever filtered by teacher success (Sec. 4.3 says no; Sec. 4.8 says failed teacher trajectories contribute no demos—confirm this holds for the released benchmark tasks).","section":"Sec. 4.3, 4.8, Table 4"}],"minor_comments":[{"comment":"Table 4 averages are unweighted means over fourteen splits of very different difficulty and size; a weighted or per-category breakdown (communication / regulated / capability) would aid interpretation.","section":"Table 4, Sec. 5.3"},{"comment":"Figure 8’s trajectory-scaling axis holds per-world mixture fixed—good—but absolute counts at each subsample point are hard to read off the prose; add exact N labels on the x-axis or a small table.","section":"Sec. 6.6, Figure 8"},{"comment":"RL uses a 50-turn train cap vs. 100-turn eval budget (Sec. 7.4); note whether any held-out gain is partly longer-horizon tolerance rather than better policy.","section":"Sec. 7.4–7.5"},{"comment":"Appendix C verifier prompts are a strength; consider reporting inter-judge agreement or a small human audit of write-diff decisions to quantify residual LLM-comparison error.","section":"Appendix C, Sec. 3.2"},{"comment":"Typos/style: “WebV oyager” spacing appears repeatedly; “aworld” → “a world” early in Sec. 1; ensure consistent πbase / πSFT notation in figures.","section":"Sec. 1–2"}],"recommendation":"major_revision","confidential_remarks":"Solid systems paper with real assets (release, grounded graders, RL substrate). The co-evolution identification gap is the main scientific soft spot; if authors add a frozen-panel experiment or clearly demote that claim, this could clear a top venue. Fit is good for a methods/systems track in agent learning. No integrity concerns from the text."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: once you can mint synthetic computer-use worlds in bulk, interior quality matters more than count. They show a 9B model going 36.5→67.1 on fourteen splits from twelve worlds, shallow clones regressing live Allrecipes while deep ones help, capability worlds transferring to held-out widgets and Online-Mind2Web, and the same worlds supporting RL (58.8→68.0 held-out) with a grounded trajectory reward plus dense step shaping.\n\nWhat is actually new is the package, not any single ingredient. Grounded DB grading already exists (τ-bench); auto-gyms exist (InfiniteWeb, CUA-Gym, etc.). Echoverse’s contribution is an operational, claim-checked notion of depth, mass-produced capability worlds that isolate failing controls across themes, a factory that repairs env/task/verifier before trusting failures as model signal, and evidence that those worlds are drop-in RL environments for login-gated workflows. The negative shallow-world result is rare and worth citing. Releasing four runnable apps with seed DBs and graders is real, not vapor.\n\nSoft spots, in proportion. Sec. 6.7’s ECHOSTAY 16.2→38.5 is a whole-loop effect; they say environment-only or verifier-only ablations are “not well defined” because repairs expand the poseable task set and re-mint references. So co-evolution is not cleanly identified as “fidelity repair sharpens a fixed measure”—it is also curriculum and factory improvement. That undercuts the strongest causal reading of the thesis, not the practical value of the loop. Deep-vs-shallow is two domains and dozens of tasks; directionally right, not a precision estimate. The dense RL judge reintroduces appearance-based signal they criticize elsewhere, though they keep validation on the grounded term and couple the two rewards carefully. Free knobs (95% claim bar, mid-band pass@4 filter, reward weights) are normal for this genre.\n\nMath and citations look fine: no load-bearing formal claims, related work is honest about Fara and the gym literature, circularity is low because outcomes tie to DBs and external live benches. Who it’s for: anyone building or training CUAs, especially enterprise/login-gated RL. I’d bring it to reading group, cite the depth/capability findings and the release, and send it to peer review. Fix the co-evolution identification story in revision if you can; don’t desk-reject over it.","headline":"Solid systems paper: depth and targeted worlds beat bulk gyms, with a real RL substrate—but the co-evolution lift is intentionally non-causal.","tokens_in":23622,"tokens_out":620,"would_cite":true,"duration_ms":15532,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Training computer-use agents improves more from deep, co-evolving worlds than from more environments.","keywords":["computer-use agents","synthetic environments","grounded verification","co-evolution","environment depth","reinforcement learning","web agents","task curriculum"],"falsifier":"Train matched models on deep versus shallow versions of the same domains and on pre- versus post-repair versions of one world, then score them on fixed live-web tasks and on a frozen task set graded by an unchanged external verifier; if shallow or unrepaired worlds match or beat deep repaired ones on those fixed measures, the interior-quality claim fails.","tokens_in":23363,"feed_emoji":"🖥️","tokens_out":955,"duration_ms":17845,"temperature":0.7,"pith_summary":"Computer-use agents only learn from actions that change real application state, yet the login-gated, stateful apps that matter cannot be trained on directly. Synthetic stand-ins therefore have to carry the load. This paper argues that once such environments can be generated in bulk, the bottleneck is no longer how many exist but what is inside each one: behavioural depth relative to target workflows, focus on the exact interactions the agent fails, and a loop that repairs the environment, tasks and verifier from the same graded rollouts used to train the model. Echoverse compiles specifications into stateful apps graded against their own databases, then co-evolves those worlds with the policy. A 9B model trained on twelve such worlds rises from 36.5% to 67.1% across fourteen splits, shallow clones can hurt live-site transfer while deep ones help, and the same worlds support reinforcement learning that lifts held-out score from 58.8% to 68.0%.","feed_headline":"Deep co-evolving worlds beat more sites for agent training","feed_subtitle":"A 9B model nearly doubles on fourteen splits; shallow clones can hurt live transfer","key_machinery":"The co-evolution loop: every graded rollout is read twice—once as repairs to the environment, its tasks and its database-grounded verifier (world first, without weakening goals), and once as training signal for the model—so a static benchmark saturates while the loop compounds.","core_discovery":"Gains for computer-use agents come less from adding more synthetic environments than from three interior properties of each world: how completely it supports the workflows it is meant to teach, whether it targets the specific interaction the agent fails, and whether environment, tasks and verifier improve alongside the model. On the same domains, shallow worlds can push live accuracy below the base model while deep ones raise it; repairing one world more than doubles the model trained on it; and the same grounded worlds serve unmodified as RL environments.","pith_inferences":["If depth and co-evolution dominate count, environment-generation pipelines should ship machine-checkable workflow claims and repair traces, not only page counts.","Live-web benchmarks that judge only from screenshots will remain weak training rewards even as synthetic grounded verifiers improve, widening the train-eval substrate split.","Capability worlds that isolate single controls may become a standard complement to full-domain clones whenever agents stall on one widget class across many sites.","Whole-loop lifts without separable ablations will keep making it hard to credit ‘better environments’ versus ‘easier or differently filtered tasks’ unless frozen external graders become standard."],"forward_implications":["Below a depth threshold, adding environments can inject noise and hurt live transfer rather than help.","Drilling one failing control across many renderings transfers to held-out widget families and to the open web.","The same database-owned worlds that supply clean supervised data also meet reset, throughput and reward needs for RL without modification.","A mid-size student can close most of the gap to its much larger teacher when trained only on deep, targeted, checkable synthetic trajectories.","Public progress should emphasize factories that find failures and repair worlds, not only larger inventories of synthetic sites."],"fun_headline_variants":["Deep co-evolving worlds beat bulk sites for agent training","Environment depth lifts 9B agents; shallow clones can hurt","Co-evolving one world more than doubles trained agent score","Stateful apps graded on their own DB nearly double 9B scores","Targeted deep worlds transfer to live sites and held-out UI"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That measured gains from repairing a world reflect genuine agent skill rather than mainly reshaping which tasks exist and how they are scored, since environment-only or verifier-only ablations are not separable in this design.","fun_headline_variants_meta":{"raw":{"variants":["Deep co-evolving worlds beat bulk sites for agent training","Environment depth lifts 9B agents; shallow clones can hurt","Co-evolving one world more than doubles trained agent score","Stateful apps graded on their own DB nearly double 9B scores","Targeted deep worlds transfer to live sites and held-out UI"]},"model":"grok-4.5","effort":"low","cost_usd":0.003533,"raw_usage":{"total_tokens":1284,"prompt_tokens":934,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":35328000,"prompt_tokens_details":{"text_tokens":934,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":279,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":934,"tokens_out":71,"duration_ms":5835,"temperature":1.0,"reasoning_tokens":279,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T18:37:44.777876+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train matched models on deep versus shallow versions of the same domains and on pre- versus post-repair versions of one world, then score them on fixed live-web tasks and on a frozen task set graded by an unchanged external verifier; if shallow or unrepaired worlds match or beat deep repaired ones on those fixed measures, the interior-quality claim fails.","supporting_citations":[],"review_version":1}