{"id":"35fc9462-df61-4613-b532-357f68a7dbca","arxiv_id":"2607.05155","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Across ~38,000 hours on 134 ultra-long real-world tasks, aggregate agent performance follows a log-sigmoid of interaction time, and measured learning speed doubles about every three months.","lead":"Agents learning from real environments improve on a precise log-sigmoid curve of interaction time, with R² about 0.998 across 134 day-long tasks. The work also reports that frontier agent learning speed has been roughly doubling every three months.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Aggregate R²≈0.998 does not establish a single intrinsic log-sigmoid law: large tmid dispersion (Fig. 5) plus near-tied S-curves (Table 1) make the fit a mixture summary, not unique evidence of one frontier process.","rationale":"The reader correctly flags population-level concentration and cut-mixing as the soft underbelly and already assigns CONDITIONAL with medium correctness risk. I agree that is the right tier: the empirical regularity of smooth, forecastable aggregate curves is real and well documented (Figs. 1, 5–8; forecasting RMSE <1). The sharper load-bearing issue for the strongest claim as written is not that the theory might be wrong in the abstract, but that the paper’s own Table 1, Appendix E, Figure 5 dispersion, and D.5 limitations already show the data do not uniquely pin down a single log-sigmoid with intrinsic parameters—only that averages are S-shaped in log time and well fit by flexible saturating forms. That underdetermination, not private-task count or the secondary three-month doubling, is what most limits reading the result as a pretraining-style scaling law. It does not justify REJECT: experience-vs-restarts (Fig. 12a), longer-horizon stability, and cross-family recurrence still make this a strong empirical contribution with clear caveats. Verdict stays CONDITIONAL; no upgrade or downgrade.","tokens_in":56815,"tokens_out":721,"duration_ms":35349,"concrete_test":"Fit every Table 1 family on the first 6.5h of the 134-task aggregate (as in Fig. 7) and compare held-out RMSE on 6.5–12h for all five models; also report the empirical distribution of per-task (or per-family) tmid and β and compute DM and BM from Assumptions D.1–D.2. If log-sigmoid is not uniquely best out-of-sample, or if DM/BM remain large, the specific single-law claim is not selected by the data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim packages high-precision evidence for a specific law S(t)=Smax/(1+(tmid/t)^β). Empirically, that precision is measured on task-averaged best-so-far curves with a free three-parameter S-curve and only a handful of time checkpoints; Table 1 shows log-probit/Gompertz/Weibull RMSE within ~3% of log-sigmoid (0.390 vs 0.398–0.404), and Appendix E states the form cannot be chosen by fit. Mechanistically, Appendix D.3–D.5 requires midpoint alignment and speed concentration (Assumptions D.1–D.2) for the average of task frontiers to collapse to one log-sigmoid; otherwise the aggregate is a convolution of shifted/heterogeneous sigmoids whose fitted (tmid, β) are window-dependent mixture summaries. Figure 5 already shows enormous dispersion (tmid from ~0.4h to 240h across families/models). Thus R²=0.998 on the 134-task mean strongly supports smooth, saturating, log-time-predictable averages, but only weakly supports a single intrinsic log-sigmoid scaling law of environment learning.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces EdgeBench, a suite of 134 ultra-long-horizon real-world agent tasks spanning six capability families, and uses roughly 38,000 hours of frontier-agent interaction to study post-deployment environment learning. The central empirical claim is that task-averaged best-so-far performance follows a three-parameter log-sigmoid in interaction time, S(t)=Smax/(1+(tmid/t)^β), with mean R²≈0.998 on the full 134-task average, similarly tight family-level and longer-horizon (28h/72h) fits, and accurate forecasts of later hours from early windows. A frontier-expansion theory on latent task graphs is offered as a sufficient mechanism; a secondary claim is that agent learning speed on a fixed 18-task slice roughly doubles every three months across recent model releases. Supporting analyses include experience-versus-restart ablations, context-length comparisons, submission-efficiency statistics, and a detailed gravitational-wave case study.","tokens_in":57246,"tokens_out":1584,"duration_ms":26612,"significance":"If the aggregate regularity holds under broader scrutiny, this is a substantial contribution: it elevates post-deployment environment learning from anecdotal trajectories to a measurable scaling object, analogous to pretraining and test-time scaling laws, and supplies a large, carefully engineered dual-loop benchmark (work/judge isolation, multilevel feedback, day-scale horizons) that the field currently lacks. Strengths include the scale of the evaluation, public release of 51 tasks plus the full harness, explicit comparison against alternative S-curves (Table 1), predictive early-window tests (Figure 7), and the experience-versus-independent-restart design (Section 5.2), which cleanly separates stateful learning from repeated sampling. The theoretical appendix is unusually thorough for an empirical agent paper and states its failure modes. Even if the unique log-sigmoid interpretation is moderated, the empirical finding of smooth, saturating, log-time-predictable aggregate learning curves across heterogeneous real tasks would remain important.","major_comments":[{"comment":"Table 1 and Appendix E: the manuscript claims a specific log-sigmoid scaling law, but full-window RMSE for log-probit (0.398), log-Gompertz (0.402), and Weibull (0.404) is within ~3% of log-sigmoid (0.390), and Appendix E explicitly states that the form cannot be chosen by fit. The abstract, title, and Section 3.2 therefore overstate uniqueness. Please reframe the central claim as high-precision evidence for smooth, saturating, log-time S-shaped aggregate curves (with log-sigmoid preferred on mechanistic grounds), report confidence intervals or bootstrap uncertainty on relative RMSE, and avoid language that implies a uniquely identified functional law from the data alone.","section":null},{"comment":"Figure 5 versus Assumptions D.1–D.2 (Appendix D.3): family-level fits show extreme tmid dispersion (e.g., ~0.4h to 240h across families/models) and heterogeneous β. The aggregate theorem requires residual midpoint alignment and speed concentration for the average of task frontiers to collapse to a single log-sigmoid rather than a convolution of shifted/heterogeneous sigmoids. The paper does not empirically test these concentration conditions (e.g., distribution of per-task tmid/β, residual mixture diagnostics, or whether fitted aggregate parameters are window-stable beyond the reported forecasts). Without that, R²=0.998 on the 134-task mean strongly supports a smooth population average but only weakly supports a single intrinsic frontier law. Add these diagnostics or qualify the theory-to-data link accordingly.","section":null},{"comment":"Section 4 and Figure 9: the ~3-month doubling claim rests on a hand-selected 18-task slice chosen for similar first-attempt performance, a rolling top-2 frontier fit, and a short calendar window of releases. Report sensitivity to slice composition, to using all models rather than rolling top-2, and to alternative learning-speed definitions (e.g., 4h or 6h gains; tmid-based speed). As written, the doubling timescale is presented as a robust generational law while remaining a free empirical summary of a small, selected panel.","section":null},{"comment":"Section 3.1 / fit protocol: aggregate curves are fit with three free parameters to a modest number of time checkpoints (roughly hourly over 12h, plus longer-horizon subsets). High R² is expected for smooth monotone averages under a flexible S-curve. Please report degrees of freedom, parameter standard errors, and leave-one-family-out or leave-one-model-out stability of (Smax, β, tmid), and clarify that Smax is an effective ceiling over the fitted regime (as Appendix D.5 notes) rather than an absolute performance bound.","section":null}],"minor_comments":[{"comment":"Only 51 of 134 tasks are publicly released. State clearly which families and difficulty strata are held back and how external groups can reproduce the aggregate scaling fits without the full suite.","section":null},{"comment":"Figure 1 and related plots: report whether scores are mean of best-so-far across three seeds or max, and whether error bands are available; several per-task tables mark * for <3 valid runs (especially GPT-5.4), which should be reflected in aggregate uncertainty.","section":null},{"comment":"Appendix B serving incidents for GPT-5.4 are important; consider a short main-text caveat near Figure 7 so forecast deviation is not read as pure model failure.","section":null},{"comment":"Notation: u = log t − log tmid is introduced cleanly in Appendix D but used earlier in Section 3.3; a one-line pointer would help non-theory readers.","section":null},{"comment":"Related work (Section 6 / Appendix F) is thorough; a compact table column for “primary reported quantity is trajectory vs endpoint” already exists—ensure AutoLab and FrontierSWE runtime comparisons are consistent with their public numbers.","section":null},{"comment":"Minor polish: “DeepSeek V4 Pro (preview)” naming, occasional future-dated system cards, and a few figure captions that restate R² without defining the score scale (0–100) could be tightened.","section":null}],"recommendation":"major_revision","confidential_remarks":"This is a high-effort, high-value empirical paper that a top venue should want after claim calibration. The main risk is rhetorical overclaim of a unique log-sigmoid law when the authors’ own Table 1 and theory assumptions already show the data support a broader S-shaped log-time regularity. I would not reject on novelty or circularity grounds; I would hold the line on reframing uniqueness and on testing midpoint/speed concentration. Scope fit for a serious ML/CL journal is strong if the revision treats EdgeBench and the aggregate regularity as co-equal contributions rather than packaging everything as discovery of one closed-form law."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing: they ran frontier agents for ~38k hours on 134 day-scale executable tasks and show that the task-averaged best-so-far score is extremely well described by a three-parameter log-sigmoid of interaction time (R² around 0.998). That is not just another leaderboard. EdgeBench itself—dual local/judge feedback, ≥12h contracts, six families, expert-built tasks—is the main engineering contribution, and the trajectory measurement is cleaner than most agent suites.\n\nWhat is actually new is treating environment-learning trajectories as a scaling object: same functional form across the full average, family averages, 28h/72h subsets, and early-window forecasts with low held-out RMSE. The experience-vs-independent-restarts ablation and the context-length check are useful; they push against “just more sampling.” Appendix D is honest about being a sufficient frontier-expansion story, not a proof that every task is logistic.\n\nSoft spots, in proportion. Table 1 and Appendix E already say log-probit/Gompertz/Weibull are nearly as good; the form is preferred on mechanism, not unique fit. Family-level tmid values scatter wildly, so the clean curve is a population average under midpoint/speed concentration—exactly what the theory needs and what the stress note flags. That does not kill the empirical claim (smooth, saturating, log-time-predictable averages), but it weakens “one intrinsic law of environment learning.” The three-month learning-speed doubling is striking and should be labeled provisional: fixed 18-task slice chosen for similar first-attempt scores. Only 51/134 tasks public and heavy API dependence limit full external replication; serving noise on GPT-5.4 is acknowledged.\n\nMath and citations look serious: scaling-law lineage, agent benchmarks, and limitations are engaged without hand-waving. Who this is for: people building long-horizon agents, eval harnesses, or post-deployment learning metrics. I would send it to peer review, bring it to reading group, and cite the benchmark plus the aggregate curve result—with the mixture/uniqueness caveat attached.","headline":"Strong empirical regularity on day-long agent learning curves, plus a real benchmark; the “single intrinsic log-sigmoid law” packaging is a bit tighter than the uniqueness evidence.","tokens_in":57988,"tokens_out":547,"would_cite":true,"duration_ms":11550,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"When AI agents learn from real-world environments, average performance rises with interaction time along a precise log-sigmoid curve.","keywords":["environment learning","scaling laws","log-sigmoid","AI agents","long-horizon benchmarks","interaction time","frontier expansion","EdgeBench"],"falsifier":"Fit the same three-parameter log-sigmoid to a larger, deliberately heterogeneous mix of tasks (or to single families kept separate) over horizons well past 12 hours; a persistent failure of R², systematic residual structure, or a better fit by a non-log-time S-curve or multi-inflection mixture would undercut the claimed law.","tokens_in":57680,"feed_emoji":"📈","tokens_out":1052,"duration_ms":11843,"temperature":0.7,"pith_summary":"Pretraining scaling laws tell us how models improve with data and compute, but what happens after deployment—when an agent must improve by interacting with a messy, feedback-rich environment—has been much less clear. This paper builds EdgeBench, a suite of 134 day-scale real-world tasks spanning science, software, optimization, professional work, formal math, and games, and runs frontier agents for roughly 38,000 hours of environment interaction. Averaging best-so-far performance over those tasks, it reports that overall progress follows a simple three-parameter log-sigmoid in interaction time with very high fit quality, and that this form holds across task families, longer horizons, and early-to-late forecasts. It also reports that measured learning speed of frontier agents has been roughly doubling every three months. The point is not just a new leaderboard: the paper treats environment learning itself as a scalable, measurable object, and argues that the regular curve is what you should expect when many tasks behave like frontier expansion on latent score graphs.","feed_headline":"Agent learning from environments follows a log-sigmoid law","feed_subtitle":"Across 134 day-long tasks and 38,000 hours, average performance tracks interaction time with R² near 0.998","key_machinery":"The log-sigmoid S(t) = Smax / (1 + (tmid/t)^β), where t is interaction time, Smax is the attainable score ceiling, tmid is the half-ceiling time, and β is the steepness in log time. The paper derives this as the many-task limit of a frontier-expansion process on latent task graphs: unlocked score mass helps unlock remaining mass, the expected growth rate is proportional to x(1−x) in a logarithmic time coordinate induced by self-similar search geometry, and averaging washes out per-task jaggedness when midpoints and speeds concentrate.","core_discovery":"Across 134 diverse real-world tasks and about 38,000 hours of agent–environment interaction, overall (task-averaged) best-so-far performance during environment learning follows a log-sigmoid scaling law of interaction time, S(t) = Smax / (1 + (tmid/t)^β), with mean R² around 0.998. The same functional form appears by task family, under 28- and 72-hour windows, and when early trajectories forecast later performance; separately, agent learning speed on a fixed slice roughly doubles every three months across recent model releases.","pith_inferences":["If the law is real, post-deployment interaction budgets may become a first-class training and product dial, analogous to how pretraining compute is planned today.","Harness design that preserves reusable state (workspace, memory, longer context) may move tmid and β as much as base model upgrades do.","Tasks engineered with strong single bottlenecks or non-scale-free feedback cycles would be natural stress tests for where the population-level log-sigmoid breaks.","A practical monitoring tool could watch live agent runs for deviation from an early-fit log-sigmoid as a signal of stuck exploration or serving failure."],"forward_implications":["Environment learning can be treated as a scaling object with its own predictable curve, not only as a pile of idiosyncratic task outcomes.","Early trajectory segments can be used to forecast later performance under a fixed interaction budget.","Progress across model generations can be tracked by learning speed (gain per fixed hours) rather than only by final score.","Long-horizon evaluation should report time-aligned improvement trajectories, not only end-state success.","Stateful continuous runs that reuse feedback should outperform equal-budget independent restarts if the frontier mechanism is right."],"fun_headline_variants":["Log-sigmoid law fits agent learning across 38k hours of tasks","Environment learning follows log-sigmoid with R² near 0.998","Agent learning speed doubles every three months on fixed tasks","First log-sigmoid scaling for real-world agent-environment learning","Task-averaged performance tracks log-sigmoid of interaction time"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"The clean law only becomes precise after averaging many tasks whose midpoints and learning speeds line up, and whose underlying task graphs mix influence across the unlocked–locked frontier without strong bottlenecks.","fun_headline_variants_meta":{"raw":{"variants":["Log-sigmoid law fits agent learning across 38k hours of tasks","Environment learning follows log-sigmoid with R² near 0.998","Agent learning speed doubles every three months on fixed tasks","First log-sigmoid scaling for real-world agent-environment learning","Task-averaged performance tracks log-sigmoid of interaction time"]},"model":"grok-4.5","effort":"low","cost_usd":0.003038,"raw_usage":{"total_tokens":1057,"prompt_tokens":778,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":30380000,"prompt_tokens_details":{"text_tokens":778,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":208,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":778,"tokens_out":71,"duration_ms":2123,"temperature":1.0,"reasoning_tokens":208,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T07:55:56.899746+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Fit the same three-parameter log-sigmoid to a larger, deliberately heterogeneous mix of tasks (or to single families kept separate) over horizons well past 12 hours; a persistent failure of R², systematic residual structure, or a better fit by a non-log-time S-curve or multi-inflection mixture would undercut the claimed law.","supporting_citations":[],"review_version":1}