{"id":"04577656-1e14-4863-8ec3-b03b37c6d655","arxiv_id":"2509.08400","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A position paper that reframes edge AI as a co-evolution loop in which wireless networks feed real-world experiences to LLMs and LLMs optimize the network in return.","lead":"This paper proposes 'ubiquitous intelligence', a vision where wireless networks and large language models (LLMs) improve each other in a continuous loop. It introduces 'experience scaling' as a new way to grow LLM capability beyond human-generated data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Experience scaling lacks a defined training mechanism; the paper never establishes that environmental interaction data can substitute for human-generated data in improving LLMs.","rationale":"The reader identified the same weakest assumption: experience data substituting for human-generated data. My analysis sharpens this into a missing-mechanism concern: the paper never defines how environmental interaction becomes a usable LLM training signal. This is more specific than 'unsupported' because it identifies a concrete gap in the argument's logic. However, this is a vision paper, not a technical contribution, and the reader's CONDITIONAL verdict already captures the appropriate status. The paper's self-listed open challenges in Section 5 do not make it internally inconsistent; they make it honest about the unproven feasibility. The citation error at [58] is real but peripheral to the central claim. I therefore recommend no change to the verdict: CONDITIONAL remains appropriate, with the condition being a proof-of-concept or formal feasibility model that demonstrates experience-scale learning works.","tokens_in":11269,"tokens_out":2686,"duration_ms":36470,"concrete_test":"Run a controlled experiment with a small LLM (e.g., 1B parameters): have an agent policy interact with a simulated environment (e.g., browser navigation or grid-world) and collect interaction trajectories containing no human-written text. Then continue pretraining or apply RL to the LLM on those trajectories and evaluate on standard benchmarks (e.g., MMLU, GSM8K) against a control trained on random or human text of comparable size. If the experience-trained model does not beat the control, the claimed experience-scaling mechanism lacks empirical support and the central claim should be treated as speculative.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that 'experience scaling'—LLMs collecting and creating data directly from their environment via wireless networks—can bypass the human-generated data bottleneck and drive continuous intelligence ascension (§3.1, §4.1.1). This requires two things: (i) a well-defined notion of 'experience' that can be converted into a training signal for LLMs, and (ii) evidence that such signal improves capabilities on tasks of interest. The paper provides neither. It asserts 'autonomous knowledge acquisition through environmental interaction' and 'continuous, multimodal experience accumulation' without specifying the learning objective, how raw or agent-generated experience is tokenized, labeled, or otherwise made usable for LLM training, or how the resulting data avoids the distributional and alignment properties that make human-generated data valuable. Section 5 explicitly lists 'scalable experience exchange' and 'global model consistency' as unsolved challenges, so even the communication and coherence layer is not established. The finite supply of human data is a real motivation, but the paper skips the decisive step: demonstrating that the proposed alternative data source is both usable and sufficient. Without that, the ubiquitous-intelligence co-evolution story rests on an unverified substitution assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces \"ubiquitous intelligence\" as a co-evolutionary paradigm in which LLMs and wireless networks jointly improve: wireless networks provide distributed, low-latency, personalized support for lifelong LLM learning, while LLMs make the network context-aware and adaptive. The central proposal is \"experience scaling\" (§3.1, §4.1.1), the idea that LLMs can bypass the human-generated-data bottleneck by collecting and creating training data directly from environmental interactions routed through wireless networks. The paper describes network-side mechanisms (edge-cloud collaborative reasoning, D2D knowledge sharing, spectrum-aware adaptation), LLM-side mechanisms (adaptive offloading, context-aware scheduling, hierarchical orchestration), and a claim of energy benefits from edge redistribution (§4.3). It closes by listing open challenges. No mathematical model, experimental evaluation, or quantitative validation is provided.","tokens_in":11540,"tokens_out":5547,"duration_ms":67723,"significance":"If the central claims were established, experience scaling would be a substantial research direction for 6G edge intelligence and continuous learning. The paper is valuable as a position statement: it synthesizes ongoing trends in edge LLM deployment, clearly names the data bottleneck, and lists important open problems (experience exchange, consistency, representation). Its strengths are the coherent conceptual framework and the explicit challenge list. However, the paper currently asserts rather than demonstrates its key premises: that environmental experience can substitute for human-generated data, that scaling laws have reached a ceiling, and that edge redistribution reduces energy. The contribution is therefore a vision paper, not a technical result.","major_comments":[{"comment":"The load-bearing claim is that LLMs can \"collect and create data directly from their environment\" and use this experience to bypass the human-data bottleneck. The manuscript never specifies how raw environmental interaction becomes a training signal: no objective, tokenization, labeling, reward model, or curation mechanism is defined. Without this, \"experience scaling\" is an assertion, not a proposal. The paper's own Section 5 lists \"scalable experience exchange\" and \"efficient knowledge representation\" as unsolved, but those are downstream of the more basic question of whether experience data can improve capabilities at all. A revision should either propose a concrete mechanism (e.g., RL from environment feedback, self-improvement loops, synthetic-data generation) and cite evidence that such signals improve LLM capabilities, or explicitly reframe experience scaling as an unproven hypoth","section":"§3.1, §4.1.1"},{"comment":"The motivation depends on the empirical claims that current scaling laws exhibit \"diminishing returns\" and that \"LLMs performance is nearing its ceiling under this paradigm.\" The only support offered is the vague statement that \"GPT-4's greatly increased scale over GPT-3 yielded diminishing per-token gains,\" with no citation or metric. The standard scaling-law metric is loss versus compute, not \"per-token gains,\" and the cited reference [25] reports power-law improvements rather than a ceiling. If this is a conjecture, it should be labeled as such; if it is an empirical claim, it needs data or references. This is load-bearing because the need for a new scaling paradigm rests on this premise.","section":"§2.1, Fig. 1"},{"comment":"The energy benefit claim is unsupported and appears physically questionable as stated. The paper asserts that redistributing computation to network edges \"harnesses wasted capacity and reduces carbon impact\" without modeling the energy efficiency of edge devices, the communication and transmission energy required for distributed training/inference, or the embodied energy of additional infrastructure. The supporting citation [58] is a black-hole physics paper (\"Hairy black holes in disformal scalar-tensor gravity theories\"), which does not address microgrids or data-center energy. A revision should either present a first-order energy model with quantitative assumptions or temper this claim to a speculative research direction with appropriate references.","section":"§4.3"},{"comment":"The conceptual framework is partly self-referential: \"ubiquitous intelligence\" is defined as the co-evolution of LLMs and wireless networks, and the same co-evolution is then presented as the source of the claimed benefits. The four principles (Continuous Intelligence Ascension, System-Orchestrated Intelligence Adaption, Permeable Semantic Networking, Ubiquitous Coherence Communications) are stated as mechanisms rather than as design goals or conjectures. As a position paper this is acceptable if clearly framed as a research vision, but the text should distinguish definitions, assumptions, and open hypotheses. For example, \"Continuous Intelligence Ascension\" is not established by any evidence; it is a target property of the proposed framework.","section":"§3.3"}],"minor_comments":[{"comment":"Caption contains typos: \"Experiece Scaling\" should be \"Experience Scaling\", and \"Compute Reource\" should be \"Compute Resource\".","section":"Fig. 1"},{"comment":"Grammar issues: \"cloud-based designs often ignores\" should be \"ignore\"; \"LLMs frameworks\" should be \"LLM frameworks\"; \"collective LLMs updates\" should be \"LLM updates\".","section":"§2.2, §3.2, §4.2.2"},{"comment":"\"Adaption\" appears in headings and principles; standard technical English uses \"Adaptation\". If intentional, please note; otherwise correct for consistency.","section":"§3.3"},{"comment":"Reference [58] is a physics paper on black holes and does not support the statement about microgrids with surplus generation. It should be replaced with a relevant energy-system or smart-grid reference, or removed.","section":"Ref. [58]"},{"comment":"The term \"experience\" is never formally defined. Does it mean raw sensor streams, interaction histories, feedback/reward signals, or synthetic data? A precise definition is needed before discussing collection, exchange, and scaling.","section":"§4.1.1"},{"comment":"The \"Efficient knowledge representation\" bullet is vague; it should explain how heterogeneous multimodal experience is intended to be tokenized, aligned, or otherwise made usable for LLM training. Currently it restates the problem without proposing a direction.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"This is a vision/position paper with no technical validation. Whether it is suitable for the journal depends on whether such papers are within scope. The main issues are fixable in revision by converting the strong claims into explicitly labeled conjectures, adding a concrete mechanism for experience scaling, and correcting the unsupported energy statement. The unrelated reference [58] should be caught in copyediting. I recommend major revision rather than rejection because the proposed research direction is timely and the open-challenge list provides a useful starting point, but the central load-bearing claims must be substantially reworked."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a vision paper, and it should be judged as one. What it does well: the writing is clear, the framing of LLM scaling as moving from pre-training to post-training to test-time and now 'experience scaling' is a useful conceptual diagram, and the paper is honest about its own gaps—Section 5 lists scalable experience exchange and global model consistency as unsolved. I also appreciate that it situates the data bottleneck against real work like Villalobos et al. and the energy argument is at least directionally motivated. As a call for a new research direction in network-driven lifelong learning for LLMs, it is a decent piece.\n\nWhere it is soft: there is no math, no data, no experiments, and—more importantly—the load-bearing assertion is never given a mechanism. The paper says LLMs should 'collect and create data directly from their environment' to bypass the human-data bottleneck, but it never specifies what 'experience' is, how it is tokenized, labeled, or converted into a training signal, or why that signal would preserve the distributional and alignment properties that make human data valuable. That is the decisive step, and it is skipped. The 'diminishing returns' of current scaling is asserted rather than quantified, and the energy-mitigation story is qualitative. The definitional circularity the reader flagged is real but not fatal: 'ubiquitous intelligence' is defined as the co-evolution of networks and models, so the touted benefits are partly built into the name. That is a rhetorical choice, not a technical flaw, but it does mean the paper's 'if true' impact is unquantified.\n\nOne concrete problem: reference [58] cites a hairy black hole paper to support the claim about microgrids with surplus generation and limited storage. That citation is wrong and should be fixed. It is the kind of error that makes you wonder how carefully the bibliography was checked, though the rest of the references look standard.\n\nMy take: this deserves peer review as a position paper for a workshop or a magazine, not as an archival technical paper. A serious referee would ask the authors to either add a proof-of-concept or feasibility model, or explicitly reframe the contribution as a research agenda. I would not cite it as a result, but I would bring it to a reading group as a starting point for a discussion on what 'experience scaling' would actually require. Send it to review, but expect heavy revision or a scope change.","headline":"A well-written vision paper that recombines known edge-AI ideas under a new label; the central claim about 'experience scaling' is plausible but lacks any mechanism or evidence, so it works as a position piece, not a technical contribution.","tokens_in":744,"tokens_out":792,"would_cite":false,"duration_ms":25481,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that LLMs can bypass the human-data bottleneck by learning directly from their environment through wireless networks, a stage it calls experience scaling.","keywords":["ubiquitous intelligence","experience scaling","LLM evolution","wireless networks","edge AI","lifelong learning","6G","distributed LLM inference"],"falsifier":"A controlled comparison: take the same base LLM, fine-tune one copy only on human-curated data and another only on environmental experience gathered by the model itself over wireless links, with matched token counts and compute. If the experience-fed model does not match or beat the human-data model on held-out tasks, or if the communication overhead grows without corresponding capability gain, the central premise of experience scaling fails.","tokens_in":11196,"feed_emoji":"📡","tokens_out":6608,"duration_ms":68842,"temperature":0.7,"pith_summary":"This paper argues that the dominant ways of making large language models smarter—pre-training on massive text, post-training on human feedback, and spending more compute at test time—are all hitting the same wall: human-produced data cannot grow fast enough to feed them. It proposes a fourth stage, called experience scaling, in which LLMs collect and create their own training signal by acting inside the real world, with wireless networks as the channel that gathers, routes, and synchronizes that experience. The result would be ubiquitous intelligence: models that keep improving continuously across edge devices and base stations instead of being frozen between centralized retraining rounds. The co-evolution cuts both ways—networks enable model growth, and models make networks semantically aware and self-organizing. The paper is a research agenda, not an experimental result, and it names communication load and model consistency as the open problems that will decide whether the vision holds.","feed_headline":"Next LLM scaling stage: learning from the environment, not human data","feed_subtitle":"The paper proposes wireless networks as the scaffolding for LLMs to gather firsthand experience and keep improving.","key_machinery":"The load-bearing object is the co-evolution loop between the wireless network and the LLM, captured by the paper's term 'ubiquitous intelligence.' On one side, network mechanisms—device-to-device links, edge caching, spectrum-aware synchronization, and hierarchical orchestration across small cells, macro cells, and cloud—collect and propagate experience for lifelong learning. On the other side, LLM-driven semantic awareness lets the network exchange meaning rather than raw data and schedule resources by context. The loop is the scaling mechanism: experience flows in, refined models flow out, and the network's orchestration determines how much capability each environment can add.","core_discovery":"The central claim is that LLM scaling has three recognized phases—pre-training, post-training, and test-time scaling—and that each depends on human-generated or human-annotated data, whose supply is running out. The paper's discovery is that a fourth phase is available in principle: let models interact directly with environments through wireless infrastructure, treat those interactions as experience, and feed that experience back into model refinement. This is 'experience scaling.' It turns the network from a pipe into a distributed learning organ: devices, base stations, and edge servers collect heterogeneous multimodal experience, share it with local and global models, and receive updated","pith_inferences":["Extension: if experience scaling is taken seriously, the scarce resource shifts from data tokens to communication bandwidth and experience quality, so network optimization should be framed as maximizing model capability gain per byte, a metric the paper does not define.","Extension: experience data will need filters or reward signals for novelty, diversity, and safety—parallel to the curation filters used for human text—otherwise uncontrolled environmental data could degrade models or amplify bias.","Extension: the cleanest near-term test is a controlled deployment where matched models are fine-tuned on wireless-collected experience versus human-curated data of equal size; if experience-trained models do not at least match, the scaling story loses its empirical basis.","Extension: the co-evolution logic transfers to other sensor-rich settings—robotics, smart cities, industrial IoT—where the same network-to-model loop could be evaluated with existing infrastructure."],"forward_implications":["LLM development could shift from discrete releases to continuous field updates, with refinements flowing over device-to-device and edge links between retraining rounds.","Wireless resource management—bandwidth, scheduling, spectrum—becomes a direct determinant of model improvement, not just a service-delivery constraint.","Personalization and privacy improve because experience is gathered and processed near users, reducing reliance on centralized cloud data collection.","Task offloading and distributed cached knowledge let heterogeneous devices share reasoning without full retraining, lowering latency for time-sensitive services.","The paradigm's own open challenges—scalable experience exchange and global model consistency—must be solved before the loop can operate at scale."],"supporting_citations":[{"why":"Supplies the evidence that human-generated data is being exhausted, which is the bottleneck motivating experience scaling.","marker":"[46]"},{"why":"Establishes the pre-training scaling-law relation among compute, data, and model size, the first scaling stage the paper argues is saturating.","marker":"[25]"},{"why":"Anchors the post-training RLHF stage and shows human feedback as a limited, labor-intensive data source.","marker":"[28]"},{"why":"Demonstrates edge inference for LLMs over wireless networks, the deployment substrate the experience-scaling loop depends on.","marker":"[14]"},{"why":"Documents cloud-centric industrial LLM deployment, the centralized baseline that the paper argues is insufficient for latency, privacy, and personalization.","marker":"[4]"},{"why":"Supports the LLMs-for-networks side by showing how generative AI can optimize next-generation wireless systems.","marker":"[17]"},{"why":"Provides the spectrum and resource management techniques that distributed learning and experience exchange would rely on.","marker":"[37]"},{"why":"Supplies the decentralized-learning-over-wireless setting and the consistency problem that the paper lists as an open challenge for global model alignment.","marker":"[60]"}],"fun_headline_variants":["Experience scaling: the fourth LLM scaling phase","Wireless networks turn LLMs into lifelong learners","LLMs learn from environment via wireless, not just data","Network-driven experience scaling for adaptive LLMs","Wireless-driven co-evolution makes LLMs self-improving"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The argument rests on the premise that experiences a model collects by acting in the environment, transmitted through wireless links, can serve as a substitute for human-generated data in making the model more capable—and that the traffic and synchronization this creates can be made tractable.","fun_headline_variants_meta":{"raw":{"variants":["Experience scaling: the fourth LLM scaling phase","Wireless networks turn LLMs into lifelong learners","LLMs learn from environment via wireless, not just data","Network-driven experience scaling for adaptive LLMs","Wireless-driven co-evolution makes LLMs self-improving"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000583,"raw_usage":{"total_tokens":2495,"prompt_tokens":575,"completion_tokens":1920,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":319,"completion_tokens_details":{"reasoning_tokens":1855}},"tokens_in":319,"tokens_out":1920,"duration_ms":12648,"temperature":1.0,"reasoning_tokens":1855,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T20:38:21.001709+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled comparison: take the same base LLM, fine-tune one copy only on human-curated data and another only on environmental experience gathered by the model itself over wireless links, with matched token counts and compute. If the experience-fed model does not match or beat the human-data model on held-out tasks, or if the communication overhead grows without corresponding capability gain, the central premise of experience scaling fails.","supporting_citations":[],"review_version":1}