{"id":"ae75019e-afb0-40bb-9542-adc8ebee2836","arxiv_id":"2508.09159","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Agoran uses AI agents with legislative, executive, and judicial branches to automatically negotiate and manage 6G network slices, achieving large performance gains on a 5G testbed.","lead":"This paper introduces Agoran, an agentic marketplace for 6G network slicing with three AI branches, and reports testbed gains in throughput, latency, and resource usage. It also shows a small fine-tuned language model can approximate a much larger model's negotiation decisions, which could make such systems practical.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full text mismatch: submitted body is a time-series paper, so Agoran's performance and LLM claims have no supporting methodology and cannot be verified.","rationale":"We read the abstract in good faith as the sole evidence of Agoran. The abstract asserts large, specific performance improvements and a surprising LLM fine-tuning result. For these claims to hold, the manuscript would need to describe the system, the testbed, the evaluation protocol, and the metrics. Instead, the supplied full text is a completely different paper on time-series analysis (JustDense). This is not a mere missing detail; it means the central claim has no evidential support whatsoever. The reader's UNVERDICTED verdict is therefore correct, and no change is needed. However, the reader's weakest_assumption singles out the single-round consensus mechanism. While that is indeed a fragile assumption—one-shot Pareto-optimal negotiation requires full preference revelation and immediate controller acceptance—it is downstream of the absence of any methodology. We partially agree with the reader because the single-round consensus would be a critical assumption if the full text were present, but the primary load-bearing problem is the mismatched manuscript. The concrete test we propose settles this by checking whether the body contains the Agoran system at all. If it does not, the debate over single-round consensus is moot; the paper cannot be verified. If a corrected full text were provided, the single-round consensus would then be the key assumption to stress-test.","tokens_in":14774,"tokens_out":7075,"duration_ms":75980,"concrete_test":"Fetch the source of arXiv:2508.09159 from arXiv and grep the full text for 'Agoran', 'SRB', 'eMBB', 'URLLC', 'PRB', '5G', 'RAN', and 'negotiation' in the main body and appendices. If any of these key terms is absent, the submission is mismatched and no claim in the abstract can be verified; the verdict must remain UNVERDICTED.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The submitted full text is 'JustDense: Just using Dense instead of Sequence Mixer for Time Series analysis' (arXiv:2508.09153v1), which does not describe Agoran, the SRB, negotiation agents, Open/AI RAN controllers, 5G testbeds, or vehicle mobility traces. The abstract's central claim—37% eMBB throughput increase, 73% URLLC latency reduction, 8.3% PRB saving, and an 80% GPT-4.1 decision-quality recovery by a 1B Llama model—appears nowhere in the body. There is no system architecture, no algorithm, no experimental protocol, no baseline definition, no error bars, and no code. Consequently, the central claim is an unsupported assertion. The reader's UNVERDICTED rating is thus not merely cautious but necessary. Among the abstract's internal assumptions, the 'consensus intent in a single round' is also unmotivated: reaching a Pareto-optimal offer via multi-objective optimization in one round requires every stakeholder to reveal complete utility information in a single message, and the abstract does not explain why such an intent would be acceptable to controllers or stakeholders. If the single-round consensus fails, the end-to-end marketplace loop breaks. But the first-order problem is that no evidence for any of this is present.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript submitted under arXiv:2508.09159 (cs.NI) presents Agoran, an agentic marketplace for 6G RAN automation, in its abstract. The abstract describes a three-branch AI architecture (Legislative, Executive, Judicial), stakeholder-side negotiation agents, a mediator agent, a multi-objective optimizer, and deployment on a private 5G testbed, with claimed gains of 37% in eMBB throughput, 73% in URLLC latency reduction, and 8.3% in PRB usage, plus a fine-tuned 1B Llama model recovering 80% of GPT-4.1's decision quality. However, the full text supplied with the submission is a different paper, 'JustDense: Just using Dense instead of Sequence Mixer for Time Series analysis' (arXiv:2508.09153v1), which contains no description of Agoran, its architecture, its experimental protocol, its baseline, or any of the claimed results. The central claims of the abstract therefore have no supporting methodology or evidence in the submitted body.","tokens_in":15095,"tokens_out":2943,"duration_ms":36385,"significance":"If the Agoran results were fully documented, the work could be of interest to the RAN automation community as a concrete proposal for stakeholder-driven slice negotiation with LLM-based compliance checking and trust management. The claimed end-to-end gains on a private 5G testbed and the resource-efficient LLM compression result would be useful if reproducible. However, in the current submission, none of this is verifiable: the full text is an unrelated time-series study, and no architecture, algorithms, experimental setup, baseline details, error bars, or code are provided for Agoran. The paper therefore cannot currently make a contribution to the literature, and the significance assessment is conditional on a complete resubmission that matches the abstract.","major_comments":[{"comment":"The submitted full text is a completely different manuscript, 'JustDense: Just using Dense instead of Sequence Mixer for Time Series analysis' (arXiv:2508.09153v1), and contains no mention of Agoran, the Service and Resource Broker, negotiation agents, Open/AI RAN controllers, 5G testbeds, vehicle mobility traces, or any of the claimed experimental results. Consequently, none of the central claims in the abstract — 37% eMBB throughput increase, 73% URLLC latency reduction, 8.3% PRB saving, and 80% GPT-4.1 decision-quality recovery — have any supporting methodology or evidence in the body. This is a load-bearing defect that prevents verification of the paper's central claims.","section":"Full Text (all sections)"},{"comment":"The abstract asserts that Negotiation Agents and the Mediator Agent 'reach a consensus intent in a single round' after multi-objective optimization, but the submission provides no mechanism, formal description, or experimental evidence for this convergence. In particular, there is no discussion of the information stakeholders must reveal, the conditions under which a Pareto-optimal offer is acceptable to all parties, or what happens if the Open and AI RAN controllers reject the negotiated intent. Since the end-to-end loop depends on this single-round consensus, this unstated assumption is load-bearing and cannot be assessed from the submitted text.","section":"Abstract (Agoran architecture)"},{"comment":"The quantitative results (37%, 73%, 8.3%, and the LLM's 80% recovery) are presented without any experimental protocol: there is no description of the private 5G testbed, the baseline configuration, the traffic models, the number of runs, the variance or confidence intervals, or the exact metrics used. The LLM claim additionally lacks details of the fine-tuning procedure, the evaluation set, the notion of 'decision quality', and the comparison to GPT-4.1. These omissions make the reported numbers neither reproducible nor statistically assessable.","section":"Abstract (experimental claims)"},{"comment":"The claim that a 1B-parameter Llama model fine-tuned for five minutes on 100 GPT-4 dialogues recovers approximately 80% of GPT-4.1's decision quality is unsupported by any methodology. No information is given about the dialogue generation process, the fine-tuning objective, the hardware used for the 1.3-second convergence figure, the 6 GiB memory measurement, or the evaluation benchmark. As stated, this is an isolated assertion with no way to verify or compare it.","section":"Abstract (LLM compression)"}],"minor_comments":[{"comment":"The phrase 'standards-aligned' is used without naming any standard or describing the alignment procedure; the reader cannot determine which 3GPP or O-RAN specifications are being referenced.","section":"Abstract"},{"comment":"The terms 'Legislative branch', 'Executive branch', and 'Judicial branch' are introduced without definitions, examples, or pseudocode, making the architecture hard to follow even before the missing full-text support.","section":"Abstract"},{"comment":"The live demo URL is mentioned but no demonstration content, duration, or behavior is described; a video link cannot substitute for an experimental section.","section":"Abstract"},{"comment":"The submitted body's title, authors, and content are inconsistent with the arXiv metadata and abstract of the Agoran paper, which is a presentation-level issue that also reflects the lack of a coherent manuscript.","section":"Full Text (JustDense)"}],"recommendation":"reject","confidential_remarks":"The mismatch between the abstract and the full text is not a normal revision issue: the body is an entirely different paper with no relation to the claimed subject. Even setting aside any concerns about submission integrity, the central claims are completely unsupported within this manuscript, so a major revision cannot repair the submission without rewriting it from scratch. I would advise the editor to return the manuscript and request a correct full-text submission before any further review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline is that the abstract describes an agentic marketplace for 6G RAN automation with real deployment numbers, but the full text is a completely different paper about dense layers in time-series models. So as submitted, the paper is not reviewable: every performance claim in the abstract—37% throughput gain, 73% latency reduction, 8.3% PRB saving, 80% GPT-4.1 decision recovery by a 1B Llama—appears nowhere in the body. There is no architecture, no experimental protocol, no baseline definition, no error bars, no code. I can't verify a single number.\n\nWhat is genuinely interesting, if the abstract is accurate, is the architectural pattern: three AI branches (Legislative, Executive, Judicial) with RAG-LLMs, a rule-based trust score, and a mediator that negotiates Pareto-optimal offers between stakeholders. That combination is new to me in RAN control. The small-LLM distillation result—fine-tuning a 1B model for five minutes on 100 GPT-4 dialogues to recover 80% of decision quality—is the kind of concrete, falsifiable claim that would matter if backed by a real testbed.\n\nThe soft spots are large. First, the full-text mismatch is fatal for peer review. A serious editor would desk reject this submission on the spot; the body has nothing to do with the abstract. Second, even taking the abstract on its own, the 'consensus intent in a single round' is an unmotivated assumption. Multi-objective negotiation with complete utility revelation in one round is a strong condition, and the abstract doesn't say what happens if the controllers reject the intent. Third, there are no references in the abstract, so novelty cannot be checked. Self-citation is not the issue; there is simply nothing to benchmark against.\n\nFor whom is this useful? For a reader working on LLM-driven RAN control, the abstract is a one-paragraph teaser that might justify a closer look at a later, complete version. But this submission, in its current form, deserves no referee time. The right move is to return it as incomplete and ask for the actual manuscript.","headline":"Interesting abstract, but the submitted body is a different paper, so the claims are unsupported and the submission is not refereable.","tokens_in":15636,"tokens_out":2350,"would_cite":false,"duration_ms":25187,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper's central claim is that Agoran's three-branch agent marketplace reconciles conflicting slice objectives at run time, and its testbed results show a 37% eMBB throughput gain, a 73% URLLC latency cut, and an 8.3%…","keywords":["agentic marketplace","6G RAN automation","network slicing","multi-objective optimization","LLM-based negotiation","Open RAN","trust scoring","service and resource broker"],"falsifier":"Re-run the same 5G testbed with deliberately conflicting slice objectives—for example, a URLLC latency target that is infeasible under current radio conditions alongside an eMBB throughput demand that exhausts nearly all physical resource blocks—and count how many negotiation rounds end without an accepted consensus intent. A substantial rejection or renegotiation rate would show the single-round mechanism is not robust.","tokens_in":14609,"feed_emoji":"📡","tokens_out":9455,"duration_ms":91272,"temperature":0.7,"pith_summary":"This paper tries to establish that the business layer of a mobile network—the service owners with conflicting goals—can be brought into the operating loop of radio access network (RAN) slicing, instead of being locked out by rigid policy-bound controllers. It proposes Agoran, an agentic marketplace with three autonomous branches: a Legislative branch that answers compliance queries with retrieval-augmented language models, an Executive branch that maintains real-time situational awareness through a watcher-updated vector database, and a Judicial branch that scores trust and arbitrates suspicious behavior. Stakeholder-side negotiation agents and a mediator use a multi-objective optimizer to produce Pareto-optimal offers and reach a single consensus intent, which is then deployed to Open and AI RAN controllers. On a private 5G testbed with realistic vehicle-mobility traces, the authors report a 37% increase in eMBB slice throughput, a 73% reduction in URLLC slice latency, and an 8.3% end-to-end saving in physical resource block usage versus a static baseline. A 1B-parameter Llama model fine-tuned for five minutes on 100 GPT-4 dialogues is reported to recover about 80% of GPT-4.1's decision quality while running within 6 GiB of memory and converging in 1.3 seconds.","feed_headline":"AI marketplace lifts 6G slice throughput 37%, cuts latency 73%","feed_subtitle":"Stakeholder agents negotiate directly with RAN controllers on a 5G testbed, saving 8.3% of physical resource blocks.","key_machinery":"The load-bearing mechanism is the Agoran Service and Resource Broker (SRB), a three-branch agentic marketplace. The Legislative branch uses retrieval-augmented large language models to answer compliance queries; the Executive branch keeps a watcher-updated vector database as its situational-awareness store; and the Judicial branch scores each agent message with a rule-based Trust Score, with arbitrating LLMs detecting malicious behavior and applying real-time incentives to restore trust. The economic core is a multi-objective optimizer that generates Pareto-optimal offers for the stakeholder-side Negotiation Agents and the SRB-side Mediator Agent, whose single-round consensus intent is the artifact actually deployed to Open and AI RAN controllers.","core_discovery":"The central claim is that conflicting service-owner objectives can be reconciled at run time through an agentic marketplace rather than through static, policy-bound slice control. In Agoran, the negotiation loop is closed by a single-round consensus intent: the mediator and the stakeholder-side agents agree on one feasible, Pareto-optimal offer, and that intent is pushed directly to Open and AI RAN controllers. The paper's evidence is the private 5G testbed evaluation—eMBB throughput up 37%, URLLC latency down 73%, and end-to-end PRB usage down 8.3% against a static baseline—together with the demonstration that a 1B-parameter Llama model can reproduce about 80% of the larger model's decision quality in 1.3 seconds.","pith_inferences":["If the single-round consensus assumption is stress-tested with intentionally incompatible slice demands, the mediator will sometimes need a second round or must drop an offer; the paper does not report those failure rates.","The same three-branch marketplace pattern could apply to other shared infrastructures—cloud schedulers, spectrum brokers, or energy grids—where multiple owners bid for hard-latency resources.","The 80% quality recovery at 1B parameters suggests frontier-model distillation may lower the operational cost of RAN automation, but the paper does not evaluate it under non-stationary traffic or adversarial agents.","Since trust scoring is rule-based, the judicial branch's incentives are only as good as those rules; a learned trust model is a natural extension the paper leaves unexplored."],"forward_implications":["Run-time reconciliation of slice owners becomes possible: the negotiation loop is designed to be fast enough for mobility-driven demand changes, with convergence reported at 1.3 seconds.","Small, cheaply trained models can substitute for frontier models in RAN arbitration, since the 1B-parameter model recovers roughly 80% of the larger model's decision quality.","The marketplace's consensus intents are deployable to Open and AI RAN controllers, so the approach is a standards-aligned evolution path rather than a proprietary overlay.","The judicial trust loop gives operators a concrete way to detect malicious agents and steer them back to cooperative behavior with incentives, which is a prerequisite for opening the RAN to outside stakeholders.","The end-to-end 8.3% PRB saving means shared spectrum and radio resources can be used more efficiently while service owners negotiate, not just at the scheduling layer."],"supporting_citations":[],"fun_headline_variants":["Agentic 6G broker boosts throughput 37%, cuts latency 73%","AI marketplace for 6G slices: +37% throughput, -73% latency","Agoran: agent-negotiated 6G slices gain 37% throughput, cut latency 73%","Negotiating 6G slices: AI broker lifts throughput 37%, cuts latency 73%","Marketplace for 6G slices: agents negotiate gains of 37% throughput, 73% lower latency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the stakeholder-side negotiation agents and the mediator always reach a single consensus intent in one round, and that Open and AI RAN controllers accept and apply that intent directly; if that handshake fails or requires revision, the marketplace loop and the reported gains break down.","fun_headline_variants_meta":{"raw":{"variants":["Agentic 6G broker boosts throughput 37%, cuts latency 73%","AI marketplace for 6G slices: +37% throughput, -73% latency","Agoran: agent-negotiated 6G slices gain 37% throughput, cut latency 73%","Negotiating 6G slices: AI broker lifts throughput 37%, cuts latency 73%","Marketplace for 6G slices: agents negotiate gains of 37% throughput, 73% lower latency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000959,"raw_usage":{"total_tokens":4154,"prompt_tokens":1080,"completion_tokens":3074,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":696,"completion_tokens_details":{"reasoning_tokens":2950}},"tokens_in":696,"tokens_out":3074,"duration_ms":19588,"temperature":1.0,"reasoning_tokens":2950,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:28:36.287069+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same 5G testbed with deliberately conflicting slice objectives—for example, a URLLC latency target that is infeasible under current radio conditions alongside an eMBB throughput demand that exhausts nearly all physical resource blocks—and count how many negotiation rounds end without an accepted consensus intent. A substantial rejection or renegotiation rate would show the single-round mechanism is not robust.","supporting_citations":[],"review_version":1}