{"id":"09f42d80-f79c-4dcc-9a52-a850f63e80ee","arxiv_id":"2501.07288","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LLM-Net, a blockchain-based network of specialized LLM providers with a text-based reputation system, is proposed and illustrated with a single truncated debate simulation that does not validate the stated quality claims.","lead":"LLM-Net is a proposed blockchain-based platform that would let many specialized AI models serve user queries, with a reputation system that records who performs well. The only experiment shown is one simple math question, which is too little to support the paper's claim that the system's quality mechanism is validated.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulation's selection signal is injected through hand-assigned intelligence labels; without independent, non-gameable quality assessment, the reputation mechanism is not validated.","rationale":"The reader's weakest_assumption correctly identifies Quality Discernment as the load-bearing premise. My stress-test confirms this and sharpens the point: the simulation does not merely rely on evaluators' discernment, it injects the quality signal directly into the respondents' prompts. The intelligence indices 0.1, 0.5, and 0.8 are not discovered by the reputation mechanism; they are inputs that cannot be observed in a real decentralized network. The coordinator's selection thus recovers the label rather than validates a mechanism. This is a stronger objection than a generic concern about evaluator reliability: it shows the single reported experiment is, at best, a prompt-following demonstration, not a test of the proposed system. The paper does have value as an architecture proposal, and the reputation mechanism could in principle work if evaluations are reliable, but no evidence is provided for that key premise. The reader's REJECT verdict is appropriate; my analysis reinforces it without changing it.","tokens_in":9986,"tokens_out":2439,"duration_ms":26118,"concrete_test":"Re-run the simulation on a benchmark with objective correctness labels (e.g., 100 GSM8K or domain QA questions) using identical underlying models for all respondents, with no explicit intelligence-index prompts. Run peer evaluation and coordinator selection exactly as in Section 3.3, then compare the accuracy of selected respondents' answers against excluded respondents and against random selection. Additionally, inject a colluding subgroup that agrees to rate each other highly regardless of quality, and check whether the coordinator's selections are robust to this manipulation. If selected-set accuracy is not meaningfully above random, or if collusion flips the ranking, the Quality Discernment premise fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central validation claim rests on Section 4's claim that the simulation 'validates the effectiveness of the reputation-based mechanism in maintaining service quality.' The experiment, however, is structured so that the outcome is determined before any evaluation occurs. Respondents are given explicit persona prompts in Section 4: 'Debater 1: 0.1 (low intelligence - basic understanding, simple logic)', 'Debater 2: 0.5 (medium intelligence - good understanding, moderate analysis)', and 'Debater 3: 0.8 (high intelligence - expert understanding, complex analysis)'. These hand-assigned labels are then reflected in the verbosity and content of the responses, and the same LLM family writes the peer-evaluation matrices (Tables 6-9) and the coordinator's summary (Table 10). The coordinator's decision to exclude Respondent 1 is therefore not evidence that the reputation mechanism can discern quality; it is evidence that a model can recognize and echo a prompted persona. This fails the Quality Discernment assumption (Section 3.3) in the only way the simulation could test it. The demonstration uses a single trivial query on which every respondent answers correctly, includes no baseline, no statistical comparison, no ablation without the persona labels, and no test of adversarial behavior such as colluding respondents inflating each other's reviews. Since the claimed mechanism must work when quality is not pre-labeled and when respondents can strategically game feedback, the simulation offers no support for the load-bearing claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes LLM-Net, a blockchain-based framework for democratizing LLMs-as-a-Service through a decentralized network of specialized LLM providers. The architecture defines requesters, coordinators, respondents, and validators, with a reputation mechanism that records qualitative interaction feedback on a blockchain and uses coordinator judgment to select high-performing respondents for future queries. The only evaluation is a Section 4 simulation of a multi-agent debate on a single question ('What is the smallest prime number after 60?') run on four LLMs, where respondents are assigned explicit intelligence indices (0.1, 0.5, 0.8). The simulation shows that the low-index respondent is excluded from future queries based on peer evaluation and coordinator summary. The paper claims this validates the reputation mechanism's effectiveness in maintaining service quality.","tokens_in":10332,"tokens_out":3015,"duration_ms":28757,"significance":"If the central claim were established, LLM-Net would contribute to a timely and important problem: maintaining quality and accountability in decentralized, blockchain-based LLM service markets. The paper is clearly written and the architecture is easy to follow, with a sensible separation of node roles and a concrete proposal for text-based reputation records. However, the evidence provided does not establish the claim. The simulation is circular because the quality signal is injected through hand-assigned persona labels, and the same LLM family generates the responses, the peer evaluations, and the coordinator's selection decision. No baseline, repeated trials, quantitative metric, or adversarial test is provided. The paper therefore currently offers a framework description and a qualitative demonstration that LLMs follow prompted personas, not a validation of the reputation mechanism.","major_comments":[{"comment":"The simulation is structured so that the outcome is predetermined by the hand-assigned intelligence indices. Respondents are explicitly prompted as 'Debater 1: 0.1 (low intelligence)', 'Debater 2: 0.5', and 'Debater 3: 0.8', and the subsequent responses, peer evaluation matrices (Tables 6-9), and coordinator summary (Table 10) simply reflect these prompted personas. The coordinator's decision to exclude Respondent 1 is therefore evidence that the LLM can recognize and echo a prompted persona, not that the reputation mechanism can discern response quality. The abstract's claim that the simulation 'validates the effectiveness of the reputation-based mechanism' is unsupported.","section":"Section 4, Tables 2-10"},{"comment":"The evaluation is entirely qualitative and lacks any statistical or comparative rigor. There is no baseline condition without the reputation mechanism, no repeated trials, no quantitative quality metric, and no error analysis. The peer evaluation matrices are anecdotal text outputs, and the coordinator's summary is a free-text narrative rather than a reproducible decision rule. Consequently, the paper does not demonstrate that the reputation mechanism improves quality over random selection, a fixed average, or any alternative selection policy.","section":"Section 4"},{"comment":"The simulation does not test the load-bearing Quality Discernment assumption. All nodes use the same LLM family, are not given incentives to disagree, and no adversarial behavior (such as colluding respondents inflating each other's reviews or validators being compromised) is considered. The claim that nodes 'possess sufficient analytical capabilities to detect and evaluate variations in response quality' is asserted but never demonstrated in a setting where quality is not pre-labeled by the experimenter. Until this assumption is tested with independent ground truth or a credible adversarial setup, the central mechanism remains unvalidated.","section":"Section 3.3, Quality Discernment assumption"}],"minor_comments":[{"comment":"The phrase 'the followings elaborate the types of nodes in detail' should be 'the following elaborates the types of nodes in detail' or similar, as 'followings' is nonstandard.","section":"Section 3.1"},{"comment":"There is a typo in the Cycle 2 entry: 'Respondet' should be 'Respondent'.","section":"Table 3"},{"comment":"The caption uses 'requestor', while the text throughout uses 'requester'; please make the terminology consistent.","section":"Figure 2 caption"},{"comment":"The sentence beginning 'This strategy has explored in various studies such as [11–13], enables multiple LLMs...' has a grammatical error; it should read 'This strategy has been explored in various studies such as [11–13] and enables multiple LLMs...'.","section":"Section 2.2"}],"recommendation":"reject","confidential_remarks":"The paper is clearly written and the framework is plausible at a high level, but the central validation is a circular demonstration. The authors would need to redesign the experiments substantially—using ground-truth quality labels, baselines, repeated trials, and adversarial scenarios—to support the claimed effectiveness of the reputation mechanism. As submitted, the manuscript is not ready for publication in a serious venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real novelty here is small but genuine: a blockchain-based LLM marketplace where reputation is kept as raw text interaction records rather than numeric scores, combined with multi-agent debate and fine-tuned expert models. The architecture is clearly laid out, the four assumptions in Section 3.3 are explicit, and the authors do admit that the simulation is a demonstration, not a deployment. That is worth something.\n\nBut the central claim — that the reputation mechanism 'validates' quality maintenance — does not survive contact with Section 4. The respondents are literally prompted to be low/medium/high intelligence (0.1, 0.5, 0.8), and the peer evaluations and coordinator summaries read back those personas. The same LLM family produces the responses, the reviews, and the selection decision, so the only thing the simulation shows is that a model can echo a prompted label. There is one trivial query, no baseline, no repeated trials, no ablation without the persona prompts, and no test of collusion or strategic gaming. The Quality Discernment assumption is the load-bearing wall, and the simulation never tests it. The stress-test note is right: the selection signal is injected before any evaluation happens.\n\nThe paper also skips prior decentralized-LLM-marketplace work, which weakens the positioning but is not itself fatal. What is fatal is the overreach in the abstract and conclusion, where the simulation is said to validate the mechanism. That is not supported by the data presented.\n\nWho gets value from this? Someone thinking through the design space of decentralized AI services will find the text-reputation idea worth considering, and the explicit assumptions give a useful checklist for what a real system would need. But the empirical section should be treated as a toy illustration, not evidence.\n\nI would send this to peer review only because the design idea is plausible enough to deserve constructive revision — not because the current claims hold. The authors need to either dramatically soften the validation language or, better, run serious experiments: multiple tasks, controlled quality, independent evaluation, and adversarial scenarios. On the current evidence, reject the claim as stated, but engage with the architecture.\n\nRecommendation: invite major revision as a position paper, and push hard for real evaluation or a humbler title.","headline":"A coherent architecture sketch undercut by a simulation whose outcome is preordained by hand-assigned intelligence labels; no real evidence for the reputation mechanism.","tokens_in":10780,"tokens_out":1138,"would_cite":false,"duration_ms":14578,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a blockchain reputation mechanism, using peer feedback recorded immutably, lets decentralized coordinators exclude weak or biased LLM providers, and a four-model simulation shows consistent exclusion of the lowest…","keywords":["LLM-Net","blockchain","decentralized AI","reputation mechanism","LLM-as-a-service","multi-agent debate","peer evaluation","quality selection"],"falsifier":"Run the same multi-agent debate with a query in a domain where the evaluator itself is weak, such as a specialized medical or legal question, and have the low-intelligence respondent give a confidently wrong answer that the evaluator rates highly; if the coordinator keeps selecting that respondent for the next query, the claim that reputation-based selection maintains quality is refuted.","tokens_in":9745,"feed_emoji":"🤖","tokens_out":6046,"duration_ms":52221,"temperature":0.7,"pith_summary":"LLM-Net is a proposed blockchain-based marketplace in which specialized language-model providers answer queries through a coordinated multi-agent debate, and the paper's central claim is that a reputation mechanism built on peer feedback and an immutable blockchain record lets coordinators keep service quality high by selecting strong respondents and excluding weak or biased ones. The authors argue this matters because LLM development today is concentrated in a few companies and limited by scarce training data, so a decentralized network of fine-tuned expert models could sustain AI progress without a central authority. Their simulation, run with Claude 3.5 Sonnet, Llama 3.1, Grok-2, and GPT-4o, shows the coordinator excluding the lowest-performing respondent in every case, which they read as validation of the reputation mechanism.","feed_headline":"Blockchain reputation system filters out weak LLM providers","feed_subtitle":"Simulation across four LLMs shows low-performing respondents being excluded without a central authority.","key_machinery":"The load-bearing mechanism is the reputation-based selection loop: respondents evaluate each other's contributions in text, validators record all interactions and feedback in immutable blockchain blocks, and the coordinator reads those records when choosing respondents for the next query. It is paired with multi-agent debate, a prompting strategy in which multiple LLM instances propose, critique, and refine answers through structured discussion. The mechanism works not by numerical reputation scores but by preserving qualitative assessments that later coordinators can interpret.","core_discovery":"On its own terms, the paper establishes that when a coordinator selects respondents based on qualitative peer evaluations stored on a blockchain, a deliberately weak respondent, prompted to reason at intelligence level 0.1, is identified as low-contributing and biased and is excluded from the next query, consistently across four different underlying LLMs. The debates over 'what is the smallest prime number after 60?' converge on the correct answer in all four models, but the peer evaluation matrices and the coordinator's summary expose Respondent 1 as making trivial or biased contributions. The authors conclude that the reputation-based selection mechanism maintains service quality by removing low-performing or biased respondents, thereby enabling decentralized expert networks to collaborate effectively.","pith_inferences":["A natural extension would be to test whether the exclusion decision holds when the evaluator LLM differs from the respondent LLMs, since in this simulation the same model family generates the responses, the peer reviews, and the coordinator's decision.","The qualitative-text reputation record could be augmented with a quantitative score for large networks, though the paper deliberately avoids scores to preserve interpretability.","The same selection loop could be applied to other decentralized AI services, such as data labeling or model routing, wherever an AI evaluator can judge output quality."],"forward_implications":["Coordinators can filter out low-performing or biased providers without a central authority, keeping future collaborations higher quality.","Blockchain records give requesters and validators an auditable history of each respondent's contributions and rewards.","Domain-specialized fine-tuned models can be offered as services and maintained collectively, reducing dependence on a single company's model and data.","Collaborative prompting among multiple respondents can reach correct consensus even when individual respondents have limited capability.","The demonstrated exclusion decision suggests reputation-based selection generalizes across different underlying LLMs."],"supporting_citations":[{"why":"Motivates the problem by estimating when public training data will be exhausted, the scarcity LLM-Net addresses.","marker":"[3]"},{"why":"Presents retrieval-augmented generation as the existing response to knowledge freshness that LLM-Net extends toward decentralized expertise.","marker":"[4]"},{"why":"Supplies the RAG architecture background the paper contrasts with its distributed expert-network approach.","marker":"[7]"},{"why":"Provides the multi-agent debate prompting strategy that the simulation uses for collaborative problem-solving.","marker":"[11-13]"},{"why":"Grounds the blockchain ledger concept that gives LLM-Net an immutable record of service delivery.","marker":"[14]"},{"why":"Establishes the smart contract mechanism used to automate agreements and reward distribution.","marker":"[15, 16]"}],"fun_headline_variants":["Blockchain peer reviews pick best LLMs, drop weak ones","Decentralized LLM network uses blockchain to oust poor performers","LLM-Net: Blockchain vetting keeps AI experts honest","Reputation on blockchain filters out weak LLM providers","Peer-evaluated LLMs: Blockchain keeps the experts in line"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole mechanism depends on the assumption that the LLM-based coordinator and validators can reliably tell good responses from bad ones and that colluding or biased respondents cannot game their judgments.","fun_headline_variants_meta":{"raw":{"variants":["Blockchain peer reviews pick best LLMs, drop weak ones","Decentralized LLM network uses blockchain to oust poor performers","LLM-Net: Blockchain vetting keeps AI experts honest","Reputation on blockchain filters out weak LLM providers","Peer-evaluated LLMs: Blockchain keeps the experts in line"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1206,"prompt_tokens":938,"completion_tokens":268,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":182}},"tokens_in":554,"tokens_out":268,"duration_ms":3602,"temperature":1.0,"reasoning_tokens":182,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:45:13.996489+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same multi-agent debate with a query in a domain where the evaluator itself is weak, such as a specialized medical or legal question, and have the low-intelligence respondent give a confidently wrong answer that the evaluator rates highly; if the coordinator keeps selecting that respondent for the next query, the claim that reputation-based selection maintains quality is refuted.","supporting_citations":[{"cited_title":"Position: Will we run out of data? limits of LLM scaling based on human-generated data,","cited_arxiv_id":null,"evidence_quote":"Motivates the problem by estimating when public training data will be exhausted, the scarcity LLM-Net addresses."},{"cited_title":"Retrieval-augmented generation for knowledge-intensive nlp tasks,","cited_arxiv_id":null,"evidence_quote":"Supplies the RAG architecture background the paper contrasts with its distributed expert-network approach."},{"cited_title":"An overview of blockchain technology: applica- tions, challenges and future trends,","cited_arxiv_id":null,"evidence_quote":"Grounds the blockchain ledger concept that gives LLM-Net an immutable record of service delivery."}],"review_version":1}