{"id":"c1e612c9-4529-4916-9fbe-dcdf6859e04d","arxiv_id":"2502.09095","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A student experiment on a blockchain marketplace found roughly a third of participants cheated or were cheated, and the authors propose a centralized credit-record authority to deter fraud.","lead":"Researchers ran a small blockchain marketplace game with 133 students and found that many reported cheating or being cheated. They argue that blockchain alone cannot prevent fraud in e-commerce and propose a trusted authority that can punish fraudsters via credit records.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 33% 'play tricks' headline is an unsupported sum of two non-exclusive survey items; the paper never reports the union or overlap.","rationale":"The reader's verdict is REJECT, and my stress-test identifies the same core problem that the reader mentioned in the rationale: the 33% figure is an arithmetic sum of overlapping self-reports. However, the reader's stated weakest_assumption was the external validity of the card-trading game as a proxy for real e-commerce fraud. My concern is more direct and internal: even granting the game as a perfect proxy, the paper's own Table 2 does not yield a 33% rate of 'playing tricks on others.' The two means are for distinct survey items, and the paper never establishes that they are disjoint or that 'got tricked' counts as 'playing tricks.' This is a correctness risk in the central quantitative claim, not just a generalizability concern. Since the reader already rejected the paper on grounds that include this arithmetic issue, my independent analysis leaves the verdict unchanged. I do not see a different, more severe load-bearing flaw that would alter the rejection; the qualitative point about blockchains' inability to verify off-chain performance is well known and not the paper's claimed novelty. A concrete re-analysis of the raw survey responses would settle whether the 33% headline can be salvaged, but absent that, the claim as presented is unsupported.","tokens_in":14858,"tokens_out":2289,"duration_ms":22350,"concrete_test":"Obtain the raw responses to the two fraud-in-experiment questions (received goods not as purchased; sent goods not as sold) and compute: (1) the proportion answering 'yes' to either item (union), (2) the proportion answering 'yes' to both (intersection), and (3) the proportion answering 'yes' to 'sent goods not as sold' alone. If the union differs from 33%, or if the headline is based on simply adding the two reported means, the 33% claim should be corrected or removed. Also report the exact question wording and response options to confirm whether 'Get tricked' was mislabeled as 'play tricks.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim in the abstract—'33% of respondents play tricks on others'—is not supported by Table 2. The table reports two separate binary items: 'Get tricked' (mean 0.17) and 'Attempt to trick' (mean 0.16). Section 4.1 says '16% ... attempt to cheat' and 'more than 17% ... are cheated. This implies that about 1/3 of subjects have the incentive to play the trick.' This adds the two percentages, but the categories are not mutually exclusive and measure different things: a respondent who was cheated is not necessarily someone who played a trick, and a respondent could appear in both categories. The abstract's phrase 'play tricks on others' matches only 'Attempt to trick,' not 'Get tricked.' Without the joint distribution or raw data, 33% is at best an upper bound under zero overlap, not an estimated proportion of trick-players. The 'Fraud in nature' items (18–29%) are hypothetical temptation questions and are also not evidence of actual misconduct. Thus the quantitative headline is internally unsupported, independent of the external-validity question about whether the card game proxies real e-commerce.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a field experiment in which 133 students in Tianjin and Trondheim traded paper letter cards on a decentralized marketplace built on Origin Protocol, followed by a questionnaire on fraud, trust, and privacy. The authors claim from the survey that 33% of respondents 'play tricks on others,' and they use this result to argue that blockchain technology alone cannot prevent fraud in e-commerce. They then propose a conceptual mechanism in which a trusted authority maintains credit records on a permissioned blockchain to deter misconduct, concluding that blockchain is an evolution, not a revolution, for e-commerce.","tokens_in":15213,"tokens_out":5903,"duration_ms":60119,"significance":"If the empirical claims were reliable, the paper would provide rare user-side evidence on fraudulent behavior in a working blockchain-based marketplace, along with cross-country comparisons of trust and privacy preferences. Strengths include the actual deployment of an Origin Protocol marketplace, the multi-trial data collection, and the authors' explicit acknowledgment that the sample is not adequate for generalization. However, the headline 33% figure is not supported by the reported data, the experimental design lacks a baseline or control condition, and the proposed punishment mechanism is asserted rather than tested. These issues undermine the paper's central empirical and policy claims, leaving it as a descriptive pilot study with limited inferential value.","major_comments":[{"comment":"The abstract's claim that '33% of respondents play tricks on others' is not supported by Table 2. The table reports two separate binary items: 'Get tricked' (mean 0.17) and 'Attempt to trick' (mean 0.16). Section 4.1 adds these percentages, but the categories are not mutually exclusive and measure different things: being cheated does not imply playing a trick, and a respondent can appear in both categories. The phrase 'play tricks on others' corresponds only to the 16% attempt item. Without reporting the joint distribution or the union of the two items, the 33% figure is at best an upper bound under the assumption of zero overlap, not an estimated proportion of trick-players. This is a load-bearing error for the paper's main empirical conclusion.","section":"§4.1, Table 2"},{"comment":"The experimental design lacks a baseline or control condition. The card-trading game with pre-determined values, hints, and prizes (steps 4-7) generates incentives to misrepresent information, but there is no comparison to a centralized platform or to a condition without blockchain. Consequently, the observed fraud cannot be attributed to the blockchain-based marketplace per se; it may reflect the game's incentive structure or general behavior in anonymous settings. The 'Fraud in nature' items in Table 2 are hypothetical temptation questions, not actual misconduct, and cannot substitute for a behavioral baseline. This limits the validity of the conclusion that blockchain technology 'hinder[s] the widespread adoption' of e-commerce.","section":"§3.2, §4.1"},{"comment":"The proposed punishment mechanism is asserted, not tested. The abstract states that 'such a punishment can effectively decrease agents' incentives to sell counterfeits and leave fake ratings,' but no experimental, simulation, or analytical evidence is provided. The mechanism is described as 'conceptual' in Section 5, and its effectiveness remains speculative. The paper should either provide evidence for the mechanism's deterrent effect or soften the causal claim in the abstract and conclusion.","section":"§4.2, Figure 2"}],"minor_comments":[{"comment":"The 'T-value' column is not defined; presumably it is a one-sample t-test of the mean against zero, but no p-values or confidence intervals are reported. The 'Dishonest when giving a rating' item has t=1.75, which is not significant at the 5% level for N=133, yet the text treats this as a finding.","section":"Table 2"},{"comment":"The literature review contains a long series of citations to reliability-engineering papers (Cheng, Elsayed, Wei, et al.) that are not integrated into the analysis and appear unrelated to blockchain e-commerce. Several cited items are listed as 'under revision' and should be identified as working papers.","section":"§2.1"},{"comment":"The statement that 'participants are advised to comment and discuss any matter regarding the concept' could introduce experimenter demand; please clarify whether the discussion was structured or recorded and how it was used.","section":"§3.2, step (8)"},{"comment":"The claim that 'more than 42% of the Chinese students are buying items online more than one time during a week' is not traceable to any reported table or appendix; provide the survey question and summary statistics.","section":"§4.4"},{"comment":"The phrase 'play tricks on others' is imprecise; the abstract should report the separate figures for 'reported being cheated' (17%) and 'reported attempting to cheat' (16%) rather than a summed 33%.","section":"Abstract"},{"comment":"There is a typo: 'we verity' should be 'we verify.'","section":"§5"},{"comment":"Reference [27] is listed as Lim et al. (2011) but cited in the text as Lim et al. (2019); please verify the correct year.","section":"References"},{"comment":"The figures need axis labels, sample sizes per country, and clearer definitions of the response categories; the text cites percentages that cannot be independently verified without this information.","section":"Figures 3-5"},{"comment":"With 75 observations and six regressors plus province fixed effects, the R-squared values (0.32-0.39) may reflect overfitting; report adjusted R-squared or F-tests of joint significance.","section":"Table 4"}],"recommendation":"reject","confidential_remarks":"The paper's headline quantitative claim is internally unsupported, and the experimental design cannot identify a blockchain-specific effect on fraud. The proposed punishment mechanism is untested, and the cross-country comparisons are based on a very small Norwegian subsample. The literature review includes a block of self-citations to reliability-engineering papers that appear tangential, which may raise concerns about citation padding. The authors are candid about some limitations, but the abstract and conclusion overstate the findings. As submitted, the central claims require new data or a fundamental reframing to be publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper has a genuine experiment and a sound qualitative point, but the headline number is an artifact. Table 2 reports 17% 'got tricked' and 16% 'attempt to trick.' The abstract's '33% of respondents play tricks' is the sum of those two non-exclusive, different questions. That isn't a proportion of trick-players; it's an upper bound assuming zero overlap. The stress test is correct.\n\nWhat's genuinely here: they actually stood up an Origin Protocol marketplace, ran five trials with 133 students, and collected post-game surveys. That's more than most blockchain-ecommerce papers do. The idea that a blockchain cannot verify off-chain goods or truthful ratings is correct, and the paper's own discussion of OpenBazaar moderators shows awareness of prior work. The cross-country trust/privacy differences are interesting descriptives, and the demographic regressions in the appendix are a good-faith attempt to rule out composition effects.\n\nSoft spots: the 33% claim is load-bearing and unsupported. The two survey items are overlapping self-reports; 'got tricked' and 'attempt to trick' are not the same behavior, and the abstract uses 'play tricks on others' as if they were one. The card game with randomly drawn letter cards and announced hints is a weak proxy for real e-commerce incentives; there's no baseline or control treatment. The proposed fix—a central authority that can downgrade a credit record on a permissioned blockchain—is asserted, not tested. The literature review carries a long block of reliability-engineering self-citations (Cheng, Wei, Li, Zhao) that are never integrated into the analysis; that's padding. The 'Fraud in nature' items are hypothetical and don't measure actual misconduct.\n\nSo the qualitative conclusion—blockchains don't solve off-chain fraud—is fine and matches the literature. But the quantitative evidence doesn't support the way it's presented. The paper is repairable: report the joint distribution, drop the 33% framing, and either add a control or explicitly label the game as a stylized exercise. As is, I wouldn't send it to a referee; I'd ask the authors to fix the measurement and reframe the claim before it deserves review.","headline":"A real experiment buried under a headline number that doesn't follow from the data; the 33% fraud claim is an unsupported sum of two overlapping self-reports.","tokens_in":15575,"tokens_out":2764,"would_cite":false,"duration_ms":28654,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A third of users cheated in a blockchain marketplace test.","keywords":["blockchain","e-commerce","fraud","trust","reputation system","decentralized marketplace","privacy","permissioned blockchain"],"falsifier":"Run the same trading game again, but verify each claimed delivery against a neutral observer and compare the verified fraud rate with the questionnaire's 33%: if verified misconduct is far lower, self-report inflated the result; if similar, the 33% is a real behavioral benchmark.","tokens_in":14663,"feed_emoji":"⚖️","tokens_out":7641,"duration_ms":73590,"temperature":0.7,"pith_summary":"This paper tests the claim that blockchains will revolutionize e-commerce by having 133 students buy and sell on a decentralized marketplace built on Origin Protocol, then surveying their behavior and attitudes. The authors report that about a third of participants tried to deceive others—shipping items that did not match the listing, attempting to cheat, or leaving dishonest ratings—and that between 18% and 29% would risk fraud for money. Their central conclusion is that a blockchain can make payments and records immutable but cannot verify off-chain delivery or truthful reviews, so blockchain technology alone does not solve e-commerce trust problems. The paper then proposes a hybrid safeguard: normal transactions remain anonymous, but in a dispute a trusted authority can downgrade the fraudulent party's real-life credit record stored on a permissioned blockchain.","feed_headline":"A third of users cheated in a blockchain marketplace test","feed_subtitle":"A 133-participant trading game shows blockchains cannot verify off-chain delivery, pushing e-commerce toward credit-backed arbitration.","key_machinery":"The mechanism that carries the argument is a physical trading game played on a real decentralized marketplace: students draw five random letter cards with pre-determined but hidden values, trade them on an Origin Protocol-based platform with Ethereum test funds, receive hints that change the cards' perceived value, and then answer a fraud questionnaire. The game makes the boundary visible: the escrow smart contract ensures payment, while delivery, item quality, and ratings are off-chain and unverifiable by the ledger. The paper's proposed remedy is a conceptual arbitration layer in which a trusted authority holds users' true identities and a permissioned blockchain stores credit records, so that fraudulent users lose real-world credit even though ordinary transactions stay anonymous.","core_discovery":"The paper's central claim is that 33% of the participants played tricks on others during the experiment, so the immutability of blockchains does not translate into honest e-commerce. Specifically, 16% of respondents said they attempted to cheat, more than 17% said they were cheated, and 2% admitted leaving dishonest ratings even without extra temptation; 18–29% said they would commit fraud for more money. Because the marketplace's smart-contract escrow guarantees payment but not the nature of delivered goods or the truthfulness of ratings, fraud moves to exactly the parts of the transaction that the blockchain cannot see. On this basis the authors argue that a decentralized marketplace needs an off-chain enforcement layer rather than purely cryptographic trust. They propose a permissioned-blockchain credit system in which a trusted authority, invoked only in disputes, can punish fraudulent users by downgrading their real-life credit records, and they report that 89% of participants would feel safer with a reputation system linked to true identity.","pith_inferences":["The 33% figure is derived by combining the 16% who admitted attempting to cheat with the 17% who said they were cheated; without participant-level linkage, the true share of deceivers could be lower.","Because the experiment relies on students playing a card game with paper cards, real e-commerce fraud could be higher or lower; a field study that independently verifies physical delivery and rating accuracy on a live decentralized marketplace would test the external validity.","The proposal's key trade-off—true identity only in disputes—depends on users actually filing cases; if the cost of arbitration is high, fraud outside the disputed transaction may go unpunished, and a low-cost automated arbitration rule may be needed.","Cross-country differences in willingness to share data suggest that a single global permissioned-credit mechanism is unlikely; region-specific trust anchors would be required."],"forward_implications":["If the 33% finding holds, blockchain-based marketplaces that rely only on escrow and immutable feedback will still face counterfeit goods, fake reviews, and cheating; adoption will be slowed until fraud is addressed.","Reputation systems linked to true identity appear acceptable to most users: 89% of participants said they would feel safer, so \"partial anonymity\" may be a viable design compromise.","A trusted authority with the power to downgrade credit records stored on a permissioned blockchain could deter fraud by making misconduct costlier than any gain, effectively importing legal enforcement into decentralization.","The observed cross-cultural differences suggest that the identity that vouches for trust—government in China, companies in Norway—should be tailored to local institutions.","Blockchain's role in e-commerce would be evolutionary (payment, escrow, record-keeping) rather than revolutionary (full replacement of intermediaries)."],"supporting_citations":[{"why":"Provides the live decentralized marketplace platform used in all five experimental trials.","marker":"Origin Protocol Inc. (Section 3.1)"},{"why":"Introduces the blockchain ledger whose immutability and lack of central authority are the premises being tested.","marker":"Nakamoto (2008)"},{"why":"Defines smart contracts, the escrow mechanism that secures payment but not delivery.","marker":"Szabo, N., 1994"},{"why":"Canonical trust game that underlies the behavioral measures the experiment's fraud results challenge.","marker":"Berg et al. (1995)"},{"why":"Prior experimental treatment claiming blockchain can promote trust, against which the paper's negative finding is a contrast.","marker":"Milosav (2019)"},{"why":"Proposal to replace human opinions with objective reputation; the paper's dishonest-rating result is evidence that this is not enough.","marker":"Dennis and Owenson (2016)"},{"why":"Anonymous fee-based reputation system whose cost-based fraud deterrence the paper argues is insufficient.","marker":"Soska and Christin (2015)"},{"why":"Existing decentralized marketplace with anonymous moderators, used as the counterexample to the paper's authority-based arbitration proposal.","marker":"OpenBazaar (OB)"}],"fun_headline_variants":["Blockchain e-commerce: 33% of users cheated in test","Evolution not revolution: blockchain users cheat in marketplace","Blockchain marketplace test shows a third of users cheat","Trust fails in blockchain e-commerce: 33% cheated","Why blockchain e-commerce is evolution, not revolution: 33% cheated"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire 33% figure rests on the assumption that a card-trading game played by students with pre-determined letter values makes people behave the way real buyers and sellers would on an actual blockchain marketplace, and that their self-reported answers describe what they actually did.","fun_headline_variants_meta":{"raw":{"variants":["Blockchain e-commerce: 33% of users cheated in test","Evolution not revolution: blockchain users cheat in marketplace","Blockchain marketplace test shows a third of users cheat","Trust fails in blockchain e-commerce: 33% cheated","Why blockchain e-commerce is evolution, not revolution: 33% cheated"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1430,"prompt_tokens":913,"completion_tokens":517,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":431}},"tokens_in":529,"tokens_out":517,"duration_ms":4948,"temperature":1.0,"reasoning_tokens":431,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T22:36:55.967430+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same trading game again, but verify each claimed delivery against a neutral observer and compare the verified fraud rate with the questionnaire's 33%: if verified misconduct is far lower, self-report inflated the result; if similar, the 33% is a real behavioral benchmark.","supporting_citations":[{"cited_title":"Bitcoin: A peer-to-peer electronic cash system","cited_arxiv_id":null,"evidence_quote":"Introduces the blockchain ledger whose immutability and lack of central authority are the premises being tested."},{"cited_title":"Smart contracts","cited_arxiv_id":null,"evidence_quote":"Defines smart contracts, the escrow mechanism that secures payment but not delivery."},{"cited_title":"and McCabe, K., 1995","cited_arxiv_id":null,"evidence_quote":"Canonical trust game that underlies the behavioral measures the experiment's fraud results challenge."},{"cited_title":"Blockchain technology: A trust or control machine? Theory and Experimental Evidence","cited_arxiv_id":null,"evidence_quote":"Prior experimental treatment claiming blockchain can promote trust, against which the paper's negative finding is a contrast."},{"cited_title":"and Owenson, G., 2016","cited_arxiv_id":null,"evidence_quote":"Proposal to replace human opinions with objective reputation; the paper's dishonest-rating result is evidence that this is not enough."},{"cited_title":"and Christin, N., 2015","cited_arxiv_id":null,"evidence_quote":"Anonymous fee-based reputation system whose cost-based fraud deterrence the paper argues is insufficient."}],"review_version":1}