{"id":"79e6d8ee-4db4-4650-aacc-630b4b9a08f4","arxiv_id":"2411.16404","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that categorizes blockchain privacy, consent, and self-sovereign identity research and finds most proposed systems lack public implementations.","lead":"This paper surveys 98 published works on blockchain-based privacy, consent management, and self-sovereign identity, sorting them by use case, technique, platform, and code availability. It is a useful map for researchers and practitioners who want to compare privacy-preserving blockchain designs at a glance.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The availability claim anchoring the 'few deployable implementations' conclusion is not visible in the paper's own tables: the promised code-availability columns are absent from Tables 2, 3, 5, and 6.","rationale":"The reader's conditional verdict is well aligned with my reading: the survey is useful but its central practical claim rests on an analysis that is not fully transparent or independently verifiable from the manuscript. I focus on a more specific and load-bearing subproblem than the reader's broad corpus-completeness concern: the paper's own displayed tables lack the promised code-availability columns, so the quantitative basis for 'few solutions provide the code on GitHub' is not evidenced in the text. This is a concrete, checkable defect rather than a purely stylistic one, because the conclusion about limited practical deployment is the survey's main actionable result. I do not think this warrants rejection: the companion GitHub dataset is a real asset, the search strings are explicitly reported, and the underlying claim may well be true. The correct response is to keep the conditional verdict and require the authors to make the availability evidence visible and reproducible. My concern is partially different from the reader's weakest assumption: the reader emphasized the title-only search and possible missed literature, while I emphasize that even the collected 98 works' availability coding is not auditable from the paper. Both point to the same underlying need for a transparent, verifiable evidence base, so I mark agreement as partial rather than full.","tokens_in":32969,"tokens_out":8583,"duration_ms":90933,"concrete_test":"Download the companion dataset at https://github.com/rodrigodg1/privacy-survey and extract the software-availability field for all 98 works. Independently enumerate repositories by (a) following every link in the paper's footnotes and dataset, (b) querying GitHub and GitLab search APIs with each paper's exact title and first author, and (c) checking the full texts of the surveyed papers for repository mentions. Recompute the counts behind 'few solutions provide the code on GitHub.' If the re-derived count is materially higher than the paper's implied count (e.g., more than 20% of the privacy and consent corpora have public code), the conclusion should be revised and the missing availability tables or columns added; if the count remains low (e.g., under 10%), the conclusion stands but the manuscript should still display the availability column to make the evidence visible and auditable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central practical conclusion—'few solutions provide the code on GitHub, while most do not provide any software availability information'—is presented as if it were read directly from Tables 2 and 3 ('As presented in Tables 2 and 3, …'). In the manuscript, however, Tables 2 and 3 contain only Work, Publication Year, Use Case, and Privacy Technique columns; the promised software-availability column is missing. The same defect affects Tables 5 and 6, whose captions promise 'software availability' but whose visible columns are Work, Year, Use Case, and Key Contribution. The only availability evidence in the body is a small number of GitHub/GitLab footnotes. Consequently, the quantitative availability claim that anchors the conclusion cannot be independently checked from the paper itself; the reader would have to trust the external GitHub dataset [41] or re-derive the counts from all 98 sources. The conclusion also states that schemes 'tend to rely on external entities and trusted third parties to protect and process sensitive information, limiting privacy' without defining or operationalizing what counts as reliance on a trusted third party. This matters because the actionable message—many proposals but few deployable, trust-minimized implementations—depends directly on these two empirical inputs. If the dataset's availability field undercounts repositories hosted on GitLab, institutional pages, or repositories mentioned in paper bodies rather than linked in footnotes, the headline result could shift from 'few implementations' to 'several implementations,' changing the paper's practical conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of blockchain-based privacy applications, with a stated focus on consent management and self-sovereign identity (SSI). The authors describe a title-based literature search across Google Scholar, Scopus, and ACM Digital Library that yields 98 works, and they organize the reviewed works into three areas: privacy-preserving data sharing, consent management, and SSI. They also review privacy-focused blockchain protocols and identity management platforms, and conclude with open challenges. The paper's central claims are that privacy must be supplied as an external layer on top of blockchain, that few proposed solutions provide publicly available implementations, and that many schemes rely on external entities and trusted third parties, limiting privacy. The contributions are primarily classificatory and expository: there are no formal derivations or empirical measurements to evaluate.","tokens_in":33189,"tokens_out":2197,"duration_ms":24272,"significance":"If the corpus is complete and the classifications are accurate, the survey is a useful reference for researchers entering the area: it provides a structured comparison of 98 works, an explicit search methodology, and a publicly listed dataset on GitHub [41]. The paper also makes a falsifiable practical claim, namely that most surveyed proposals lack accessible implementations, which is a valuable observation if it is supported by the data. At the same time, the survey's value depends entirely on the correctness and completeness of its literature corpus and on the accuracy of the summary tables, since there are no formal results to verify. The absence of machine-checked proofs, quantitative evaluations, or systematic quality appraisal means the contribution is an organizing and synthesizing one rather than a technical one.","major_comments":[{"comment":"The central availability claim is not verifiable from the manuscript's own tables. The text in Section 5.2 states, \"As presented in Tables 2 and 3 ... few solutions provide the code on GitHub, while most do not provide any software availability information,\" and the table captions promise a code/software availability column. However, Tables 2 and 3 contain only Work, Publication Year, Use Case, and Privacy Technique columns; Tables 5 and 6 contain only Work, Publication Year, Use Case, and Key Contribution columns. The claimed software availability information appears only in a small number of footnotes and in the external GitHub dataset [41]. As a result, the quantitative basis for the conclusion that \"few solutions provide the code on GitHub\" cannot be independently checked from the paper. The authors should either add the promised availability columns to the tables, or present an explicit count of available/unavailable implementations in the text with a precise definition of what counts as \"available\" (e.g., linked public repository versus code mentioned in the paper body), and clearly state which of the 98 sources were assessed and how.","section":"Section 5.2, Tables 2 and 3; Section 5.3, Tables 5 and 6"},{"comment":"The conclusion that surveyed schemes \"tend to rely on external entities and trusted third parties to protect and process sensitive information, limiting privacy\" is not operationalized or directly supported by the presented tables. The tables classify privacy techniques but do not include a column or coding for trust assumptions, such as whether a scheme requires a centralized proxy re-encryption server, a trusted authority for key management, or an off-chain storage provider. Without a definition of what counts as reliance on a trusted third party, and without a per-work classification, this statement is an impression rather than a survey result. The authors should either add a trust-assumption dimension to the classification or revise the conclusion to state this as an interpretive observation with the specific supporting examples from the surveyed works.","section":"Section 9, Conclusion; Section 5.2"},{"comment":"The completeness claims of the survey are stronger than the search methodology justifies. The selection procedure uses only title-field keyword searches in three databases, with no backward or forward citation chasing, no grey literature or preprint coverage, and no independent verification of the 98 retrieved summaries. Given this, the claim in Section 2 that \"no existing study connects these aspects\" is too strong: it is a claim about the entire literature, not just about the title-matched corpus. The authors should soften this claim to \"no study found by our search connects these aspects\" and add an explicit limitations paragraph discussing recall risk, including the fact that relevant work with different title wording would be missed.","section":"Section 4, Survey Methodology"}],"minor_comments":[{"comment":"The search-string numbering is inconsistent: the privacy strings are labeled S01, S02, and S03, and then the consent strings reuse S03, followed by S04 and S05, while the identity strings jump to S07, S08, and S09, skipping S06. This should be renumbered for clarity.","section":"Section 4, search string listing"},{"comment":"The sentence \"Sections 6, 7. Section 8 identifies the opportunities...\" is grammatically incomplete; it should read something like \"Section 6 reviews privacy-focused blockchain platforms, Section 7 reviews identity management platforms, and Section 8 identifies opportunities for further research.\"","section":"Section 1, last paragraph of the introduction"},{"comment":"The first author of reference [49] is listed only as \"Raghav\" without a surname; this should be corrected to the full author name as it appears in the source publication.","section":"References, entry [49]"},{"comment":"The table lists both \"Anonymization\" and \"K-anonymity\" as separate techniques, and also lists \"Hash Anonymous Identity\" separately; since several works use multiple techniques, the table would be clearer if the relationship between these categories and the individual works' primary classification were explained in a note.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a survey venue and the authors have made a good-faith effort to document their corpus. The main concern is that the headline practical conclusion about code availability is not auditable from the manuscript's own tables, and the trusted-thrid-party claim is not coded systematically. Both issues are fixable and do not require new experiments. I would also note that the authors' own prior work is cited as [10] and appears in the discussion of privacy-enhancing mechanisms; this is not circular, but the relationship should be transparent in the revision if any of the survey's classifications rely on [10]."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a competent survey that adds a legitimate joint lens—privacy, consent, and self-sovereign identity in blockchain applications—with an explicit code-availability analysis. The Table 1 comparison against prior surveys is fair and shows that earlier work covers these topics only in pairs. For a newcomer or a practitioner scoping the field, the structured tables of 98 works by use case, privacy technique, and platform are genuinely useful, and the GitHub dataset is a nice touch. The search methodology is described transparently, even if it is limited to title-field queries in three databases.\n\nThe soft spots are real but not fatal. The title-only search, with no citation chasing or quality appraisal, constrains completeness; the claim that \"no existing study connects these aspects\" is too strong for a survey genre that evolves quickly. More importantly, the paper's central practical conclusion—that few solutions provide code on GitHub—is presented as if read directly from Tables 2 and 3, but those tables have no software-availability column. The same is true for Tables 5 and 6, whose captions promise \"software availability\" but whose visible columns stop at key contribution. Only a few footnotes point to GitHub repositories. So the availability claim cannot be independently checked from the manuscript itself; a reader would need to re-derive the counts from the external dataset. That is a presentation and verifiability defect, not a fabricated result, but it matters because the paper's actionable message depends on it. Similarly, the conclusion that schemes \"tend to rely on external entities and trusted third parties\" is asserted without defining what counts as such reliance, so the reader cannot tell how that judgment was reached.\n\nThere is also a structural oddity: Sections 6 and 7 in the text appear out of order relative to the stated organization, and Section 5.3 mixes narrative text with table material—minor but worth cleaning up. The self-citation [10] is not a problem here; the survey's classifications do not depend on that work.\n\nWho is this for? Someone entering the blockchain-privacy area who wants a broad map of the literature and a snapshot of platform options. It does not change practice or offer new technical results, but it is a fair synthesis. I would send it to peer review, though not without requiring the authors to add the missing availability columns (or explicitly citing the dataset rows), soften the uniqueness claim, and operationalize \"reliance on trusted third parties.\" With those revisions, it would be a serviceable reference.","headline":"A useful but overclaiming survey: the joint privacy/consent/SSI synthesis is legitimate, yet the headline code-availability result is not verifiable from the paper's own tables.","tokens_in":33773,"tokens_out":1714,"would_cite":true,"duration_ms":18226,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 98 blockchain privacy, consent, and identity papers concludes that privacy is not inherent to blockchain and most proposed solutions rely on external trusted entities, with few public implementations.","keywords":["blockchain","privacy","consent management","self-sovereign identity","data sharing","survey","GDPR","decentralized identity"],"falsifier":"A repeatable search that adds backward and forward citation chasing, preprint sources, and direct code checks, and that finds a substantial set of surveyed-period systems offering full privacy, no trusted third parties, and public software, would overturn the paper's conclusion that the field is mostly proposals with limited implementations.","tokens_in":32767,"feed_emoji":"🔐","tokens_out":6090,"duration_ms":57271,"temperature":0.7,"pith_summary":"The paper is a survey of 98 blockchain-based works on privacy, consent management, and self-sovereign identity, organized around three research questions about privacy mechanisms, identity control, and supporting platforms. It argues that blockchain offers transparency, auditability, and immutability for multi-stakeholder data sharing, but privacy is not built in and must be added through external schemes. Reviewing the corpus by use case, privacy technique, platform, and software availability, the survey concludes that most proposed schemes use cryptographic tools such as encryption and zero-knowledge proofs yet tend to depend on external entities and trusted third parties to protect and process sensitive information, which limits the privacy they can actually provide. It also finds that few works make their code publicly available. The survey positions itself as the first to connect privacy, consent, and identity management together with implementation availability, and it draws a list of open research opportunities from that gap.","feed_headline":"Privacy must be bolted onto blockchain, 98-paper survey finds","feed_subtitle":"Existing schemes lean on trusted third parties and few ship code, so real-world privacy remains limited.","key_machinery":"The argument is carried by a classification grid rather than by a single theorem. Each of the 98 selected works is sorted by use case, by the privacy technique it employs, by whether it addresses consent or self-sovereign identity, by the blockchain platform used for prototyping, and by whether its software is available. A second organizing device is the three-layer privacy taxonomy: Layer-0 covers network-level tools, Layer-1 covers on-chain protocol techniques from homomorphic encryption to confidential transactions, and Layer-2 covers off-chain proofs such as zk-SNARKs, zk-STARKs, and Bulletproofs. These two grids generate the survey's conclusions: privacy is always an add-on, and most add-ons route sensitive material through external parties.","core_discovery":"The central claim is that blockchain cannot deliver privacy by itself: because transactions and smart-contract inputs are visible to consensus nodes, any privacy guarantee must come from mechanisms layered around the ledger. Analyzing 40 privacy works, 33 consent works, and 25 self-sovereign identity works, the survey finds that the dominant privacy techniques are data encryption, homomorphic encryption, access control, and zk-SNARK-related proofs, organized by network-layer, on-chain, and off-chain strategies. Its cross-cutting finding is that these schemes tend to rely on external entities and trusted third parties to protect and process sensitive information, and that most do not provide software, so the literature is rich in proposals but thin in deployable, trust-minimized implementations. The paper further claims that earlier surveys treat privacy, consent, or identity separately and that no existing study combines these dimensions with an analysis of open-source availability.","pith_inferences":["A consequence the authors leave implicit is that the field's bottleneck is deployment, not cryptographic invention: with most schemes never shipped, the next useful step is reference implementations and benchmarks rather than new schemes.","The survey's finding about trusted third parties suggests a testable taxonomy: rank the 98 works by the number and role of external trust assumptions and see whether privacy guarantees weaken as trust assumptions grow.","The layer table hints at a possible design rule: Layer-2 zero-knowledge proofs can give transactional privacy without trusted setups, so combining them with self-sovereign identity-based consent could yield fully trust-minimized systems; that combination is not yet evaluated in the surveyed corpus."],"forward_implications":["Applications that need both GDPR-style consent and blockchain auditability must combine on-chain access-control records with off-chain encrypted storage; neither blockchain alone nor off-chain storage alone suffices.","Practitioners should treat a system's privacy claim as limited whenever a trusted third party or external entity can see or process the sensitive data.","Researchers evaluating these schemes should measure not only security properties but also whether working code exists, since the survey finds most works lack software availability.","Privacy-focused platforms such as the ones surveyed often trade away smart-contract support or introduce trusted setups, so choosing a platform means accepting one of these limitations.","Open problems identified include on-chain encryption performance, post-quantum NTRU schemes, differential privacy versus encryption, cross-chain transactions, confidential assets, and decentralized identity interoperability."],"supporting_citations":[{"why":"supplies the working definition of consent management and explains its regulatory role.","marker":"[6]"},{"why":"defines self-sovereign identity and the verifiable-credential model the survey analyzes.","marker":"[7]"},{"why":"prior privacy-protection survey whose cryptography categories the survey extends.","marker":"[12]"},{"why":"earlier survey that already linked privacy-preserving techniques with SSI concepts, used as a comparison baseline.","marker":"[13]"},{"why":"earlier identity-management survey whose IdM components frame the SSI analysis.","marker":"[20]"},{"why":"introduces the transparent blockchain design that creates the privacy gap motivating the survey.","marker":"[22]"},{"why":"baseline decentralized personal-data system showing blockchain privacy via access control and off-chain storage.","marker":"[42]"},{"why":"influential blockchain consent-management system that anchors the healthcare consent analysis.","marker":"[83]"}],"fun_headline_variants":["Blockchain has no native privacy; 98-paper survey shows add-ons","Survey: Blockchain privacy relies on trusted third parties","Few blockchain privacy tools ship code; most are just proposals","98-paper survey: Privacy must be layered onto blockchain","Blockchain privacy gap: schemes lean on external trust, little code"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's conclusions depend on the corpus being complete and accurately summarized, yet the search covered only titles in three bibliographic databases, without citation chasing, grey literature, or independent verification of the 98 summaries.","fun_headline_variants_meta":{"raw":{"variants":["Blockchain has no native privacy; 98-paper survey shows add-ons","Survey: Blockchain privacy relies on trusted third parties","Few blockchain privacy tools ship code; most are just proposals","98-paper survey: Privacy must be layered onto blockchain","Blockchain privacy gap: schemes lean on external trust, little code"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000718,"raw_usage":{"total_tokens":3214,"prompt_tokens":921,"completion_tokens":2293,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":2217}},"tokens_in":537,"tokens_out":2293,"duration_ms":14415,"temperature":1.0,"reasoning_tokens":2217,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:08:38.163226+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A repeatable search that adds backward and forward citation chasing, preprint sources, and direct code checks, and that finds a substantial set of surveyed-period systems offering full privacy, no trusted third parties, and public software, would overturn the paper's conclusion that the field is mostly proposals with limited implementations.","supporting_citations":[{"cited_title":"A systematic review of blockchain for consent management,","cited_arxiv_id":null,"evidence_quote":"supplies the working definition of consent management and explains its regulatory role."}],"review_version":1}