{"id":"b668eecd-9061-4d8f-8e94-d526ed90a9a5","arxiv_id":"2509.09351","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A full-year census of top security papers plus 24 interviews shows ethics reporting is inconsistent and skewed toward compliance, with little deep harm-benefit deliberation.","lead":"This paper analyzes how computer security researchers report and reason about ethics, examining all 1,154 papers from the top four security conferences in 2024 and interviewing 24 researchers. It finds inconsistent ethics reporting, focused on approval and disclosure, and suggests community-level changes to support deeper ethical deliberation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's central claim is internally inconsistent: 'lack' of balancing harms/benefits contradicts Table 3, where it is the most common ethics topic.","rationale":"The reader's verdict identifies the absence of inter-rater reliability as the weakest assumption. That is a legitimate methodological concern, but I find a more direct threat to the central claim: the abstract mischaracterizes the paper's own results. Section 4.1 explicitly states that balancing/mitigating risks and benefits is the most common ethical discussion (159), exceeding the counts for the three 'strong focus' items listed in the abstract (IRB/ERB approval = 121, human subjects = 116, responsible disclosure = 257). While disclosure is indeed most frequent overall, IRB approval and human subjects are each less frequent than balancing. Therefore the abstract's juxtaposition of a 'strong focus' on those compliance topics with a 'lack of discussion of balancing' is not supported by the reported frequencies. This is an internal inconsistency, not an external judgment. If the authors intend 'lack' to mean shallowness of discussion, they need to provide an additional measure of depth; the current coding only records presence/absence. Thus the central claim as stated is not fully supported. The IRR concern remains relevant for the precision of the counts, but even with perfect coding the abstract's characterization would still be misleading. Hence I recommend conditional acceptance pending a clarified, data-consistent abstract (or added depth analysis).","tokens_in":29406,"tokens_out":14534,"duration_ms":160880,"concrete_test":"Analytical check: compare the abstract's wording with the frequencies in Table 3 and the statement in Section 4.1. Specifically, verify that 'Balancing benefits and harm' = 159 > 'Review board approved' = 121 and > 'Human Subjects' = 116. If this inequality holds, the abstract's 'lack of discussion of balancing harms and benefits' is contradicted by the paper's own metric. To settle whether the findings are merely poorly summarized or substantively wrong, authors could also report the proportion of ethics-discussing papers (315) that contain a substantive (not just keyword-level) risk-benefit discussion; but the immediate test is the internal consistency check.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In Section 4.1 and Table 3, the authors report that 'the most common discussion of ethics was on the principle of beneficence through balancing/mitigating risks and benefits (159).' However, the abstract states the meta-analysis found 'a lack of discussion of balancing harms and benefits,' while simultaneously claiming a 'strong focus' on IRB/ERB approval (121 papers) and human subjects protection (116 papers). The 159 balancing mentions exceed both of these 'strong focus' categories. Thus, as written, the abstract's central claim is not supported by the paper's own frequency data. If 'lack' is intended to mean insufficient depth or quality of such discussions, that dimension was not coded or reported; the paper only records presence/absence of balancing mentions. Consequently, the headline finding is misleading and needs either a revision of the abstract or an additional depth-based analysis to justify the characterization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a data-driven assessment of ethics reporting and decision-making in computer security research. The authors manually coded all 1,154 papers published at the four top security conferences (CCS, IEEE S&P, NDSS, USENIX Security) in 2024 for the presence and type of ethics discussions, using a Menlo Report-derived codebook. They supplement this meta-analysis with 24 semi-structured interviews with security and privacy researchers, including authors, reviewers, ethics-committee members, and program chairs. The paper's central claims are that ethics reporting in published papers is inconsistent and compliance-oriented (e.g., review-board approval, human-subjects protection, responsible disclosure), that deeper ethical reasoning reported by authors is often absent from papers, and that the community lacks shared standards and guidance. The paper also proposes recommendations for conferences, review boards, and ethics education.","tokens_in":29631,"tokens_out":2896,"duration_ms":34176,"significance":"If the headline findings are accurate, this is a valuable empirical contribution to the ongoing discussion of research ethics in computer security. The manuscript makes its codebook, interview guide, and extended appendix available, and it provides a rare corpus-level snapshot of ethics reporting across a full publication year. The interview data add a useful qualitative layer, and the authors are transparent about many limitations, including sample bias and the non-generalizability of the interview findings. The main quantitative claim in the abstract, however, is not supported by the paper's own frequency data, which is a load-bearing inconsistency that must be addressed before the paper can be accepted.","major_comments":[{"comment":"The abstract states the meta-analysis found 'a lack of discussion of balancing harms and benefits,' but Table 3 reports that 'balancing/mitigating risks and benefits' is the most commonly coded Menlo principle, appearing in 159 papers—more than IRB/ERB approval (121) or human-subjects protection (116). Section 4.1 also explicitly says this was 'the most common discussion of ethics.' If 'lack' is meant to refer to depth or quality rather than presence, that dimension was not coded in the meta-analysis; the paper only records presence/absence. The abstract should be revised to match the data, or the authors should add a depth-based analysis to justify the current wording.","section":"Abstract; Section 4.1, Table 3"},{"comment":"The quantitative claims (e.g., 839 papers with no ethics discussion, 159 with risk-benefit discussion) rest entirely on manual coding by three researchers, yet no inter-rater reliability statistic is reported for the meta-analysis. The paper reports only a 10% spot-check by the lead author and team discussions for ambiguous cases. For the interview coding, Krippendorff's alpha > 0.80 is reported; an analogous reliability measure (or a detailed justification of the consensus process) should be provided for the paper coding, or at least for a reliability subsample, so readers can assess the stability of the headline numbers.","section":"Section 3.1.2"}],"minor_comments":[{"comment":"The phrase 'all 1154 top-tier security papers published in 2024' is imprecise: the dataset covers papers from four selected venues, not all top-tier security venues. The Limitations section correctly notes this, but the abstract should say 'all papers at the top four security conferences' to avoid overstatement.","section":"Abstract; Section 3.1"},{"comment":"The claim that 'published research seems to view the Menlo Report principles as a useful framework' is an inference from the coders' mapping, not from authors' explicit use: only 36 papers explicitly referenced the Menlo Report. This sentence should be softened or clarified.","section":"Section 4.1"},{"comment":"The table is dense and the row/column totals are not always immediately checkable. Adding a 'Total' column or explicit totals in the table header would improve readability.","section":"Appendix F, Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is an extended version of a CCS 2025 publication, and the topic is well within the venue's scope. The main concern is the abstract's contradiction of the paper's own Table 3; this is fixable but requires either a substantial rewording of the central claim or an additional depth-based coding. The missing inter-rater reliability for the meta-analysis is a separate but important rigor issue that the authors should address."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should read this paper—it is the most complete census of ethics reporting in top security venues that I know of, and the interviews add a layer the quantitative data can't reach. The claim that the de-facto standard is compliance-oriented and uneven is well supported.\n\nWhat is actually new: they manually coded all 1154 papers from CCS, IEEE S&P, NDSS, and USENIX Security 2024 for ethics presence, Menlo principles, IRB/ERB approval, and responsible disclosure, with a published codebook and keyword list. They also interviewed 24 security researchers, including reviewers and ethics committee members, about how they reason about ethics in their own work and during peer review. The finding that researchers' private ethical reasoning is much richer than what appears in published papers—and that much ethics writing is boilerplate—is a useful contribution. The materials on Zenodo are a real plus.\n\nThe soft spots are real but manageable. First, the abstract's central claim is internally inconsistent with the paper's own data. It says there is 'a lack of discussion of balancing harms and benefits,' but Table 3 shows balancing/mitigating risks and benefits is the most commonly coded ethics topic (159 papers), ahead of IRB/ERB approval (121) and human subjects protection (116). In context, the real story is that 839 of 1154 papers have no ethics discussion at all; when ethics is discussed, balancing is actually frequent. The abstract needs rewording or a depth-based analysis to say what 'lack' means. Second, the 1154-paper coding has no inter-rater reliability statistic—only a 10% spot check by the lead author. That is a moderate limitation, not a fatal one; the central qualitative patterns would probably survive, but the headline numbers could shift. Third, the interview sample is 24 self-selected participants, mostly from the US and Germany, and the authors are appropriately cautious about generalization.\n\nThe central argument holds. The 2024 snapshot may already be dated because USENIX Security's 2025 ethics guidelines changed the landscape, but the paper acknowledges that and frames it as a baseline.\n\nThis paper deserves a serious referee. I would send it to peer review and likely accept after an abstract revision that matches the data. I'd bring it to our reading group.","headline":"The first full-year census of ethics reporting in top-4 security venues, with a strong interview study—but the abstract overstates a 'lack' of risk–benefit discussion that its own tables contradict.","tokens_in":30103,"tokens_out":3345,"would_cite":true,"duration_ms":35700,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ethics reporting in top security research is inconsistent and compliance-focused, with deep ethical reasoning rarely making it into the written record.","keywords":["security ethics","research ethics","meta-analysis","ethics reporting","Menlo Report","IRB","peer review","usable security and privacy"],"falsifier":"Recode a random sample of 200 papers from the corpus with two independent coders using the paper's own codebook; if inter-rater agreement (e.g., Cohen's kappa) falls below about 0.7, or if the recoded share of papers with no ethics discussion moves more than a few percentage points away from the reported 839/1154, the quantitative core of the claim is not stable. A complementary check is to run the paper's own keyword list over the full text of the 839 'no ethics' papers; if a substantial share contain terms like 'consent' or 'harm,' the manual screen may have missed discussions.","tokens_in":29346,"feed_emoji":"⚖️","tokens_out":6486,"duration_ms":61875,"temperature":0.7,"pith_summary":"This paper claims that ethics reporting in top-tier computer security research is inconsistent and tilted toward formal compliance, with an emphasis on institutional approval, human subjects protection, and responsible disclosure, while the harder question of balancing harms and benefits is rarely discussed. The claim rests on a manual review of all 1,154 papers published in 2024 at the four main security conferences, plus interviews with 24 security researchers. The interviews show that authors do engage in substantive ethical reasoning, but that reasoning is mostly invisible in the published papers. A sympathetic reader would care because the paper reframes ethics in security research from a matter of individual conscience to a measurable, systemic feature of the publication pipeline, one that could be improved with clearer community standards and earlier, proactive guidance.","feed_headline":"Three of four top security papers skip ethics altogether","feed_subtitle":"Authors say they reason deeply about ethics; the written record is mostly compliance talk—IRB, disclosure, human subjects.","key_machinery":"The central analytical instrument is the Menlo Report's four principles—respect for persons, beneficence, justice, and respect for law and public interest—used as a codebook to classify every paper in the corpus. Each paper was skimmed and keyword-checked for indicators of each principle, for IRB/ERB status, consent, deception, and vulnerability disclosure. The second mechanism is a two-stage qualitative study: a semi-structured interview guide covering decision-making, peer review, and hopes for the future, analyzed with open coding and affinity diagramming. Together these let the paper measure both the written record (what gets reported) and the lived process (what authors and reviewers ac","core_discovery":"By coding every paper published in 2024 at the big four security venues against the Menlo Report's four ethical principles—respect for persons, beneficence, justice, and respect for law and public interest—the authors find that 839 of 1,154 papers (about 73%) contain no discussion of ethics at all. Among the 315 papers that do engage with ethics, the dominant themes are institutional approval (121 board-approved, 19 exempt), confidentiality (110 papers), informed consent (89), and vulnerability disclosure (257 papers); the explicit balancing of risks and benefits appears in only 159 papers. Interviews with 24 authors, reviewers, ethics committee members, and program chairs reveal a community","pith_inferences":["If the gap between interview reasoning and written reporting is real and caused partly by space and incentive constraints, then venue policies that allocate dedicated space for ethics reasoning (a change the paper notes one major conference has recently introduced) could measurably increase the depth of published ethics sections; this is a testable prediction.","The paper's focus on the Menlo Report may itself shape what is counted as 'ethics'; a complementary analysis coding for consequentialist vs. deontological reasoning, or for dual-use and environmental harms, could reveal whether risk-benefit balancing is genuinely rare or simply phrased in language the codebook missed.","The authors' deliberate choice not to release the full paper-level coding means the quantitative results cannot be independently audited; a privacy-preserving release of aggregated codes by venue and paper type would let the community verify the headline rates without identifying individual authors.","If the 'hidden curriculum' finding generalizes, formalizing ethics mentorship—for example, pairing junior researchers with ethics-experienced reviewers before submission—could reduce the inconsistency the paper documents."],"forward_implications":["If the findings are right, the de facto ethics standard in top security venues is compliance-oriented: getting the board approval and writing the disclosure paragraph matters more than demonstrating that harms were weighed against benefits.","The written literature systematically underrepresents the ethical reasoning of authors, so readers—especially junior researchers—cannot learn the field's actual ethical practice from published papers alone.","Wide variation in calls for papers across venues means authors and reviewers face different and sometimes conflicting ethics expectations for essentially similar work.","Research ethics committees and peer reviewers operate without a shared framework, making ethics review subjective, reactive, and inconsistent across subdisciplines and geographies.","Because ethics is largely learned informally from advisors and colleagues, researchers without access to experienced mentors are at a structural disadvantage."],"fun_headline_variants":["73% of top security papers omit ethics entirely","Only a quarter of security papers address ethics at all","Security ethics reporting: heavy on compliance, light on reasoning","Top security papers skip ethics in 3 of 4 cases","Security researchers want ethics, but papers rarely show it"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline numbers depend on the reliability of three coders' manual classification of 1,154 papers; the paper reports only a 10% spot-check by the lead author and no formal inter-rater reliability statistic, and the authors themselves note they may have missed ethics discussions written in language they did not associate with ethics.","fun_headline_variants_meta":{"raw":{"variants":["73% of top security papers omit ethics entirely","Only a quarter of security papers address ethics at all","Security ethics reporting: heavy on compliance, light on reasoning","Top security papers skip ethics in 3 of 4 cases","Security researchers want ethics, but papers rarely show it"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000472,"raw_usage":{"total_tokens":2184,"prompt_tokens":747,"completion_tokens":1437,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":1373}},"tokens_in":491,"tokens_out":1437,"duration_ms":11985,"temperature":1.0,"reasoning_tokens":1373,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T19:12:31.244831+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recode a random sample of 200 papers from the corpus with two independent coders using the paper's own codebook; if inter-rater agreement (e.g., Cohen's kappa) falls below about 0.7, or if the recoded share of papers with no ethics discussion moves more than a few percentage points away from the reported 839/1154, the quantitative core of the claim is not stable. A complementary check is to run the paper's own keyword list over the full text of the 839 'no ethics' papers; if a substantial share contain terms like 'consent' or 'harm,' the manual screen may have missed discussions.","supporting_citations":[],"review_version":1}