{"id":"6dfb7a6e-117c-4775-b9a8-9c4881d88e01","arxiv_id":"2601.16513","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"OpenAI's public discourse overwhelmingly emphasizes safety and risk, using explicit 'ethics' vocabulary rarely and superficially, which the authors interpret as evidence of ethics-washing.","lead":"This paper counts how often OpenAI's public posts and papers use the words 'ethics', 'safety', and 'alignment' from 2015 to 2025. It finds safety and risk dominate while 'ethics' rarely appears, and argues this signals 'ethics-washing' — ethics as reputation rather than a real constraint.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corpus selection may not support the generalization: excluding Charter, safety pages, and testimony could reverse the ethics-omission finding; publication corpus size is also inconsistent (30 vs 180).","rationale":"Reader's conditional verdict already identified the corpus sampling as the weak point; I agree. The descriptive numbers within the chosen sections are well-documented and the code release is a positive reproducibility step. However, the paper's own language overreaches: it speaks of 'OpenAI's public discourse' and 'OpenAI increasingly omits an ethics vocabulary' while limiting the corpus to two website sections and a filtered publication subset. The authors' supplementary site-search only tested the literal word 'ethics,' not conceptual equivalents, and by their own admission was not exhaustive. Thus the possibility remains that substantive ethics language lives on the Charter, preparedness, safety, or policy pages. In addition, there is an unresolved internal inconsistency in the corpus size: the text says both '30 publications' and '180 machine-readable publications' in the corpus, which should be resolved by re-reading the released code. Neither issue requires rejecting the paper; rather, it requires rewriting the conclusion as a scoped claim about News/Research discourse and either running the expanded corpus check or toning down the general 'ethics-washing' assertion. Verdict remains conditional.","tokens_in":11607,"tokens_out":10597,"duration_ms":116064,"concrete_test":"Scrape the complete openai.com domain (including /charter, /safety, /alignment, system cards, and blog subdomains), collect all OpenAI-authored arXiv preprints regardless of ethics keyword, and obtain OpenAI's public testimony and policy papers; lemmatize and count occurrences of 'ethic*' plus a validated academic-ethics lexicon (e.g., fairness, justice, accountability, transparency, non-maleficence, beneficence) across channels. Then recompute the annual frequency ratios. If ethics-related terms or substantive ethical concepts appear at rates comparable to safety/risk on excluded channels, the omission conclusion fails; if they remain rare, the finding survives as a scoped result about News/Research only.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is whether the sampled corpus supports the unrestricted conclusion about 'OpenAI's public discourse.' The paper deliberately restricts the corpus to OpenAI.com News and Research sections and to ethics-containing publications (30 of 180 reviewed), explicitly conceding this is 'a snapshot versus a complete archival record.' Absent from the sampling frame are foundational channels where OpenAI makes normative commitments: the OpenAI Charter, the Preparedness Framework, dedicated Safety & Alignment pages, system cards, and the CEO's Congressional testimony. If those sources contain substantive ethical reasoning under vocabularies like fairness, justice, accountability, or even 'ethics' itself, the central finding that OpenAI 'omits an ethics vocabulary' would be an artifact of section selection. The exploratory 'site:openai.com ethics' audit only searched the literal token and did not quantify conceptual equivalents, and the audit is described as 'not exhaustive.' The paper's conclusion that 'OpenAI increasingly omits an ethics vocabulary' therefore extends beyond the defensible claim that its News and Research sections rarely use the word 'ethics.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a mixed-methods case study of OpenAI's public discourse from December 2015 to July 2025. The authors assembled a corpus of 424 web articles from OpenAI's News and Research sections and 30 publications (from 180 screened) that contain 'ethic-' terms. They combine keyword-frequency analysis, HDBSCAN/PCA topic modeling, and a qualitative audit to show that safety, risk, alignment, and governance vocabularies dominate while explicit 'ethics' language is rare. The paper concludes that OpenAI increasingly omits an ethics vocabulary and that this constitutes 'significant ethics washing,' with implications for governance and accountability.","tokens_in":11858,"tokens_out":3430,"duration_ms":40907,"significance":"If the core finding is valid, the paper offers a useful longitudinal documentation of how a frontier AI developer's public communications privilege safety/risk framings over explicit ethics vocabularies, and it connects this to the ethics-washing literature. The manuscript has clear strengths: it is transparent about its mixed-method design, it releases code for reproducibility, and it includes a qualitative audit that goes beyond simple counts. The temporal coverage is substantial. However, the manuscript's central claim is currently broader than the evidence: the corpus is restricted to two website sections and linked arXiv preprints, yet conclusions are phrased in terms of OpenAI's public discourse as a whole. The quantitative analysis also has unresolved internal inconsistencies and limited validation. These issues are fixable within the manuscript's scope, but they need substantive revision.","major_comments":[{"comment":"The central conclusion—that 'OpenAI increasingly omits an ethics vocabulary' and engages in 'significant ethics washing'—is stated for OpenAI's public discourse generally, but the corpus is deliberately limited to OpenAI.com News/Research sections and linked arXiv preprints. The OpenAI Charter, Preparedness Framework updates, dedicated safety pages, many system cards, policy papers, and congressional testimony are excluded. The exploratory 'site:openai.com ethics' audit is described as 'not exhaustive' and only searches the literal token. As a result, the evidence supports a claim about the sampled sections, not about OpenAI's public discourse as a whole. Please either expand the sampling frame or narrow the conclusion and title accordingly.","section":"Corpus Construction; Results; Discussion; Conclusion"},{"comment":"There is an internal inconsistency in the corpus description. The text states that the total corpus is 454 documents (424 web articles + 30 publications), but later says 'Considering the entire corpus (424 web articles and 180 machine-readable publications)'. The publication-level frequency results and Figure 3 depend on which of these denominators is used. The reader cannot determine whether quantitative analyses were run on the 30 ethics-screened publications or on all 180 machine-readable publications. This must be clarified and made consistent throughout.","section":"Corpus Construction; Corpus-wide structure; Figure 3"},{"comment":"All keyword frequencies are reported as raw annual counts with no normalization for corpus size, document length, or total words. The headline comparison (safety ≈ 687 vs. ethics ≈ 7 in 2024) is likely robust in this corpus, but the temporal trends and cross-corpora comparisons could be artifacts of the number of documents published per year. The qualitative audit also reports no inter-coder reliability metric. Please report rates or per-document counts, and provide reliability estimates or a clear rationale for their absence.","section":"Results; Figure 1; Figure 2"},{"comment":"The HDBSCAN/PCA clustering is used to support substantive claims such as 'safety is the gravitational hub' and that policy undergoes a semantic drift from RL to governance. The manuscript gives no validation of the clusters—no stability checks, coherence scores, or sensitivity analysis for the chosen hyperparameters. PCA plots can suggest structure that is not statistically strong. These analyses should be presented as exploratory, or the clustering should be validated and the thresholds for cluster assignment reported.","section":"Quantitative Analysis; Figure 4; Discursive Pivots"},{"comment":"The claim that OpenAI communicates 'without applying academic and advocacy ethics frameworks or vocabularies' rests on the low frequency of the literal 'ethic-' stem. However, academic AI ethics frameworks prominently include concepts such as fairness, justice, accountability, transparency, and non-maleficence. The paper reports a 75-concept keyword library but the results focus on a few terms. If these principle-level concepts are in the library, their counts should be reported; if not, the conclusion about the absence of ethics vocabularies is under-supported. The qualitative audit of adjacent moral language helps, but the quantitative evidence needs to match the breadth of the claim.","section":"Method; Results; Discussion"}],"minor_comments":[{"comment":"The phrase 'significant ethics washing' appears in the conclusion but is presented as a direct finding rather than an interpretive label. Earlier in the paper the authors are appropriately careful about the distinction between data and interpretation; the conclusion should preserve that nuance.","section":"Abstract and Conclusion"},{"comment":"The counts for the publications corpus are inconsistent in places: the text mentions '31 contained a variation of the term ethic' and '30 items' in the final corpus, and later 'n=30' appears. Please reconcile these numbers and define the exact inclusion rule.","section":"Corpus Construction"},{"comment":"The PCA plots have no axis labels or explained-variance annotations. As presented, the visual clusters are difficult to interpret and the claimed structural divergence between web articles and publications would be easier to assess with better metadata.","section":"Figure 4"},{"comment":"There are minor typographical issues, e.g., 'faming' in the Method section and inconsistent reference formatting for the Information Research volume. The paper also cites 'Hao (2025)' as a book but the entry could be clarified with publisher information.","section":"General"},{"comment":"The manuscript acknowledges its corpus is 'a snapshot versus a complete archival record,' which is good. A dedicated limitations subsection that explicitly states which OpenAI communication channels are excluded would help prevent over-interpretation by readers.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"The core frequency result is plausibly robust within the sampled News/Research corpus, and the reproducible code is a plus. The main risk is that the paper overstates the generalizability of its corpus to 'OpenAI's public discourse.' I recommend major revision rather than rejection because the necessary changes—narrowing claims, fixing the denominator inconsistency, and reporting normalized or validated quantities—are achievable within the manuscript's scope. If the authors cannot or will not narrow the conclusion, the paper may need to be rejected as it stands."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Core finding is real and useful: the authors assembled a decade of OpenAI.com News and Research posts plus arXiv preprints and show the gap between safety/risk and ethics is stark—hundreds of mentions vs single digits. The qualitative audit adds texture: when 'ethics' appears, it is usually a stock phrase or a side reference. That descriptive result is a legitimate empirical contribution, and the code release helps.\n\nCredit is due for the method's self-restraint: the paper explicitly calls the corpus a snapshot, separates web from academic registers, and shows authorship patterns matter—co-authored academic pieces use ethics more than in-house web copy. The keyword library is large, and the manual audit gives me confidence the frequency gap is not a preprocessing artifact.\n\nThe soft spot is that the corpus is narrower than the conclusions. The paper claims to characterize 'OpenAI's public discourse,' but it draws only from News, Research, and arXiv. That excludes the Charter, the Preparedness Framework, safety policy pages, system cards, and Congressional testimony—places where OpenAI makes normative commitments. If those sources use 'fairness,' 'accountability,' or even 'ethics' substantively, the claim that OpenAI 'omits an ethics vocabulary' becomes an artifact of section selection. The text itself concedes the snapshot in the method but drops that caution in the conclusion, landing on 'significant ethics washing' and 'increasingly omits an ethics vocabulary.' The data support a narrower claim: in these sampled channels, ethics terminology is rare and largely rhetorical. The leap to ethics-washing imports a motive and a totality the corpus cannot prove.\n\nSmaller issues: the HDBSCAN/PCA clusters are offered without validation; raw frequencies have no variance or inter-coder reliability; and the released code is not accompanied by the dataset or keyword library, so exact reproduction is not possible. The 30 vs 180 publication numbers are just a sampling-frame distinction, not a contradiction.\n\nThis paper deserves a serious referee. The descriptive finding is important enough to matter, and the overreach is correctable with revised framing. Recommend major revision, not rejection.","headline":"A genuinely useful longitudinal count of OpenAI's web corpus showing 'ethics' is rare while safety/risk dominate; the paper then overclaims ethics-washing beyond what the corpus supports.","tokens_in":12303,"tokens_out":2794,"would_cite":true,"duration_ms":28419,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Over a decade of OpenAI's public communications, mentions of 'safety' reached hundreds per year while 'ethics' rarely exceeded seven, a disparity the paper reads as evidence of ethics-washing.","keywords":["AI ethics","ethics-washing","OpenAI","discourse analysis","safety","risk","content analysis","corporate communication"],"falsifier":"A concrete check: use the Wayback Machine to archive every page under openai.com that contains the word 'ethics' across all subdomains (including /charter, /safety, /global-affairs, and blog posts), and code each mention for substantive engagement (e.g., definitional discussion, trade-off analysis, or governance commitments) versus passing use. If the full-site audit yields dozens of documents with substantive ethics analysis, the paper's claim that ethics functions only rhetorically in OpenAI's public discourse fails.","tokens_in":11508,"feed_emoji":"⚖️","tokens_out":4736,"duration_ms":117931,"temperature":0.7,"pith_summary":"The paper tries to establish that OpenAI's public-facing discourse has systematically replaced an ethics vocabulary with safety, risk, and alignment language. Over 2015–2025, web articles and research publications mention safety and risk hundreds of times per year while explicit 'ethics' appears only a handful of times, and mostly in passing phrases. A qualitative audit of all 16 web articles and 30 publications that use ethics terms finds the references are peripheral—headers, stock phrases, external citations—rather than sustained ethical analysis. The paper concludes that this pattern constitutes 'ethics-washing': ethics functions rhetorically as reputational cover while actual governance is framed around compliance, risk, and technical safety. If correct, this matters because it suggests OpenAI's public commitments do not translate into an ethics-grounded governance framework, and external, exogenously imposed accountability is needed.","feed_headline":"OpenAI says 'safety' 700 times, 'ethics' seven","feed_subtitle":"A decade of OpenAI's own posts shows safety and risk crowding out ethics, a pattern the paper calls ethics-washing.","key_machinery":"The argument is carried by a comparative discourse analysis of OpenAI's own public materials: a corpus of 424 web articles from OpenAI's News and Research sections and 30 OpenAI-authored or co-authored publications (25 linked academic preprints and 5 site-hosted items). Two analytic tools do the work: a keyword-concept frequency analysis tracking 75 concepts (ethics, safety, risk, alignment, governance, etc.) over time, and a qualitative audit of every document containing 'ethic-' that logs where and how the term is used. The central discriminator is the ratio between ethics mentions and safety/risk mentions, together with the location and framing of ethics in the text; the paper uses this r","core_discovery":"The paper's central claim is that OpenAI's public discourse, as expressed through its own website articles and its linked academic preprints, is dominated by safety and risk vocabulary rather than ethics vocabulary. Across the corpus, 'safety' peaks at nearly 700 mentions in a year and 'risk' at 386, while 'ethics' never exceeds seven annual mentions; in the 424 web articles, only 16 (3.8%) contain any variant of 'ethics.' When ethics does appear—in publications more than on the website—it is typically in footnotes, appendices, policy copy, or stock phrases such as 'ethical standards,' not in analytic engagement. The authors interpret this asymmetry as evidence of ethics-washing: the organiz","pith_inferences":["If the finding holds, a testable follow-up is whether other frontier AI labs show the same ethics-to-safety substitution in their public communications; a cross-company comparison would show whether this is an OpenAI-specific pattern or an industry-wide one.","The paper's reliance on archived web sections means the claim may undercount ethics discourse on other OpenAI channels, such as the charter page, safety documentation, blog posts, or congressional testimony; checking those would sharpen or soften the ethics-washing conclusion.","One implicit consequence the authors do not develop is that the absence of ethics vocabulary may itself be a rational response to legal and reputational exposure: explicit ethics commitments are harder to litigate against than safety claims, so the marginalization of 'ethics' could be a strategic feature rather than an oversight.","The findings suggest a concrete audit strategy for regulators: instead of counting ethics statements, track whether ethics terms ever appear in the decision-relevant parts of system cards and preparedness frameworks, not just in headers and footnotes."],"forward_implications":["If OpenAI's public discourse is representative, then the organization's own communication has shifted from ethics to safety, risk, and alignment, and this shift privileges technical staff and compliance roles over social-science and community perspectives.","Because the paper finds ethics almost absent from web articles and only slightly more present in academic publications, it implies that OpenAI's public ethics discourse is activated mainly in collaborative, academic contexts, not in its general-audience voice.","The paper's conclusion is that AI governance and accountability for OpenAI should be exogenously imposed—by regulation, multistakeholder dialogue, and transparency requirements—rather than left to internal, voluntary ethics statements.","The paper argues that the observed discourse is not neutral but reallocates epistemic authority, defining what counts as responsible AI and who counts as an expert."],"fun_headline_variants":["OpenAI's own posts: safety 700, ethics 7","Safety trumps ethics in OpenAI's discourse","Ethics-washing? OpenAI's risk talk drowns ethics","OpenAI's ethics gap: safety 700 vs ethics 7","In OpenAI's words, safety eclipses ethics 100:1"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the corpus—articles from OpenAI.com's News and Research sections plus linked preprints—captures OpenAI's substantive public ethical discourse; if substantive ethics discussions live on other OpenAI channels like the charter page, safety documents, or policy papers, the conclusion that OpenAI 'omits' ethics vocabulary would be overstated.","fun_headline_variants_meta":{"raw":{"variants":["OpenAI's own posts: safety 700, ethics 7","Safety trumps ethics in OpenAI's discourse","Ethics-washing? OpenAI's risk talk drowns ethics","OpenAI's ethics gap: safety 700 vs ethics 7","In OpenAI's words, safety eclipses ethics 100:1"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000316,"raw_usage":{"total_tokens":1615,"prompt_tokens":723,"completion_tokens":892,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":467,"completion_tokens_details":{"reasoning_tokens":807}},"tokens_in":467,"tokens_out":892,"duration_ms":6806,"temperature":1.0,"reasoning_tokens":807,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T08:31:37.514610+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: use the Wayback Machine to archive every page under openai.com that contains the word 'ethics' across all subdomains (including /charter, /safety, /global-affairs, and blog posts), and code each mention for substantive engagement (e.g., definitional discussion, trade-off analysis, or governance commitments) versus passing use. If the full-site audit yields dozens of documents with substantive ethics analysis, the paper's claim that ethics functions only rhetorically in OpenAI's public discourse fails.","supporting_citations":[],"review_version":1}