{"id":"fdd2e722-5081-43a4-9278-e419c147b788","arxiv_id":"2505.07468","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A systematic comparison of 44 primary U.S. and Chinese AI governance documents finds strong overlap on transparency, reliability, and pro-safety AI use, suggesting concrete topics for bilateral dialogue.","lead":"This paper analyzed 44 U.S. and Chinese AI policy and corporate documents and found that the two countries share strong concerns about transparency, reliability, and using AI for safety despite geopolitical tensions. It identifies specific topics for bilateral AI governance dialogue, which could help de-risk advanced AI through cooperation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corpus representativeness is the load-bearing risk: the overlap findings depend heavily on the now-rescinded Biden AI EO, and the paper never tests whether strong/moderate grades survive its removal.","rationale":"The reader's weakest assumption is corpus representativeness, and I agree that this is the most load-bearing premise. The empirical claim is not that individual documents mention similar words; it is that U.S. and Chinese positions overlap in ways that make bilateral dialogue promising. That inference requires the documents to stand in for the countries' positions. The paper's own Note, §2.2, §3.4, and §7 all concede the fragility: the U.S. picture is concentrated in an executive order that has since been rescinded, and the corpus is asymmetric by design. The internal coding procedure (two authors independently categorizing overlap cases, consensus resolution) is transparent but does not address external validity; no inter-coder reliability metric is reported, and no sensitivity analysis is provided. I would not reject the paper: the taxonomy is explicit, the qualitative comparisons are nuanced, and the authors flag the temporal limitation. But the conditional verdict is appropriate because a single concrete re-analysis—removing the EO or adding successor documents—could materially change the Table 5 grades. This concern is about external validity and robustness, not about internal inconsistency or author conduct.","tokens_in":16589,"tokens_out":6672,"duration_ms":67785,"concrete_test":"Re-run the coding with the Biden AI EO removed from the U.S. corpus while keeping the Chinese corpus fixed, and recompute the Table 5 overlap grades; additionally, if feasible, add the January 2025 Trump executive order and other post-rescission U.S. policy documents. If any currently strong category (especially convening or pro-safety use of AI) drops to moderate/weak or becomes unsupported because its evidence came mostly from the EO, the central claim is not robust to the paper's own acknowledged temporal fragility. A secondary check: recompute overlap based on the number of distinct issuing institutions per category rather than total quotes, to see whether the results are driven by one or two heavily quoted documents.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a systematic reading of 44 primary U.S. and Chinese documents reveals several areas of strong and moderate overlap in risk perception and governance approaches. That claim would be secure only if the selected corpus approximates each country's actual policy positions closely enough that the overlap grades are not artifacts of document choice. Two features make this condition fragile. First, the corpus is deliberately asymmetric: 30 Chinese vs. 14 U.S. documents, and the U.S. side is dominated by the now-rescinded Biden AI EO—161 of 189 U.S. quotes come from executive-branch documents, with the EO the largest single source (§3.4, Table 2). Second, the paper never tests robustness to alternative selections: the initial Note and §7 acknowledge the EO's rescission and the asymmetry, and §2.2 argues for policy continuity, but no analysis shows which Table 5 categories remain strong or moderate if the EO is dropped or if post-rescission U.S. documents (e.g., the January 2025 'Removing Barriers to American Leadership in AI' EO) are substituted. Because 'strong' governance overlaps such as convening and 'use of AI for pro-safety purposes' are illustrated mainly with EO provisions, their persistence is an empirical question the paper leaves open. If those categories collapse or drop to weak without the EO, the headline 'promising topics' result would be an artifact of a single now-invalid document, not a robust finding about U.S.-China common ground.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 44 primary U.S. and Chinese AI policy and corporate documents using an adapted version of the AGORA taxonomy, coding quoted spans for sociotechnical risks and governance approaches, then grading the degree of U.S.–China overlap as strong, moderate, or weak. It finds strong overlap on limited user transparency, poor reliability, the pro-safety role of AI, and stakeholder convening, and moderate overlap on robustness, bias, interpretability, dangerous capabilities, cybersecurity, external auditing, licensing/registration, watermarking, model evaluations, and adversarial testing. Based on these grades, the paper recommends specific topics for bilateral dialogue, including national-security-relevant evaluation standards, technical standardization, content provenance, and Track II discussions. The paper explicitly acknowledges that the analysis was completed before the rescission of the Biden AI Executive Order, that the corpus is asymmetric (about twice as many Chinese as U.S. documents), and that the coding is qualitative without formal inter-coder reliability metrics.","tokens_in":16800,"tokens_out":5155,"duration_ms":46707,"significance":"If the overlap findings are robust, the paper is a valuable and timely contribution: it provides systematic, original-language evidence for concrete common ground in U.S.–China AI governance, an area where claims of convergence or divergence are often made but rarely documented at the level of specific policy provisions. The use of fluent researchers for Chinese-language documents, the grounding in the AGORA taxonomy, and the explicit separation of risk perception from governance approaches are notable strengths. The paper is also appropriately cautious, noting that overlap does not imply endorsement or guaranteed dialogue value. However, the significance is conditional on the representativeness of the corpus, particularly the heavy reliance on the now-rescinded Biden AI EO, and on the reliability of the qualitative overlap grading. These issues are fixable but currently leave the headline result more fragile than the authors' framing suggests.","major_comments":[{"comment":"The U.S. side of the corpus is dominated by the rescinded Biden AI EO: 161 of 189 U.S. quotes come from executive-branch documents, with the EO the largest single source. The paper acknowledges this in the initial Note, §2.2, and §7, but it never tests how the overlap grades in Table 5 change if the EO is removed or if post-rescission U.S. documents (e.g., the January 2025 executive order cited as [51]) are substituted. Because several categories graded 'strong' or 'moderate'—notably convening, the pro-safety role of AI, licensing/registration, external auditing, and watermarking—are illustrated primarily with EO provisions, the headline result could be an artifact of a single now-invalid document. I request a robustness analysis: re-grade the overlap categories excluding the EO and any other rescinded provisions, and report which grades persist, weaken, or disappear. This is essential for the paper's central claim that these topics are promising for ongoing dialogue.","section":"§3.4, Table 2; §7"},{"comment":"The strong/moderate/weak classification criteria in Table 4 are defined qualitatively ('most, if not all' for strong; 'some overlap... key differences' for moderate), and the paper does not operationalize the thresholds or report inter-coder reliability metrics. The text states that two authors independently categorized each case of overlap and reached consensus, but no agreement statistics or per-category evidence counts are given. Without a table showing, for each category, how many U.S. and Chinese documents/quotes were coded under that category, the reader cannot assess whether a grade of 'strong' rests on broad cross-document support or on a small number of quotes. Please provide a supplemental evidence table (or an appendix) with quote counts per category per country, and clarify the decision rules used to map the presence of quotes to the three overlap grades.","section":"§3.2, §3.5, Table 4"},{"comment":"The corpus asymmetry (30 Chinese vs. 14 U.S. documents) and the differing sampling strategies—broad inclusion of Chinese corporate whitepapers versus focused sampling of model cards from three U.S. frontier labs—may bias the comparison toward convergence. A narrower U.S. set will tend to find fewer divergent positions, making 'overlap' easier to achieve. The paper acknowledges the numerical asymmetry but does not test how sensitive the Table 5 results are to alternative document selections, such as adding more U.S. corporate governance documents or excluding Chinese state/party documents that have no U.S. counterpart. At minimum, the authors should discuss whether the current design under- or over-estimates overlap, and ideally should run a sensitivity check with a more balanced corpus to show that the strong and moderate grades are not an artifact of selection.","section":"§3.1, §3.4"}],"minor_comments":[{"comment":"The phrase 'appears oto be different' should read 'appears to be different'.","section":"§5.2.3"},{"comment":"The phrase 'cloud compute U.S.ge' should read 'cloud compute usage'.","section":"§5.2.6"},{"comment":"The FLOPS thresholds are rendered as '1026' and '1023'; they should be typeset as 10^26 and 10^23 with proper superscripts.","section":"§5.2.6"},{"comment":"Table 5 lists 'Pro-safety role of AI systems' while the section heading is 'Use of AI for pro-safety purposes'; harmonize the terminology across the paper.","section":"Table 5 vs §5.1.1"},{"comment":"The text says the analysis is restricted to 'areas of high and moderate overlap', but the grading scheme in Table 4 uses 'strong' rather than 'high'; use consistent labels.","section":"§3.5"},{"comment":"The acronym list includes UNCLOS and EEZ, which are not used in the body of the paper; either remove them or explain their relevance.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The central concern—corpus representativeness and the dependency of the overlap grades on the rescinded Biden AI EO—is fixable with a robustness analysis, which I would recommend making a condition of acceptance. The paper is well-written, honest about its limitations, and addresses a timely and practically important question. I see no citation-pattern or scope issues; the authors' use of translations by a co-author is transparent and appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read if you work on AI governance or US-China relations. The paper does something genuinely useful: it codes 44 primary documents in the original languages and maps where US and Chinese actors converge on risk perception and governance approaches, using an adapted AGORA taxonomy. That granularity is new relative to earlier comparative work, and the authors are careful to distinguish strong, moderate, and weak overlap rather than forcing a binary. The quotes are documented, the coding framework is in the appendix, and the limitations section is honest about the fast-moving policy landscape.\n\nThe main empirical claim—that there are several areas of strong and moderate overlap—holds up better than the stress-test note suggests. The two 'strong' risk categories, limited user transparency and poor reliability, rest on NIST, Chinese standards, and MIIT, not just the Biden EO. The governance overlaps (pro-safety use of AI, convening) do lean more heavily on the EO, but they also have support from Tencent, Zhipu, CAICT, and Chinese standards. So the rescission of the EO would weaken but not collapse the headline result.\n\nThe real soft spots are two. First, the corpus asymmetry (30 Chinese vs 14 US documents) and the heavy weight on the Biden EO mean that the overlap grades could shift with a different document set. The authors acknowledge this but do not test robustness—a sensitivity analysis dropping the EO or substituting the January 2025 Trump order would have been straightforward and would have made the claims much stronger. Second, the coding is qualitative with no reported inter-coder reliability; the authors describe a collaborative review process but do not quantify agreement. That is a limitation, not a fatal flaw, for a scoping study of this kind.\n\nThere is also a minor issue: the paper was written before the EO rescission and the initial Note is a bit awkward, but Section 2.2 and Section 7 deal with it directly. The citation pattern looks clean; the self-citations (Ding's translations) support access, not the central claim.\n\nWho is this for? AI governance researchers, diplomats, and people running Track II processes. It is a solid scoping study, not a definitive measurement. I would send it to peer review—it already is peer-reviewed, but if this were a new submission I would not desk-reject it. The claims are appropriately qualified and the method is transparent. My recommendation: engage with it, but ask for a sensitivity analysis and clearer coding reliability if it goes through another round.","headline":"A transparent, original-language comparison of US and Chinese AI governance documents that delivers a plausible map of common ground, with a real but non-fatal weakness in corpus representativeness.","tokens_in":17405,"tokens_out":2131,"would_cite":true,"duration_ms":19779,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a systematic comparison of 44 primary U.S. and Chinese AI governance documents reveals strong or moderate overlap in most major risk and governance categories, giving concrete topics for bilateral dialogue.","keywords":["AI governance","US-China relations","risk perception","governance approaches","document analysis","bilateral dialogue","AI policy"],"falsifier":"A direct check would be to re-run the same coding on a balanced post-2025 corpus: equal numbers of U.S. and Chinese government documents, current U.S. policy statements in place of the rescinded order, and updated Chinese standards. If the strong overlaps shrink to moderate or the moderate list changes substantially, the original classification was sample-driven; if the same categories persist, the finding is stable.","tokens_in":16331,"feed_emoji":"🤝","tokens_out":6544,"duration_ms":58518,"temperature":0.7,"pith_summary":"The paper asks whether the United States and China, despite strategic rivalry, have enough shared concerns about AI to make bilateral governance dialogue worthwhile. It answers by reading 44 primary government and corporate documents from both countries in their original languages, coding them for how they describe AI risks and what governance tools they favor. The central claim is that risk perceptions converge strongly on limited user transparency and poor reliability, and governance approaches converge strongly on using AI for safety and on multi-stakeholder convening. A broader set of areas—bias, robustness, dangerous capabilities, weak cybersecurity, external auditing, licensing or registration, watermarking, model evaluations, and adversarial testing—shows moderate overlap. If this mapping is right, the two governments have concrete, usable topics for talks even where their overall regulatory philosophies differ.","feed_headline":"44 AI documents show where US and China can talk","feed_subtitle":"Government and corporate texts converge on transparency, reliability, safety uses, and convening, suggesting concrete dialogue topics.","key_machinery":"The load-bearing instrument is an adapted coding taxonomy that classifies each passage of a governance document into named risk categories and named governance approaches, together with a three-level grading rule that labels cross-national agreement strong, moderate, or weak. The taxonomy's definitions determine what counts as a match: for example, transparency is compared by whether both sides require disclosure of how a system works to users, and adversarial testing is compared by whether red-team obligations exist and who bears them. This machinery lets the authors convert qualitative documents into comparable positions and then see where the two countries land in the same cell.","core_discovery":"The discovery is empirical: a systematic side-by-side reading of U.S. and Chinese AI policy texts finds more common ground than the conflict narrative suggests. On risk perception, the authors classify transparency and reliability as strong overlap, meaning U.S. and Chinese sources understand the problem in similar terms; they classify robustness, bias, interpretability, dangerous capabilities, and cybersecurity as moderate overlap, with real agreement but differing frames—for example, Chinese documents often treat chemical and biological risks under 'content security' while U.S. documents talk about capability proliferation. On governance, both sides strongly endorse using AI for safety and convening stakeholders; they moderately converge on external auditing, licensing and registration, watermarking, model evaluations, and adversarial testing. Privacy is the one risk category judged weakly aligned. The authors take these overlaps as evidence that dialogue topics can be selected pragmatically rather than waiting for agreement on first principles.","pith_inferences":["A balanced re-test with equal document counts per country and with current U.S. policy texts replacing the rescinded executive order would likely preserve some overlaps but could downgrade several moderate categories; the strong categories are the most sample-sensitive because they rest heavily on that single U.S. document.","The overlap on 'dangerous capabilities' may be shallower than the moderate label suggests: Chinese sources largely frame the issue as prohibited content while U.S. sources frame it as capability access, so agreeing on red lines would require resolving that frame difference first.","A testable extension would be to code post-2025 U.S. executive actions and China's updated generative-AI safety standard with the same taxonomy and check whether the strong-overlap set shrinks, which would locate how much of the finding was tied to the 2023–2024 document window.","Because companies on both sides read each other's technical reports, evaluation practices such as dangerous-capability testing could spread across borders even if formal regulation diverges, making corporate documents a leading indicator for future government alignment."],"forward_implications":["Dialogues can open with the strongly aligned topics—transparency, reliability, AI-for-safety, and multi-stakeholder convening—where documents on both sides already speak in compatible terms.","Technical standards-setting talks could make early progress on commercial product safety, where reliability, robustness, and adversarial testing overlap, without touching national-security topics.","Content watermarking and provenance coordination is a concrete candidate for shared industry engagement, since both governments and leading firms support labeling mechanisms.","Model evaluations and red-teaming are a middle ground: both sides require them, but talks would need to bridge the U.S. focus on dangerous capabilities with China's focus on content security.","Privacy is the weak spot in risk perception, so bilateral dialogue should not begin there if the goal is quick common ground."],"supporting_citations":[{"why":"Supplies most U.S. government positions on risks and governance; its later rescission is the paper's stated temporal fragility.","marker":"[2]"},{"why":"Provides the U.S. definitions of transparency, reliability, and robustness that anchor the strong-overlap comparisons.","marker":"[35]"},{"why":"Gives China's technical safety requirements for generative AI, used heavily in transparency, content safety, and monitoring.","marker":"[39]"},{"why":"Sets China's baseline generative AI service obligations, shaping the moderate-overlap governance findings.","marker":"[21]"},{"why":"Supports the Chinese side of transparency and algorithm-registration comparisons.","marker":"[19]"},{"why":"Represents the U.S. corporate framing of dangerous capabilities and pre-deployment evaluation.","marker":"[41]"},{"why":"Supplies Chinese corporate positions on cybersecurity, model-weight protection, and red-teaming.","marker":"[49]"},{"why":"Provides Chinese corporate views on robustness, external auditing, and privacy-preserving technical measures.","marker":"[3]"}],"fun_headline_variants":["US-China AI documents show shared risk perceptions","Common ground on AI governance found in US-China papers","US and China align on AI safety and transparency","AI policy overlap suggests US-China dialogue paths"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the assumption that the chosen 44 documents—about twice as many Chinese as U.S. texts, with the U.S. side heavily represented by a 2023 executive order that was later rescinded—reflect each country's actual governance positions closely enough that the measured overlaps are not artifacts of the sample.","fun_headline_variants_meta":{"raw":{"variants":["US-China AI documents show shared risk perceptions","Common ground on AI governance found in US-China papers","US and China align on AI safety and transparency","AI policy overlap suggests US-China dialogue paths"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1361,"prompt_tokens":923,"completion_tokens":438,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":377}},"tokens_in":539,"tokens_out":438,"duration_ms":4712,"temperature":1.0,"reasoning_tokens":377,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:15:28.387586+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check would be to re-run the same coding on a balanced post-2025 corpus: equal numbers of U.S. and Chinese government documents, current U.S. policy statements in place of the rescinded order, and updated Chinese standards. If the strong overlaps shrink to moderate or the moderate list changes substantially, the original classification was sample-driven; if the same categories persist, the finding is stable.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies most U.S. government positions on risks and governance; its later rescission is the paper's stated temporal fragility."},{"cited_title":"Artificial intelligence risk management framework (AI RMF 1.0)","cited_arxiv_id":null,"evidence_quote":"Provides the U.S. definitions of transparency, reliability, and robustness that anchor the strong-overlap comparisons."},{"cited_title":"Basic safety requirements for Generative Artificial Intelligence services","cited_arxiv_id":null,"evidence_quote":"Gives China's technical safety requirements for generative AI, used heavily in transparency, content safety, and monitoring."},{"cited_title":"Interim measures for the management of Generative Artificial Intelligence services, 7 2023","cited_arxiv_id":null,"evidence_quote":"Sets China's baseline generative AI service obligations, shaping the moderate-overlap governance findings."},{"cited_title":"Provisions on the management of algorithmic recommendations in Internet information services","cited_arxiv_id":null,"evidence_quote":"Supports the Chinese side of transparency and algorithm-registration comparisons."},{"cited_title":"Ten- cent large model security and safety report","cited_arxiv_id":null,"evidence_quote":"Supplies Chinese corporate positions on cybersecurity, model-weight protection, and red-teaming."},{"cited_title":"White paper on the governance and use of Generative Artificial Intelligence","cited_arxiv_id":null,"evidence_quote":"Provides Chinese corporate views on robustness, external auditing, and privacy-preserving technical measures."}],"review_version":1}