{"id":"0286cb5c-6476-42c3-ac33-ec713c51f2ca","arxiv_id":"2608.10329","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Obligation-level auditing of 70,075 EPA comments shows comments are modestly linked to rule revisions, while organization-heavy dockets skew toward editorial rather than substantive changes.","lead":"This paper measures whether EPA rulemakers respond to public comments by tracking changes to individual regulatory duties, not whole rules. It finds engagement is modestly linked to revisions, while the mix of commenters helps explain which rules get engaged at all.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Finding 3 rests on an unaudited submitter-type classifier whose permissive rule likely labels many individuals as organizations; the stricter ternary version leaves only 4 org-majority obligations, so the headline equity contrast is not yet established.","rationale":"The manuscript is unusually transparent: it reports blind audits, flags the submitter-type proxy as a dependency (§4.1, §8.2), and presents Finding 3 as conditional. The extraction and matching audits are genuine independent support, and Findings 1 and 2 do not turn on the editorial-versus-substantive boundary. However, the single load-bearing point for the paper's equity claim is Finding 3, and every estimate in Finding 3 is conditional on an unaudited binary Title-pattern classifier that is plausibly biased in the direction of the result: comments without one of three phrases are all coded organizational. The paper's own stricter ternary classifier reduces the organizational rate from 28.6% to 14.4% and leaves only 4 of 111 engaged obligations organizational-majority, so the contrast is underpowered rather than confirmed. Because the audit-corrected SE-not-org cell has n=2 and the CI lower bound is about 1.0, small label corrections could erase the association. This does not change the reader's conditional verdict, but it is precisely why the paper should not move toward acceptance without the planned blind validation audit.","tokens_in":20833,"tokens_out":5197,"duration_ms":50972,"concrete_test":"Have two independent coders blind-label submitter type (organization vs. individual vs. ambiguous) for every comment associated with the 116 F3-determining obligations, using comment body text and attachment headers rather than the Title field; recompute per-obligation organizational-majority status and re-run the audit-corrected 2×2 Fisher's exact test with these human labels. If the OR falls toward 1 (or p > 0.05), Finding 3 is a classifier artifact; also report permissive-classifier agreement with the blind labels so the direction and size of the misclassification are explicit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Finding 3, the paper's equity headline, is conditional on a Title-field classifier (§4.1) that is not blind-audited anywhere in the manuscript. The permissive rule tags a comment as individual only if its title contains \"submitted by,\" \"on behalf of,\" or \"written by\"; all other comments are labeled organizational. A comment titled with a personal name or with no recognizable phrase is therefore assigned to the organizational category, plausibly inflating the organizational-majority set in exactly the direction of the reported association. The paper's own stricter ternary classifier drops the organizational rate from 28.6% to 14.4% and leaves only 4 of 111 engaged obligations organizational-majority, rendering the Finding 3 contrast underpowered rather than confirmed (§6.4, §8.2). The authors explicitly state that a blind validation audit on 200–300 stratified submitter labels is still required (§8.2). Given that the audit-corrected contingency has an SE-not-org cell of n=2 and a confidence interval whose lower bound is near 1.0, even modest submitter-label error could eliminate the association. This is the weakest load-bearing link in the paper: F1 and F2 do not depend on submitter type, and the extraction/matching audits are strong, but the central equity claim cannot be evaluated until the submitter-type reconstruction itself is validated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an obligation-level framework for auditing responsiveness in EPA notice-and-comment rulemaking, using LLM-assisted extraction of regulatory obligations, comment-obligation matching, and proposed-final outcome classification, applied to 70,075 comments across 36 anchor dockets. It reports three descriptive findings: (F1) engagement is associated with revision at a modest within-docket magnitude (OR = 1.93, 95% CI [1.04, 3.57]); (F2) support-versus-opposition direction does not significantly differentiate revision outcomes; and (F3) under a permissive submitter-type reconstruction, organizational-majority engagement concentrates in editorial-refinement rather than substantive-modification outcomes at the cross-docket level (audit-corrected Fisher's exact OR = 4.90, p = 0.044, n = 115). The paper interprets this pattern as upstream structural inequity in which different commenter populations select into different types of rulemakings, rather than as evidence of within-docket differential treatment.","tokens_in":21137,"tokens_out":4244,"duration_ms":37988,"significance":"The obligation-level unit of analysis and the explicit, audited validation hierarchy are genuine contributions to computational administrative-law and accountability research. The paper is unusually transparent: blind dual-annotator audits are reported for extraction, matching, and outcome classification; robustness checks include leave-one-docket-out tests and alternative classifier specifications; and the authors repeatedly flag which components remain unvalidated. If the findings hold, F1 and F2 provide useful, robust descriptive evidence about the limits of preference-aggregation accounts of agency responsiveness. The measurement-validity result that cosine similarity is essentially invalid for the editorial-versus-substantive distinction (kappa = 0.137) is a valuable caution for the field. The central equity headline, F3, is, however, conditional on an unaudited submitter-type proxy and on a small, fragile audit-corrected contingency; that is the principal weakness preventing the paper from making a fully supported claim at this stage.","major_comments":[{"comment":"Finding 3, the paper's equity headline, depends on a Title-field submitter-type classifier that is not validated anywhere in the manuscript. The permissive rule classifies any Title not containing 'submitted by', 'on behalf of', or 'written by' as organizational, which likely mislabels many individual submissions; the stricter ternary classifier yields only 4 organizational-majority obligations among 111 engaged obligations, leaving the contrast underpowered rather than confirmed. The audit-corrected contingency in Table 3 has an SE-not-org cell of 2 and a Fisher's exact p of 0.044 with an approximate CI lower bound near 1.0. The authors state in §8.2 that a blind validation audit on 200–300 stratified submitter labels is still required. Until that audit is performed, the cross-docket equity contrast is not established, and this is the weakest load-bearing link in the paper.","section":"4.1, 6.4, 8.2"},{"comment":"The obligation-extraction pipeline is validated for precision on a single development rule (n = 66) with no recall measurement, and the obligation-comment matcher's retrieval step is not audited for recall (the prefilter uses cosine >= 0.3 and a 50-character minimum comment length; see §8.1). Because the engagement counts underlying F1 and F3 are conditional on retrieved comment-obligation pairs, a systematic recall deficit could alter the estimated engagement-revision association in unknown directions. The authors acknowledge this limitation in §8.1, but a below-threshold sampling check or a recall audit is needed to support the framework's core reliability claim.","section":"5, 8.1"},{"comment":"The blind audit shows that the cosine-similarity outcome classifier agrees with human raters at only kappa = 0.137 on the SE/MO contrast, meaning the classifier is invalid for the very distinction that drives F3. The paper correctly substitutes audit-corrected labels for the headline F3 estimate, but this leaves the only defensible F3 evidence as a single small contingency with n = 115 and one cell containing 2 observations. The pre-audit OR of 3.07 should not be presented as a classifier-based estimate of the effect; the audit-corrected estimate, while directionally consistent, is fragile and requires additional validation before the finding can be regarded as robust.","section":"6.5, 6.4"}],"minor_comments":[{"comment":"The text refers to '116 engaged obligations' in several places while the audit-corrected contingency uses n = 115 after one obligation is reclassified; please standardize the denominator across the text, Table 3, and Figure 4 captions.","section":"6.4, Table 3, Figure 4"},{"comment":"The abstract states that the blind human audit 'preserves this third finding under corrected labels,' but the audit corrects outcome labels only, not the submitter-type proxy; consider rewording to make this distinction explicit.","section":"Abstract"},{"comment":"The figure legend refers to green components as audited, but the figure is not visible in the text version; please confirm that the rendered figure matches the caption and that the color coding is accessible.","section":"Figure 1"},{"comment":"The relationship between the 12,730 total obligations and 12,243 proposed-side obligations is clear from the NEW category, but a brief parenthetical would avoid reader confusion about the decomposition.","section":"6.3"}],"recommendation":"major_revision","confidential_remarks":"The paper's honest disclosure of its own limitations is commendable, but the central equity claim is conditional on a proxy classifier that the authors themselves identify as requiring validation. I would advise the editor to require the submitter-type validation audit, or a clearly labeled downgrade of Finding 3 to a hypothesis-generation result, before acceptance. The framework itself, and Findings 1 and 2, are otherwise sound and suitable for publication after the load-bearing validation gap is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe genuinely new thing here is the unit of analysis: discrete regulatory obligations rather than whole rules or comment corpora. That is not cosmetic. The paper shows that obligation-level measurement surfaces an engagement-revision association, a directional null, and a cross-docket organizational/individual outcome pattern that rule-level work would miss. The extraction and matching audits are strong on their face: K=1.0 on extraction, K=1.0/0.953 on matching, on samples of 66 and 100. The authors are unusually transparent about which labels are audited, descriptive, or deferred; the limitations section reads like an honest audit log rather than a hedge.\n\nFindings 1 and 2 hold up. F1's within-docket OR of 1.93 (95% CI 1.04–3.57) is modest and survives leave-one-docket-out. F2 is a null, carefully interpreted as informative rather than as a power failure, with the overlapping-subsets caveat acknowledged.\n\nF3 is the soft spot, and the stress-test note lands on the right place. The equity contrast depends on a submitter-type classifier reconstructed from Title-field patterns: titles containing 'submitted by', 'on behalf of', or 'written by' are individual; everything else is organizational. That permissive rule almost certainly labels many unnamed individuals as organizational. The stricter ternary classifier leaves only 4 of 111 engaged obligations organizational-majority, so the contrast is underpowered rather than confirmed. The audit-corrected Fisher's exact has one cell with n=2, p=0.044, and the authors explicitly state that a blind validation of 200–300 stratified submitter labels is still required. Modest label error could flip the association. The outcome-state audit is a real contribution—the cosine classifier's K=0.137 on the editorial/substantive boundary is a useful measurement lesson—but the audit was performed by the two authors using their own rubric, so shared blind spots cannot be ruled out.\n\nMinor issues: extraction precision is validated on one rule (n=66) with no recall measurement; matcher recall is unmeasured; anchor selection is biased toward high-attention dockets; mass-comment campaigns are not deduplicated. The authors flag all of these. Also, code and data were not accessible from the preprint, which makes the reproducibility claims hard to check.\n\nBottom line: the framework and Findings 1 and 2 deserve serious engagement; F3 is suggestive, not established. This deserves peer review, with submitter-type validation as a condition for the equity claim to stand. I'd bring it to reading group and would cite it for the obligation-level audit method.","headline":"A genuinely new obligation-level audit framework with solid Findings 1 and 2, but the equity headline (Finding 3) rests on an unaudited submitter-type proxy that the paper itself concedes needs blind validation.","tokens_in":21620,"tokens_out":1675,"would_cite":true,"duration_ms":16580,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper measures responsiveness at the level of individual regulatory obligations and finds that public engagement tracks revision only modestly.","keywords":["notice-and-comment rulemaking","regulatory responsiveness","obligation-level auditing","EPA rulemaking","public comments","LLM-assisted measurement","equity in administrative systems","algorithmic accountability"],"falsifier":"Run the planned blind validation audit on 200-300 stratified submitter labels: if the permissive Title-pattern classifier's organizational and individual assignments are substantially wrong, rebuild the audit-corrected 2x2 contingency with validated labels and check whether organizational-majority obligations still concentrate in editorial-refinement outcomes.","tokens_in":20596,"feed_emoji":"⚖️","tokens_out":7023,"duration_ms":59765,"temperature":0.7,"pith_summary":"Notice-and-comment rulemaking grants everyone the same formal right to comment, but this paper argues that formal access is not the same as substantive capacity to change regulation. It introduces obligation-level responsiveness auditing, which tracks the fate of each discrete regulatory duty from proposed to final rule and matches comments to the specific obligation they address. Applied to 70,075 comments across 36 EPA rulemakings, it finds that heavily addressed obligations are revised more often but only modestly, that support versus opposition does not clearly predict revision, and that organizational-majority engagement concentrates in editorial refinements rather than substantive modifications across dockets. The authors place the equity asymmetry upstream of agency response: in differential capacity to identify, interpret, and contest specific legal obligations, not simply in whether comments are heeded.","feed_headline":"Commented-on EPA provisions see nearly double revision odds","feed_subtitle":"An obligation-level audit of 70,075 EPA comments shows engagement helps modestly and locates the equity gap upstream.","key_machinery":"The central object is the regulatory obligation, defined as an actor-modal-action-object tuple describing who must do what in binding rule text. The framework has three load-bearing components: a structural deontic parser plus large-language-model verifier that extracts obligations from proposed and final rule text; an LLM-judged matcher that classifies whether a comment addresses a specific obligation and with what stance; and an outcome-state classifier that assigns proposed obligations to survived-unchanged, survived-edited, modified, dropped, or new using cosine similarity thresholds. Each component is evaluated against blind dual-annotator human judgment, and the audit of the outcome-state classifier is what drives the correction to Finding 3.","core_discovery":"The central discovery is that measuring responsiveness at the level of the discrete regulatory obligation rather than the rule, docket, or comment makes visible a pattern that coarser units obscure. Across 12,730 obligations from 29 rulemakings, obligations addressed by five or more commenters are revised 19.5 percentage points more often than less-addressed ones (chi-square = 24.53, p < 0.001), and a docket-fixed-effects logistic regression puts the within-docket odds ratio at 1.93 (95% CI [1.04, 3.57]). Direction of engagement does not separate outcomes: obligations with at least one opposing commenter were revised 62.3% of the time versus 68.3% for those with at least one supporting commenter (p = 0.155), a null the paper reads as inconsistent with simple preference aggregation. On the 115 obligations with sufficient engagement, a blind human audit of outcome labels preserves a cross-docket association between organizational-majority commenter composition and editorial-refinement outcomes (Fisher's exact OR = 4.90, p = 0.044), though the estimate is conditional on a permissive title-pattern reconstruction of commenter identity. The audit also shows that cosine text similarity agrees with human raters only at kappa = 0.137 on the editorial-versus-substantive boundary, so the framework rather than the classifier is the contribution.","pith_inferences":["If the framework were extended to other agencies, obligation-level auditing might reveal different responsiveness patterns; the EPA-specific high-attention sample limits generalization.","A validated submitter-type label set, such as the paper's planned 200-300-sample audit, would tell whether Finding 3's cross-docket pattern is an artifact of the permissive title classifier; until then the equity claim should be read as conditional.","Deduplicating mass-comment campaigns could shrink or erase the engagement-revision association; treating each comment as an independent voice may overstate Finding 1.","Replacing the cosine outcome classifier with an LLM-based classifier audited on the same infrastructure could sharpen editorial-versus-substantive estimates and make the methodology reusable."],"forward_implications":["Agencies and researchers should report responsiveness at the obligation level, since rule-level summaries collapse the unit where change actually occurs.","Strong-version responsiveness accounts that treat comments as primary drivers of provision-level outcomes are inconsistent with the modest within-docket odds ratio.","Simple preference-aggregation models, in which opposition should track with revision, are not supported by the null directional finding.","The equity asymmetry in rulemaking is better characterized as upstream engagement capacity, which populations can identify and contest specific legal duties, than as differential treatment within a shared docket.","Text-shape similarity alone cannot separate editorial refinement from substantive compliance change; responsiveness measures need human-validated outcome labels at that boundary."],"supporting_citations":[{"why":"Supplies the discrete-deontic-instruction unit of analysis, the structural extraction approach, and the three-category audit protocol the pipeline extends.","marker":"[20]"},{"why":"Defines the strong-version responsiveness account and preference-aggregation expectation that Findings 1 and 2 test against.","marker":"[21]"},{"why":"Provides the organized-interest influence baseline and business-bias account that Finding 3's equity comparison engages.","marker":"[34]"},{"why":"Operationalizes the procedural-versus-substantive responsiveness distinction and supplies the human-LLM kappa anchor for matcher validation.","marker":"[22]"},{"why":"Provides the sentence-transformer cosine similarity used to assign proposed-final outcome states, the classifier the audit shows is brittle.","marker":"[26]"},{"why":"Supplies the kappa agreement benchmarks used to interpret inter-rater and classifier-versus-human agreement in the audits.","marker":"[19]"}],"fun_headline_variants":["Obligation-level audit finds comment focus boosts EPA revisions","EPA rule revisions double when comments concentrate","Comment direction doesn't sway EPA rule changes, audit shows","Organizational comments lead to editorial tweaks in EPA rules","New audit pinpoints where public comments change EPA regulations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Finding 3 stands or falls on the assumption that the permissive Title-field pattern classifier, which tags titles containing 'submitted by,' 'on behalf of,' or 'written by' as individuals and everything else as organizations, reconstructs organizational versus individual commenter identity reliably enough to compare across dockets.","fun_headline_variants_meta":{"raw":{"variants":["Obligation-level audit finds comment focus boosts EPA revisions","EPA rule revisions double when comments concentrate","Comment direction doesn't sway EPA rule changes, audit shows","Organizational comments lead to editorial tweaks in EPA rules","New audit pinpoints where public comments change EPA regulations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000255,"raw_usage":{"total_tokens":1670,"prompt_tokens":1142,"completion_tokens":528,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":758,"completion_tokens_details":{"reasoning_tokens":452}},"tokens_in":758,"tokens_out":528,"duration_ms":5607,"temperature":1.0,"reasoning_tokens":452,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:21:57.710540+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the planned blind validation audit on 200-300 stratified submitter labels: if the permissive Title-pattern classifier's organizational and individual assignments are substantially wrong, rebuild the audit-corrected 2x2 contingency with validated labels and check whether organizational-majority obligations still concentrate in editorial-refinement outcomes.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the discrete-deontic-instruction unit of analysis, the structural extraction approach, and the three-category audit protocol the pipeline extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the strong-version responsiveness account and preference-aggregation expectation that Findings 1 and 2 test against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Operationalizes the procedural-versus-substantive responsiveness distinction and supplies the human-LLM kappa anchor for matcher validation."}],"review_version":1}