REVIEW 3 major objections 4 minor 35 references
Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper measures responsiveness at the level of individual regulatory obligations and finds that public engagement tracks revision only modestly.
desk verdict A genuinely new obligation-level audit framework with solid Findings 1 and 2, but the equity headline (Finding 3) rests on an unaudited submitter-type proxy that the paper itself concedes needs blind validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the regulatory obligation, defined as an actor-modal-action-object tuple describing who must do what in binding rule text. The framework has three load-bearing components: a structural deontic parser plus large-language-model verifier that extracts obligations from proposed and final rule text; an LLM-judged matcher that classifies whether a comment addresses a specific obligation and with what stance; and an outcome-state classifier that assigns proposed obligations to survived-unchanged, survived-edited, modified, dropped, or new using cosine similarity thresholds. Each component is evaluated against blind dual-annotator human judgment, and the audit of the outcome-state classifier is what drives the correction to Finding 3.
What would settle it
Run the planned blind validation audit on 200-300 stratified submitter labels: if the permissive Title-pattern classifier's organizational and individual assignments are substantially wrong, rebuild the audit-corrected 2x2 contingency with validated labels and check whether organizational-majority obligations still concentrate in editorial-refinement outcomes.
Extended reading notes
Core claim
The central discovery is that measuring responsiveness at the level of the discrete regulatory obligation rather than the rule, docket, or comment makes visible a pattern that coarser units obscure. Across 12,730 obligations from 29 rulemakings, obligations addressed by five or more commenters are revised 19.5 percentage points more often than less-addressed ones (chi-square = 24.53, p < 0.001), and a docket-fixed-effects logistic regression puts the within-docket odds ratio at 1.93 (95% CI [1.04, 3.57]). Direction of engagement does not separate outcomes: obligations with at least one opposing commenter were revised 62.3% of the time versus 68.3% for those with at least one supporting commenter (p = 0.155), a null the paper reads as inconsistent with simple preference aggregation. On the 115 obligations with sufficient engagement, a blind human audit of outcome labels preserves a cross-docket association between organizational-majority commenter composition and editorial-refinement outcomes (Fisher's exact OR = 4.90, p = 0.044), though the estimate is conditional on a permissive title-pattern reconstruction of commenter identity. The audit also shows that cosine text similarity agrees with human raters only at kappa = 0.137 on the editorial-versus-substantive boundary, so the framework rather than the classifier is the contribution.
Load-bearing premise
Finding 3 stands or falls on the assumption that the permissive Title-field pattern classifier, which tags titles containing 'submitted by,' 'on behalf of,' or 'written by' as individuals and everything else as organizations, reconstructs organizational versus individual commenter identity reliably enough to compare across dockets.
Editorial extensions
If this is right
- Agencies and researchers should report responsiveness at the obligation level, since rule-level summaries collapse the unit where change actually occurs.
- Strong-version responsiveness accounts that treat comments as primary drivers of provision-level outcomes are inconsistent with the modest within-docket odds ratio.
- Simple preference-aggregation models, in which opposition should track with revision, are not supported by the null directional finding.
- The equity asymmetry in rulemaking is better characterized as upstream engagement capacity, which populations can identify and contest specific legal duties, than as differential treatment within a shared docket.
- Text-shape similarity alone cannot separate editorial refinement from substantive compliance change; responsiveness measures need human-validated outcome labels at that boundary.
Reading between the lines
- If the framework were extended to other agencies, obligation-level auditing might reveal different responsiveness patterns; the EPA-specific high-attention sample limits generalization.
- A validated submitter-type label set, such as the paper's planned 200-300-sample audit, would tell whether Finding 3's cross-docket pattern is an artifact of the permissive title classifier; until then the equity claim should be read as conditional.
- Deduplicating mass-comment campaigns could shrink or erase the engagement-revision association; treating each comment as an independent voice may overstate Finding 1.
- Replacing the cosine outcome classifier with an LLM-based classifier audited on the same infrastructure could sharpen editorial-versus-substantive estimates and make the methodology reusable.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces an obligation-level framework for auditing responsiveness in EPA notice-and-comment rulemaking, using LLM-assisted extraction of regulatory obligations, comment-obligation matching, and proposed-final outcome classification, applied to 70,075 comments across 36 anchor dockets. It reports three descriptive findings: (F1) engagement is associated with revision at a modest within-docket magnitude (OR = 1.93, 95% CI [1.04, 3.57]); (F2) support-versus-opposition direction does not significantly differentiate revision outcomes; and (F3) under a permissive submitter-type reconstruction, organizational-majority engagement concentrates in editorial-refinement rather than substantive-modification outcomes at the cross-docket level (audit-corrected Fisher's exact OR = 4.90, p = 0.044, n = 115). The paper interprets this pattern as upstream structural inequity in which different commenter populations select into different types of rulemakings, rather than as evidence of within-docket differential treatment.
Significance. The obligation-level unit of analysis and the explicit, audited validation hierarchy are genuine contributions to computational administrative-law and accountability research. The paper is unusually transparent: blind dual-annotator audits are reported for extraction, matching, and outcome classification; robustness checks include leave-one-docket-out tests and alternative classifier specifications; and the authors repeatedly flag which components remain unvalidated. If the findings hold, F1 and F2 provide useful, robust descriptive evidence about the limits of preference-aggregation accounts of agency responsiveness. The measurement-validity result that cosine similarity is essentially invalid for the editorial-versus-substantive distinction (kappa = 0.137) is a valuable caution for the field. The central equity headline, F3, is, however, conditional on an unaudited submitter-type proxy and on a small, fragile audit-corrected contingency; that is the principal weakness preventing the paper from making a fully supported claim at this stage.
major comments (3)
- [4.1, 6.4, 8.2] Finding 3, the paper's equity headline, depends on a Title-field submitter-type classifier that is not validated anywhere in the manuscript. The permissive rule classifies any Title not containing 'submitted by', 'on behalf of', or 'written by' as organizational, which likely mislabels many individual submissions; the stricter ternary classifier yields only 4 organizational-majority obligations among 111 engaged obligations, leaving the contrast underpowered rather than confirmed. The audit-corrected contingency in Table 3 has an SE-not-org cell of 2 and a Fisher's exact p of 0.044 with an approximate CI lower bound near 1.0. The authors state in §8.2 that a blind validation audit on 200–300 stratified submitter labels is still required. Until that audit is performed, the cross-docket equity contrast is not established, and this is the weakest load-bearing link in the paper.
- [5, 8.1] The obligation-extraction pipeline is validated for precision on a single development rule (n = 66) with no recall measurement, and the obligation-comment matcher's retrieval step is not audited for recall (the prefilter uses cosine >= 0.3 and a 50-character minimum comment length; see §8.1). Because the engagement counts underlying F1 and F3 are conditional on retrieved comment-obligation pairs, a systematic recall deficit could alter the estimated engagement-revision association in unknown directions. The authors acknowledge this limitation in §8.1, but a below-threshold sampling check or a recall audit is needed to support the framework's core reliability claim.
- [6.5, 6.4] The blind audit shows that the cosine-similarity outcome classifier agrees with human raters at only kappa = 0.137 on the SE/MO contrast, meaning the classifier is invalid for the very distinction that drives F3. The paper correctly substitutes audit-corrected labels for the headline F3 estimate, but this leaves the only defensible F3 evidence as a single small contingency with n = 115 and one cell containing 2 observations. The pre-audit OR of 3.07 should not be presented as a classifier-based estimate of the effect; the audit-corrected estimate, while directionally consistent, is fragile and requires additional validation before the finding can be regarded as robust.
minor comments (4)
- [6.4, Table 3, Figure 4] The text refers to '116 engaged obligations' in several places while the audit-corrected contingency uses n = 115 after one obligation is reclassified; please standardize the denominator across the text, Table 3, and Figure 4 captions.
- [Abstract] The abstract states that the blind human audit 'preserves this third finding under corrected labels,' but the audit corrects outcome labels only, not the submitter-type proxy; consider rewording to make this distinction explicit.
- [Figure 1] The figure legend refers to green components as audited, but the figure is not visible in the text version; please confirm that the rendered figure matches the caption and that the color coding is accessible.
- [6.3] The relationship between the 12,730 total obligations and 12,243 proposed-side obligations is clear from the NEW category, but a brief parenthetical would avoid reader confusion about the decomposition.
Circularity Check
No significant circularity: the empirical associations are not derived from the measurement choices, and the outcome audit corrects rather than confirms the brittle classifier.
full rationale
Walking the derivation chain, each load-bearing step is an empirical measurement rather than a definitional reduction. Finding 1 compares matcher-derived engagement counts with independently extracted proposed-final outcomes; the logistic within-docket OR is estimated, not imposed by construction. Finding 2 compares stance labels from the audited matcher with revised-versus-not-revised outcomes; the null is data-driven. Finding 3 combines a separately measured submitter-type proxy with outcome labels; the audit-corrected contingency uses human judgments that replaced, rather than were derived from, the cosine classifier, whose poor agreement (kappa = 0.137 on the SE/MO contrast) is reported openly. The paper explicitly conditions Finding 3 on the permissive submitter-type reconstruction and lists a blind validation audit as required future work; that is an acknowledged validity limitation, not a claim that the reconstruction itself is the evidence. There are no self-citations by the authors, and the external benchmarks (Leahey's FRTracker protocol, Hansen's validation threshold) are used transparently rather than invoked to forbid alternatives. The mild concern that the authors designed and conducted their own audit rubric is a matter of rater independence, not circularity by construction. No equation or fitted parameter is renamed as a prediction; no central claim reduces to its input. Score 0.
Assumptions & free parameters
free parameters (5)
- engagement threshold =
>= 5 addressing commenters
- outcome-state cosine thresholds =
0.95 / 0.85 / 0.55
- matcher prefilter =
cosine >= 0.3 and comment length >= 50 characters
- cfr_section gate drop =
text-only similarity, no section co-occurrence gate
- permissive submitter-type classifier =
individual if title contains 'submitted by', 'on behalf of', or 'written by'; else organizational
assumptions (6)
- standard math Statistical test assumptions hold (chi-square, logistic regression with cluster-robust SE, Fisher exact).
- domain assumption regulations.gov Organization Name is universally PII-redacted, forcing submitter-type reconstruction from Title field.
- ad hoc to paper Title-field patterns separate individual from organizational submitters well enough for cross-docket comparisons.
- domain assumption The LLM verifier with closed enums produces valid obligation extractions beyond the single audited development rule.
- ad hoc to paper The SURVIVED-edited versus MODIFIED boundary captures procedural versus substantive regulatory change.
- domain assumption The 36 anchor rulemakings support cross-docket generalization of engagement patterns.
Cite this review
Pith. "Pith review of Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking." pith.science (2026). https://pith.science/paper/OEQXGMUQ
@misc{pith2026260810329,
author = {Pith},
title = {Pith review of: Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking},
year = {2026},
howpublished = {\url{https://pith.science/paper/OEQXGMUQ}},
note = {Machine review of arXiv:2608.10329}
}
read the original abstract
Notice-and-comment rulemaking gives any affected party the same formal right to influence federal regulation, but formal access is not substantive capacity to shape rule text. Existing strategies operate at the rule or aggregate-corpus level, too coarse to capture the discrete regulatory obligations where commenters seek change. We introduce obligation-level responsiveness auditing, an auditable, AI-assisted framework for measuring whether public-comment engagement co-occurs with changes to specific regulatory duties. The framework extracts proposed and final-rule obligations, matches comments to the obligations they address, and classifies proposed-final outcomes; each load-bearing component is evaluated against blind human judgment. We apply the framework to 70,075 comments across 36 EPA anchor rulemakings, drawn from a corpus of 786,197 comments across 6,145 dockets from 2010-2022. Three descriptive findings emerge. First, engagement is associated with revision at a modest within-docket magnitude. Second, support-versus-opposition direction does not clearly differentiate outcomes, an informative null inconsistent with simple preference-aggregation. Third, under a permissive reconstruction of commenter type, organizational-majority engagement concentrates in editorial-refinement rather than substantive-modification outcomes at the cross-docket level. A blind human audit of the load-bearing outcome contrast preserves this third finding under corrected labels and reveals that text-similarity methods are insufficient for distinguishing editorial from substantive regulatory change, a measurement-validity lesson we treat as a supporting methodological contribution. Together, these findings locate the equity asymmetry upstream of agency response: in differential capacity across commenter populations to identify, interpret, and contest specific legal obligations.
Figures
Reference graph
Works this paper leans on
-
[1]
Omar Al-Ubaydli and Patrick A. McLaughlin. 2017. RegData: A numerical data- base on industry-specific regulations for all United States industries and federal regulations, 1997–2012.Regulation & Governance11, 1 (2017), 109–123
work page 2017
-
[2]
Steven J. Balla, Reeve T. Bull, Bridget C.E. Dooling, Emily Hammond, Michael Herz, Michael A. Livermore, and Beth Simone Noveck. 2021.Mass, Computer- Generated, and Fraudulent Comments: Report for the Administrative Conference of the United States. Technical Report. Administrative Conference of the United States (ACUS), Washington, DC
work page 2021
-
[4]
A. Colin Cameron, Jonah B. Gelbach, and Douglas L. Miller. 2008. Bootstrap- based improvements for inference with clustered errors.Review of Economics and Statistics90, 3 (2008), 414–427
work page 2008
- [5]
-
[6]
Sarah H. Cen and Rohan Alur. 2024. From Transparency to Accountability and Back: A Discussion of Access and Evidence in AI Auditing. InProceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO ’24). doi:10.1145/3689904.3694711
arXiv 2024
-
[7]
Benjamin M. Chen and Brian Libgober. 2023. Do Administrative Procedures Fix Cognitive Biases?Journal of Public Administration Research and Theory34 (2023), 105–121. doi:10.1093/jopart/muac054
-
[8]
Cary Coglianese, Gabriel Scheffler, and Daniel E. Walters. 2021. Unrules.Stanford Law Review73 (2021), 885–967
work page 2021
-
[9]
Eric Corbett, Emily Denton, and Sheena Erete. 2023. Power and Public Participa- tion in AI. InProceedings of the 3rd ACM Conference on Equity and Access in Algo- rithms, Mechanisms, and Optimization (EAAMO ’23). doi:10.1145/3617694.3623228
arXiv 2023
Show all 35 references
-
[10]
Desmarais, and John A
Mia Costa, Bruce A. Desmarais, and John A. Hird. 2019. Public Comments’ Influence on Science Use in U.S. Rulemaking: The Case of EPA’s National Emission Standards.American Review of Public Administration49, 1 (2019), 36–50. doi:10. 1177/0275074018795287
2019
-
[11]
Mariano-Florentino Cuéllar. 2005. Rethinking Regulatory Democracy.Adminis- trative Law Review57, 2 (2005), 411–499
2005
-
[12]
Weiguo Dong, Ronghua Zhu, Michael Jensen, et al. 2026. An LLM-based NLP pipeline to assist government agencies in digesting massive public comments and mitigating spam.Information Processing and Management63 (2026), 104821. EAAMO ’26, November 5–7, 2026, Munich, Germany Fan and Yao
2026
-
[13]
Vlad Eidelman and Brian Grom. 2019. Argument Identification in Public Com- ments from eRulemaking. InProceedings of the 17th International Conference on Artificial Intelligence and Law (ICAIL ’19). arXiv:1905.00572
2019 arXiv
-
[14]
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatGPT outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences120, 30 (2023), e2305016120
2023
-
[15]
2026.Validating Large Language Model Annotations
Anne Lundgaard Hansen. 2026.Validating Large Language Model Annotations. Technical Report FEDS Working Paper 2026-020. Board of Governors of the Federal Reserve System. doi:10.17016/FEDS.2026.020
2026 doi
-
[16]
Moynihan
Pamela Herd and Donald P. Moynihan. 2018.Administrative Burden: Policymaking by Other Means. Russell Sage Foundation, New York
2018
-
[17]
Michael Heseltine and Bernhard Clemm von Hohenberg. 2024. Large language models as a substitute for human experts in annotating political text.Research & Politics(2024). doi:10.1177/20531680241236239
2024 doi
-
[18]
Janssen, Jiawei Wang, Anna Min-Venditti, Niranjan Karanjia, and John M
Sungsoo Kim, Marijn A. Janssen, Jiawei Wang, Anna Min-Venditti, Niranjan Karanjia, and John M. Anderies. 2026. All Public Voices Are Equal, But Are Some More Equal Than Others to LLMs? arXiv:2604.17247
2026 arXiv
-
[19]
Richard Landis and Gary G
J. Richard Landis and Gary G. Koch. 1977. The Measurement of Observer Agree- ment for Categorical Data.Biometrics33, 1 (1977), 159–174. doi:10.2307/2529310
1977 doi
-
[20]
Andrew Leahey. 2026. Only One-Third of Proposed Regulatory Obligations Survive to the Final Rule.Yale Journal on Regulation, Notice & Comment (2026). https://www.yalejreg.com/nc/only-one-third-of-proposed-regulatory- obligations-survive-to-the-final-rule/
2026
-
[21]
Brian Libgober and Steven Rashin. 2023. What Public Comments During Rulemaking Do (and Why).American Politics Research51, 6 (2023), 715–730. doi:10.1177/1532673X231175686
2023 doi
-
[22]
Aileen Love. 2026. Hearing, not heeding: procedural acknowledgment and substantive influence in rulemaking.Journal of Public Administration Research and Theory(2026). doi:10.1093/jopart/muag003
2026 doi
-
[23]
Nathan Mantel and William Haenszel. 1959. Statistical aspects of the analysis of data from retrospective studies of disease.Journal of the National Cancer Institute 22, 4 (1959), 719–748
1959
-
[24]
Paul Mohai and Robin Saha. 2007. Racial inequality in the distribution of haz- ardous waste: a national-level reassessment.Social Problems54, 3 (2007), 343–370
2007
-
[25]
Nicholas Pangakis, Samuel Wolken, and Neil Fasching. 2023. Automated Anno- tation with Generative AI Requires Validation.arXiv preprint arXiv:2306.00176 (2023)
2023 arXiv
-
[26]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of EMNLP 2019. arXiv:1908.10084
2019 arXiv
-
[27]
Mona Sloane, Emanuel Moss, Olaitan Awomolo, and Laura Forlano. 2022. Partic- ipation Is not a Design Fix for Machine Learning. InProceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO ’22). doi:10.1145/3551624.3555285
2022
-
[28]
Ray Smith. 2007. An Overview of the Tesseract OCR Engine. InProceedings of ICDAR 2007, Vol. 2. 629–633
2007
-
[29]
Tessum, Joshua S
Christopher W. Tessum, Joshua S. Apte, Andrew L. Goodkind, Nicholas Z. Muller, Kimberley A. Mullins, David A. Paolella, Stephen Polasky, Nathaniel P. Springer, Sumil K. Thakrar, Julian D. Marshall, and Jason D. Hill. 2019. Inequity in con- sumption of goods and services adds t...
2019
-
[30]
Government Accountability Office
U.S. Government Accountability Office. 2019.Federal Rulemaking — Selected Agen- cies Should Clearly Communicate Practices Associated with Identity Information in the Public Comment Process. Technical Report GAO-19-483. U.S. Government Accountability Office, Washington, DC
2019
-
[31]
Government Accountability Office
U.S. Government Accountability Office. 2020.Federal Rulemaking: Information on Selected Agencies’ Management of Public Comments. Technical Report GAO- 20-383R. U.S. Government Accountability Office, Washington, DC
2020
-
[32]
Government Accountability Office
U.S. Government Accountability Office. 2021.Federal Rulemaking: Selected Agen- cies Should Fully Describe Public Comment Data and Their Limitations. Technical Report GAO-21-103181. U.S. Government Accountability Office, Washington, DC
2021
-
[33]
Edwin B. Wilson. 1927. Probable Inference, the Law of Succession, and Statistical Inference.J. Amer. Statist. Assoc.22, 158 (1927), 209–212. doi:10.1080/01621459. 1927.10502953
1927
-
[34]
Jason Webb Yackee and Susan Webb Yackee. 2006. A Bias Towards Business? Assessing Interest Group Influence on the U.S. Bureaucracy.Journal of Politics 68, 1 (2006), 128–139. doi:10.1111/j.1468-2508.2006.00375.x
2006
-
[35]
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024. Can Large Language Models Transform Computational Social Sci- ence?Computational Linguistics50, 1 (2024), 237–291
2024
-
[6006]
doi:10.1073/pnas.1818859116
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.