{"id":"a11073eb-16bf-4321-873e-1010f4675641","arxiv_id":"2502.14869","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A two-survey participatory pipeline turns laypeople's brainstormed stakeholder-action pairs for generative AI harms in the news environment into prioritized, LLM-generated policy fact sheets.","lead":"This paper develops a participatory method for AI governance: laypeople read short fictional scenarios about generative AI harms in news, propose who should act and what they should do, then a second group ranks those proposals by priority and agreement. The authors argue this gives policymakers an empirical, democratic complement to expert-only risk assessments.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lay-input enrichment claim rests on an unmeasured comparison: without an expert-generated baseline for the same impact types, the paper cannot show that its participatory SAPs add value over expert-centered methods.","rationale":"The reader's weakest assumption identifies the absence of a comparison against expert-generated mitigation lists, and my reading converges on that as the most load-bearing gap. The paper is a clearly described proof of concept, and its internal mechanics appear coherent, but the central value proposition is that lay input enriches expert-centered approaches. Without a baseline, the paper cannot distinguish between 'lay stakeholders add distinctive knowledge' and 'lay stakeholders reproduce a subset of standard expert recommendations.' The single-expert validation of fact sheets is a related but secondary weakness: it affects the policy-delivery component rather than the core participatory-mapping claim. I agree with the reader's CONDITIONAL assessment: the concern does not invalidate the method, but it makes the central claim depend on future comparative validation. Therefore the verdict should remain unchanged rather than being raised to ACCEPT or lowered to REJECT.","tokens_in":26951,"tokens_out":2850,"duration_ms":31786,"concrete_test":"Run a matched expert baseline: recruit 10–15 AI governance and media policy experts and ask them to generate SAPs for the same ten impact types using the same scenario prompts from Section 3.2 (or a standard expert-focused elicitation format). Have blind coders measure the overlap between the expert SAPs and the 228 lay SAPs, and rate both sets on specificity, actionability, and policy feasibility. If the expert lists cover nearly all lay-unique SAPs or are rated more actionable, the central 'enrichment' claim is not supported; if the lay SAPs contain materially novel, feasible actions or responsibility allocations that experts omit, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this participatory approach enriches expert-centered risk mitigation. To support \"enriches,\" the paper needs evidence that lay-generated stakeholder-action pairs contain something beyond what experts would produce. Survey 1 asks lay participants to brainstorm SAPs after reading LLM-written scenarios, and Survey 2 ranks those SAPs, but there is no expert condition and no comparison to expert-generated mitigation lists. The resulting SAPs largely align with obvious expert categories: government, technology companies, news publishers, fact-checking, transparency, and education. Those categories also appear in standard regulatory frameworks and industry guidance, so the distinctive contribution—lay blind spots, novel responsibility allocations, or materially different action proposals—is asserted rather than demonstrated. Section 5.4 candidly acknowledges the single-expert validation and the lack of policymaker testing, which weakens the downstream fact-sheet claim, but the more load-bearing gap is the missing baseline: if lay SAPs are mostly redundant with existing expert checklists, the democratic-enrichment justification for the method collapses, even though the pipeline itself may still function as a proof of concept. This is not an internal inconsistency; it is an absent comparison needed to support the paper's stated value proposition.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a participatory, forward-looking method for eliciting lay-stakeholder input to inform AI policy. Using ten negative impact types of generative AI in the media environment (selected from prior scenario-based work), the authors run two Prolific surveys: Survey 1 asks forty participants to brainstorm stakeholder-action pairs (SAPs) after reading LLM-written scenarios; Survey 2 asks eighty-six participants to rate a subset of the resulting 228 SAPs for agreement and priority. The ranked SAPs are then synthesized into one-page policy fact sheets using GPT-4o with prompt engineering. The authors report descriptive findings: government, technology companies, and news publishers are most frequently named as responsible actors; fact-checking is the most highly valued action; and actions involving bans or limits are generally rated lower in priority and agreement. They frame the work as a proof-of-concept to enrich expert-centered risk mitigation with democratic, participatory input.","tokens_in":27183,"tokens_out":5648,"duration_ms":49877,"significance":"If the method is taken as a proof of concept rather than as an effectiveness demonstration, it makes a useful contribution: it shows a concrete, transparent pipeline for turning brief, scenario-based stimuli into a structured map of lay-proposed responsibilities and actions, and it makes all ranked tables and stimuli available in the appendix. The paper is unusually candid about its limitations, including non-representative sampling, single-expert validation of the fact sheets, and the absence of policy-maker testing. It also ships reproducible materials, fair-pay recruitment details, and a positionality statement, which are strengths. The main significance hinges on whether the 'enrichment' claim over expert-centered methods is actually demonstrated; as it stands, the paper establishes feasibility of the pipeline but not the added value relative to existing expert-generated mitigation frameworks.","major_comments":[{"comment":"The central claim that the approach 'enriches the development of risk mitigation strategies' is not supported by a comparison against expert-generated mitigation strategies. The two surveys elicit lay SAPs and rank them, but there is no expert baseline for the same ten impact types, so the paper cannot show that lay input contributes non-redundant actions, novel responsibility allocations, or blind spots beyond expert-centered checklists. I recommend adding an expert or literature-baseline condition (e.g., having policy/domain experts produce SAPs for the same scenarios, or coding the lay SAPs against existing risk frameworks such as the EU AI Act or NIST AI RMF) and reporting overlap and novelty. If such a comparison is outside the scope, the abstract and Section 5 should be reworded to claim feasibility of eliciting and ranking lay input rather than 'enrichment.'","section":"Abstract; §5"},{"comment":"The selection of the ten impact types is not fully data-driven. The paper first computes a weighted average of severity, plausibility, magnitude, and specificity (with severity weighted 0.4), but then one author replaces three of the top ten ('Labor: Changing Job Roles', 'Media Quality: Clickbait', 'Political: Opinion Monopoly') with three from the top 20 based on the question 'Which impact types ... are likely receptive to policy intervention?'. This expert overlay changes the set that all downstream results depend on, and the rationale for the replacements is not justified beyond a one-sentence question. Please report the full ranked list, the criteria used for the replacements, and a sensitivity analysis (e.g., whether the main findings about government/tech/news-publisher responsibility allocations and fact-checking value hold under alternative impact-type selections).","section":"§3.1"},{"comment":"The consensus threshold is arbitrary and partly circular. The paper states that 'the average standard deviation of the agreement scores was 1.5, so we designate any score with a standard deviation less than or equal to 1.5 as higher consensus and those above 1.5 as lower consensus.' Because 1.5 is the average of the sample, this rule mechanically labels roughly half the SAPs as lower consensus regardless of the actual degree of agreement. This categorization feeds into Section 4.2 and 5.2's claims about 'contested' actions. Either justify a substantive threshold (e.g., SD < 1 on a 7-point scale) or report the continuous SD values and use a regression or correlation analysis for agreement 'controversiality', and test robustness of the conclusions to the threshold.","section":"§3.3"},{"comment":"The paper generalizes from small, non-representative Prolific samples (N=40 and N=86) to 'lay stakeholders' and 'the public'. While Section 5.4 acknowledges non-representativeness, Section 4.2 and 5.1 use phrases like 'in the eyes of the public' and 'in the eyes of laypeople' without qualification. For a policy-facing method, this is more than a presentational issue: the estimates of priority and agreement are sample-specific and likely influenced by the US-context scenarios. Recommend consistently framing all results as 'in this sample' and adding an explicit statement that the method's output is illustrative for policy exploration, not an estimate of population preferences.","section":"§5.1; §5.4"}],"minor_comments":[{"comment":"The paper says the policy fact sheets were 'validated by an expert' and 'guide the usefulness,' while Section 5.4 states validation with policy makers was not performed. The wording should be consistent; suggest 'reviewed by one policy expert' rather than 'validated.'","section":"§3.5 / §5.3"},{"comment":"The claim that human-written scenarios 'tend to be more complex, often intertwining several impact types' is presented without evidence or citation; please soften or support it.","section":"§3.2"},{"comment":"Table 8 skips the number 15 in the SAP numbering, Table 9 has two items numbered 19, and Table 4 has a row with the stakeholder label split by the table line break. Please correct numbering and formatting for readability.","section":"Appendix A.3"},{"comment":"Some percentages in Figure 2/Figure 4 are not fully defined (e.g., how the 228 SAPs distribute across the 12 stakeholders and 41 actions). Adding a counts table would improve transparency.","section":"§4.1"},{"comment":"There is a grammatical typo: 'lay people perceptions' should be 'lay people's perceptions.'","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is readable, transparent, and unusually honest about its limitations. My main concern is that the 'enrichment' claim is not tested, and the arbitrary consensus threshold and non-representative samples could be addressed with sensitivity analyses and careful wording. The paper fits a participatory-governance venue and could become a solid proof-of-concept contribution after the comparison/reframing and the methodological fixes above."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, honest methods paper with new empirical data. The main claim—that lay input enriches expert-centered mitigation—is plausible but not yet demonstrated, because there's no expert baseline. It deserves a serious referee.\n\nThe genuinely new thing is the pipeline: LLM-written scenarios, lay brainstorming of stakeholder-action pairs, a second survey ranking those pairs by agreement and priority, then LLM-generated policy fact sheets. The integration is new, and the empirical core—228 SAPs across ten impact types, plus priority rankings from a separate sample—is new and descriptively useful. The paper is also unusually candid: the limitations section admits the sample is not representative, the fact sheets are validated by one author, and the recommendations are raw and untested against existing law. That honesty buys credit.\n\nThe descriptive findings are worth keeping. Lay stakeholders assign responsibility beyond the usual regulatory suspects—schools, unions, local communities—and they rank outright bans and limits low while favoring fact-checking and transparency. Those patterns are not simply predictable from the EU AI Act's provider/deployer framing.\n\nThe soft spots, in proportion. The main one is exactly what your stress-test flags: the paper claims to 'enrich' expert-centered mitigation but never compares lay-generated SAPs to an expert-generated list for the same impact types. Without that baseline, 'enrich' is asserted. I don't think the claim collapses—the school, union, and anti-ban findings are non-obvious—but it's the gap to fix. Secondary weaknesses: impact-type selection mixes a weighted score with one author's manual replacement of three types; the SD 1.5 consensus cutoff is arbitrary; the samples are small and non-representative; and fact-sheet usefulness rests on one expert. The paper acknowledges most of this, and none of it breaks the proof-of-concept.\n\nWho it's for: people building participatory foresight methods for AI governance, and researchers wanting a worked example of SAP elicitation. It's a methods paper, not a policy paper. I'd bring it to reading group, and I'd cite it if I were working on participatory AI assessment. Recommend conditional acceptance, with the expert-baseline comparison as the key revision.","headline":"A solid, honest methods paper with new empirical data; the \"enriches expert mitigation\" claim needs an expert baseline before it's demonstrated.","tokens_in":27694,"tokens_out":2838,"would_cite":true,"duration_ms":25401,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new participatory method lets laypeople map who should fix AI harms, producing 228 stakeholder-action pairs that prioritize fact-checking over outright bans.","keywords":["participatory governance","stakeholder-action pairs","generative AI risk mitigation","scenario-based surveys","anticipatory governance","policy fact sheets","lay stakeholder input","AI impact assessment"],"falsifier":"Run the same two-survey pipeline on a nationally representative sample and have professional policy analysts work with the resulting fact sheets alongside expert-authored mitigation summaries; if the lay-derived sheets change no decisions or simply duplicate the expert lists, the method's added value is contradicted.","tokens_in":26746,"feed_emoji":"🗳️","tokens_out":4199,"duration_ms":35970,"temperature":0.7,"pith_summary":"The paper tries to establish a participatory method for AI risk mitigation that starts from lay stakeholders rather than experts. The authors show that when ordinary people are given short fictional scenarios of generative AI harms in the news environment, they can brainstorm mitigation ideas and prioritize them, yielding 228 stakeholder-action pairs across ten impact types. These pairs name who should act (most often government, technology companies, and news publishers) and what they should do, with fact-checking emerging as the most valued action. The authors argue this approach fills blind spots in expert-centered impact assessments and provides empirically grounded input for policymakers, delivered as LLM-generated one-page fact sheets.","feed_headline":"Public maps 228 fixes for AI harms, with fact-checking on top","feed_subtitle":"A two-survey method turns ordinary people's priorities into policy fact sheets.","key_machinery":"The central object is the stakeholder-action pair (SAP): an assignment of a specific mitigating or preventing action to a specific actor, e.g., 'news publishers should fact check AI-generated content.' The method pairs these with GPT-4-generated narrative scenarios that contextualize each impact for lay readers, a two-survey workflow (brainstorming, then agreement/priority ranking), and a final GPT-4o prompt step that converts ranked SAPs into one-page policy fact sheets under an expert's guidance. The machinery carries the burden of showing that non-expert input can be both broad and prioritized, and that it can be translated into a usable policy artifact.","core_discovery":"Lay stakeholders can produce a usable map of responsibility for AI harm mitigation. Survey 1 had 40 US participants brainstorm, after reading GPT-4-written scenarios, stakeholder-action pairs (SAPs) for ten negative impacts such as fake news, manipulation, and addiction; after consolidation this produced 228 SAPs, 12 stakeholder types, and 41 actions. Survey 2 had 86 different participants rate each SAP for agreement and priority. The rankings show fact-checking as the highest priority across impact types, assigned mainly to news publishers, technology companies, and social media platforms; government was the most frequently named actor in brainstorming yet fell to fourth in priority rankings; and outright bans or heavy restrictions consistently ranked low. The paper claims this demonstrates that a broader participative base can enrich expert-driven mitigation strategies and inform policy in a format policymakers can actually use.","pith_inferences":["The paper does not compare its SAP output against expert-generated mitigation lists, so the strongest version of its claim—that lay input adds information experts would miss—remains untested; a direct comparison would settle it.","Because the sample is small and US-only, the low support for bans may reflect the American self-regulatory context, and a European replication under the AI Act would likely produce different priority rankings.","The fact-sheet validation rests on one author's expertise, so the final step of the pipeline should be treated as a proof of concept until tested with actual policy staff; a randomized usefulness trial would be the natural next test.","The use of LLM-written scenarios and LLM-generated summaries creates a possible risk of circularity: the same technology being governed is also the instrument that elicits and condenses public input."],"forward_implications":["Lay brainstorming surfaces actors that expert-dominated lists tend to miss, such as schools, unions, and local communities, and assigns them concrete responsibilities.","Fact-checking and transparency emerge as the public's top mitigation priorities, pointing policymakers toward investment in verification infrastructure rather than outright restrictions.","The consistently low priority given to bans suggests public opinion in the US context favors targeted oversight over prohibitive AI regulation.","The fact-sheet format offers a template for converting crowdsourced survey data into one-page briefs that can enter existing policy workflows.","The approach is portable: the same scenario-plus-ranking pipeline can be applied to other emerging technologies where harms are still being anticipated."],"supporting_citations":[{"why":"Supplies the GPT-4 scenario-writing method and the pre-policy impact evaluations used to select the ten impact types.","marker":"[6]"},{"why":"Provides the impact typology and prior scenario-based work on generative AI in the news environment that this study extends.","marker":"[34]"},{"why":"Establishes the value of stakeholder inclusion in scenario planning, the methodological basis for lay involvement.","marker":"[3]"},{"why":"G rounds the claim that algorithmic impacts are co-constructed and require participatory assessment rather than purely expert judgment.","marker":"[38]"},{"why":"Articulates the accountability frame for algorithmic impact assessment in the public interest that the SAP mapping operationalizes.","marker":"[42]"},{"why":"Provides the critique of expert-centered risk-based regulation that motivates the participatory approach.","marker":"[61]"},{"why":"Models risk-scenario-based auditing approaches that the stakeholder-action mapping extends to mitigation strategies.","marker":"[39]"}],"fun_headline_variants":["Public input ranks fact-checking top AI harm fix","Laypeople prioritize fact-checking over bans for AI risks","Participatory study: 228 AI harm fixes, fact-checking wins","Citizen survey maps 228 AI harm mitigations for policymakers","AI policy: public prefers fact-checking, not heavy bans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a small, non-representative group of US crowdworkers, reacting to short AI-written stories, produces stakeholder-action pairs that are meaningful and useful for real policy, and that a single expert's judgment is enough to validate the policy fact sheets built from them.","fun_headline_variants_meta":{"raw":{"variants":["Public input ranks fact-checking top AI harm fix","Laypeople prioritize fact-checking over bans for AI risks","Participatory study: 228 AI harm fixes, fact-checking wins","Citizen survey maps 228 AI harm mitigations for policymakers","AI policy: public prefers fact-checking, not heavy bans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000643,"raw_usage":{"total_tokens":2945,"prompt_tokens":918,"completion_tokens":2027,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":1941}},"tokens_in":534,"tokens_out":2027,"duration_ms":14535,"temperature":1.0,"reasoning_tokens":1941,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:44:52.222763+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same two-survey pipeline on a nationally representative sample and have professional policy analysts work with the resulting fact sheets alongside expert-authored mitigation summaries; if the lay-derived sheets change no decisions or simply duplicate the expert lists, the method's added value is contradicted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the GPT-4 scenario-writing method and the pre-policy impact evaluations used to select the ten impact types."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the impact typology and prior scenario-based work on generative AI in the news environment that this study extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the value of stakeholder inclusion in scenario planning, the methodological basis for lay involvement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Articulates the accountability frame for algorithmic impact assessment in the public interest that the SAP mapping operationalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the critique of expert-centered risk-based regulation that motivates the participatory approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Models risk-scenario-based auditing approaches that the stakeholder-action mapping extends to mitigation strategies."}],"review_version":1}