{"id":"e0c088f8-d3b7-4382-a7d6-040a13f54580","arxiv_id":"2502.05324","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The Atlas of AI Risks, built from crowdsourced design requirements and LLM-generated content, outperformed a baseline dashboard on usability, aesthetics, and perceived balanced assessment in a 140-participant evaluation.","lead":"This paper presents an interactive Atlas of AI Risks, a narrative-style map showing the many uses, risks, benefits, and mitigations of facial recognition in plain language for non-experts. A 140-person study suggests it is more usable and more balanced than a baseline dashboard, supporting public deliberation about AI policy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The between-group comparison relies on a baseline deliberately stripped of the exact features being measured, so the headline superiority may reflect the control's construction rather than a real advance over state-of-the-art practice.","rationale":"The reader's conditional verdict is appropriate. The Atlas contribution is real: a deployed public tool, a formative study with 40 participants, an evaluation with 140 demographically matched participants, attention checks, and documented LLM content validation. However, the headline comparison is not between the Atlas and an independently existing expert dashboard; it is between the Atlas and a control the authors constructed to lack the very features measured. The paper's own baseline description admits the difference on R2/R4/R5, so the main effect is partly a consequence of the control's design. This does not invalidate the usability or aesthetics findings as descriptive results, but it does mean the 'more effective than baseline' claim is conditional on the baseline being representative. A direct re-run with the actual AIID spatial view, plus a content-controlled arm, would settle whether the advantage is due to the visualization design or to richer, more curated content. Thus the verdict should remain conditional pending that check.","tokens_in":20483,"tokens_out":4681,"duration_ms":49506,"concrete_test":"Re-run the evaluation with the actual, unmodified AI Incident Database spatial view (the artifact cited as the state of the art) as the control in the same task, with n=70 per arm and pre-registered tests for SUS, aesthetics, and balanced-assessment agreement. To separate design from content, add a crossed control: present the Atlas's 138 uses and impact cards in the baseline dashboard layout; if that condition performs as well as the Atlas, the advantage is the content, not the visualization. If the Atlas still beats the real AIID view and its own content in a conventional layout, the central claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is that the baseline is a fair representative of existing AI-risk visualizations. The paper states in the Baseline paragraph that the control 'differed in balancing risks and benefits (R2), reducing complexity (R4), and appealing to audience (R5)' — the same three dimensions later used as outcome measures. The baseline is a custom, stripped-down dashboard, not the unmodified AI Incident Database spatial view it claims to mimic, and it presents different content (incident descriptions and news articles) than the Atlas (138 LLM-generated uses with bespoke illustrations and impact cards). Consequently, the observed advantages on perceived balanced assessment (53% vs 32%), SUS (68 vs 50), and aesthetics could be produced by content volume, illustration quality, or the control's information poverty, rather than by the Atlas's visualization design. The central claim compares against 'a baseline visualization' without establishing that this control represents state-of-the-art practice, so the comparative conclusion is not yet supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Atlas of AI Risks, a narrative-style interactive visualization that maps uses of facial-recognition technology together with risks, benefits, and mitigations, aimed at non-experts. The authors first run a 40-participant crowdsourced formative study to derive six design requirements (R1-R6), then use LLM and generative-image prompts to create 138 uses and associated content, and then build the Atlas using the requirements. A second study with 140 US-demographic-matched participants compares the Atlas with a custom-built baseline that mimics an AI Incident Database spatial view, using the System Usability Scale, perceived visual aesthetics scales, a single self-report item on balanced assessment, exploration time, and open-ended questions. The paper reports that the Atlas outperforms the baseline on all quantitative metrics, and additionally demonstrates that the underlying format can be populated with 379 uses derived from the AI Incident Database.","tokens_in":20664,"tokens_out":5100,"duration_ms":51500,"significance":"If the comparative results held, the paper would provide a concrete, reusable design pattern for making AI risk information accessible to non-experts, a useful contrast to expert-oriented incident databases. The study is real: it uses a careful crowdsourcing setup with attention checks, demographic matching, and a publicly available artifact. The formative-to-evaluation pipeline and the explicit publication of prompts and supplementary materials are strengths. However, the headline between-group comparison is currently undermined by the construction of the baseline condition and by the absence of inferential statistics; these issues must be resolved before the central claim can be accepted.","major_comments":[{"comment":"The control condition is constructed so that it differs from the Atlas on exactly the dimensions later reported as outcomes. The Baseline paragraph states that the baseline 'displayed various uses (R1), categorized them (R2), and allowed exploration (R6)' but 'differed in balancing risks and benefits (R2), reducing complexity (R4), and appealing to audience (R5)'; Figure 4 then reports advantages on balanced assessment, SUS, and aesthetics. As a result, the 53% vs. 32% difference in balanced assessment and the SUS/aesthetics gaps may reflect the intentional information poverty of the control rather than the Atlas's visualization design. The paper needs to show that the baseline is a fair representative of current state-of-the-art practice, for example by using the unmodified AI Incident Database spatial view (or an independently existing tool) as the control, or by adding a third condition that holds content constant and varies only the visualization design; otherwise the comparative claim is confounded.","section":"Evaluation: Baseline"},{"comment":"No inferential statistics are reported for any of the headline comparisons. Figure 4 marks differences as 'statistically significant' but the text reports only percentages and means (SUS 68 vs. 50, classic aesthetics 3.77 vs. 2.98, etc.) without p-values, test statistics, effect sizes, or confidence intervals. The paper must report the tests used (e.g., t-test/Mann-Whitney for SUS and aesthetics, chi-square/Fisher for the balanced-assessment proportion), with effect sizes and CIs, and should state the pre-specified significance threshold; without these, the claims of superiority are not quantitatively supported.","section":"Quantitative results (Figure 4)"},{"comment":"The central construct of 'balanced assessment' (R2) is measured with a single self-report item ('The tool helped me to understand both the risks and benefits of facial recognition'), reported as a binary percentage. This is a weak measure for the paper's key outcome, and it is especially problematic because the two conditions presented different content (138 LLM-generated uses with illustrations and impact cards versus a stripped dashboard with news-article pop-ups). The authors already collected pre- and post-task emails from participants; analyzing these emails for the number/balance of risks and benefits mentioned would provide a direct behavioral outcome and would strengthen the claim considerably. At minimum, the item's wording and threshold should be justified and the content confound acknowledged.","section":"Metrics: quantitative metrics (Q2)"},{"comment":"The section claims 'successful testing' of generalizability, but what is actually demonstrated is a technical repopulation of the Atlas with 379 uses derived from the AI Incident Database; no user study or expert evaluation is reported for this generalized version. The claim that the design 'allows it to generalize across the diverse range of technology applications' is therefore supported only by a feasibility demonstration, not by testing. The manuscript should either add an evaluation of the generalized Atlas (even a small usability or comprehension check) or rephrase this contribution as a technical generalization demo rather than successful testing.","section":"Demonstrating the Generalizability of the Tool"}],"minor_comments":[{"comment":"In the Baseline paragraph, 'categorized them (R2)' appears to be a typo: categorization is requirement R3, while R2 is balanced assessment; this should be corrected to avoid confusion about which dimensions the baseline actually matched.","section":"Evaluation: Baseline"},{"comment":"Because the two conditions differ in the amount and type of content and in the number of interactive affordances, longer exploration time should be interpreted as engagement only after controlling for content volume; as reported, it is a descriptive difference rather than a controlled measure of engagement.","section":"Quantitative results: exploration time"},{"comment":"The paper does not report reliability statistics (e.g., Cronbach's alpha) for the SUS and Perceived Visual Aesthetics scales in this sample; adding these would support the claim that the instruments performed acceptably in the crowdsourced setting.","section":"Metrics: reliability"},{"comment":"The demographic matching covers age, sex, and ethnicity only, while recruitment additionally required interest in technology and English fluency; the text should state this limitation more explicitly when describing the sample as representative of the US population.","section":"Execution: participants"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real contribution and I'd send it out for review, but the headline comparison is weaker than it looks. The Atlas is a substantial tool, the six design requirements come from a real formative study, and the 140-person evaluation is run carefully in several respects: attention checks, demographic quotas, and a public artifact. The paper also names its limitations honestly in the discussion.\n\nThe main problem is the baseline. It was built as a stripped-down dashboard that deliberately omits the very features the Atlas is measured on: balanced presentation of risks and benefits, reduced complexity, and broad appeal. So finding that the Atlas wins on those dimensions is partly built into the study design. The paper calls the control 'state-of-the-art' because it mimics the AI Incident Database's spatial view, but it is a custom, simplified version, not the actual AIID interface. The two conditions also differ in content — the Atlas shows LLM-generated use cases with illustrations and impact cards, while the baseline shows incident descriptions and news articles — so the comparison mixes content and design. The claim that the Atlas beats existing practice needs the assumption that this control fairly represents current tools, and that assumption isn't established. That said, the paper is transparent about the difference, and the fix is straightforward: compare against the real AIID view or another existing tool.\n\nThe statistical reporting is thin. The 53% vs 32% difference and the SUS 68 vs 50 are given without p-values, confidence intervals, or effect sizes. The inter-rater reliability for the LLM-generated risk/benefit annotations is low (ICC around 0.23–0.47, called 'fair' but on the low side). The content also comes from the authors' own prompts and in-house expert validation, so the input side is somewhat self-grounded, though the tool itself is evaluated externally.\n\nThe generalizability demonstration with the AI Incident Database is useful as a deployment test, but it is not a user evaluation, so it does not show that the usability results carry over to other AI domains.\n\nOverall: the design requirements and the Atlas itself are valuable, and the evaluation is competent but not decisive. The central claim about superiority over existing practice is plausible but not yet proven. I'd accept this for peer review with a request for a stronger baseline comparison and fuller statistics. The paper is aimed at HCI researchers, visualization people, and anyone working on public AI risk communication.","headline":"A genuinely useful HCI contribution whose headline comparative claim is partly an artifact of the baseline design, but it deserves peer review.","tokens_in":21169,"tokens_out":3179,"would_cite":true,"duration_ms":32508,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A narrative map of AI risks can help non-experts weigh facial recognition's benefits against its harms, the paper reports.","keywords":["AI risk visualization","facial recognition","narrative visualization","public understanding of AI","human-computer interaction","crowdsourced design requirements","AI Incident Database","risk communication"],"falsifier":"Run the same user study with a control dashboard that presents the identical 138 uses, risks, benefits, and mitigations in a conventional expert-dashboard layout; if users rate the two tools equally on balanced assessment and usability, the Atlas's reported advantage comes from its content, not from its map-narrative design.","tokens_in":20320,"feed_emoji":"🗺️","tokens_out":6647,"duration_ms":63764,"temperature":0.7,"pith_summary":"The paper presents the Atlas of AI Risks, a narrative-style interactive visualization meant to give ordinary, non-technical people a balanced picture of what AI technologies do and what can go wrong. It argues that existing AI risk tools concentrate on technical failures such as data bias and are built for experts, leaving regular users unable to see societal risks like surveillance or job loss. To build and test the Atlas, the authors first ran a 40-person formative study that produced six design requirements, then used a large language model and an image model to expand a facial-recognition case study to 138 uses with associated risks, benefits, and mitigations. In a 140-person evaluation matching the US population by age, sex, and ethnicity, the Atlas outperformed a baseline dashboard in usability, in classic and expressive aesthetics, and in helping users understand both benefits and risks of facial recognition. If the result holds, narrative, map-like visualization could be a viable way to support informed public judgment and debate about AI regulation.","feed_headline":"Narrative atlas beats dashboard for public AI risk literacy","feed_subtitle":"In a 140-person test, users found it easier to weigh risks and benefits of facial recognition.","key_machinery":"The load-bearing mechanism is the pairing of six crowdsourced design requirements (multiple uses, balanced assessment, structured uses, reduced complexity, broad appeal, engaging exploration) with a concrete visualization grammar: sentence embeddings and t-SNE place uses on a map; a Martini Glass narrative structure with progressive disclosure reveals complexity in stages; colour coding marks daily versus non-daily uses and risk categories; and impact assessment cards provide both a tooltip and a detailed profile. These components together are what the paper credits for making risk information approachable and for the measured usability and balance advantages.","core_discovery":"On the paper's own terms, the central discovery is that a map-atlas metaphor with narrative sections and progressive disclosure can communicate broad AI risk information to people who are not AI experts. Each technology use is one dot on a two-dimensional map, positioned by semantic similarity; users can split dots by risk level, read brief tooltips or detailed profiles, and follow a five-section Martini Glass story that moves from what facial recognition is, to daily uses, to risk categories, to common harms, to a mitigation dashboard. The paper reports that 53% of participants using the Atlas said it helped them understand both risks and benefits, against 32% for the baseline; the Atlas's System Usability Scale score was 68 versus 50; and it scored higher on classic aesthetics, expressive aesthetics, and pleasurable interaction. The authors also extended the same five-component use format to 379 AI uses derived from incident reports to show the tool is not limited to facial recognition.","pith_inferences":["If the advantage is causal rather than content-driven, the same narrative atlas pattern could be tried for other contested technologies such as generative media, autonomous vehicles, and workplace surveillance, using the same five-component use format and a similar evaluation.","A sharper test would give the control dashboard the Atlas's own content and balanced framing, so only the visual narrative differs; the paper's comparison does not isolate the narrative design from the richer content it generated.","The shorter task-completion time for Atlas users (about 6.5 minutes versus 10.3 for the baseline) hints that the tool may increase comprehension efficiency, but the paper labels this as longer exploration, so the interpretation is ambiguous and worth re-measuring with explicit engagement metrics.","The design pattern suggests a possible plain-language AI risk label standard, in which every high-risk AI use carries an atlas-style card that separates purpose, capability, user, subject, and domain; that is an extension the paper gestures toward but does not develop."],"forward_implications":["If the Atlas's advantage is real, a non-expert can weigh trade-offs of facial recognition without first becoming an AI specialist, which is what the paper aimed to enable.","The same design could be populated from existing incident records: the authors show that the five-component format covers 379 AI uses drawn from a database of 649 incident reports, so the tool can shift from a facial-recognition case study to a general AI risk atlas.","Regulators and municipalities could embed the atlas format into public AI registers, consumer AI databases, and classroom materials, the three deployment scenarios the paper proposes.","Because SUS scores stayed similar across self-reported low, average, and high technological knowledge, the tool's usability appears not to depend on technical background, a direct corollary of the study's reported results."],"supporting_citations":[{"why":"Supplies the System Usability Scale that produced the Atlas's headline usability score of 68 versus 50.","marker":"(Brooke 1996)"},{"why":"Supplies the classic/expressive/pleasurable aesthetics dimensions on which the Atlas was rated higher.","marker":"(Lavie and Tractinsky 2004)"},{"why":"Provides the spatial AI Incident Database dashboard that the study used as the baseline comparison.","marker":"(CSET 2024b)"},{"why":"Provides the AI Incident Database whose incident reports the authors converted into 379 uses to test generalizability.","marker":"(McGregor 2021)"},{"why":"Supplies the Martini Glass narrative structure used to order the Atlas's five story sections.","marker":"(Segel and Heer 2010)"},{"why":"Supplies the five-component definition of an AI use that structures each Atlas entry and makes the tool generalizable.","marker":"(Golpayegani, Pandit, and Lewis 2023)"},{"why":"Supplies the capability/human-interaction/systemic-impact layers used to organize risks, benefits, and mitigations.","marker":"(Weidinger et al. 2023)"},{"why":"Supplies evidence-based visual communication techniques such as groupings and icon arrays that justify the design choices.","marker":"(Franconeri et al. 2021)"}],"fun_headline_variants":["AI risk atlas outshines dashboards for non-experts","Narrative maps beat dashboards for AI risk understanding","Atlas of AI risks lifts comprehension by 21 points for laypeople","Story-driven atlas helps non-tech users weigh AI pros and cons","Map-metaphor tool improves AI risk literacy in 140-person test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the baseline dashboard is a fair representative of existing expert-oriented AI risk visualization, because the baseline differed from the Atlas on exactly the dimensions the study then measured: balance of risks and benefits, reduced complexity, and broad appeal.","fun_headline_variants_meta":{"raw":{"variants":["AI risk atlas outshines dashboards for non-experts","Narrative maps beat dashboards for AI risk understanding","Atlas of AI risks lifts comprehension by 21 points for laypeople","Story-driven atlas helps non-tech users weigh AI pros and cons","Map-metaphor tool improves AI risk literacy in 140-person test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1351,"prompt_tokens":962,"completion_tokens":389,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":315}},"tokens_in":578,"tokens_out":389,"duration_ms":4987,"temperature":1.0,"reasoning_tokens":315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T19:45:51.387791+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same user study with a control dashboard that presents the identical 138 uses, risks, benefits, and mitigations in a conventional expert-dashboard layout; if users rate the two tools equally on balanced assessment and usability, the Atlas's reported advantage comes from its content, not from its map-narrative design.","supporting_citations":[],"review_version":1}