{"id":"370e1f63-8da7-4696-94bd-3c51e7b39e35","arxiv_id":"2502.00637","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A human-AI collaboration framework using five story elements is applied to a real Uber self-driving incident to produce an AI ethics comic narrative.","lead":"This paper proposes a framework for humans and generative AI to collaborate on making visual stories from real AI news events, and demonstrates it with a comic about a 2016 Uber self-driving incident. It aims to give the public factual AI ethics narratives instead of science fiction scenarios.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'real-world data' claim is not auditable: the 10 source reports are unlisted and AI-generated panels contain admitted inaccuracies, so the narrative's authenticity cannot be independently verified.","rationale":"The reader's verdict is CONDITIONAL, and the identified weakest assumption—faithful representation of real-world events from unlisted news reports—is indeed the load-bearing point for the paper's core claim. My reading reinforces rather than redirects that concern: the absence of source lists and the admitted AI-generated visual inaccuracies make it impossible to independently check whether the comic is 'authentic' rather than a researcher-curated reconstruction. The framework itself is plausible and described in sufficient detail to be a useful proof-of-concept, so REJECT would be too strong; the condition should be to release the source data and trace each panel. Thus the verdict remains UNCHANGED with the stated conditions.","tokens_in":15574,"tokens_out":4747,"duration_ms":49729,"concrete_test":"Publish or otherwise provide the 10 source news reports (e.g., URLs or archived PDFs). Then build a traceability matrix: for each textual claim in Figure 9, identify the exact sentence in a listed report from which it was derived; for each AI-generated panel, list which visual elements (vehicle type, signage, colors, street layout, people) are supported by any source, and which came solely from the image generator. If any central factual assertion cannot be mapped to a source, or if prominent visual elements are unsupported by the sources, the 'data-driven / authentic' claim fails and the paper should be revised to either add the missing evidence or weaken the claim to 'researcher-generated narrative with AI illustration.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the five-element framework produces authentic AI ethics narratives from real-world cases. The load-bearing premise is that the Figure 9 comic faithfully represents the 10 news reports collected in Section 3.3.1. This premise is currently unverifiable: the 10 reports are not listed, cited, or quoted with attributions, so no reader can trace any panel's assertion (e.g., 'Uber argued pedestrians are not within its scope of protection', 'former employees revealed red-light running was common') back to a documented source. The few quoted texts in Section 3.3.2 are not source-referenced. The paper also admits in Section 3.3.2 that AI-generated images include mistaken details such as traffic-light colors and car text, yet those images constitute the visual story. Because the paper's motivation is to replace fictional or corporate narratives with authentic real-world narratives, an unauditable chain from source reports to comic leaves the central claim unsupported as presented. This is an evidence gap rather than an internal contradiction, and it is exactly what the conditional verdict should hinge on.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a conceptual framework for human-AI collaboration in data-driven visual storytelling about AI ethics, organizing the story model into five elements: structure, time, opinion, reality, and yearning. Humans are described as directing overall narrative, theme, and emotional depth, while generative AI enriches details, expands scenes, and advances plot progression. The framework is implemented on a 2016 Uber self-driving vehicle incident drawn from the AI Incident Database: the authors collected news reports, coded them into stakeholder perspectives, used a GPT-based image generator with iterative prompt engineering to create comic panels, and presented the result as a comicboarding narrative. The paper claims this demonstrates a feasible, repeatable approach to constructing authentic AI ethics narratives based on real-world events, in contrast to science-fiction or corporate narratives.","tokens_in":15745,"tokens_out":3014,"duration_ms":33969,"significance":"If the central claim is supported, the framework offers a concrete workflow for moving AI ethics storytelling from speculative fiction toward documented events, and it usefully assigns complementary roles to humans and generative AI in visual storytelling. The strengths of the paper include its honest acknowledgment of limitations (single case, reliance on expert interpretation), its detailed walk-through of prompt engineering and image correction, and its choice of a multi-stakeholder, open-ended narrative format that is well suited to contested AI ethics incidents. However, the load-bearing claims about 'authenticity' and 'accurate public understanding' are not empirically tested: the story artifact is produced and evaluated by the same researchers, no independent raters or audience study is provided, and the evidentiary chain from news reports to comic panels is not auditable. The contribution is therefore best read as a proof-of-concept design case study, not as a validated method.","major_comments":[{"comment":"The 10 news reports that form the data basis are never listed, cited, or quoted with attribution. The text states that 'we collected all news articles related to this incident' but provides no URLs, report titles, or publication dates, so no reader can trace assertions in Figure 9 (e.g., 'Uber argued that pedestrians are not its customers' or 'former employees revealed that Uber's autonomous vehicles running red lights were not uncommon') back to a documented source. Because the paper's central claim is that the resulting narrative is authentic relative to real-world data, this unauditable chain leaves the central claim unsupported as presented. Please include a full source list with per-panel sourcing, an appendix with the structured texts and their provenance, or a clear reclassification of the comic as author interpretation rather than a documentary reconstruction.","section":"Section 3.3.1 and Section 4"},{"comment":"The paper admits that AI-generated images contain mistaken or blurred details, such as 'directions and colors of traffic lights, texts on the car,' yet these images are retained as satisfactory and constitute the visual story. Since the paper's motivation is to give the public an accurate understanding of AI ethics issues, unverified and potentially misleading visual details directly undercut the authenticity claim. The authors should either verify the factual accuracy of every visual element against the source reports, explain why the admitted inaccuracies are immaterial to the ethical content, or modify the claim to acknowledge that the images are illustrative rather than factually accurate reconstructions.","section":"Section 3.3.2"},{"comment":"The framework's role division is derived from a literature review and the coding of two researchers, and the demonstration is built by the same authors using that framework; there is no independent evaluation of whether the resulting story is authentic, engaging, or informative. This creates a circularity in the evaluation sense: the artifact is judged satisfactory by its creators using criteria they also defined. The paper explicitly defers user research to future work, but then the abstract's claims about 'shaping public accurate understanding' and 'promoting active public engagement' are not backed by evidence in this manuscript. Please either add an evaluation component (independent raters, audience study, or expert assessment against pre-defined criteria) or explicitly scope the paper as a proof-of-concept that does not yet validate those outcome claims.","section":"Section 3.2 and Section 5"}],"minor_comments":[{"comment":"The phrase 'shape the public accurate understanding' should be revised, for example to 'shape the public's accurate understanding' or 'shape accurate public understanding.'","section":"Abstract"},{"comment":"The keyword list mixes semicolons and commas inconsistently and repeats 'Human-AI' with different hyphenation; please standardize the punctuation and style.","section":"Keywords"},{"comment":"The event summary says the vehicle 'suddenly accelerated against a red light and sped toward the opposite junction,' while the later example text says the car 'ran a red light ... and almost hit a pedestrian who was running the red light'; please reconcile these descriptions or clarify that they come from different reports.","section":"Section 3.3.1"},{"comment":"The text says the story is presented 'from the perspectives of four primary stakeholders: eyewitnesses, Uber as a corporation, Uber employees, and the government,' but Section 3.3.1 earlier says the narrative is structured 'from the perspectives of three key stakeholders'; please make the number of stakeholders consistent.","section":"Section 3.3.3"},{"comment":"The phrase 'The key characteristic of this new case' and later 'our new case' appear to be typographical errors for 'news case'; please correct them throughout.","section":"Section 3.3.3"},{"comment":"The text uses 'comicboarding' in some places and 'comicboardings' in others; please choose one consistent term and use it throughout.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the paper is honest about its limitations and presents a clear proof-of-concept, but its headline claims about authenticity and public understanding exceed what the evidence supports. The most pressing issue is the unauditable source chain from 10 unnamed news reports to the comic panels; this is fixable with an appendix or supplementary material. The paper's positioning relative to prior data-storytelling frameworks (e.g., Li et al.'s review of human-AI collaboration tools) should also be sharpened, since the five-element framework is presented as new but draws heavily on existing story models. The self-citation [65] is a related earlier abstract and is not problematic, but the incremental contribution over that abstract should be made explicit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the five-element framework (structure, time, opinion, reality, yearning) is the real contribution. It gives practitioners a concrete way to divide storytelling labor between humans and generative AI, and the paper applies it to a real AI ethics incident rather than another sci-fi scenario. That is a legitimate extension of existing human-AI storytelling work.\n\nWhat the paper does well: the role division is concrete and illustrated with a detailed prompt-engineering workflow (Figures 4-8). The authors are unusually honest in the limitations: they admit single-case validation, dependence on expert interpretation, and lack of granular step-by-step detail. The multi-stakeholder comic structure is a sensible way to present competing narratives in an open-ended case.\n\nWhere it gets soft: the paper's load-bearing claim is that the comic is an authentic narrative built from real-world data. That claim is not auditable as presented. The 10 news reports collected in Section 3.3.1 are never listed, cited, or quoted with attribution. A reader cannot trace any panel assertion — 'Uber argued pedestrians are not within its scope of protection,' 'former employees revealed red-light running was common' — back to a documented source. The few quoted texts in Section 3.3.2 have no source references. Separately, the authors keep AI-generated images they know contain wrong details (traffic-light colors, car text) because the scenes are otherwise acceptable. For a paper whose motivation is replacing misleading fictional/corporate narratives with authentic ones, those admitted inaccuracies sit in tension with the authenticity claim. The tension is acknowledged by the authors, and it is not fatal to the framework, but it does reduce the central claim to a plausibility argument.\n\nThe circularity the reader flagged is real but mild: the framework comes from the authors' coding of the literature, and the validating artifact is a story the authors built with that same framework. That is normal for a proof-of-concept; it means the evaluation is demonstration, not independent validation. No user data either, so statements about public understanding and engagement are aspirations, not findings.\n\nBottom line: this is a useful reference workflow for HCI/visualization people working on human-AI co-creation and data comics, and for AI-ethics educators who want a concrete alternative to fiction-based narratives. It deserves a serious referee — the framework is clearly described and the authors have thought about limitations. But it needs revision before publication: list the source reports, add traceability from text to panel, and either soften the authenticity claim or test it with readers.","headline":"A clear, honest proof-of-concept for a five-element human-AI storytelling framework; the 'authentic real-world' claim is undercut by unlisted sources and admitted AI image errors.","tokens_in":16256,"tokens_out":1873,"would_cite":false,"duration_ms":19392,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that authentic AI ethics narratives can be built from documented real-world events by dividing the work between humans and generative AI across five story elements, and it demonstrates this workflow on the 2016 Uber…","keywords":["AI ethics narratives","data-driven visual storytelling","human-AI collaboration","generative AI","comicboarding","AI incident database","autonomous vehicles","prompt engineering"],"falsifier":"For the 2016 Uber red-light incident, checking the original police report, court record, or full news archive against the comic's key claims would settle authenticity; if the documentary record contradicts any central panel, for example that the vehicle accelerated on its own or that the company's first public statement blamed the driver, then the workflow has not produced a faithful narrative.","tokens_in":15383,"feed_emoji":"🤖","tokens_out":8014,"duration_ms":73034,"temperature":0.7,"pith_summary":"This paper tries to establish that AI ethics narratives can be grounded in documented events rather than science fiction, and that a human-AI division of labor can make that grounding practical. It proposes a five-element story model—structure, time, opinion, reality, yearning—with humans setting direction, theme, and emotional depth while generative AI enriches details and scenes. The authors implement the model on the 2016 Uber autonomous-vehicle red-light incident, piecing together ten news reports into a comic told from four stakeholder perspectives. If the framework works beyond this case, it gives storytellers and policy communicators a repeatable workflow for turning fragmented incident reports into accessible visual narratives.","feed_headline":"Human-AI team turns a real Uber incident into an ethics comic","feed_subtitle":"People set theme and emotion; AI fills in scenes; the format was tested on the 2016 self-driving red-light case.","key_machinery":"The central object is the five-element data-driven visual storytelling model (structure, time, opinion, reality, yearning) used as a division-of-labor scheme between human and AI. Humans act as directors for each element, deciding narrative framework, timeline, theme, emotional depth, and positive values, while a prompt-based image generator expands scenes and fills in concrete details. The machinery also includes a prompt-engineering loop in which structured news texts become image prompts and are iteratively refined through conversational edits until the generated panel matches the documented event; the comicboarding format, a multi-panel sequential-art presentation, then carries the story across panels to a broad audience.","core_discovery":"The central claim is that data-driven visual storytelling with human-AI collaboration can construct authentic AI ethics narratives from real-world news data, and that the five-element model is a valid way to organize that collaboration. In the demonstrated workflow, humans control the overall structure, the timeline, the thematic stance, and the emotional arc; generative AI supplies scene detail, visual realism, and plot progression. The resulting comic about the Uber red-light incident presents the event from the viewpoints of the company, eyewitnesses, former employees, and the government, leaving the question of responsibility open. The authors argue that this open, multi-perspective format counters the misleading effect of a single corporate press release and shifts autonomous-driving ethics discourse away from trolley-problem speculation toward documented events.","pith_inferences":["This suggests a concrete stress test: applying the same five-element workflow to other AI incident types, such as bias, privacy, or automation failures, would show whether the human-direction/AI-detail split is general or specific to driving incidents.","A reader could go further and audit authenticity by attaching a source sentence from the news reports to each comic panel; the paper does not provide that provenance, but the framework would be stronger for it.","Because the paper does not run a user study, the claimed public-engagement benefit remains open: a controlled comparison of the comic against a single news article and a science-fiction scenario could measure shifts in understanding or policy preference.","As image generation improves, the iterative prompt-editing bottleneck should shrink, which would shift the human role toward theme selection and factual verification—a change the framework could absorb without redrawing its five elements."],"forward_implications":["The workflow can turn a pile of fragmented news reports into a coherent visual account without a professional illustrator, because the human side needs only direction-setting and prompt refinement.","An open-ended, multi-stakeholder comic can expose contradictions between a company's first public statement and later eyewitness and employee accounts, giving readers a reason to question any single report.","The five-element division of labor, where humans set structure, theme, timeline, and emotional depth while AI enriches details and scenes, offers a template for other data-storytelling tasks beyond AI ethics.","Because the story ends with unresolved responsibility and a government investigation, it supports public debate rather than a manufactured conclusion, which matches the unresolved state of autonomous-vehicle policy.","Using documented incidents re-centers autonomous-vehicle ethics on mundane regulatory questions, such as who is accountable when a self-driving car runs a red light, instead of hypothetical trolley problems."],"supporting_citations":[{"why":"Documents expert concern that AI narratives are missing and dominated by fiction or corporate marketing, motivating the need for authentic real-world narratives.","marker":"[15]"},{"why":"Supplies the incident-data source the paper draws its real-world case from and supports using real AI harms to raise awareness.","marker":"[18]"},{"why":"Warns that fictional AI narratives distort public benchmarks for evaluating real technology, justifying the shift to documented events.","marker":"[29]"},{"why":"Argues that science-fiction-based AI policy discussion overlooks mundane real ethical issues, supporting the paper's grounding in news cases.","marker":"[32]"},{"why":"Provides the three-stage data-storytelling process (data exploration, story creation, storytelling) that the case-study protocol adapts.","marker":"[37]"},{"why":"Reviews human-AI collaboration patterns in data storytelling tools and serves as the tool-centric view the paper contrasts with its element-level framework.","marker":"[38]"},{"why":"Establishes the design space of narrative visualization and storytelling with data, providing the conceptual foundation for data-driven visual storytelling.","marker":"[49]"},{"why":"Earlier presentation of the same ethically sensitive self-driving event that supplies the comicboarding demonstration and its choice of medium.","marker":"[65]"}],"fun_headline_variants":["Real Uber crash becomes an ethics comic via human-AI collab","Multi-view comic on Uber case leaves responsibility open","Data-driven ethics comic: humans set story, AI draws scenes","Real-incident AI ethics comic: humans lead, AI illustrates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The narrative stays authentic only if the ten news reports are themselves complete and accurate; if those reports omit or distort what happened, the comic cannot be a faithful account no matter how well the five-element framework is executed.","fun_headline_variants_meta":{"raw":{"variants":["Real Uber crash becomes an ethics comic via human-AI collab","Multi-view comic on Uber case leaves responsibility open","Data-driven ethics comic: humans set story, AI draws scenes","Real-incident AI ethics comic: humans lead, AI illustrates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00088,"raw_usage":{"total_tokens":3778,"prompt_tokens":895,"completion_tokens":2883,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":2814}},"tokens_in":511,"tokens_out":2883,"duration_ms":20614,"temperature":1.0,"reasoning_tokens":2814,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T18:13:30.654231+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For the 2016 Uber red-light incident, checking the original police report, court record, or full news archive against the comic's key claims would settle authenticity; if the documentary record contradicts any central panel, for example that the vehicle accelerated on its own or that the company's first public statement blamed the driver, then the workflow has not produced a faithful narrative.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents expert concern that AI narratives are missing and dominated by fiction or corporate marketing, motivating the need for authentic real-world narratives."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the incident-data source the paper draws its real-world case from and supports using real AI harms to raise awareness."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Warns that fictional AI narratives distort public benchmarks for evaluating real technology, justifying the shift to documented events."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues that science-fiction-based AI policy discussion overlooks mundane real ethical issues, supporting the paper's grounding in news cases."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reviews human-AI collaboration patterns in data storytelling tools and serves as the tool-centric view the paper contrasts with its element-level framework."},{"cited_title":"In International Conference on Human-Computer Interaction","cited_arxiv_id":null,"evidence_quote":"Establishes the design space of narrative visualization and storytelling with data, providing the conceptual foundation for data-driven visual storytelling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier presentation of the same ethically sensitive self-driving event that supplies the comicboarding demonstration and its choice of medium."}],"review_version":1}