{"id":"bce11ac4-1dc0-453c-9806-9d5db8d3df76","arxiv_id":"2607.10523","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"TIC is a process-oriented taxonomy of recurring issues in data communication, refined on 700 real-world narratives and mapped onto analysis, construction, and reception stages.","lead":"The paper introduces TIC, a six-dimension taxonomy of how data narratives break down, built from literature plus annotation of 700 real-world cases. It gives authors, editors, and fact-checkers a shared map of where quantitative claims go wrong across data, analysis, visuals, text, reasoning, and audience interpretation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged corpus/IRR limits.","rationale":"The strongest claim is integrative and formative: a process-oriented taxonomy plus annotated corpus and leverage-point framework. The manuscript already scopes itself as multi-label analytic lenses rather than mutually exclusive ground truth, documents pilot/open/main coding with reconciliation and spot-checks, and lists domain, modality, and IRR limits. The reader's CONDITIONAL verdict with high confidence already tracks that soft spot without over-penalizing a qualitative HCI taxonomy paper. Stress-testing does not surface a more central failure mode (e.g., process framework contradicting cases, or taxonomy collapsing under its own examples). Therefore the verdict should remain CONDITIONAL; no upgrade or downgrade is warranted from this pass.","tokens_in":32519,"tokens_out":445,"duration_ms":6079,"concrete_test":"Have two independent coders re-annotate a stratified random sample of 100 items from the released corpus using only Table II; report multi-label agreement (e.g., Jaccard/Krippendorff) and list any leaf types that systematically fail to stabilize. If agreement is low and new high-frequency types emerge outside TIC, the generality claim weakens; if agreement is moderate-to-high and types map cleanly, the reader's concern does not move the verdict.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that TIC (six dimensions + process mapping) is a useful structured diagnostic lens, not a prevalence-estimated or exhaustive universal classification. That claim is supported by the 34-paper synthesis, directed content analysis of 700 cases, multi-label examples, Table II, Figure 7, and explicit Limitations (§VII.A–B) that already concede domain/modality bias and formative coding without full-corpus formal IRR. The reader's weakest_assumption correctly identifies the main soft spot for generality claims, but it does not undercut the stated contribution as an analytical lens refined from literature and practice. No internal inconsistency, hidden mathematical assumption, or unacknowledged circularity is load-bearing for the strongest claim as written.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces TIC, a six-dimensional taxonomy of issues in data-driven communication (data, quantitative analysis, visual encoding, text/rhetoric, reasoning, and audience interpretation), refined from a 34-paper literature synthesis and directed content analysis of 700 real-world narratives from fact-checking APIs, prior research datasets, and controversial websites. TIC is multi-label and process-oriented rather than mutually exclusive; the authors situate it in an analysis–construction–consumption framework (Figure 7) drawing on Hall’s encoding/decoding model, report issue distributions across corpora (Figure 6), release an annotated corpus with rationales and a browsing interface, and discuss validation as a whole package, authorial intent, scrutiny calibration, and design implications for authoring, editorial, and fact-checking support.","tokens_in":32706,"tokens_out":733,"duration_ms":8070,"significance":"If accepted as an analytical lens rather than a prevalence estimate, TIC is a useful integrative contribution for visualization, HCI, and data journalism. Prior work is largely siloed by modality (statistics, chart design, claim verification); the paper’s main value is connecting issue types to process stages, actors, and leverage points, with concrete extensions such as “scope dilution” and an openly browsable multi-label case corpus. The design-implications section is actionable for linters, claim-review tools, and fact-checking aids. Strengths include transparent methodology (PRISMA-style review, pilot/open/main coding with reconciliation), explicit multi-label framing, and candid Limitations on domain/modality bias and formative coding without full-corpus formal IRR.","major_comments":[{"comment":"§III.B.2 and §VII.B: The taxonomy is the central contribution, yet formal inter-coder reliability is not reported for the full 700-item set; only pilot calibration, 15% dual open coding with reconciliation, lead-author main coding, and spot-checks are described. For a formative taxonomy this is defensible, but the manuscript should either (a) report agreement metrics on a held-out dual-coded subset for middle-layer categories, or (b) more tightly bound claims of stability/extensibility to “analytic lens refined from literature and practice,” and state how boundary cases (e.g., selective reasoning vs. evidence distortion vs. visual–text mismatch) were resolved in the codebook.","section":null},{"comment":"§IV.C / Figure 6 and Abstract/§I: Distribution percentages are corpus-conditioned (Fact-Check text-heavy; Prior Work visualization-centered; Controversial sites climate/vaccines). The text mostly treats them as descriptive, but phrasing such as “most prevalent issue types” can be read as general prevalence. Explicitly frame Figure 6 as corpus-specific descriptive patterns, not estimates of issue rates in data communication at large, and avoid language that implies representativeness beyond the curated sources.","section":null},{"comment":"§V / Figure 7 vs. §IV.A.6: Interpretation biases are defined as reception-side and excluded from artifact annotation, yet the process framework places them as a full TIC dimension with leverage points. Clarify operational status: which categories are artifact-codable vs. hypothesized reception mechanisms, and what evidence (beyond literature) supports mapping interpretation biases onto consumption-stage interventions. Without this, Dimension 6 risks reading as a literature appendix rather than an empirically grounded taxonomy layer.","section":null}],"minor_comments":[],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a careful qualitative taxonomy paper, not a new theory of deception. TIC’s value is the cross-modal package—data, analysis, visual encoding, text, reasoning, interpretation—mapped onto analysis–construction–consumption, plus a 700-case annotated set and a named “scope dilution” pattern that prior chart taxonomies mostly miss.\n\nWhat is actually new is the synthesis and the process situating, not most of the leaf categories. Huff, Campbell, McNutt, Lo et al., Lan et al., and the fact-checking literature already cover large pieces. The authors are honest about that. What they do well is the method: PRISMA-style lit pass, directed content analysis with pilot/open coding/reconciliation/spot-check, multi-label coding that matches how narratives actually break, clear examples (causal co-trending, inverted axes, geologic baselines), and design implications that track the leverage points. The annotated corpus and browser are genuine research products, not window dressing. Circularity is low; categories come from literature plus cases, not from defining the conclusion into existence.\n\nSoft spots, in proportion: the corpus is skewed to climate/COVID/politics and mostly static text+chart forms, and they never report formal IRR on the full set. Both are already in Limitations. That weakens any claim of prevalence or universal coverage, but it does not undercut the stated contribution as a diagnostic lens. I would not treat the distribution percentages as population estimates. No load-bearing math, no hidden circularity, citation pattern looks appropriate for an HCI/vis taxonomy paper.\n\nWho it is for: people building authoring linters, fact-checking tools, or teaching data narrative critique. A serious editor should send it to peer review. I would cite the taxonomy and process figure when framing multimodal narrative failures, and I would bring it to reading group if the group cares about trustworthy data communication rather than pure perception experiments.\n\nRecommendation: engage; accept-shaped with the usual ask to tighten generality language and, if possible, add reliability evidence or a clearer statement that TIC is formative.","headline":"Useful integrative taxonomy that stitches known failure modes into a process lens; soft on generality, but the claim as written holds and the corpus work is real.","tokens_in":33318,"tokens_out":526,"would_cite":true,"duration_ms":11264,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Data narratives fail as a whole package, not as isolated bad charts or wrong numbers, and a six-part taxonomy maps where they break.","keywords":["data narratives","misleading visualization","taxonomy","fact-checking","data communication","reasoning fallacies","audience interpretation","sociotechnical support"],"falsifier":"Independent multi-annotator coding of a broader, modality-balanced sample (including dashboards, video, and scrollytelling outside climate, COVID, and politics) that either fails to recover TIC's dimensions and co-occurrence patterns or shows that many real-world failures fall outside the six dimensions and process stages.","tokens_in":33404,"feed_emoji":"📊","tokens_out":616,"duration_ms":7717,"temperature":0.7,"pith_summary":"This paper argues that failures in data-driven public communication rarely reduce to a single wrong number or a deceptive chart. Instead, they emerge as quantitative evidence is turned into claims, visuals, text, and arguments across a longer meaning-making process. The authors introduce TIC, a taxonomy of issues in data communication built from prior literature and refined by qualitative annotation of 700 real-world narratives from fact-checking sites, research datasets, and controversial media. TIC groups recurring breakdowns into six dimensions: data, quantitative analysis, visual encoding, text and rhetoric, reasoning, and audience interpretation, and places those dimensions inside a process framework of analysis, narrative construction, and audience reception. The contribution is a shared diagnostic language for authors, editors, fact-checkers, and tool builders who need to see how issues arise, propagate, and compound rather than treating statistics, charts, and wording as separate problems.","feed_headline":"Data stories fail as whole packages, not lone bad charts","feed_subtitle":"A six-part taxonomy maps how numbers become misleading claims across analysis, design, and reading","key_machinery":"TIC (Taxonomy of Issues in Data Communication): a multi-layer scheme of six dimensions, middle-layer categories, and leaf subtypes, mapped onto an analysis–construction–consumption process framework that links issue types to activities, actors, and leverage points.","core_discovery":"Problematic data narratives are best diagnosed with TIC, a six-dimensional, process-oriented taxonomy of issues spanning data integrity, quantitative analysis, visual encoding, textual and rhetorical framing, reasoning fallacies, and audience interpretation biases, situated across analysis, construction, and consumption so that breakdowns can be located where they enter and how they compound.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Data narratives fail as processes, not lone bad charts","TIC maps six ways data stories break from numbers to readers","How data claims go wrong across analysis, design, and reception","Taxonomy tracks data narrative failures that compound over stages","Data stories mislead via data, visuals, text, and interpretation"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The claim that a literature synthesis plus 700 curated narratives from fact-check APIs, prior research sets, and three controversial websites, coded without formal full-corpus reliability, is representative enough to define a general taxonomy of data-communication failures.","fun_headline_variants_meta":{"raw":{"variants":["Data narratives fail as processes, not lone bad charts","TIC maps six ways data stories break from numbers to readers","How data claims go wrong across analysis, design, and reception","Taxonomy tracks data narrative failures that compound over stages","Data stories mislead via data, visuals, text, and interpretation"]},"model":"grok-4.5","effort":"low","cost_usd":0.003928,"raw_usage":{"total_tokens":1180,"prompt_tokens":740,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":39280000,"prompt_tokens_details":{"text_tokens":740,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":374,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":740,"tokens_out":66,"duration_ms":3679,"temperature":1.0,"reasoning_tokens":374,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T11:04:05.192901+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Independent multi-annotator coding of a broader, modality-balanced sample (including dashboards, video, and scrollytelling outside climate, COVID, and politics) that either fails to recover TIC's dimensions and co-occurrence patterns or shows that many real-world failures fall outside the six dimensions and process stages.","supporting_citations":[],"review_version":1}