{"id":"1ca79b66-aba3-4457-953a-a045b17653d3","arxiv_id":"2502.07649","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Crowd-sourced Community Notes on social media can function as a sociotechnical linter for data visualizations, extending the linting metaphor from code to human computation.","lead":"This paper proposes that crowd-sourced notes on social media, such as Community Notes on X, can act as a 'linter' for data visualizations, flagging misleading charts through human judgment rather than automated rules. It introduces a playful taxonomy of linters and argues that human computation adds social context that automated tools miss.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim fails the paper's own 'fixable' rung: Section 4.2.2 concedes Community Notes offer no repair mechanism, yet concludes they satisfy the ladder.","rationale":"The reader identified the provenance of the linting ladder as the weakest assumption. I agree that this is a vulnerability, but I found a more concrete and internal problem: even accepting the ladder as the relevant standard, the paper's own application of it is inconsistent. The 'fixable' rung is asserted and then immediately conceded away in the same subsection. This matters because the central claim ('Community Notes are linters') is established by mapping the ladder onto Community Notes; if one rung is missing, the mapping is incomplete. A design provocation can tolerate a loose metaphor, but the paper's argumentative structure is a checklist, and the checklist has a broken item. The proposed test would settle the issue by coding the examples against the rungs. If the coding shows no fixable content, the paper should either soften its claim (e.g., 'Community Notes are linter-like but lack the fixable dimension') or amend the ladder. This does not warrant rejection of the provocation, but it does warrant a conditional verdict that requires the authors to resolve the contradiction.","tokens_in":16624,"tokens_out":3973,"duration_ms":43859,"concrete_test":"Independently code the three Figure 3 examples against the four ladder rungs, using the paper's definitions. For 'fixable,' record whether the note text or the Community Notes system offers a concrete change to the visualization (e.g., rescale axes, add missing models, change aggregation). If none of the three notes offers any repair, the §4.2.2 claim fails its own evidentiary standard. Also state in the coding protocol whether the ladder requires all four rungs; if it does, the classification collapses; if it does not, the paper must explain why the missing rung is optional.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing support for 'Community Notes are linters' is the claim in §4.2.2 that they 'conform to the expected behaviors of linters, as outlined in the linting ladder.' The paper's own text undermines the fourth rung. After stating Community Notes 'can Support Fixable Changes,' it immediately qualifies: the only mechanism offered is social shaming/deletion, and 'a standard linter would not suggest you should delete your entire code.' Editing capabilities are 'inconsistent,' and no note in Figure 3 proposes a repair. So if the ladder is a checklist (all rungs required), Community Notes do not qualify. If it is not a checklist, then the §4.2.2 demonstration that they satisfy all four rungs is overstated, and the paper has not defined which rungs are necessary. Either way, the central classification rests on an unresolved internal contradiction in the evaluative standard, not merely on a questionable external definition. This is more specific than a general worry about the ladder's provenance: the claimed fit fails on the paper's own criteria.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that Community Notes on social media can be understood as a form of human computation that performs 'linting' of data visualizations. It extends a 'linting ladder' from the authors' prior work, consisting of the properties checkable, customizable, blamable, and fixable, and applies this ladder to Community Notes, using three annotated examples of misleading charts. The paper also situates linters on a puritanical/neutral/rebellious alignment chart and discusses implications for AI-powered linters, uncertainty, and the philosophy of linting. The authors explicitly frame the contribution as a provocation rather than an empirical study.","tokens_in":16808,"tokens_out":4341,"duration_ms":39062,"significance":"If the classification holds, the paper offers visualization researchers a generative lens: crowd moderation of charts as a sociotechnical complement to automated linting. The alignment-chart framing is thought-provoking, the three examples are real and instructive, and the discussion of implied truth effects and unresolvable lints opens concrete future work. The paper is also honest in its limitations section, acknowledging that it does not explore how general users understand the content. Its main weakness is that the central claim rests on a checklist application of the ladder despite admitted gaps in the 'fixable' rung, and the word 'demonstrate' in the abstract is stronger than the illustrative evidence provided.","major_comments":[{"comment":"The 'fixable' rung is not satisfied by the paper's own account. The paragraph 'Community Notes can Support Fixable Changes' concedes that the only mechanisms are social shaming/deletion and inconsistent editing, and even states that 'a standard linter would not suggest you should delete your entire code.' The subsequent conclusion nevertheless states that Community Notes 'can even guide fixable improvements.' If the ladder is a checklist, Community Notes do not qualify; if it is not a checklist, the paper needs to say which rungs are necessary and why. As written, the central claim that Community Notes 'conform to the expected behaviors of linters' is internally inconsistent.","section":"§4.2.2"},{"comment":"The evidence for the central claim consists of three hand-picked Community Notes with no selection criteria, no corpus, and no inter-rater reliability or alternative sampling. The abstract's 'We demonstrate' is accordingly stronger than the support provided; the paper itself labels the work a provocation and its limitations section says it does not explore how general users understand the content. Please either reframe the abstract and Section 4 as an illustrative analysis, or add a transparent and systematic selection procedure for the examples.","section":"§4.2.1, Figure 3"},{"comment":"The status of the linting ladder's rungs is underspecified. The ladder is introduced as 'a collection of four properties that broadly describe the behaviors expectable from a linter,' and then used as a checklist in §4.2.2. Yet earlier examples, such as StackOverflow downvotes and architectural review boards, lack the fixable property or apply it loosely. The paper should state whether the ladder is a set of necessary conditions, a set of common features, or a similarity metric; otherwise the classification of Community Notes as linters cannot be distinguished from a partial match.","section":"§3"},{"comment":"The claim that human computation 'enhances traditional linting by offering social insights' is not supported by any comparative evaluation. The paper explains how Community Notes could surface contextual or cultural knowledge, but it does not show that they do so more effectively or differently than existing automated linters in a way that would justify 'enhances.' Please soften this to a proposal or provide evidence of the claimed enhancement.","section":"Abstract, §4.3"}],"minor_comments":[{"comment":"Typo: 'the issue switch a data visualization' should read 'the issue with a data visualization.'","section":"§4.2.2"},{"comment":"Typo: 'and and can even guide fixable improvements' contains a doubled 'and.'","section":"§4.2.2"},{"comment":"Typo: 'but that household incomehas' is missing a space between 'income' and 'has.'","section":"§2.1"},{"comment":"References [36] and [37] are duplicate entries for the same Lazar et al. book chapter; please consolidate them.","section":"References"},{"comment":"The word 'nonuplets' in the text preceding Figure 2 is unclear; if it is intended as a term of art, it should be defined.","section":"§3, Figure 2"}],"recommendation":"major_revision","confidential_remarks":"This is a position/provocation paper, and under a strict empirical journal standard its evidence is thin. The framing and examples are valuable for the community, but the internal inconsistency on the 'fixable' rung undermines the central classification and should be resolved before publication. If the venue is an extended-abstract workshop track, the revision burden would be lighter, but the abstract should still not overclaim by saying 'demonstrate.'"},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a good alt.chi paper that deserves a serious referee, but it has a specific, fixable overreach that the authors nearly admit. The novel piece is the alignment chart and the move to treat Community Notes as sociotechnical linting powered by human computation. The three examples are well chosen and really do look like lints. Credit where due: the paper is honestly framed as a provocation, the related work is appropriate, and the writing is clear.\n\nThe soft spot is exactly what the stress-tester flags. In §4.2.2 the authors walk the linting ladder rung by rung, and for 'fixable' they concede that the only mechanism is social shaming or deletion, that a standard linter wouldn't suggest deleting your entire code, and that editing capabilities are inconsistent. Then the paragraph concludes that Community Notes 'can even guide fixable improvements.' That does not follow. If the ladder is a checklist, the rung fails. If it is a spectrum, the demonstration that all four rungs are satisfied is overstated. The paper would be stronger if it openly said that the fixable rung is only partially met, and then explored the design space for that rung. That would be consistent with the provocation genre.\n\nOther soft spots are minor for the genre: the examples are hand-picked (no systematic corpus), and the ladder comes from the authors' own prior work, though that is not circular—it is a reasonable reuse. The paper's self-aware limitation statement in §5.0.4 is accurate.\n\nNet: I'd send this to review. It is the kind of work that will get cited for the framing even if the detailed claim is debated. A good referee can ask for a modest revision to bring the conclusion in line with the analysis.","headline":"A lively, clearly-written provocation whose central analogy works better than its own conclusions admit; the fixable-rung overreach is real but fixable.","tokens_in":17349,"tokens_out":2502,"would_cite":true,"duration_ms":23343,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Community Notes are a linter for misleading charts","keywords":["Community Notes","human computation","linting","data visualization","sociotechnical linters","crowdsourced moderation","misinformation","alt.chi"],"falsifier":"Collect a corpus of Community Notes attached to data visualizations on X and check whether any original chart is edited, deleted, or corrected after a note appears; finding zero fixable changes would weaken the ladder case. Separately, run a survey or field experiment comparing trust in un-noted misleading charts with noted ones; if viewers trust them equally, the implied-truth-effect implication drawn from treating notes as lints does not hold.","tokens_in":16362,"feed_emoji":"📊","tokens_out":6433,"duration_ms":51109,"temperature":0.7,"pith_summary":"The paper argues that Community Notes—the crowd-sourced context boxes attached to posts on X—are effectively linters for data visualizations: they flag misleading or error-prone charts, point to the specific problem, and are approved or rejected by a community rating process. The authors extend the linter idea beyond code and beyond automated checks to human computation, then test the metaphor against the four-rung \"linting ladder\" of checkable, customizable, blamable, and fixable. They examine three real Community Notes that catch an inflation-unadjusted rent-versus-income chart, a cherry-picked leaderboard of language models, and a misleading choice of job-growth aggregation, showing that the same categories of issues automated visualization linters detect arise organically from the crowd. If the claim holds, crowd moderation becomes a legitimate form of visualization evaluation that adds social insight—cultural norms and community standards—that rule-based tools miss.","feed_headline":"Community Notes are a linter for misleading charts","feed_subtitle":"If true, crowd moderation becomes a legitimate complement to automated chart checkers.","key_machinery":"The load-bearing mechanism is the \"linting ladder\" from the authors' prior work, a four-property model of the behavior expected from a linter: checkable, customizable, blamable, and fixable. The paper applies that model as a checklist to Community Notes and also situates notes on a two-axis alignment chart that classifies linters by input type (code, text, anything) and evaluation mode (static analysis, computer, human). The other piece of machinery is the Community Notes platform itself: notes are written only after enough people request them, rated for helpfulness by the community, and surfaced through a ranking algorithm, and this structured crowd process is what earns the label \"human computation.\"","core_discovery":"In the paper's own framing, the central discovery is that Community Notes satisfy every rung of the linting ladder and therefore count as a genuine, human-powered linter of data visualizations. A note is checkable because it tests a chart against community norms about what counts as misleading; customizable because the community's ratings and the ranking algorithm determine whether a note appears at all; blamable because it names the specific failing, such as unadjusted scales, missing models, or a debatable aggregate; and fixable because the visible social pressure of a note can lead the original author to edit or delete the post. Because evaluation is done by people rather than static analysis, the paper classes Community Notes in the \"evaluation rebel\" cell of its alignment chart, with text as the input. The work is presented as a provocation rather than a user study: it reframes crowd moderation as linting and draws out consequences, including the implied truth effect, unresolvable conflicting lints, and the risks of non-expert or adversarial note-writers.","pith_inferences":["A test the authors do not run: measure whether charts that receive Community Notes are subsequently edited or deleted, which would operationalize the fixable rung and separate social shaming from actual repair.","The same ladder-plus-alignment analysis could be applied to other community-moderated artifacts, such as Wikipedia edit reviews or mapping-platform change discussions, where crowds enforce norms on mutable objects.","If the implied truth effect applies to charts, a randomized field experiment comparing trust in un-noted misleading charts against noted ones would give a direct welfare test of the provocation.","The note-ranking algorithm gives the platform, not the individual viewer, control over the customizable rung; a user-level \"hide note\" option would be a concrete product experiment on how individual customization changes trust."],"forward_implications":["Crowd moderation of visualizations becomes a legitimate branch of visualization evaluation, complementing automated linters such as VizLinter.","Visualization tool builders can use the linting ladder as a checklist for whether their own feedback mechanisms are checkable, customizable, blamable, and fixable.","If the comparison holds, the absence of a note will be read as a signal of correctness, so platforms must design for note absence, not only note presence.","Human-powered linting catches issues that automated tools miss, including cherry-picked comparisons and aggregation choices that are socially misleading rather than formally wrong.","AI- or LLM-based linters should be judged by whether they reproduce the ladder and the social context that makes notes meaningful, not merely by whether they automate the checking step."],"supporting_citations":[{"why":"Supplies the linting ladder (checkable, customizable, blamable, fixable) that the paper uses as the model against which Community Notes are tested.","marker":"[46]"},{"why":"Describes the Community Notes ranking algorithm and note-generation rules, the platform behavior the paper argues constitutes a linter.","marker":"[57]"},{"why":"VizLinter is the automated visualization linter whose detected issue types (scale, missing data, aggregation) are compared with what Community Notes catch organically.","marker":"[10]"},{"why":"Provides the survey and taxonomy of human computation used to argue Community Notes qualify as a human computation platform.","marker":"[53]"},{"why":"The implied truth effect is invoked to draw out the consequence that un-noted misleading charts may be perceived as accurate.","marker":"[51]"},{"why":"Surfacing visualization mirages supplies the \"manipulated scales\" example and the statistically mediated evaluation mode used to position the notes.","marker":"[44]"}],"fun_headline_variants":["Crowd moderation as a chart linter","Linting goes human: Community Notes for visuals","People-powered linting for data charts","Community Notes check charts like code linters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion rests on accepting the four-rung linting ladder as the definitive model of what a linter is; if that prior-work checklist is not the right standard, the claim that Community Notes are linters loses its foundation.","fun_headline_variants_meta":{"raw":{"variants":["Crowd moderation as a chart linter","Linting goes human: Community Notes for visuals","People-powered linting for data charts","Community Notes check charts like code linters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000333,"raw_usage":{"total_tokens":1831,"prompt_tokens":905,"completion_tokens":926,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":870}},"tokens_in":521,"tokens_out":926,"duration_ms":9062,"temperature":1.0,"reasoning_tokens":870,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T12:00:29.642708+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a corpus of Community Notes attached to data visualizations on X and check whether any original chart is edited, deleted, or corrected after a note appears; finding zero fixable changes would weaken the ladder case. Separately, run a survey or field experiment comparing trust in un-noted misleading charts with noted ones; if viewers trust them equally, the implied-truth-effect implication drawn from treating notes as lints does not hold.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the Community Notes ranking algorithm and note-generation rules, the platform behavior the paper argues constitutes a linter."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The implied truth effect is invoked to draw out the consequence that un-noted misleading charts may be perceived as accurate."}],"review_version":1}