{"id":"b01f0213-2f7a-432d-8611-3f29662f2bbc","arxiv_id":"1908.08893","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Researchers can publish replication work by embedding it into studies with added novelty, and such embedded replications typically take the form of re-evaluation, expansion, or specialization.","lead":"This position paper argues that replication studies are rarely published because they lack novelty, and offers three strategies (re-evaluation, expansion, specialization) for embedding replications into novel work. It analyzes eight replications of Cleveland and McGill's classic study on graphical perception as case examples.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that strict replication studies are effectively unpublishable rests on a non-exhaustive survey and critiques, not on acceptance data; if that premise fails, the 'best way' recommendation becomes one strategy among several.","rationale":"The reader's conditional verdict is appropriate. I searched for a more damaging internal inconsistency and did not find one: the taxonomy is clearly presented, the case-study papers are real and mostly fit the stated categories, and the practical advice is actionable. The single load-bearing point is the premise that strict replications are unpublishable. The paper provides opinion and anecdote, not systematic evidence. This matters because the prescriptive conclusion ('the best way... is to distinguish the work with added novelty') is only necessary if publishing a strict replication is not a viable alternative. The paper itself qualifies the premise ('incredibly difficult', 'some may exist'), which is honest but also signals that the empirical boundary is unknown. A corpus re-analysis of Hornbaek et al. would at least establish whether strict replications are ever published in the venues discussed. The taxonomy's circularity—built and validated on the same eight examples—is a secondary weakness, but it does not undermine the central recommendation as much as the unsupported premise does. Therefore I agree with the reader's weakest-assumption identification and recommend no change to the conditional verdict.","tokens_in":8945,"tokens_out":6874,"duration_ms":75160,"concrete_test":"Re-analyze the Hornbaek et al. [13] corpus of 891 HCI papers: extract every paper they coded as a replication, classify each as strict, partial, or conceptual using their own definitions, and tabulate which were published in peer-reviewed venues (CHI, TVCG, EuroVis, etc.). If any strict replication appears in the published set—or if Dragicevic and Jansen's 'A Replication Study' [6] qualifies under that definition—Section 5.1's 'incredibly difficult' claim is overstated for those venues. If the published set contains zero strict replications, the premise survives this dataset, though rejection pressure is still not directly measured.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.1 asserts that as long as novelty is mandatory, 'it will be incredibly difficult to publish strict replication studies,' and Section 1 motivates this with citations to prior critiques and workshop reports. The only original evidence is the statement that a non-exhaustive survey found no strict replications of Cleveland and McGill (Sections 4-5.1). Absence in a hand-picked set of follow-ups does not demonstrate rejection: the set is self-selected for published, novel-embedded work, and no data are given on papers submitted but rejected for lacking novelty. The cited works [8,13,43] are position papers and a literature review of attempted replications, not acceptance statistics. The paper's own title is stronger than its body: 'some may exist' and 'incredibly difficult' already concede exceptions. If a strict replication can be published in a mainstream venue, or if reviewers treat novelty as one criterion rather than a gate, the central recommendation in Section 5.2 is still useful but is no longer the 'best way' grounded in a demonstrated constraint; it becomes one viable strategy among several.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper addresses the difficulty of publishing replication studies in information visualization and vision science. The authors argue that strict replication studies—those that only re-conduct and confirm an earlier experiment—lack the novelty required by publication venues and are therefore effectively unpublishable. They propose that researchers can successfully publish replication work by embedding it in studies that carry additional novelty, and they define a three-category taxonomy of such embedded replications: re-evaluation (testing an original finding under new conditions or participant pools), expansion (broadening the original conclusions with additional experimental conditions), and specialization (transferring the original knowledge to a specialized domain). The taxonomy is illustrated with a non-exhaustive case study of eight papers that replicate aspects of Cleveland and McGill's seminal graphical perception work. The paper concludes with practical advice for vision scientists who wish to contribute to visualization research via replication studies, recommending specialization as a particularly promising route, and suggests that reviewers could encourage replication by rewarding papers that include small replication components.","tokens_in":9285,"tokens_out":3633,"duration_ms":34474,"significance":"If the central claim is accepted, the paper contributes a useful vocabulary for describing replication-embedded contributions and offers concrete guidance for researchers at the vision science/visualization interface. The taxonomy is simple and memorable, and the case study of Cleveland and McGill replications gives the advice a concrete grounding. The authors are honest about the non-exhaustive nature of their survey and explicitly frame the paper as a position piece. The practical recommendation—embed replication in a novel contribution—is likely sound regardless of whether strict replication is truly unpublishable, because it aligns with observed publication practices and gives novice researchers a low-risk entry strategy. However, the paper's title and framing depend on an empirical premise that is not systematically demonstrated, and the taxonomy is validated only on the same examples from which it was induced. These issues do not destroy the paper's practical value but do require careful reframing and additional evidence before the strong claims can stand.","major_comments":[{"comment":"The central premise that strict replication studies are effectively unpublishable is not supported by the evidence provided. The claim relies on citations to position papers and a literature review [8,13,43] rather than systematic acceptance or rejection data, and the authors' own survey (Section 4.2) only examines eight published papers that embed replication, which cannot demonstrate that strict replications are rejected or unpublishable. The text itself concedes that 'some may exist' and that it is 'incredibly difficult' rather than impossible. Please either soften the claim to reflect the available evidence (e.g., 'rarely published' or 'currently disincentivized') or supply systematic evidence on submission and rejection outcomes to justify the strong framing.","section":"Section 5.1 and Section 1"},{"comment":"The taxonomy is circular in construction: Section 3.1 states that the authors' 'evaluation of prior work shows that the vast majority of replicated work in information visualization falls within one of the three following categories,' and then Section 4.2 uses the same body of work to demonstrate that taxonomy. There is no independent validation, no a priori definition of the categories, and no test against negative cases or a broader systematically sampled corpus. Please clarify whether the taxonomy is intended as a descriptive framework derived from the examples (which would require acknowledging its exploratory nature more explicitly) or as an empirical claim about the distribution of replication styles (which would require a more rigorous survey method).","section":"Section 3.1 and Section 4.2"},{"comment":"The case study's methodology is not described with enough detail to assess its evidentiary weight. The authors do not report the search strategy, inclusion criteria, screening process, or the number of papers examined before arriving at the eight in Table 1. Without this information, the reader cannot judge whether the eight examples are representative or whether the absence of strict replication studies reflects reality or selection bias. In addition, some classifications in Table 1 appear to blur the category boundaries: for example, Heer and Bostock [11] are labeled 'Re-evaluate' but the text describes them as also extending the original study to new encodings, which would overlap with 'Expand.' Please operationalize the categories and provide transparency about how each paper was assigned.","section":"Section 4.2 and Table 1"}],"minor_comments":[{"comment":"There are several wording and typographical errors that should be corrected, including 'Simple put' (should be 'Simply put'), 'Y ou' at the start of the title, 'wishing contribute' (missing 'to'), 'cite' instead of 'cited' in Section 4, 'shined new light' (should be 'shed new light' or 'shed light'), and 'in deep knowledge' (should be 'deep knowledge').","section":"Abstract and throughout"},{"comment":"The discussion of prior replication classifications would benefit from a table or figure that explicitly compares Hornbaek et al.'s strict/partial/conceptual categories with Kosara and Haroz's reanalysis/direct/conceptual categories and the authors' new re-evaluate/expand/specialize taxonomy, as this would help readers see the intended complementarity more clearly.","section":"Section 2"},{"comment":"The sentence 'The main novelty of this paper was not to validate the findings of Cleveland and McGill, but to test the viability of online user study like crowdsourcing' is clear in intent but the phrase 'online user study like crowdsourcing' should read 'online user studies such as crowdsourcing' for grammatical correctness.","section":"Section 4.2.1"},{"comment":"The claim that 'we believe specializing represents the best opportunity' is presented without a supporting argument for why specialization is superior to re-evaluation or expansion for vision scientists; a short justification would strengthen the practical advice.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"This is a clearly written position paper that will likely be useful to the community, but its empirical claims are currently stronger than the evidence. The authors should decide whether they want to present this as a deliberately provocative opinion piece (in which case the wording should be softened to match the admitted uncertainty in Section 5.1) or as a more rigorous empirical analysis (in which case the survey needs to be expanded and systematized). The circularity of the taxonomy is a separate concern that is fixable by repositioning the taxonomy as a proposed framework rather than a discovered empirical regularity. Given the practical value of the advice, I believe these issues are addressable within a major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful, practical position paper, not a rigorous empirical study. The three-way taxonomy (re-evaluate, expand, specialize) is a legitimate re-grouping of existing replication categories, but the paper's title overstates the claim: the body only says strict replication studies are 'incredibly difficult' to publish, and the evidence is a non-exhaustive case study of eight papers.\n\nThe genuinely new thing is the classification axis. Hornbæk et al. and Kosara & Haroz classify by study similarity; this paper classifies by the type of novel contribution attached to the replication. That's a useful frameshift for practitioners, and the Cleveland & McGill case study is well chosen. The eight examples are clearly described and the taxonomy maps onto them without obvious strain. The paper is also honest about its limits: it says the analysis is non-exhaustive and that 'some may exist' of strict replications. It's an easy read and would be genuinely helpful for, say, a vision scientist wanting to enter visualization research.\n\nThe soft spots are the ones the reader flagged. First, the title and abstract say 'can't publish replication studies,' but the actual claim is narrower—strict replication studies are hard to publish. That's an inconsistency, and it matters because the paper's central advice ('best way... is to distinguish the work with added novelty') only has teeth if the premise is true. The support for that premise is thin: no acceptance/rejection data, just citations to critiques and a self-selected survey. So the 'best way' is better stated as 'one sensible strategy' unless someone does the systematic study. Second, the taxonomy in Section 3.1 is induced from the same papers that form the case study in Section 4.2. That's circular validation, and the paper doesn't try to address it. Third, the eight papers are all published in top venues, which makes sense for a how-to guide but doesn't test whether the strategy fails elsewhere. These are real limitations, but they are typical of a position paper, not disqualifying.\n\nThe paper is for vision scientists, visualization researchers, and probably anyone in HCI thinking about replication. I'd bring it to a reading group, and I'd cite it if I write about replication in visualization. I'd send it to peer review: it deserves serious referees. I'd ask for a toned-down title and a clearer admission that the taxonomy is a heuristic, not a validated framework.","headline":"A practical, clearly written position paper with a useful taxonomy, but the title overclaims and the evidence base is a self-selected case study rather than systematic data.","tokens_in":9649,"tokens_out":2136,"would_cite":true,"duration_ms":20150,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This position paper argues that strict replication studies are effectively unpublishable under current novelty requirements, but researchers can publish replication value by embedding it in studies that re-evaluate, expand, or specialize…","keywords":["replication studies","novelty","visualization research","graphical perception","position paper","research publication","reproducibility","vision science"],"falsifier":"A systematic survey of visualization and HCI venues that finds a published strict replication of a well-known result, or review data showing reviewers rate strict replication as equal in value to novel contributions, would cast doubt on the claimed dominance of the novelty requirement.","tokens_in":8736,"feed_emoji":"🔁","tokens_out":5938,"duration_ms":50539,"temperature":0.7,"pith_summary":"This position paper argues that strict replication studies are effectively unpublishable in visualization and human-computer interaction because reviewers require novelty, so replication value must be embedded in studies that make new contributions. The authors identify three forms of embedded replication: re-evaluation (same objective, new environment or population), expansion (broaden conclusions with new conditions), and specialization (apply conclusions to a specific domain). They ground the taxonomy in a case study of replications of the classic graphical perception ranking, showing how influential follow-ups combined confirmation with added novelty. The paper is addressed to vision scientists seeking to contribute to visualization research, and recommends specialization as the strongest path. The claim matters because it offers a concrete route to producing validated, publishable research in a publication system that penalizes pure replication.","feed_headline":"Three ways to publish replication studies despite novelty pressure","feed_subtitle":"A position paper shows how re-evaluation, expansion, and specialization carry replication value into print.","key_machinery":"The central mechanism is the taxonomy of three embedded-replication forms. Re-evaluation repeats an earlier experiment's objective in a different setting, such as a crowdsourced subject pool. Expansion adds new experimental conditions to generalize or deepen earlier conclusions. Specialization transfers the original study's conclusions into a specific domain or application, often leveraging domain expertise. These categories are complementary to prior similarity-based classifications (strict, partial, conceptual) and are defined by the kind of novelty each carries.","core_discovery":"The paper's central discovery is a taxonomy of how successful replication studies in information visualization actually carry novelty. Rather than classifying replication by similarity to the original (as prior work does), the authors classify by the type of novel contribution attached: re-evaluation, expansion, and specialization. Using the seminal graphical perception study as a case study, they show that widely-cited replication-like studies all fit one of these three patterns, and none is a strict replication. The conclusion is that researchers should not attempt strict replication; they should design studies that both re-confirm prior findings and advance a new objective, environment, or domain.","pith_inferences":["The taxonomy likely generalizes beyond visualization to other empirical fields where novelty is required, such as human-computer interaction or cognitive science.","If the novelty requirement weakens through new journal policies or pre-registration, strict replication may become publishable, making the paper's strategic advice time-bound.","The paper's own literature search found no strict replications of the graphical perception study; a larger systematic review could test whether strict replications ever appear in other venues or subfields.","The bonus-points mechanism could be evaluated empirically by comparing review scores of submissions that include a replication component versus those that do not."],"forward_implications":["Vision scientists can enter visualization research by replicating a known perceptual study and specializing it to a visualization context.","Re-evaluation studies using new participant pools, such as crowdsourced subjects, can confirm that older lab-based findings still hold in modern settings.","Expansion studies can generalize perceptual laws, such as modeling correlation perception, to new chart types or tasks.","Strict replication alone should be avoided if the goal is publication under current novelty standards.","Reviewers and venues could encourage embedded replications by explicitly rewarding replication components with bonus points, analogous to data-availability incentives."],"supporting_citations":[{"why":"Supplies the prescriptive definition of replication and the finding that only 3% of HCI papers attempt replication.","marker":"[13]"},{"why":"Establishes the replication crisis in visualization and the threats to validity that replication addresses.","marker":"[22]"},{"why":"The seminal graphical perception study that serves as the paper's case study for embedded replications.","marker":"[4]"},{"why":"Example of a re-evaluation replication that moved the original experiment to a crowdsourced subject pool.","marker":"[11]"},{"why":"Example of an expansion replication that tests whether affective priming changes graphical perception judgments.","marker":"[9]"},{"why":"Example of a specialization replication that narrows the original study to bar charts.","marker":"[40]"},{"why":"Supports the claim that reviewers do not value replication work, motivating the need for embedded novelty.","marker":"[8]"}],"fun_headline_variants":["Publish replication studies by adding novelty: 3 ways","Replication studies only publish with embedded novelty","Three ways replication studies sneak into journals","Re-evaluate, expand, specialize: publish replications anyway","Novelty needed: three methods to publish replication studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the premise that the novelty requirement is so dominant across publication venues that strict replication studies are effectively unpublishable; the paper supports this only with anecdotal and secondary evidence rather than systematic acceptance data.","fun_headline_variants_meta":{"raw":{"variants":["Publish replication studies by adding novelty: 3 ways","Replication studies only publish with embedded novelty","Three ways replication studies sneak into journals","Re-evaluate, expand, specialize: publish replications anyway","Novelty needed: three methods to publish replication studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000255,"raw_usage":{"total_tokens":1533,"prompt_tokens":870,"completion_tokens":663,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":588}},"tokens_in":486,"tokens_out":663,"duration_ms":7128,"temperature":1.0,"reasoning_tokens":588,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:25:40.168318+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic survey of visualization and HCI venues that finds a published strict replication of a well-known result, or review data showing reviewers rate strict replication as equal in value to novel contributions, would cast doubt on the claimed dominance of the novelty requirement.","supporting_citations":[{"cited_title":"Hornbæk, S","cited_arxiv_id":null,"evidence_quote":"Supplies the prescriptive definition of replication and the finding that only 3% of HCI papers attempt replication."},{"cited_title":"Kosara and S","cited_arxiv_id":null,"evidence_quote":"Establishes the replication crisis in visualization and the threats to validity that replication addresses."},{"cited_title":"Cleveland and R","cited_arxiv_id":null,"evidence_quote":"The seminal graphical perception study that serves as the paper's case study for embedded replications."},{"cited_title":"Heer and M","cited_arxiv_id":null,"evidence_quote":"Example of a re-evaluation replication that moved the original experiment to a crowdsourced subject pool."},{"cited_title":"Harrison, D","cited_arxiv_id":null,"evidence_quote":"Example of an expansion replication that tests whether affective priming changes graphical perception judgments."},{"cited_title":"Talbot, V","cited_arxiv_id":null,"evidence_quote":"Example of a specialization replication that narrows the original study to bar charts."},{"cited_title":"Greenberg and B","cited_arxiv_id":null,"evidence_quote":"Supports the claim that reviewers do not value replication work, motivating the need for embedded novelty."}],"review_version":1}