Pith. sign in

REVIEW 3 major objections 4 minor 44 references

You Can't Publish Replication Studies (and How to Anyways)

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This position paper argues that strict replication studies are effectively unpublishable under current novelty requirements, but researchers can publish replication value by embedding it in studies that re-evaluate, expand, or specialize…

desk verdict A practical, clearly written position paper with a useful taxonomy, but the title overclaims and the evidence base is a self-selected case study rather than systematic data. read the letter →

arxiv 1908.08893 v1 pith:P5BLQWRD submitted 2019-08-23 cs.HC cs.GR

classification cs.HCcs.GR
keywords replicationstudiesnoveltyvisualizationresearchgraphicalperceptionpositionpaperpublicationreproducibilityvisionscience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This position paper argues that strict replication studies are effectively unpublishable in visualization and human-computer interaction because reviewers require novelty, so replication value must be embedded in studies that make new contributions. The authors identify three forms of embedded replication: re-evaluation (same objective, new environment or population), expansion (broaden conclusions with new conditions), and specialization (apply conclusions to a specific domain). They ground the taxonomy in a case study of replications of the classic graphical perception ranking, showing how influential follow-ups combined confirmation with added novelty. The paper is addressed to vision scientists seeking to contribute to visualization research, and recommends specialization as the strongest path. The claim matters because it offers a concrete route to producing validated, publishable research in a publication system that penalizes pure replication.

What carries the argument

The central mechanism is the taxonomy of three embedded-replication forms. Re-evaluation repeats an earlier experiment's objective in a different setting, such as a crowdsourced subject pool. Expansion adds new experimental conditions to generalize or deepen earlier conclusions. Specialization transfers the original study's conclusions into a specific domain or application, often leveraging domain expertise. These categories are complementary to prior similarity-based classifications (strict, partial, conceptual) and are defined by the kind of novelty each carries.

What would settle it

A systematic survey of visualization and HCI venues that finds a published strict replication of a well-known result, or review data showing reviewers rate strict replication as equal in value to novel contributions, would cast doubt on the claimed dominance of the novelty requirement.

Watch

Extended reading notes

Core claim

The paper's central discovery is a taxonomy of how successful replication studies in information visualization actually carry novelty. Rather than classifying replication by similarity to the original (as prior work does), the authors classify by the type of novel contribution attached: re-evaluation, expansion, and specialization. Using the seminal graphical perception study as a case study, they show that widely-cited replication-like studies all fit one of these three patterns, and none is a strict replication. The conclusion is that researchers should not attempt strict replication; they should design studies that both re-confirm prior findings and advance a new objective, environment, or domain.

Load-bearing premise

The argument depends on the premise that the novelty requirement is so dominant across publication venues that strict replication studies are effectively unpublishable; the paper supports this only with anecdotal and secondary evidence rather than systematic acceptance data.

Editorial extensions

If this is right

  • Vision scientists can enter visualization research by replicating a known perceptual study and specializing it to a visualization context.
  • Re-evaluation studies using new participant pools, such as crowdsourced subjects, can confirm that older lab-based findings still hold in modern settings.
  • Expansion studies can generalize perceptual laws, such as modeling correlation perception, to new chart types or tasks.
  • Strict replication alone should be avoided if the goal is publication under current novelty standards.
  • Reviewers and venues could encourage embedded replications by explicitly rewarding replication components with bonus points, analogous to data-availability incentives.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy likely generalizes beyond visualization to other empirical fields where novelty is required, such as human-computer interaction or cognitive science.
  • If the novelty requirement weakens through new journal policies or pre-registration, strict replication may become publishable, making the paper's strategic advice time-bound.
  • The paper's own literature search found no strict replications of the graphical perception study; a larger systematic review could test whether strict replications ever appear in other venues or subfields.
  • The bonus-points mechanism could be evaluated empirically by comparing review scores of submissions that include a replication component versus those that do not.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This position paper addresses the difficulty of publishing replication studies in information visualization and vision science. The authors argue that strict replication studies—those that only re-conduct and confirm an earlier experiment—lack the novelty required by publication venues and are therefore effectively unpublishable. They propose that researchers can successfully publish replication work by embedding it in studies that carry additional novelty, and they define a three-category taxonomy of such embedded replications: re-evaluation (testing an original finding under new conditions or participant pools), expansion (broadening the original conclusions with additional experimental conditions), and specialization (transferring the original knowledge to a specialized domain). The taxonomy is illustrated with a non-exhaustive case study of eight papers that replicate aspects of Cleveland and McGill's seminal graphical perception work. The paper concludes with practical advice for vision scientists who wish to contribute to visualization research via replication studies, recommending specialization as a particularly promising route, and suggests that reviewers could encourage replication by rewarding papers that include small replication components.

Significance. If the central claim is accepted, the paper contributes a useful vocabulary for describing replication-embedded contributions and offers concrete guidance for researchers at the vision science/visualization interface. The taxonomy is simple and memorable, and the case study of Cleveland and McGill replications gives the advice a concrete grounding. The authors are honest about the non-exhaustive nature of their survey and explicitly frame the paper as a position piece. The practical recommendation—embed replication in a novel contribution—is likely sound regardless of whether strict replication is truly unpublishable, because it aligns with observed publication practices and gives novice researchers a low-risk entry strategy. However, the paper's title and framing depend on an empirical premise that is not systematically demonstrated, and the taxonomy is validated only on the same examples from which it was induced. These issues do not destroy the paper's practical value but do require careful reframing and additional evidence before the strong claims can stand.

major comments (3)
  1. [Section 5.1 and Section 1] The central premise that strict replication studies are effectively unpublishable is not supported by the evidence provided. The claim relies on citations to position papers and a literature review [8,13,43] rather than systematic acceptance or rejection data, and the authors' own survey (Section 4.2) only examines eight published papers that embed replication, which cannot demonstrate that strict replications are rejected or unpublishable. The text itself concedes that 'some may exist' and that it is 'incredibly difficult' rather than impossible. Please either soften the claim to reflect the available evidence (e.g., 'rarely published' or 'currently disincentivized') or supply systematic evidence on submission and rejection outcomes to justify the strong framing.
  2. [Section 3.1 and Section 4.2] The taxonomy is circular in construction: Section 3.1 states that the authors' 'evaluation of prior work shows that the vast majority of replicated work in information visualization falls within one of the three following categories,' and then Section 4.2 uses the same body of work to demonstrate that taxonomy. There is no independent validation, no a priori definition of the categories, and no test against negative cases or a broader systematically sampled corpus. Please clarify whether the taxonomy is intended as a descriptive framework derived from the examples (which would require acknowledging its exploratory nature more explicitly) or as an empirical claim about the distribution of replication styles (which would require a more rigorous survey method).
  3. [Section 4.2 and Table 1] The case study's methodology is not described with enough detail to assess its evidentiary weight. The authors do not report the search strategy, inclusion criteria, screening process, or the number of papers examined before arriving at the eight in Table 1. Without this information, the reader cannot judge whether the eight examples are representative or whether the absence of strict replication studies reflects reality or selection bias. In addition, some classifications in Table 1 appear to blur the category boundaries: for example, Heer and Bostock [11] are labeled 'Re-evaluate' but the text describes them as also extending the original study to new encodings, which would overlap with 'Expand.' Please operationalize the categories and provide transparency about how each paper was assigned.
minor comments (4)
  1. [Abstract and throughout] There are several wording and typographical errors that should be corrected, including 'Simple put' (should be 'Simply put'), 'Y ou' at the start of the title, 'wishing contribute' (missing 'to'), 'cite' instead of 'cited' in Section 4, 'shined new light' (should be 'shed new light' or 'shed light'), and 'in deep knowledge' (should be 'deep knowledge').
  2. [Section 2] The discussion of prior replication classifications would benefit from a table or figure that explicitly compares Hornbaek et al.'s strict/partial/conceptual categories with Kosara and Haroz's reanalysis/direct/conceptual categories and the authors' new re-evaluate/expand/specialize taxonomy, as this would help readers see the intended complementarity more clearly.
  3. [Section 4.2.1] The sentence 'The main novelty of this paper was not to validate the findings of Cleveland and McGill, but to test the viability of online user study like crowdsourcing' is clear in intent but the phrase 'online user study like crowdsourcing' should read 'online user studies such as crowdsourcing' for grammatical correctness.
  4. [Section 5.2] The claim that 'we believe specializing represents the best opportunity' is presented without a supporting argument for why specialization is superior to re-evaluation or expansion for vision scientists; a short justification would strengthen the practical advice.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an explicitly positioned argument whose recommendation follows from a stated premise; the taxonomy and Cleveland-McGill case study are illustrative, not a derived-and-tested prediction.

full rationale

The paper makes no formal derivation, fits no parameters, and reports no predictive claim. The central recommendation (Section 5.2: 'the best way for individual researchers to publish replication studies is to distinguish the work with the help of added novelty') follows from the explicitly stated premise in Section 5.1 that novelty is a necessary review criterion and that strict replication supplies no novelty by definition. That premise is supported by citations to prior critiques and by the authors' self-described 'non-exhaustive' survey, which is weak empirical evidence but not circular. The three-category taxonomy in Section 3.1 is introduced as 'our evaluation of prior work' and then illustrated in Section 4.2 with the same Cleveland-McGill examples; the paper explicitly labels this a 'non-exhaustive case study' and never claims the examples independently validate the taxonomy. This is an interpretive framework, not a derivation whose output is equivalent to its input. There is no load-bearing self-citation chain and no imported uniqueness theorem. Accordingly, no circular step meeting the evidence standard is present.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper introduces a classification framework but no free parameters or entities. It relies on the empirical premise that the novelty requirement blocks strict replications, and on the exhaustiveness of its self-generated taxonomy, which is validated on the same examples used to construct it.

assumptions (2)
  • domain assumption Publication venues require novelty as a necessary contribution, making strict replication studies effectively unpublishable.
    The paper supports this with citations to prior critiques and its own survey, but does not provide systematic acceptance-rate data; Section 1 and Section 5.1.
  • ad hoc to paper Every successful embedded replication in visualization falls into one of the three categories: re-evaluation, expansion, or specialization.
    The authors state this based on 'our evaluation of prior work' (Section 3.1) and then use the same eight papers as evidence, without independent validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of You Can't Publish Replication Studies (and How to Anyways)." pith.science (2026). https://pith.science/paper/P5BLQWRD

@misc{pith2026190808893,
  author       = {Pith},
  title        = {Pith review of: You Can't Publish Replication Studies (and How to Anyways)},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P5BLQWRD}},
  note         = {Machine review of arXiv:1908.08893}
}
read the original abstract

Reproducibility has been increasingly encouraged by communities of science in order to validate experimental conclusions, and replication studies represent a significant opportunity to vision scientists wishing contribute new perceptual models, methods, or insights to the visualization community. Unfortunately, the notion of replication of previous studies does not lend itself to how we communicate research findings. Simple put, studies that re-conduct and confirm earlier results do not hold any novelty, a key element to the modern research publication system. Nevertheless, savvy researchers have discovered ways to produce replication studies by embedding them into other sufficiently novel studies. In this position paper, we define three methods -- re-evaluation, expansion, and specialization -- for embedding a replication study into a novel published work. Within this context, we provide a non-exhaustive case study on replications of Cleveland and McGill's seminal work on graphical perception. As it turns out, numerous replication studies have been carried out based on that work, which have both confirmed prior findings and shined new light on our understanding of human perception. Finally, we discuss how publishing a true replication study should be avoided, while providing suggestions for how vision scientists and others can still use replication studies as a vehicle to producing visualization research publications.

Figures

Figures reproduced from arXiv: 1908.08893 by the authors.

Figure 1
Figure 1. Reproduction of the Cleveland and McGill’s graphical en [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Reproduction of Mackinlay graphical encoding rankings [25]. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 44 canonical work pages

  1. [11]

    Heer and M

    J. Heer and M. Bostock. Crowdsourcing Graphical Perception: Using Mechanical Turk to Assess Visualization Design. In ACM SIGCHI: on Human Factors in Computing Systems, pp. 203–212, 2010

  2. [1]

    Changing Order: Replication and Induction in Scientific Practice

    C. Bazerman. Book Review of “Changing Order: Replication and Induction in Scientific Practice” by H. M. Collins. Philosophy of the Social Sciences/Philosophie des Sciences Sociales, 19(1):115, 1989

  3. [2]

    Bertini, N

    E. Bertini, N. Elmqvist, and T. Wischgoll. Judgment Error in Pie Chart Variations. In Proceedings of the Eurographics/IEEE VGTC Conference on Visualization: Short Papers, pp. 91–95, 2016

  4. [3]

    Bronstein

    R. Bronstein. Publication Politics, Experimenter Bias and the Replica- tion Process in Social Science Research. Journal of Social Behavior and Personality, 5(4):71, 1990

  5. [4]

    Cleveland and R

    W. Cleveland and R. McGill. Graphical Perception: Theory, Experi- mentation, and Application to the Development of Graphical Methods. Journal of the American Statistical Association , 79(387):531–554, 1984

  6. [5]

    Demiralp, M

    C ¸. Demiralp, M. Bernstein, and J. Heer. Learning Perceptual Kernels for Visualization Design. IEEE Transactions on Visualization and Computer Graphics, 20(12):1933–1942, 2014

  7. [6]

    Dragicevic and Y

    P. Dragicevic and Y . Jansen. Blinded with Science or Informed by Charts? A Replication Study. IEEE Transactions on Visualization and Computer Graphics, 24(1):781–790, 2017

  8. [7]

    Gramazio, K

    C. Gramazio, K. Schloss, and D. Laidlaw. The Relation Between Visu- alization Size, Grouping, and User Performance. IEEE Transactions on Visualization and Computer Graphics, 2014

Show all 44 references
  1. [8]

    Greenberg and B

    S. Greenberg and B. Buxton. Usability Evaluation Considered Harmful (Some of the Time). In ACM SIGCHI Conference on Human Factors in Computing Systems, pp. 111–120, 2008

  2. [9]

    Harrison, D

    L. Harrison, D. Skau, S. Franconeri, A. Lu, and R. Chang. Influenc- ing Visual Judgment through Affective Priming. In ACM SIGCHI Conference on Human Factors in Computing Systems, pp. 2949–2958, 2013

  3. [10]

    Harrison, F

    L. Harrison, F. Yang, S. Franconeri, and R. Chang. Ranking Visual- izations of Correlation Using Weber’s Law. IEEE Transactions on Visualization and Computer Graphics, 20(12):1943–1952, 2014

  4. [12]

    J. Heer, N. Kong, and M. Agrawala. Sizing the Horizon: The Effects of Chart Size and Layering on the Graphical Perception of Time Series Visualizations. In ACM SIGCHI Conference on Human Factors in Computing Systems, 2009

  5. [13]

    Hornbæk, S

    K. Hornbæk, S. Sander, J. A. Bargas-Avila, and J. Grue Simonsen. Is Once Enough?: On the Extent and Content of Replications in Human- Computer Interaction. In ACM SIGCHI Conference on Human Factors in Computing Systems, pp. 3523–3532, 2014

  6. [14]

    Hullman, E

    J. Hullman, E. Adar, and P. Shah. The Impact of Social Information on Visual Judgments. In ACM SIGCHI Conference on Human Factors in Computing Systems, pp. 1461–1470, 2011

  7. [15]

    Jakobsen and K

    M. Jakobsen and K. Hornbæk. Interactive Visualizations on Large and Small Displays: The Interrelation of Display Size, Information Space, and Scale. IEEE Transactions on Visualization and Computer Graphics, 19(12):2336–2345, 2013

  8. [16]

    Jasny, G

    B. Jasny, G. Chin, L. Chong, and S. Vignieri. Again, and Again, and Again... American Association for the Advancement of Science, 2011

  9. [17]

    Jones, P

    K. Jones, P. Derby, and E. Schmidlin. An Investigation of the Preva- lence of Replication Research in Human Factors. Human Factors, 52(5):586–595, 2010

  10. [18]

    Kay and J

    M. Kay and J. Heer. Beyond Weber’s Law: A Second Look at Ranking Visualizations of Correlation. IEEE Transactions on Visualization and Computer Graphics, 22(1):469–478, 2015

  11. [19]

    S.-H. Kim, Z. Dong, H. Xian, B. Upatising, and J.-S. Yi. Does an Eye Tracker Tell the Truth About Visualizations?: Findings While Investigating Visualizations for Decision Making. IEEE Transactions on Visualization and Computer Graphics, 18(12):2421–2430, 2012

  12. [20]

    N. Kong, J. Heer, and M. Agrawala. Perceptual Guidelines for Creat- ing Rectangular Treemaps. IEEE Transactions on Visualization and Computer Graphics, 16(6):990–998, 2010

  13. [21]

    R. Kosara. An Empire Built on Sand: Reexamining What We Think We Know About Visualization. In Beyond Time and Errors on Novel Evaluation Methods for Visualization, pp. 162–168, 2016

  14. [22]

    Kosara and S

    R. Kosara and S. Haroz. Skipping the Replication Crisis in Visualiza- tion: Threats to Study Validity and How to Address Them: Position Paper. In IEEE Evaluation and Beyond-Methodological Approaches for Visualization (BELIV), pp. 102–107, 2018

  15. [23]

    Kosara and C

    R. Kosara and C. Ziemkiewicz. Do Mechanical Turks Dream of Square Pie Charts? In Beyond Time and Errors on Novel Evaluation Methods for Visualization, pp. 63–70, 2010

  16. [24]

    Lallemand, V

    C. Lallemand, V . Koenig, and G. Gronier. Replicating an International Survey on User Experience: Challenges, Successes and Limitations. In ACM SIGCHI Conference on Human Factors in Computing Systems, 2013

  17. [25]

    Mackinlay

    J. Mackinlay. Automating the Design of Graphical Presentations of Re- lational Information. ACM Transactions On Graphics (TOG), 5(2):110– 141, 1986

  18. [26]

    Neumann, K

    L. Neumann, K. Matkovic, and W. Purgathofer. Perception Based Color Image Difference. Computer Graphics Forum, 17(3):233–241, 1998

  19. [27]

    W. Newman. A Preliminary Analysis of the Products of HCI Research, Using Pro Forma Abstracts. In ACM SIGCHI Conference on Human Factors in Computing Systems, vol. 94, pp. 278–284, 1994

  20. [28]

    Ottley, E

    A. Ottley, E. M. Peck, L. Harrison, D. Afergan, C. Ziemkiewicz, H. Tay- lor, P. Han, and R. Chang. Improving Bayesian Reasoning: The Effects of Phrasing, Visualization, and Spatial Ability. IEEE Transactions on Visualization and Computer Graphics, 22(1):529–538, 2015

  21. [29]

    R. Peng. Reproducible Research and Biostatistics. Biostatistics, 10(3):405–408, 2009

  22. [30]

    S. Pinker. A Theory of Graph Comprehension. Artificial Intelligence and the Future of Testing, pp. 73–126, 1990

  23. [31]

    K. Reda, P. Nalawade, and K. Ansah-Koi. Graphical Perception of Continuous Quantitative Maps: The Effects of Spatial Frequency and Colormap Design. In ACM SIGCHI Conference on Human Factors in Computing Systems, p. 272, 2018

  24. [32]

    R. A. Rensink and G. Baldridge. The Perception of Correlation in Scatterplots. Computer Graphics Forum, 29(3):1203–1210, 2010

  25. [33]

    Rosenthal

    R. Rosenthal. Replication in Behavioral Research. Journal of Social Behavior and Personality, 5(4):1, 1990

  26. [34]

    Saket, A

    B. Saket, A. Srinivasan, E. D. Ragan, and A. Endert. Evaluating Inter- active Graphical Encodings for Data Visualization. IEEE Transactions on Visualization and Computer Graphics, 24(3):1316–1330, 2018

  27. [35]

    Sedlmair, A

    M. Sedlmair, A. Tatu, T. Munzner, and M. Tory. A Taxonomy of Visual Cluster Separation Factors. Computer Graphics Forum, 31(3pt4):1335– 1344, 2012

  28. [36]

    Skau and R

    D. Skau and R. Kosara. Arcs, Angles, or Areas: Individual Data Encod- ings in Pie and Donut Charts. Computer Graphics Forum, 35(3):121– 130, 2016

  29. [37]

    Sukumar and R

    P. Sukumar and R. Metoyer. Towards Designing Unbiased Replication Studies in Information Visualization. In IEEE Evaluation and Beyond- Methodological Approaches for Visualization (BELIV), pp. 93–101, 2018

  30. [38]

    P. T. Sukumar and R. Metoyer. Replication and Transparency of Qual- itative Research from a Constructivist Perspective. OSF Preprints, 2019

  31. [39]

    D. Szafir. Modeling Color Difference for Visualization Design. IEEE Transactions on Visualization and Computer Graphics, 24(1):392–401, 2018

  32. [40]

    Talbot, V

    J. Talbot, V . Setlur, and A. Anand. Four Experiments on the Perception of Bar Charts. IEEE Transactions on Visualization and Computer Graphics, 20(12):2152–2160, 2014

  33. [41]

    A. C. Valdez, A. K. Schaar, J. R. Hildebrandt, and M. Ziefle. Re- quirements for Reproducibility of Research in Situational and Spatio- Temporal Visualization: Position Paper. In IEEE Evaluation and Beyond-Methodological Approaches for Visualization (BELIV) , pp. 53–59, 2018

  34. [42]

    C. Ware. Information Visualization: Perception for Design. Elsevier, 2012

  35. [43]

    Wilson, E

    M. Wilson, E. Chi, S. Reeves, and D. Coyle. RepliCHI: The Workshop II. In ACM SIGCHI Conference on Human Factors in Computing Systems (Extended Abstracts), pp. 33–36, 2014

  36. [44]

    F. Yang, L. Harrison, R. A. Rensink, S. Franconeri, and R. Chang. Cor- relation Judgment and Visualization Features: A Comparative Study. IEEE Transactions on Visualization and Computer Graphics, 2018

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.