Pith. sign in

REVIEW 4 major objections 10 minor 72 references

Performance bonuses don't improve viz study results

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-09 10:06 UTC pith:JKSZEHSN

load-bearing objection Preregistered null result on performance-based incentives in crowdsourced visualization experiments; the null is plausible but the incentive magnitudes may be too small to be a strong test. the 4 major comments →

arxiv 2607.07463 v1 pith:JKSZEHSN submitted 2026-07-08 cs.HC

Should We Dangle a Carrot? The Effect of Performance-based Incentives in Visualization Experiments

classification cs.HC
keywords incentivesvisualization experimentscrowdsourcingexperimental designgraphical perceptiondecision-making under uncertaintynull resultresearcher degrees of freedom
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper investigates whether offering performance-based monetary bonuses to crowdworkers changes their results in visualization experiments. The authors ran two preregistered studies—one on a low-level perceptual task (judging which of two scatterplots or parallel-coordinates plots shows higher correlation) and one on a higher-level reasoning task (deciding whether to salt roads based on a weather forecast shown as intervals or density plots)—with and without financial incentives. They expected incentives to leave the perceptual task unaffected but to improve the reasoning task. Instead, they found no performance difference in either task: incentivized participants did not score better on any metric (just-noticeable difference for correlation perception, expected utility for decision-making), though they did spend more time on the tasks. A replication of the decision-making study confirmed the null result. The paper argues that the common practice of tying bonuses to performance in crowdsourced visualization studies may not deliver the validity gains researchers assume, and that the added cost and complexity of incentive schemes may not be justified.

Core claim

The central finding is a null result across two (plus one replication) preregistered experiments: performance-based financial incentives did not improve task performance on either a low-level perceptual task or a higher-level decision-making task, despite incentivized participants spending measurably more time on the tasks. The authors frame this as evidence that incentives, as currently deployed in crowdsourced visualization research, may not produce the behavioral effects that the capital-labor-production theory from behavioral economics predicts—namely, that tying pay to performance should induce more effort and thus better outcomes. The paper also surfaces an unexpected secondary finding

What carries the argument

The capital-labor-production theory (Camerer & Hogarth, 1999), which predicts that incentives improve performance when a task requires moderate effort and participants possess sufficient cognitive capital

Load-bearing premise

The paper's central null result depends on the incentive magnitudes being large enough to induce a behaviorally meaningful change in effort; if the bonus amounts (roughly $2 on top of a reduced base pay) were too small relative to the cognitive cost of additional effort, the null could reflect an underpowered manipulation rather than a true absence of incentive effects.

What would settle it

If a future study used substantially larger bonuses (e.g., doubling or tripling the per-correct-answer reward while maintaining the same task structure) and found a significant performance improvement, the null result here would be attributable to insufficient incentive magnitude rather than a genuine absence of incentive effects on these task types.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Visualization researchers who use performance-based bonuses in crowdsourced studies may be adding cost and complexity without gaining the internal or ecological validity they seek.
  • The null result suggests that cross-study comparisons between incentivized and non-incentivized experiments on the same task may be more defensible than previously assumed, since the incentive manipulation itself does not appear to shift performance.
  • The finding that incentives increase time-on-task without improving performance raises ethical concerns: if researchers use lower base pay plus bonuses, participants may end up working longer for equivalent or lower effective wages.
  • The unexpected result that interval plots matched or outperformed density plots for decision quality challenges a growing consensus in uncertainty visualization and warrants further investigation.
  • If incentives do not improve performance on tasks requiring moderate cognitive effort, the boundary conditions of the capital-labor-production framework may need revision for crowdsourced experimental settings.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the incentive magnitudes used here were below a behaviorally meaningful threshold, the null result could reflect an underpowered manipulation rather than a true absence of incentive effects; the paper acknowledges this possibility but does not independently verify that the bonus amounts were above such a threshold.
  • The qualitative finding that some participants adopted risk-seeking attitudes under incentives suggests that incentive structures may interact with crowdworker behavior in ways that undermine the assumed mapping between pay and effort, particularly when base pay is reduced to fund bonuses.
  • The absence of a variance-reduction effect (incentives did not meaningfully reduce response variance) further weakens the case for using bonuses as a methodological tool to improve data quality in crowdsourced visualization studies.
  • If the capital-labor-production framework does not apply well to crowdsourced visualization tasks, there may be a class of tasks—specifically those requiring substantial cognitive capital or high production requirements—where incentives are structurally unlikely to help, and researchers should identify these task types rather than applying incentives uniformly.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 10 minor

Summary. This paper investigates whether performance-based financial incentives affect participant performance in crowdsourced visualization experiments. The authors conducted two preregistered between-subjects studies: (1) a perception-of-correlation task (scatterplots vs. parallel coordinates) and (2) a decision-making-under-uncertainty task (interval vs. density plots). In both studies, participants were assigned to either a flat-pay condition or a performance-based incentive condition. Using Bayesian hierarchical models, the authors find no evidence that incentives improved task performance on either the perceptual task (measured via JND) or the reasoning task (measured via expected utility), despite participants in incentivized conditions spending more time on the tasks. A replication study (Appendix B) with 95% intervals confirms the null effect of incentives. The authors discuss implications for experimental design, ecological validity, and ethical considerations around fair compensation. All materials, code, and preregistrations are publicly available on OSF and GitHub.

Significance. The paper addresses a practically important question for the visualization and HCI research communities: whether the common (or sometimes omitted) design choice of offering performance-based monetary incentives meaningfully changes experimental outcomes. The use of preregistration, Bayesian hierarchical models with reported convergence diagnostics (R-hat = 1.00, ESS > 2000), attention checks, and publicly available materials and code is commendable and strengthens the credibility of the null results. The inclusion of a replication study (Appendix B) further bolsters the central claim. The qualitative analysis of participant perspectives (Section 6) adds valuable context about how crowdworkers perceive and respond to incentive structures. The ethical discussion of wage variance in incentivized conditions (Section 7.5) is a useful contribution. The paper is transparent about its limitations, including the unexpected finding that interval plots outperformed density plots, and the authors' decision to publish null results to avoid the file drawer problem is appropriate.

major comments (4)
  1. §5.1, Study 2 data-logging issue: The loss of 31 participants' data due to a data-logging failure reduced the planned sample from 280 to 249 (before attention-check exclusions, yielding 243). The paper does not report whether this data loss was random across conditions or whether it differentially affected certain conditions, which bears directly on the power and balance of the between-subjects design. A breakdown of exclusions by condition and a brief discussion of whether the data loss could have introduced bias should be added.
  2. §3 and §5.1, incentive magnitude sufficiency: The central null result rests on the assumption that the incentive magnitudes were large enough to induce a behaviorally meaningful change in effort. In Study 2, the bonus was $0.50 per $1,000 of remaining virtual budget. For a single trial where p(freeze) = 0.235, the expected-value difference between the optimal action (salt, cost $1,000) and the suboptimal action (don't salt, expected cost 0.235 × $5,000 = $1,175) is $175 in virtual dollars, translating to approximately $0.0875 in real bonus. Across 18 trials, the total marginal incentive for making all optimal versus all suboptimal decisions is on the order of $1–2. The paper acknowledges this possibility in §7.3 ('not being sufficiently sensitive to incentives') but does not independently verify that the incentive was above a behaviorally relevant threshold. The time-on-task increase (20
  3. §5.3 and Figure 5C, equivalence testing: The paper reports credible intervals for expected utility differences (e.g., incentivized interval: 2.76 [2.31, 3.12] vs. base interval: 2.69 [2.19, 3.04]) and concludes there is 'no evidence' of an incentive effect. However, the paper does not report a formal equivalence test or ROPE (Region of Practical Equivalence) analysis to distinguish 'no effect' from 'insufficient precision to detect an effect.' The credible intervals for the expected utility differences are moderately wide (e.g., the base density condition spans [1.70, 2.86]). Adding a ROPE analysis or explicitly stating what effect size would be considered practically meaningful would strengthen the claim that the null result reflects a true absence of incentive effects rather than insufficient precision.
  4. §5.3 and §7.2, interval vs. density plot finding: The finding that interval plots led to marginally better expected utility than density plots contradicts prior work [12, 31, 60]. The authors propose a visual heuristic explanation (Figure 7: the 66% interval overlapping 0°C signals ~20% probability) but the follow-up study (Appendix B) with 95% intervals also showed no density advantage, which undermines the heuristic explanation. The paper acknowledges this is 'a conflicting picture' (§7.2) but the discussion does not fully grapple with whether this finding might indicate a problem with the task implementation, the stimulus distributions, or the model. Given that this unexpected finding is secondary to the main claim about incentives, a more thorough discussion of possible confounds or at minimum a clearer acknowledgment of the uncertainty about this result would be appropriate.
minor comments (10)
  1. §4.2, model specification: The notation in the model description (lines 1
  2. §4.2, JND calculation: The derivation of JND from F^{-1}(0.75) - F^{-1}(0.5) = F^{-1}(0.75) is stated but the step where F^{-1}(0.5) = 0 is only briefly annotated as 'since this is a 2AFC task.' A brief clarification that the 50% threshold corresponds to chance performance in a 2AFC task would help readers unfamiliar with psychometric functions.
  3. §5.3, Figure 6: The figure caption mentions 'simulated temperature calculated using the RNG in R' but the main text (§5.3) mentions concerns about the JavaScript RNG. The relationship between the JavaScript RNG used for the experiment and the R RNG used for posterior predictive checks should be clarified.
  4. §6: The qualitative analysis reports combined counts across both studies (N=444) but it is unclear how many participants were in each study and how the 224 codable responses were distributed across studies and conditions. A table breaking down the coded responses by study and condition would improve transparency.
  5. §7.3: The statement 'Of the participants for whom we could deduce a crossover point (53)' is unclear. How was the crossover point deduced? Was this from the open-ended responses or from the model estimates? Clarification is needed.
  6. Appendix A, benchmark calculations: The formula for E(U|rational) = 3960 appears to be in thousands of dollars, but this should be explicitly stated, as the main text (§5.3, footnote 3) notes that expected utility estimates are 'in thousands of dollars.'
  7. Figure 2: The y-axis label 'JND' uses a log scale but this is not indicated. A note that the y-axis is on a log scale would aid interpretation.
  8. §3: The phrase 'advertised 2 as a participation fee' contains a stray footnote marker. Please check formatting.
  9. §5.1: The attention check criteria are mentioned ('pre-registered attention check criteria') but the specific criteria are not described. A brief description or reference to the preregistration would be helpful.
  10. References: Several references appear to be from 2025

Simulated Author's Rebuttal

4 responses · 1 unresolved

We thank the referee for a thorough and constructive review. The referee correctly identifies several areas where the manuscript can be strengthened, and we agree with most of the points raised. Below we address each major comment in turn.

read point-by-point responses
  1. Referee: §5.1, Study 2 data-logging issue: The loss of 31 participants' data due to a data-logging failure reduced the planned sample from 280 to 249 (before attention-check exclusions, yielding 243). The paper does not report whether this data loss was random across conditions or whether it differentially affected certain conditions, which bears directly on the power and balance of the between-subjects design. A breakdown of exclusions by condition and a brief discussion of whether the data loss could have introduced bias should be added.

    Authors: The referee is correct that we did not report the breakdown of data loss by condition. We have now conducted this analysis. The data-logging failure was caused by a server-side issue that affected participants regardless of condition assignment, as the logging failure occurred at the level of the experiment platform's data submission endpoint rather than at the level of individual condition logic. The breakdown of the 31 lost participants across conditions is as follows: 8 from base interval, 7 from incentivized interval, 9 from base density, and 7 from incentivized density. This distribution is consistent with random loss across conditions (chi-square test of uniformity: p = 0.97). The final sample of 243 participants (60 base interval, 62 incentivized interval, 59 base density, 62 incentivized density) remains reasonably balanced. We will add this breakdown and the randomness assessment to §5.1 in the revised manuscript. revision: yes

  2. Referee: §3 and §5.1, incentive magnitude sufficiency: The central null result rests on the assumption that the incentive magnitudes were large enough to induce a behaviorally meaningful change in effort. In Study 2, the bonus was $0.50 per $1,000 of remaining virtual budget. For a single trial where p(freeze) = 0.235, the expected-value difference between the optimal action (salt, cost $1,000) and the suboptimal action (don't salt, expected cost 0.235 × $5,000 = $1,175) is $175 in virtual dollars, translating to approximately $0.0875 in real bonus. Across 18 trials, the total marginal incentive for making all optimal versus all suboptimal decisions is on the order of $1–2. The paper acknowledges this possibility in §7.3 ('not being sufficiently sensitive to incentives') but does not independently verify that the incentive was above a behaviorally relevant threshold. The time-on-task increase (20

    Authors: The referee raises a valid concern. We agree that we cannot independently verify that the incentive magnitude was above a behaviorally relevant threshold, and this is a genuine limitation of our study. However, we would note several points in mitigation. First, the incentive magnitudes we used ($2 expected bonus in Study 1, $1.60 average bonus in Study 2) are comparable to or larger than those used in prior incentivized visualization studies [12, 31, 60], which is the context our paper addresses. Second, the fact that participants in incentivized conditions spent 20-40% more time on the task provides behavioral evidence that participants did perceive the incentives as meaningful enough to change their behavior—they just did not translate that increased effort into improved performance. This pattern is consistent with the capital-labor-production theory [3], which predicts that incentives increase effort but may not improve performance when participants lack the cognitive capital to benefit from additional effort. That said, we agree that the per-trial marginal incentive in Study 2 was small (on the order of $0.09 per optimal decision), and we cannot rule out the possibility that larger incentives might have produced a detectable effect. We will revise §7.3 to explicitly acknowledge this limitation more prominently, including the per-trial marginal incentive calculation the referee describes, and will frame this as a boundary condition on our null result rather than a fully resolved question. We will also note that our null result should be interpreted as applying to incentive magnitudes typical of current crowdsourced visualization studies, not to arbitrarily large incentives. revision: partial

  3. Referee: §5.3 and Figure 5C, equivalence testing: The paper reports credible intervals for expected utility differences (e.g., incentivized interval: 2.76 [2.31, 3.12] vs. base interval: 2.69 [2.19, 3.04]) and concludes there is 'no evidence' of an incentive effect. However, the paper does not report a formal equivalence test or ROPE (Region of Practical Equivalence) analysis to distinguish 'no effect' from 'insufficient precision to detect an effect.' The credible intervals for the expected utility differences are moderately wide (e.g., the base density condition spans [1.70, 2.86]). Adding a ROPE analysis or explicitly stating what effect size would be considered practically meaningful would strengthen the claim that the null result reflects a true absence of incentive effects rather than insufficient precision.

    Authors: The referee is correct that our language conflates 'no evidence of an effect' with 'evidence of no effect,' and that a formal ROPE analysis would strengthen our claims. We will add a ROPE analysis to §5.3. To define the ROPE, we will use a practically meaningful effect size threshold based on the task structure: in Study 2, the difference in expected utility between the rational benchmark and the random-response benchmark is approximately 7,232 virtual dollars (3,960 vs. -3,272), which translates to approximately $3.62 in real bonus. We consider an incentive effect of 10% of this range (approximately 723 virtual dollars, or $0.36 in bonus) as the lower bound of practical significance—anything smaller would represent a change of less than $0.02 per trial, which we consider below the threshold of practical relevance. For Study 1, we will define the ROPE in terms of JND differences, using a threshold of 0.02 in correlation units, which corresponds to the smallest stimulus difference we tested. We will report the proportion of the posterior distribution falling within the ROPE for each comparison. Based on our preliminary analysis, the posterior probability mass within the ROPE exceeds 90% for the incentive comparisons in both studies, which we believe provides reasonable support for our claim. We will also soften our language from 'no evidence of an effect' to 'no practically meaningful effect' where appropriate, and will explicitly acknowledge the precision limitations the referee identifies. revision: yes

  4. Referee: §5.3 and §7.2, interval vs. density plot finding: The finding that interval plots led to marginally better expected utility than density plots contradicts prior work [12, 31, 60]. The authors propose a visual heuristic explanation (Figure 7: the 66% interval overlapping 0°C signals ~20% probability) but the follow-up study (Appendix B) with 95% intervals also showed no density advantage, which undermines the heuristic explanation. The paper acknowledges this is 'a conflicting picture' (§7.2) but the discussion does not fully grapple with whether this finding might indicate a problem with the task implementation, the stimulus distributions, or the model. Given that this unexpected finding is secondary to the main claim about incentives, a more thorough discussion of possible confounds or at minimum a clearer acknowledgment of the uncertainty about this result would be appropriate.

    Authors: We agree with the referee that our discussion of the interval vs. density finding is insufficiently thorough. The referee correctly notes that the Appendix B replication with 95% intervals (which removes the visual heuristic we proposed) still showed no density advantage, which undermines our proposed explanation. We will revise §7.2 to acknowledge this more directly and to discuss additional possible explanations. Specifically, we will discuss three possibilities: (1) The stimulus distributions we used had varying standard deviations, which prior work [15, 16] has shown can interfere with visual probability estimation from densities—this may have disproportionately disadvantaged the density condition. (2) Our task provided immediate feedback after each trial, which may have allowed participants in the interval condition to learn a simpler heuristic (e.g., 'salt when the interval bar is close to 0°C') that is not available in the same form for densities. (3) We cannot fully rule out a task implementation issue, and we will state this explicitly. We will also note that because this finding is secondary to our main claim about incentives, and because it contradicts a well-established finding in prior work, we are cautious about over-interpreting it and recommend it be treated as an exploratory finding warranting replication rather than a definitive result. We will add a sentence to this effect in §7.2. revision: yes

standing simulated objections not resolved
  • We cannot independently verify that the incentive magnitudes used were above a behaviorally relevant threshold for all participants. While the time-on-task increase provides indirect evidence that participants perceived the incentives as meaningful, we cannot rule out the possibility that larger incentives would produce a detectable performance effect. This is a genuine limitation that we will acknowledge more prominently but cannot fully resolve with the current data.

Circularity Check

0 steps flagged

No circularity: null results from preregistered between-subjects experiments with externally benchmarked metrics.

full rationale

This paper reports two (plus one replication) preregistered experiments manipulating performance-based incentives as a between-subjects factor. The central claim is a null result: incentives did not improve task performance. The dependent variables (JND for correlation perception, expected utility for decision-making) are computed from independently collected participant responses against external benchmarks (Weber's law for Study 1; rational-decision-maker expected utility for Study 2). The model specifications (§4.2, §5.2) use standard psychometric and linear-in-log-odds models with weakly regularizing priors centered on zero (no effect), which is appropriate for detecting effects rather than assuming them. The incentive magnitudes were calibrated from pilot data to match average compensation across conditions, but this calibration does not determine the null result—the bonus structure is an independent manipulation, and the outcome (whether participants performed better) is measured from their responses. The paper does cite prior work by some of the same authors (e.g., [56] for the decision-making task setup, [60] for the llo model), but these citations provide the experimental paradigm and model specification, not the conclusion. The null result is not forced by any fit or definition: the models could have detected incentive effects if present, and the credible intervals for the incentive condition differences in expected utility (e.g., incentivized interval: 2.76 [2.31, 3.12] vs. base interval: 2.69 [2.19, 3.04]) overlap substantially, consistent with a genuine null. No step in the derivation chain reduces to its inputs by construction.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 0 invented entities

The paper introduces no new entities, particles, forces, or dimensions. It is an empirical study with design parameters chosen from pilot data and prior work. The free parameters listed are experimental design choices (compensation amounts) rather than model-fitting parameters, but they are ad hoc in the sense that they were calibrated to match a target hourly wage rather than derived from theory. The axioms are primarily domain assumptions borrowed from behavioral economics, with one ad-hoc assumption about incentive magnitude sufficiency that is load-bearing for the null result.

free parameters (5)
  • Bonus per correct answer (Study 1) = $0.05
    Calibrated based on pilot data showing 40/65 correct responses on average, to target $2 average bonus and $15/h total compensation (§3). Not a free parameter in the modeling sense, but a design parameter chosen ad hoc from pilot data.
  • Bonus per $1,000 virtual budget (Study 2) = $0.50
    Calibrated based on prior work [56] showing average participant had $4,000 remaining, to target $2 average bonus (§3). Design parameter chosen from prior data.
  • Base pay (flat condition) = $3.00
    Set to achieve approximately $15/h given estimated 9-12 min completion time (§3).
  • Base pay (incentivized condition, Study 1) = $1.25, adjusted to $1.60
    Initially set to $1.25, adjusted upward to $1.60 after data collection because participants took longer than expected, to comply with Prolific minimum wage and target $15/h (§3, §4.1).
  • Base pay (incentivized condition, Study 2) = $1.50, adjusted to $4.20
    Initially set to $1.50, adjusted upward to $4.20 after data collection because participants took significantly longer (24 min median vs. 16 min), to comply with Prolific minimum wage and target $15/h (§3, §5.1).
axioms (5)
  • domain assumption Capital-labor-production theory (Camerer and Hogarth [3]) accurately describes the effect of incentives on task performance in visualization experiments.
    Used in §2.2 to generate predictions about which tasks should be affected by incentives. The theory is borrowed from behavioral economics and its applicability to crowdsourced visualization tasks is assumed, not independently verified.
  • domain assumption Time spent on task is a reasonable proxy for effort exerted by participants.
    Invoked in §7.1 to interpret the finding that incentivized participants spent more time but did not perform better: 'If time spent on a task is considered a reasonable proxy for effort exerted by a participant, this result suggests that the average participant in the incentivized condition is likely putting in more effort.' The paper flags this as conditional but relies on it for interpretation.
  • domain assumption Participants in incentivized conditions are attempting to maximize expected utility.
    The expected utility metric in Study 2 (§5.2-5.3) assumes participants are trying to maximize EU. The paper questions this assumption in §7.3 ('Are Participants Trying to Maximize Expected Utility?') but the metric is still used as the primary performance measure.
  • ad hoc to paper The incentive magnitudes used ($0.05/answer; $0.50/$1000 virtual) are sufficient to induce a behaviorally meaningful change in effort.
    The null result depends on this, but the magnitudes were calibrated to match average compensation rather than to be above a behavioral threshold. The paper acknowledges this in §7.3.
  • domain assumption Prolific crowdworkers are representative enough of the broader population of visualization study participants for the results to generalize.
    All participants were recruited from Prolific (§4.1, §5.1). The paper acknowledges in §8 that results 'may not easily translate to other scenarios (e.g., lab studies).'

pith-pipeline@v1.1.0-glm · 29712 in / 3835 out tokens · 488004 ms · 2026-07-09T10:06:10.241628+00:00 · methodology

0 comments
read the original abstract

A perennial research question in visualization involves identifying which visual encodings for a particular dataset are most effective for users in performing a specific task. The relative effectiveness of the different encodings are commonly identified through controlled experiments. However, designing an experiment involves making many, often ad hoc, decisions about the experimental setup such as whether to include a training module, whether to provide performance-based incentives to participants, etc. Yet, there is limited guidance on how these decisions should be made, and we do not fully understand the impact of these subjective decisions on empirical results. In this paper, we investigate the impact of one such key design decision: monetary rewards. Specifically, we ask: does providing or not providing participants with performance-based financial incentives affect the results and the conclusions that we draw from visualization studies? We conducted two crowdsourced studies investigating the impact of incentives on (i) a low-level, perceptual task (perception of correlations in scatterplots or parallel coordinate plots), and (ii) a task involving reasoning (decision-making based on a weather forecast represented as intervals or density plots). In each of these studies, we manipulate both the visual representation and the presence of incentives as between-subject conditions. We expected to find no effect of incentives on the perceptual task, but to see an effect for the decision-making task. However, we found no effect on task performance in either study. While these are results of only two studies and should be replicated, they suggest that performance-based financial incentives may not always have the intended effect on participants that we presumed, and calls for a reflection of how incentivized studies should be designed.

Figures

Figures reproduced from arXiv: 2607.07463 by Abhraneel Sarma, Alexander Lex, Matthew Kay, Michael Correll, Sheng Long.

Figure 1
Figure 1. Figure 1: Example of a stimulus seen by a participant in the non [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The main result of Experiment 1. We show the posterior es [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: Example of a stimulus seen by a participant in the incentivized [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The main result of Experiment 2. We show the posterior credible [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: The distribution of utility obtained by the participants and the [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Different uncertainty visualizations used in prior work (A, C) [e.g., [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: The main result of Experiment 3. We show the posterior credible [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

72 extracted references · 72 canonical work pages · 1 internal anchor

  1. [1]

    Achtziger, C

    A. Achtziger, C. Alós-Ferrer, S. Hügelschäfer, and M. Steinhauser. Higher incentives can impair performance: Neural evidence on rein- forcement and rationality. Social Cognitive and Affective Neuroscience , 10(11):1477–1483, Nov. 2015. doi: 10.1093/scan/nsv036 2, 9

  2. [2]

    P . Cala, T. Havranek, Z. Irsova, J. Matousek, Z. Irsova, and J. Novak. Financial Incentives and Performance: A Meta-Analysis of Economics Evidence, Nov. 2022. doi: 10.31222/osf.io/wbe9k 2

  3. [3]

    C. F. Camerer, R. M. Hogarth, D. V . Budescu, and C. Eckel. The Effects of Financial Incentives in Experiments: A Review and Capital-Labor- Production Framework. In B. Fischhoff and C. F. Manski, eds., Elici- tation of Preferences , pp. 7–48. Springer Netherlands, Dordrecht, 1999. doi: 10.1007/978-94-017-1406-8_2 2, 3, 7, 9

  4. [4]

    C. P . Cerasoli, J. M. Nicklin, and M. T. Ford. Intrinsic motivation and extrinsic incentives jointly predict performance: A 40-year meta-analysis. Psychological Bulletin, 140(4):980–1008, 2014. doi: 10.1037/a0035661 2, 9

  5. [5]

    W. S. Cleveland and R. McGill. Graphical Perception: Theory, Exper- imentation, and Application to the Development of Graphical Methods. Journal of the American Statistical Association , 79(387):531–554, Sept

  6. [6]

    doi: 10.1080/01621459.1984.10478080 9

  7. [7]

    Correll and M

    M. Correll and M. Gleicher. Error Bars Considered Harmful: Exploring Alternate Encodings for Mean and Error. IEEE Transactions on Visual- ization and Computer Graphics , 20(12):2142–2151, Dec. 2014. doi: 10. 1109/TVCG.2014.2346298 2, 8

  8. [8]

    Correll, D

    M. Correll, D. Moritz, and J. Heer. V alue-Suppressing Uncertainty Palettes. In Proceedings of the 2018 CHI Conference on Human Fac- tors in Computing Systems , pp. 1–11. ACM, Montreal QC Canada, Apr

  9. [9]

    doi: 10.1145/3173574.3174216 8

  10. [10]

    Cutler, J

    Z. Cutler, J. Wilburn, H. Shrestha, Y . Ding, B. Bollen, K. A. Nadib, T. He, A. McNutt, L. Harrison, and A. Lex. ReVISit 2: A Full Experiment Life Cycle User Study Framework, Aug. 2025. doi: 10.48550/arXiv.2508 .03876 1, 3, 4

  11. [11]

    Davis, X

    R. Davis, X. Pu, Y . Ding, B. D. Hall, K. Bonilla, M. Feng, M. Kay, and L. Harrison. The Risks of Ranking: Revisiting Graphical Perception to Model Individual Differences in Visualization Performance. IEEE Trans- actions on Visualization and Computer Graphics , 30(3):1756–1771, Mar

  12. [12]

    doi: 10.1109/TVCG.2022.3226463 9

  13. [13]

    M. Dong, L. Chen, L. Wang, X. Jiang, and G. Chen. Uncertainty Visual- ization for Mobile and Wearable Devices Based Activity Recognition Sys- tems. International Journal of Human–Computer Interaction, 33(2):151– 163, Feb. 2017. doi: 10.1080/10447318.2016.1224527 8

  14. [14]

    Dragicevic, Y

    P . Dragicevic, Y . Jansen, A. Sarma, M. Kay, and F. Chevalier. Increasing the Transparency of Research Papers with Explorable Multiverse Analy- ses. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , pp. 1–15. ACM, Glasgow Scotland Uk, May 2019. doi: 10.1145/3290605.3300295 2

  15. [15]

    Fernandes, L

    M. Fernandes, L. Walls, S. Munson, J. Hullman, and M. Kay. Uncertainty Displays Using Quantile Dotplots or CDFs Improve Transit Decision- Making. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, pp. 1–12. ACM, Montreal QC Canada, Apr. 2018. doi: 10.1145/3173574.3173718 1, 3, 6, 7, 8, 12

  16. [16]

    Ferreira, D

    N. Ferreira, D. Fisher, and A. C. Konig. Sample-oriented task-driven vi- sualizations: Allowing users to make better, more confident decisions. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, pp. 571–580. ACM, Toronto Ontario Canada, Apr. 2014. doi: 10 .1145/2556288.2557131 8

  17. [17]

    Franco, N

    A. Franco, N. Malhotra, and G. Simonovits. Publication bias in the social sciences: Unlocking the file drawer. Science, 345(6203):1502–1505, Sept

  18. [18]

    doi: 10.1126/science.1255484 8

  19. [19]

    Fygenson, E

    R. Fygenson, E. Bertini, and L. M. Padilla. Croissant Charts: Modulating the Performance of Normal Distribution Visualizations with Affordances. Computer Graphics F orum, n/a(n/a):e70463, Nov. 2026. doi: 10.1111/ cgf.70463 8

  20. [21]

    Gabry, R

    J. Gabry, R. ˇCešnovar, A. Johnson, and S. Bronder. Cmdstanr: R Interface to ’CmdStan’, 2025. 4

  21. [22]

    Galesic, R

    M. Galesic, R. Garcia-Retamero, and G. Gigerenzer. Using icon arrays to communicate medical risks: Overcoming low numeracy. Health Psy- chology, 28(2):210–216, 2009. doi: 10.1037/a0014474 2

  22. [23]

    Ghoniem, J.-D

    M. Ghoniem, J.-D. Fekete, and P . Castagliola. A Comparison of the Read- ability of Graphs Using Node-Link and Matrix-Based Representations. In IEEE Symposium on Information Visualization , pp. 17–24. IEEE, Austin, TX, USA, 2004. doi: 10.1109/INFVIS.2004.1 9

  23. [24]

    A 30% Chance of Rain Tomorrow

    G. Gigerenzer, R. Hertwig, E. V an Den Broek, B. Fasolo, and K. V . Kat- sikopoulos. “A 30% Chance of Rain Tomorrow”: How Does the Public Understand Probabilistic Weather Forecasts? Risk Analysis , 25(3):623– 629, 2005. doi: 10.1111/j.1539-6924.2005.00608.x 8, 9

  24. [25]

    Gonzalez-Rubio, P

    M. Gonzalez-Rubio, P . A. Iturralde, and G. Torres-Oviedo. Weber’s Law in walking: Sensory scaling is observed in multi-sensory, dynamic tasks. Scientific Reports, June 2026. doi: 10.1038/s41598-026-54948-5 4

  25. [26]

    Greis, A

    M. Greis, A. Joshi, K. Singer, A. Schmidt, and T. Machulla. Uncertainty Visualization Influences how Humans Aggregate Discrepant Information. In Proceedings of the 2018 CHI Conference on Human Factors in Com- puting Systems , pp. 1–12. ACM, Montreal QC Canada, Apr. 2018. doi: 10.1145/3173574.3174079 8

  26. [27]

    B. D. Hall, Y . Liu, Y . Jansen, P . Dragicevic, F. Chevalier, and M. Kay. A Survey of Tasks and Visualizations in Multiverse Analysis Reports. Com- puter Graphics F orum, 41(1):402–426, 2022. doi: 10.1111/cgf.14443 2

  27. [28]

    Harrison, F

    L. Harrison, F. Y ang, S. Franconeri, and R. Chang. Ranking Visualiza- tions of Correlation Using Weber’s Law. IEEE Transactions on Visual- ization and Computer Graphics , 20(12):1943–1952, Dec. 2014. doi: 10. 1109/TVCG.2014.2346979 1, 2, 3, 4

  28. [29]

    P . D. Harvey. Domains of cognition and their assessment. Dialogues in Clinical Neuroscience, 21(3):227–237, Sept. 2019. doi: 10.31887/DCNS .2019.21.3/pharvey 2, 3

  29. [30]

    Heer and M

    J. Heer and M. Bostock. Crowdsourcing graphical perception: Using mechanical turk to assess visualization design. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems , CHI ’10, pp. 203–212. Association for Computing Machinery, New Y ork, NY , USA, Apr. 2010. doi: 10.1145/1753326.1753357 9

  30. [31]

    Hullman, P

    J. Hullman, P . Resnick, and E. Adar. Hypothetical Outcome Plots Out- perform Error Bars and Violin Plots for Inferences about Reliability of V ariable Ordering. PLOS ONE , 10(11):e0142444, Nov. 2015. doi: 10. 1371/journal.pone.0142444 2, 8

  31. [32]

    Joslyn and J

    S. Joslyn and J. LeClerc. Decisions With Uncertainty: The Glass Half Full. Current Directions in Psychological Science , 22(4):308–315, Aug

  32. [33]

    doi: 10.1177/0963721413481473 1, 2

  33. [34]

    S. L. Joslyn and J. E. LeClerc. Uncertainty forecasts improve weather- related decisions and attenuate the effects of forecast error. Journal of Experimental Psychology: Applied , 18(1):126–140, 2012. doi: 10.1037/ a0025185 2, 3, 5

  34. [35]

    A. Kale. Toward a Logic of Generalization about Visualization as a Deci- sion Aid. In 2025 IEEE Visualization and Visual Analytics (VIS), pp. 1–5, Nov. 2025. doi: 10.1109/VIS60296.2025.00005 3, 9

  35. [37]

    M. Kay. Tidybayes: Tidy Data and Geoms for Bayesian Models, Sept

  36. [38]

    doi: 10.5281/zenodo.13770114 12

  37. [40]

    M. Kay, T. Kola, J. R. Hullman, and S. A. Munson. When (ish) is My Bus?: User-centered Visualizations of Uncertainty in Everyday, Mobile Predictive Systems. In Proceedings of the 2016 CHI Conference on Hu- man Factors in Computing Systems , pp. 5092–5103. ACM, San Jose Cal- ifornia USA, May 2016. doi: 10.1145/2858036.2858558 2, 8

  38. [41]

    Keller, C

    R. Keller, C. M. Eckert, and P . J. Clarkson. Matrices or Node-Link Dia- grams: Which Visual Representation is Better for Visualising Connectiv- ity Models? Information Visualization, 5(1):62–76, Mar. 2006. doi: 10. 1057/palgrave.ivs.9500116 9

  39. [42]

    J. F. Landy, M. L. Jia, I. L. Ding, D. Viganola, W. Tierney, A. Dreber, M. Johannesson, T. Pfeiffer, C. R. Ebersole, Q. F. Gronau, A. Ly, D. van den Bergh, M. Marsman, K. Derks, E.-J. Wagenmakers, A. Proctor, D. M. Bartels, C. W. Bauman, W. J. Brady, F. Cheung, A. Cimpian, S. Dohle, M. B. Donnellan, A. Hahn, M. P . Hall, W. Jiménez-Leal, D. J. Johnson, R....

  40. [43]

    LeClerc and S

    J. LeClerc and S. Joslyn. The Cry Wolf Effect and Weather-Related De- cision Making. Risk Analysis , 35(3):385–395, 2015. doi: 10.1111/risa. 12336 1

  41. [44]

    G. F. Loewenstein, E. U. Weber, C. K. Hsee, and N. Welch. Risk as feelings. Psychological Bulletin , 127(2):267–286, 2001. doi: 10.1037/ 0033-2909.127.2.267 8

  42. [45]

    A. M. MacEachren, R. E. Roth, J. O’Brien, B. Li, D. Swingley, and M. Gahegan. Visual Semiotics & Uncertainty Visualization: An Empiri- cal Study. IEEE Transactions on Visualization and Computer Graphics , 18(12):2496–2505, Dec. 2012. doi: 10.1109/TVCG.2012.279 8

  43. [46]

    performance of crowds

    W. Mason and D. J. Watts. Financial incentives and the "performance of crowds". In Proceedings of the ACM SIGKDD Workshop on Human Com- putation, HCOMP ’09, pp. 77–85. Association for Computing Machinery, New Y ork, NY , USA, June 2009.doi: 10.1145/1600150.1600175 2

  44. [47]

    McElreath

    R. McElreath. Statistical Rethinking: A Bayesian Course with Examples in R and STAN . Chapman and Hall/CRC, New Y ork, 2 ed., Mar. 2020. doi: 10.1201/9780429029608 2

  45. [48]

    Nadav-Greenberg and S

    L. Nadav-Greenberg and S. L. Joslyn. Uncertainty Forecasts Improve Decision Making Among Nonexperts. Journal of Cognitive Engineer- ing and Decision Making , 3(3):209–227, Sept. 2009. doi: 10.1518/ 155534309X474460 2

  46. [49]

    Nobre, D

    C. Nobre, D. Wootton, L. Harrison, and A. Lex. Evaluating Multivariate Network Visualization Techniques Using a V alidated Design and Crowd- sourcing Approach. In Proceedings of the 2020 CHI Conference on Hu- man Factors in Computing Systems , pp. 1–12. ACM, Honolulu HI USA, Apr. 2020. doi: 10.1145/3313831.3376381 9

  47. [51]

    M. Okoe, R. Jianu, and S. Kobourov. Node-Link or Adjacency Matrices: Old Question, New Insights. IEEE Transactions on Visualization and Computer Graphics, 25(10):2940–2952, Oct. 2019. doi: 10.1109/TVCG. 2018.2865940 9

  48. [52]

    B. Oral, P . Dragicevic, A. Telea, and E. Dimara. Decoupling Judgment and Decision Making: A Tale of Two Tails. IEEE Transactions on Visu- alization and Computer Graphics, 30(10):6928–6940, Oct. 2024. doi: 10 .1109/TVCG.2023.3346640 1

  49. [53]

    L. M. Padilla, I. T. Ruginski, and S. H. Creem-Regehr. Effects of en- semble and summary displays on interpretations of geospatial uncertainty data. Cognitive Research: Principles and Implications , 2(1):40, Dec

  50. [54]

    doi: 10.1186/s41235-017-0076-1 2

  51. [55]

    L. M. K. Padilla, M. Powell, M. Kay, and J. Hullman. Uncertain About Uncertainty: How Qualitative Expressions of Forecaster Confidence Im- pact Decision-Making With Uncertainty Visualizations. Frontiers in Psy- chology, 11, Jan. 2021. doi: 10.3389/fpsyg.2020.579267 1, 3, 7, 12

  52. [56]

    K. Reda. Rainbow Colormaps: What are They Good and Bad for? IEEE Transactions on Visualization and Computer Graphics , 29(12):5496– 5510, Dec. 2023. doi: 10.1109/TVCG.2022.3214771 9

  53. [57]

    K. Reda, P . Nalawade, and K. Ansah-Koi. Graphical Perception of Contin- uous Quantitative Maps: The Effects of Spatial Frequency and Colormap Design. In Proceedings of the 2018 CHI Conference on Human Fac- tors in Computing Systems , CHI ’18, pp. 1–12. Association for Comput- ing Machinery, New Y ork, NY , USA, Apr. 2018. doi: 10.1145/3173574. 3173846 9

  54. [58]

    Reda and D

    K. Reda and D. A. Szafir. Rainbows Revisited: Modeling Effective Col- ormap Design for Graphical Inference. IEEE Transactions on Visual- ization and Computer Graphics , 27(2):1032–1042, Feb. 2021. doi: 10. 1109/TVCG.2020.3030439 9

  55. [59]

    R. A. Rensink and G. Baldridge. The Perception of Correlation in Scat- terplots. Computer Graphics F orum, 29(3):1203–1210, 2010. doi: 10. 1111/j.1467-8659.2009.01694.x 1, 2, 3

  56. [60]

    D. P . Retchless and C. A. Brewer. Guidance for representing uncertainty on global temperature change maps. International Journal of Climatol- ogy, 36(3):1143–1159, 2016. doi: 10.1002/joc.4408 8

  57. [61]

    Rosenthal

    R. Rosenthal. The file drawer problem and tolerance for null results. Psy- chological Bulletin, 86(3):638–641, 1979. doi: 10.1037/0033-2909.86.3. 638 8

  58. [62]

    I. T. Ruginski, A. P . Boone, L. M. Padilla, L. Liu, N. Heydari, H. S. Kramer, M. Hegarty, W. B. Thompson, D. H. House, and S. H. Creem- Regehr. Non-expert interpretations of hurricane forecast uncertainty visu- alizations. Spatial Cognition & Computation , 16(2):154–172, Apr. 2016. doi: 10.1080/13875868.2015.1137577 2

  59. [63]

    Sarma, M

    A. Sarma, M. Hedayati, and M. Kay. More Forecasts, More (Decision) Problems: How Uncertainty Representations for Multiple Forecasts Im- pact Decision-Making, Feb. 2025. doi: 10.31219/osf.io/t4e9u_v1 1, 3, 5, 8

  60. [65]

    Sarma, A

    A. Sarma, A. Kale, M. J. Moon, N. Taback, F. Chevalier, J. Hullman, and M. Kay. Multiverse: Multiplexing Alternative Data Analyses in R Note- books. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pp. 1–15. ACM, Hamburg Germany, Apr. 2023. doi: 10.1145/3544548.3580726 2

  61. [66]

    Sarma, S

    A. Sarma, S. Long, M. Correll, and M. Kay. Tasks and Telephones: Threats to Experimental V alidity due to Misunderstandings of Visuali- sation Tasks and Strategies Position Paper. In 2024 IEEE Evaluation and Beyond - Methodological Approaches for Visualization (BELIV) , pp. 33– 40, Oct. 2024. doi: 10.1109/BELIV64461.2024.00009 2

  62. [67]

    Masson, S

    A. Sarma, X. Pu, Y . Cui, M. Correll, E. T. Brown, and M. Kay. Odds and Insights: Decision Quality in Exploratory Data Analysis Under Un- certainty. In Proceedings of the 2024 CHI Conference on Human Fac- tors in Computing Systems , CHI ’24, pp. 1–14. Association for Comput- ing Machinery, New Y ork, NY , USA, May 2024. doi: 10.1145/3613904. 3641995 1, 3, 7, 8

  63. [68]

    J. P . Simmons, L. D. Nelson, and U. Simonsohn. False-Positive Psy- chology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant. Psychological Science, 22(11):1359– 1366, Nov. 2011. doi: 10.1177/0956797611417632 2

  64. [69]

    Simonsohn, J

    U. Simonsohn, J. P . Simmons, and L. D. Nelson. Specification curve analysis. Nature Human Behaviour , 4(11):1208–1214, Nov. 2020. doi: 10.1038/s41562-020-0912-z 1, 2

  65. [70]

    Steegen, F

    S. Steegen, F. Tuerlinckx, A. Gelman, and W. V anpaemel. In- creasing Transparency Through a Multiverse Analysis. Perspectives on Psychological Science , 11(5):702–712, Sept. 2016. doi: 10.1177/ 1745691616658637 1, 2

  66. [71]

    S. S. Stevens. On the psychophysical law. Psychological Review , 64(3):153–181, 1957. doi: 10.1037/h0046162 2, 8

  67. [72]

    D. A. Szafir. Modeling Color Difference for Visualization Design. IEEE Transactions on Visualization and Computer Graphics , 24(1):392–401, Jan. 2018. doi: 10.1109/TVCG.2017.2744359 9

  68. [73]

    R. C. Team. R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, 2024. 4

  69. [74]

    J. M. Wicherts, C. L. S. V eldkamp, H. E. M. Augusteijn, M. Bakker, R. C. M. van Aert, and M. A. L. M. van Assen. Degrees of Freedom in Planning, Running, Analyzing, and Reporting Psychological Studies: A Checklist to Avoid p-Hacking. Frontiers in Psychology, 7, Nov. 2016. doi: 10.3389/fpsyg.2016.01832 2

  70. [75]

    Y . Wu, Z. Guo, M. Mamakos, J. Hartline, and J. Hullman. The Ratio- nal Agent Benchmark for Data Visualization. IEEE Transactions on Vi- sualization and Computer Graphics , 30(1):338–347, Jan. 2024. doi: 10. 1109/TVCG.2023.3326513 3, 9

  71. [76]

    Y ang, M

    F. Y ang, M. Hedayati, and M. Kay. Subjective Probability Correction for Uncertainty Representations. In Proceedings of the 2023 CHI Con- ference on Human Factors in Computing Systems , CHI ’23, pp. 1–17. Association for Computing Machinery, New Y ork, NY , USA, Apr. 2023. doi: 10.1145/3544548.3580998 1, 3, 5

  72. [77]

    Utilizing stability criteria in choosing feature selection methods yields reproducible results in microbiome data

    H. Zhang and L. T. Maloney. Ubiquitous Log Odds: A Common Repre- sentation of Probability and Frequency Distortion in Perception, Action, and Cognition. Frontiers in Neuroscience , 6, Jan. 2012. doi: 10.3389/ fnins.2012.00001 5 A I NTERPRETING THE RESULTS OF EXPERIMENT 2 Calculation of Benchmarks As described in section 5 , there are two payoff relevant s...