REVIEW 3 major objections 6 minor 3 references
An open-source Modular Online Psychophysics Platform (MOPP)
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read An open-source web platform, MOPP, lets researchers build and run psychophysics experiments online without writing code, and a pilot study finds the resulting data match published lab results.
desk verdict MOPP is a genuinely useful open-source platform paper whose infrastructure contribution stands, but the 'similar to lab' validation claim overshoots what a 13–17-participant pilot with Bayes factors near 0.5 can establish. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery that carries the argument is the three-part platform stack — a Node.js server that runs experiments, a MongoDB database for storage, and a React.js client that serves both researcher and participant interfaces — together with four built-in quality-control tools: email/IP authentication, reCAPTCHA, a virtual chinrest that computes viewing distance from the retinal blind spot, and a Taylor's E staircase test of visual acuity. The validation machinery is the comparison of each pilot task's summary statistics to published lab values, using t-tests and Bayes factors.
What would settle it
A preregistered equivalence study with a larger sample would settle it: for example, collect Mooney face d' from 60 online MOPP participants and test whether the 95% confidence interval falls inside a pre-specified equivalence margin around the Schwiedrzik et al. benchmark of 1.19; if the interval excludes that margin or the Bayes factor favors a meaningful difference, the central validation claim fails for that task.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that a single modular web platform can cover the full lifecycle of an online psychophysics experiment — building, hosting, launching, participant screening, visual calibration, data collection, and download — and that the five preloaded tasks yield behavioral patterns comparable to published laboratory benchmarks. The length task reproduced the classic contraction toward the mean; the numerosity task reproduced sublinear power-law estimation; biological-motion and Mooney-face tasks gave d' values that did not differ significantly from published online and offline samples; and key-tapping rates matched the original study. The authors therefore conclude that MOPP's task implementations are validated and that the platform can help researchers collect large psychophysics datasets online in a standardized manner.
Load-bearing premise
The load-bearing premise is that 'no significant difference' from published lab means, assessed with 13–17 participants per task, is enough to establish that online MOPP data are equivalent to laboratory data.
Editorial extensions
If this is right
- Researchers who cannot program can assemble experiments from preloaded tasks, clone existing experiments, and modify task parameters such as distributions, trial counts, and stimulus durations.
- Each participant's viewing distance and visual acuity are measured automatically, so stimulus sizes can be corrected in real time and data can be filtered post hoc for participants who moved.
- Public mode with email/IP authentication and reCAPTCHA reduces duplicate and bot responses, while the completion code integrates with crowdsourcing platforms.
- The five validated preloaded tasks provide ready-to-run benchmarks for line length, numerosity, biological motion, Mooney faces, and motor tapping, with new tasks added through jsPsych plugins.
- Because the same experiment can run in supervised mode in the lab and public mode online, MOPP enables direct offline-online comparisons within one platform.
Reading between the lines
- The paper leaves implicit that its validation template — compare pilot data against published lab benchmarks — could be reused as a general onboarding test for any new task added to MOPP.
- A stronger equivalence claim would require specifying a tolerance margin in advance; the reported Bayes factors of 0.27–0.53 are anecdotal, so the platform's standardization benefit is better supported by its built-in calibration than by this pilot alone.
- Because viewing distance is recorded twice per participant, MOPP opens a practical post-hoc quality metric — excluding sessions where distance shifted — that the paper describes but does not formalize into a recommendation.
- If the installation barrier (manual AWS setup) is lowered by a future installer, the modular sharing model could make this platform a common interchange format for lab-specific visual tasks, a consequence the authors mention only as a possibility.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MOPP, an open-source web-based platform for building and hosting psychophysics experiments. It describes the three-component architecture (Node.js server, MongoDB database, React client), a workflow for researchers and participants, and five preloaded tasks: length estimation, numerosity, biological motion discrimination, Mooney face detection, and key tapping. To evaluate the platform, the authors ran a small online pilot (17 participants) and compared the results with published laboratory benchmarks, concluding that 'in all five tasks, the data yielded results similar to previous publications collected in laboratory settings, validating our task implementations.' The code and pilot data are publicly available.
Significance. If the platform performs as described, MOPP addresses a real need: lowering the technical barrier for non-programmers, providing integrated viewing-distance and visual-acuity calibration, and building in participant credibility checks. The open-source release, documented workflow, and public pilot data are concrete strengths that will benefit the community. However, the validation evidence is the weakest link: the pilot is small, the statistical comparisons are not designed to establish equivalence, and two of the five tasks are compared only qualitatively. The platform’s architectural and usability claims are plausible, but the strong wording of the validation claim goes beyond what the data support.
major comments (3)
- [§C.3, §C.4, §C.5] The comparisons with laboratory benchmarks rely on non-significant t-tests and Bayes factors between 0.27 and 0.53 with N=13–17 per task. For example, the biological motion d' comparison reports BF10 = 0.40 and 0.50 (vs. N=189 and N=19), and the Mooney face comparison reports BF10 = 0.53. The text states that the values 'did not differ significantly' and presents this as evidence of similarity. With such small samples, these tests have very low power, and a null result is not evidence of equivalence. No pre-specified equivalence bound, TOST procedure, or region of practical equivalence is provided. The Bayes factors are anecdotal and prior-dependent. This is the load-bearing step for the central validation claim, and it is statistically insecure.
- [§C.1, §C.2] The length and numerosity tasks are compared to previous findings only in terms of the direction of the effect—'regression to the mean' and 'underestimation'—rather than to specific published effect sizes. In the numerosity task, the regression is fit only for stimuli exceeding ten dots, with no justification for this restriction. The reported slope (β1 = 0.88, 95% CI [0.79, 0.96]) is not contrasted with a quantitative benchmark from the cited studies, so the claim of similarity to laboratory results is supported only at a qualitative level for these two tasks.
- [Discussion, paragraph 2] The statement 'In all five tasks, the data yielded results similar to previous publications collected in laboratory settings, validating our task implementations' overstates what the pilot can establish. Three of the five tasks have direct statistical comparisons, and those are statistically weak; two tasks have only directional qualitative matches. The term 'validating' is too strong for a feasibility pilot of this size and with these analyses. I recommend softening the claim to 'consistent with' and explicitly describing the results as a preliminary demonstration rather than a validation.
minor comments (6)
- [Section B.1, Figure 2] The text contains missing placeholders for icons, e.g., 'using the button' and 'drag-and-drop using the cursor'. The icons appear to be absent from the manuscript; please include them or describe them in words.
- [References and text] The citation 'Weil et al (2018)' in §C.3 is inconsistent with the reference list entry 'Weil, R. S., Schwarzkopf, D. S., ...' Please standardize the author-year format. Similarly, 'Schwiedrzik (2018)' in §C.4 should be 'Schwiedrzik et al. (2018)'.
- [Abstract] The abstract contains a formatting error: 'i ii)' should be 'iii)'. Please correct.
- [§C.5] For the key-tapping comparison, the manuscript reports the pooled single-hand mean (55.4 ± 7.0 taps) versus the literature value (60.3 ± 1.4 taps), but it does not report the per-hand means or the exact pooling method for the pilot data. Since the excluded participants differ by condition, this would help the reader interpret the comparison.
- [§C.4] When citing the Mooney face benchmark, the phrase 'their N = 18' is ambiguous because the reference includes several experiments; please specify that this is the number of participants in the relevant condition of Schwiedrzik et al. (2018).
- [General] The manuscript would benefit from a version number or archived DOI for the GitLab repository to improve reproducibility and allow readers to cite a stable snapshot of the code.
Circularity Check
No significant circularity: the pilot validation compares MOPP data against external published laboratory benchmarks, and no equation or fitted parameter is presented as a prediction derived from those benchmarks.
full rationale
The paper's central claim is that MOPP produces data 'similar to those reported in laboratory settings.' This is supported by pilot comparisons against externally published studies (e.g., Weil et al., 2018; Schwiedrzik et al., 2018; Noyce et al., 2014), which are independent of the authors' own fitting procedures. The regression fits in the length and numerosity tasks (slope 0.51 for length, slope 0.88 for numerosity) are descriptive summaries of the pilot data, not quantities derived from the benchmark values; the benchmarks are used only as qualitative or statistical comparisons. The Mooney and biological motion stimuli are adapted from the same public stimulus sets used in the benchmark papers, which makes the comparison appropriate rather than circular. The key-tapping comparison uses the original study's published mean as an external benchmark. None of the paper's equations reduces a claimed prediction to a fitted input, and no load-bearing premise is justified solely by a self-citation or by an imported uniqueness theorem. Concern that non-significance or anecdotal Bayes factors do not establish equivalence is a statistical validity issue, not circularity. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Length task linear regression parameters =
beta0 = 3.68, beta1 = 0.51
- Numerosity task log-log regression parameters =
beta0 = 0.20, beta1 = 0.88
- Pilot stimulus design settings =
Uniform 1-18 for length; Gaussian mean 22, SD 5 for numerosity; 15/10/7 dots for biological motion; 24 trials per task
assumptions (5)
- domain assumption Published lab results used as benchmarks (Weil et al., 2018; Schwiedrzik et al., 2018; Noyce et al., 2014) are valid and comparable to the MOPP pilot measurements.
- domain assumption The virtual chinrest (Li et al., 2020) accurately infers viewing distance from blind spot geometry in home environments.
- domain assumption Taylor's E test (Bach, 2006) measures visual acuity validly when administered online.
- domain assumption Email-based authentication, unique IP checks, and reCAPTCHA sufficiently establish participant credibility.
- domain assumption Browser-based stimulus presentation preserves the timing and appearance needed for these visual tasks.
Cite this review
Pith. "Pith review of An open-source Modular Online Psychophysics Platform (MOPP)." pith.science (2026). https://pith.science/paper/KO3A5U4K
@misc{pith2026250523137,
author = {Pith},
title = {Pith review of: An open-source Modular Online Psychophysics Platform (MOPP)},
year = {2026},
howpublished = {\url{https://pith.science/paper/KO3A5U4K}},
note = {Machine review of arXiv:2505.23137}
}
read the original abstract
In recent years, there is a growing need and opportunity to use online platforms for psychophysics research. Online experiments make it possible to evaluate large and diverse populations remotely and quickly, complementing laboratory-based research. However, developing and running online psychophysics experiments poses several challenges: i) a high barrier-to-entry for researchers who often need to learn complex code-based platforms, ii) an uncontrolled experimental environment, and iii) questionable credibility of the participants. Here, we introduce an open-source Modular Online Psychophysics Platform (MOPP) to address these challenges. Through the simple web-based interface of MOPP, researchers can build modular experiments, share them with others, and copy or modify tasks from each others environments. MOPP provides built-in features to calibrate for viewing distance and to measure visual acuity. It also includes email-based and IP-based authentication, and reCAPTCHA verification. We developed five example psychophysics tasks, that come preloaded in the environment, and ran a pilot experiment which was hosted on the AWS (Amazon Web Services) cloud. Pilot data collected for these tasks yielded similar results to those reported in laboratory settings. MOPP can thus help researchers collect large psychophysics datasets online, with reduced turnaround time, and in a standardized manner.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[10]
Just Another Tool for Online Studies
https://doi.org/10.1111/j.1467-9450.1961.tb01215.x Grootswagers, T. (2020). A primer on running human behavioural experiments online. Behavior Research Methods, 52(6), 2283–2286. https://doi.org/10.3758/s13428-020-01395-3 Indow, T., & Ida, M. (1977). Scaling of dot numerosity. Perception & Psychophysics, 22(3), 265–276. https://doi.org/10.3758/BF03199689 ...
-
[188]
https://doi.org/10.1177/0963721414531598 Peer, E., Rothschild, D., Gordon, A., & Damer, E. (2022). Erratum to Peer et al. (2021) Data quality of platforms and panels for online behavioral research. Behavior Research Methods, 54(5), 2618–2620. https://doi.org/10.3758/s13428-022- 01909-1 Peirce, J., Gray, J. R., Simpson, S., MacAskill, M., Höchenberger, R.,...
-
[428]
https://doi.org/10.1016/j.cub.2008.02.052 Chandler, J., Mueller, P., & Paolacci, G. (2014). Nonnaïveté among Amazon Mechanical Turk workers: Consequences and solutions for behavioral researchers. Behavior Research Methods, 46(1), 112–130. https://doi.org/10.3758/s13428-013-0365-7 Chmielewski, M., & Kucker, S. C. (2020). An MTurk Crisis? Shifts in Data Qua...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.