REVIEW 3 major objections 4 minor 39 references
Characterising Volunteers' Task Execution Patterns Across Projects on Multi-Project Citizen Science Platforms
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read On multi-project citizen science platforms, volunteers who already contribute to one project tend to do more tasks in a new project than volunteers recruited from outside, even though outside recruitment delivers more people.
desk verdict A useful descriptive case study of cross-project engagement on citizen science platforms, but its central claim about inherited volunteers' higher engagement is confounded by selection and should not drive design recommendations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument turns on a metric the paper calls balance in computing, $(t - m)/\min(t, m)$, where $t$ is the average number of tasks performed by volunteers inherited from other projects on the platform and $m$ is the average for volunteers recruited from outside; a positive value means inherited volunteers do more work per person. This metric is the direct evidence for the paper's main claim about the value of cross-project engagement. It sits inside a Goal-Question-Metric (GQM) framework that also defines the exploration rate ($p/a$, projects tried over projects available), the engagement rate ($g/a$, projects contributed to on at least two days over projects available), relative activity duration, balance in recruitment, and two Gini inequality coefficients. The qualitative half uses the Semiotic Inspection Method, a protocol for reading the designer-to-user messages encoded in interface signs, to classify the cross-project features platforms actually communicate.
What would settle it
Recompute the balance-in-computing metric on a platform that logs referral sources or first page views before the first task. If volunteers who verifiably arrived from outside the platform perform as many or more tasks per person than volunteers who verifiably came from other projects, the paper's central claim fails.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is a consistent asymmetry in how volunteers reach new projects. On Crowdcrafting, Socientize, and GeoTag-X, most projects recruit more volunteers from outside the platform than from its other projects, yet the volunteers inherited from other projects perform more tasks per person; the paper's 'balance in computing' is positive for most projects. At platform level, between 13 and 26 per cent of volunteers explore multiple projects, at most 6 per cent become regular contributors to multiple projects, and a few flagship projects capture most recruitment and task execution, with Gini coefficients between 0.47 and 0.95. The interface inspection adds a design finding: the platforms offer project search, featured-project lists, and lists of the volunteer's own projects, but no personalised or explainable recommendation of which new project to try next.
Load-bearing premise
The load-bearing assumption is that a volunteer was recruited by the first project they contributed to after registering; if volunteers browse or are directed to other projects before that first contribution, the split between 'inherited' and 'recruited outside' is misclassified, and the higher productivity of inherited volunteers could be an artifact of that misclassification.
Editorial extensions
If this is right
- Platform managers can treat cross-project movement as a retention lever: volunteers who regularly engage with multiple projects show longer relative activity duration than those who stay in one project, so encouraging exploration may keep people on the platform.
- New projects should expect recruitment campaigns to bring many one-time helpers and should not mistake headcount for engagement; the smaller group of inherited volunteers is the more productive segment per person.
- A concrete design gap is the absence of personalised, explainable project recommendations; filling it could raise the small fraction of volunteers who become multi-project regulars.
- The Gini coefficients give platform managers a monitoring instrument: when a few projects capture nearly all recruitment and task execution, the platform may need to intervene to keep less visible projects viable.
Reading between the lines
- Our inference: the balance-in-computing ratio is scale-free but denominator-sensitive; on a project with very few inherited volunteers, a single highly active volunteer can make the metric look strongly positive, so the aggregate result should be read alongside counts, not just ratios.
- Our inference: the paper's design recommendation assumes that multi-project participation causes longer retention, but the data are correlational; a randomised trial that prompts some explorers with personalised recommendations and not others would test whether the recommendations themselves extend engagement.
- Our inference: the extreme Gini inequality suggests an early-visibility advantage; if that dynamic holds, recommendation algorithms that favour new or small projects could be a testable way to flatten the distribution of attention.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a Goal-Question-Metric (GQM) framework plus the Semiotic Inspection Method (SIM) to characterize volunteers' cross-project task-execution behavior on multi-project citizen science platforms. It applies the quantitative metrics to task-execution logs from Crowdcrafting, Socientize, and GeoTag-X and the qualitative inspection to Crowdcrafting, Zooniverse, and CitSci.org. The main reported findings are that only a minority of volunteers explore multiple projects, very few become regular multi-project contributors, attention and recruitment are highly concentrated in a small number of projects, and volunteers inherited from other projects on the platform tend to perform more tasks than volunteers recruited from outside. The paper also identifies three classes of cross-project interface signs and derives design recommendations, most notably that platforms should offer personalized, explainable project recommendations.
Significance. If the quantitative claims were causally supported, the paper would be a useful empirical contribution to citizen-science HCI: it addresses a real gap regarding cross-project dynamics, its metrics are defined transparently from log data, and its use of actual task-execution data from three PyBossa-based platforms is a strength. The SIM-based inspection is a credible qualitative complement, and the resulting design recommendations are actionable. However, the headline claim that inherited volunteers are more engaged than outside-recruited volunteers is currently supported only by a confounded comparison, which substantially lowers the significance of the paper's central contribution unless the analysis is redone or the claim is reframed.
major comments (3)
- [Section 4.2, Figure 3, Section 4.3] The 'balance in computing' comparison is confounded by selection. A volunteer is classified as 'inherited' only after having performed tasks in another project on the platform, whereas outside-recruited volunteers are all volunteers whose first task was in the project, including the 67%–93% platform transients documented in Table 2. Inherited volunteers are therefore, by construction, a self-selected subset of returning volunteers, and a positive mean difference in task counts is expected even if cross-project features add no value. The conclusion in Section 4.3 that 'inherited volunteers perform more tasks than the recruited ones' and the design recommendation to invest in cross-project engagement rest directly on this comparison. The paper should compare inherited volunteers with outside-recruited volunteers who have comparable platform tenure or activity (for example, via matching on number of active days or number of prior projects), or perform a within-volunteer comparison of task counts in the first versus subsequent projects, or explicitly reframe the result as a descriptive statement about group composition rather than as evidence for the value of cross-project engagement.
- [Section 4.1] The recruitment attribution assumption is untestable with the available data. The paper estimates that each volunteer was recruited by the project in which she/he performed the first task after registering, but volunteers may browse the platform, see recommendations, or perform tasks in a project that is not the one that actually brought them to the platform. Because the balance-in-recruitment and balance-in-computing metrics depend entirely on this binary classification, the potential direction and magnitude of the resulting misclassification bias should be discussed, and the conclusions should be explicitly conditioned on the assumption. At minimum, the paper should consistently describe the comparison as being between first-project contributors and later-project contributors rather than between 'outside-recruited' and 'inherited' volunteers.
- [Section 4.2, Table 2] The definition of the 'multi-project regular' class is internally inconsistent. The text states that this class consists of volunteers who 'executed tasks on least two different projects', which is the same condition used for the 'multi-project explorer' class, yet the two classes are said to be mutually exclusive and have different reported percentages. If 'regular' is intended to require a minimum number of active days per project, as suggested by the engagement-rate metric in Section 3.2, that condition should be stated explicitly and applied consistently. This distinction matters because Figure 2 and the claim that multi-project regulars exhibit longer relative activity duration rely on it.
minor comments (4)
- [Section 3.2] The balance-in-recruitment and balance-in-computing formulas use min(n,u) and min(t,m) in the denominator, which are undefined when either value is zero. Please specify how projects with no inherited or no outside-recruited volunteers are handled in the analysis.
- [Section 4.2] There are several typographical and formatting errors, including 'on least two different projects', 'ananalysis', 'Crowcrafting' (in the Figure 3 caption), and 'e.g. of for example' in Section 4.3. Please proofread the text and correct these issues.
- [Figure 3] The caption refers to the balance in recruitment as the 'right' plot and the balance in computing as the 'left' plot, while the surrounding text discusses them in the reverse order. Please verify that the panels and captions correspond to the intended metrics.
- [Table 3] The Gini coefficient values have inconsistent spacing (for example, '0 .95' in the Crowdcrafting row). Please standardize the formatting.
Circularity Check
No significant circularity: the paper is an empirical measurement study whose metrics are defined from task logs and evaluated on external platform data.
full rationale
The paper's load-bearing empirical claim—that inherited volunteers perform more tasks than outside-recruited ones (Section 4.2, Figure 3)—is not equivalent by construction to its input. The balance-in-computing metric is defined directly as (t-m)/min(t,m), where t and m are averages computed from task execution logs; no parameter is fitted, and the sign of the balance is not forced by the definition. The classification of a volunteer as 'recruited' by the project of first task and 'inherited' by later projects (Section 4.1) is an attribution assumption, and the comparison does contain a selection confound—inherited volunteers are, by definition, a subset of returning volunteers, whereas outside-recruited volunteers include the full inflow of mostly transient participants. This is a validity threat to the interpretive claim about cross-project features, but it is not circular: the paper does not define 'inherited volunteers are more engaged' into the metric, and the observed distribution (Figure 3) could in principle have been negative for many projects. The only overlap with the authors' prior work is the 'relative activity duration' metric, cited to Ponciano and Brasileiro [19]; that citation supplies a definition, not the paper's conclusions, and the multi-project comparison using it is computed from the present dataset. No self-citation is used to forbid alternatives or to import a uniqueness claim. The limitations section acknowledges task-complexity and generalizability issues but does not assert that the analysis proves a causal effect; the absence of a discussion of the selection confound is a correctness concern outside the circularity definition. Accordingly, the derivation chain is self-contained against external task-log data, and the circularity score is 0.
Assumptions & free parameters
free parameters (2)
- Regular-volunteer threshold =
at least 2 days of task execution
- Multi-project threshold =
at least 2 projects
assumptions (4)
- domain assumption Task execution logs from a platform's API contain all relevant volunteer activity for the studied period.
- domain assumption The first project in which a volunteer performs a task is the project that recruited her/him.
- domain assumption Volunteer join dates are available and correctly recorded in the platform data.
- domain assumption The platforms studied quantitatively (all PyBossa-based) are representative of multi-project citizen science platforms.
Cite this review
Pith. "Pith review of Characterising Volunteers' Task Execution Patterns Across Projects on Multi-Project Citizen Science Platforms." pith.science (2026). https://pith.science/paper/XJDCBD66
@misc{pith2026190801344,
author = {Pith},
title = {Pith review of: Characterising Volunteers' Task Execution Patterns Across Projects on Multi-Project Citizen Science Platforms},
year = {2026},
howpublished = {\url{https://pith.science/paper/XJDCBD66}},
note = {Machine review of arXiv:1908.01344}
}
read the original abstract
Citizen science projects engage people in activities that are part of a scientific research effort. On multi-project citizen science platforms, scientists can create projects consisting of tasks. Volunteers, in turn, participate in executing the project's tasks. Such type of platforms seeks to connect volunteers and scientists' projects, adding value to both. However, little is known about volunteer's cross-project engagement patterns and the benefits of such patterns for scientists and volunteers. This work proposes a Goal, Question, and Metric (GQM) approach to analyse volunteers' cross-project task execution patterns and employs the Semiotic Inspection Method (SIM) to analyse the communicability of the platform's cross-project features. In doing so, it investigates what are the features of platforms to foster volunteers' cross-project engagement, to what extent multi-project platforms facilitate the attraction of volunteers to perform tasks in new projects, and to what extent multi-project participation increases engagement on the platforms. Results from analyses on real platforms show that volunteers tend to explore multiple projects, but they perform tasks regularly in just a few of them; few projects attract much attention from volunteers; volunteers recruited from other projects on the platform tend to get more engaged than those recruited outside the platform. System inspection shows that platforms still lack personalised and explainable recommendations of projects and tasks. The findings are translated into useful claims about how to design and manage multi-project platforms.
Figures
Reference graph
Works this paper leans on
-
[1]
Ricardo Matsumura Araujo. 2013. 99designs: An analysis of creative competition in crowdsourced design. In First AAAI conference on Human computation and crowdsourcing. AAAI, Palo Alto, US, 17–24
work page 2013
-
[2]
Elizabeth H Boakes, Gianfranco Gliozzo, Valentine Seymour, Martin Harvey, Chloë Smith, David B Roy, and Muki Haklay. 2016. Patterns of contribution to citizen science biodiversity projects increase understanding of volunteersâĂŹ recording behaviour. Scientific reports 6 (2016), 33051
work page 2016
-
[3]
Victor R Basili1 Gianluigi Caldiera and H Dieter Rombach. 1994. The goal question metric approach. Encyclopedia of software engineering 2 (1994), 528–532
work page 1994
-
[4]
E Gil Clary, Mark Snyder, Robert D Ridge, John Copeland, Arthur A Stukas, Julie Haugen, and Peter Miene. 1998. Understanding and assessing the motivations of volunteers: a functional approach. Journal of personality and social psychology 74, 6 (1998), 1516
work page 1998
-
[5]
Joe Cox, Eun Young Oh, Brooke Simmons, Chris Lintott, Karen Masters, Anita Greenhill, Gary Graham, and Kate Holmes. 2015. Defining and measuring success in online citizen science: A case study of Zooniverse projects. Computing in Science & Engineering 17, 4 (2015), 28–41
work page 2015
-
[6]
Clarisse Sieckenius De Souza. 2005. The semiotic engineering of human-computer interaction. MIT press, Cambridge, Massachusetts, US
work page 2005
-
[7]
Clarisse Sieckenius De Souza, Carla Faria Leitão, Raquel Oliveira Prates, and Elton José da Silva. 2006. The semiotic inspection method. In Proceedings of VII Brazilian symposium on Human factors in computing systems . ACM, NY, US, 148–157
work page 2006
-
[8]
Djellel Difallah, Elena Filatova, and Panos Ipeirotis. 2018. Demographics and Dynamics of Mechanical Turk Workers. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining (WSDM ’18) . ACM, New York, NY, USA, 135–143
work page 2018
Show all 39 references
-
[9]
Melissa Eitzel, Jessica Cappadonna, Chris Santos-Lang, Ruth Duerr, Sarah Eliza- beth West, Arika Virapongse, Christopher Kyba, Anne Bowser, Caren Cooper, Andrea Sforzi, Anya Metcalfe, Edward Harris, Martin Thiel, Mordechai Haklay, Lesandro Ponciano, Joseph Roche, Luidi Ceccaro...
2017
-
[10]
Daniel Lombrana González et al. 2018. Scifabric/pybossa: v2.10.0. https://doi. org/10.5281/zenodo.1402409
2018 doi
-
[11]
Alexandra Eveleigh, Charlene Jennett, Ann Blandford, Philip Brohan, and Anna L Cox. 2014. Designing for dabblers and deterring drop-outs in citizen science. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . ACM, NY, US, 2985–2994
2014
-
[12]
Muki Haklay. 2013. Citizen science and volunteered geographic information: Overview and typology of participation. In Crowdsourcing geographic knowledge. Springer, Netherlands, 105–122
2013
-
[13]
Alan Irwin. 1995. Citizen Science: A Study of People, Expertise and Sustainable Development. Psychology Press, London, UK
1995
-
[14]
Edith Law and Luis von Ahn. 2011. Human computation. Synthesis lectures on artificial intelligence and machine learning 5, 3 (2011), 1–121
2011
-
[15]
Chris Lintott and Jason Reed. 2013. Human computation in citizen science. In Handbook of human computation . Springer, NY, US, 153–162
2013
-
[16]
Pietro Michelucci and Janis L Dickinson. 2016. The power of crowds. Science 351, 6268 (2016), 32–33
2016
-
[17]
Heather L O’Brien and Elaine G Toms. 2008. What is user engagement? A conceptual framework for defining user engagement with technology. Journal of the American society for Information Science and Technology 59, 6 (2008), 938–955
2008
-
[18]
Lesandro Ponciano and Nazareno Andrade. 2018. Perspectivas em Computação Social. In Computação Brasil, Raquel Prates and Thais Castro (Eds.). Vol. 36. Sociedade Brasileira de Computação, Porto Alegre, Brasil, 30–33
2018
-
[19]
Lesandro Ponciano and Francisco Brasileiro. 2014. Finding Volunteers’ Engage- ment Profiles in Human Computation for Citizen Science Projects. Human Computation 1, 2 (2014), 245–264
2014
-
[20]
Lesandro Ponciano and Francisco Brasileiro. 2018. Agreement-based credibil- ity assessment and task replication in human computation systems. Future Generation Computer Systems 87 (2018), 159–170
2018
-
[21]
Lesandro Ponciano, Francisco Brasileiro, Nazareno Andrade, and Lívia Sampaio
-
[22]
Lesandro Ponciano, Francisco Brasileiro, Robert Simpson, and Arfon Smith. 2014. Volunteers’ engagement in human computation for astronomy projects. Com- puting in Science & Engineering 16, 6 (2014), 52–59
2014
-
[23]
Jennifer Preece. 2016. Citizen science: New research challenges for human– computer interaction. International Journal of Human-Computer Interaction 32, 8 (2016), 585–612
2016
-
[24]
M Jordan Raddick, Georgia Bracey, Pamela L Gay, Chris J Lintott, Phil Murray, Kevin Schawinski, Alexander S Szalay, and Jan Vandenberg. 2010. Galaxy Zoo: Exploring the Motivations of Citizen Science Volunteers. Astronomy Education Characterising Volunteers’ Task Execution Patt...
2010
-
[25]
Christine Robson, Marti Hearst, Chris Kau, and Jeffrey Pierce. 2013. Comparing the use of social networking and traditional media channels for promoting citizen science. In Proceedings of the 2013 conference on Computer supported cooperative work. ACM, NY, US, 1463–1468
2013
-
[26]
Dana Rotman, Jenny Preece, Jen Hammock, Kezee Procita, Derek Hansen, Cynthia Parr, Darcy Lewis, and David Jacobs. 2012. Dynamic changes in motivation in collaborative citizen-science projects. In Proceedings of the ACM 2012 conference on computer supported cooperative work . A...
2012
-
[27]
Robert Simpson, Kevin R Page, and David De Roure. 2014. Zooniverse: observing the world’s largest citizen science platform. InProceedings of the 23rd international conference on world wide web . ACM, NY, US, 1049–1054
2014
-
[28]
Ianna Sodré and Francisco Brasileiro. 2017. An analysis of the use of qualifications on the Amazon mechanical Turk online labor market. Computer Supported Cooperative Work (CSCW) 26, 4-6 (2017), 837–872
2017
-
[29]
Jessica Suzuki and Edna Dias Canedo. 2018. Interaction Design Process Oriented by Metrics. In International Conference on Human-Computer Interaction . Springer, Cham, Switzerland, 290–297
2018
-
[30]
DM Rini van Solingen and Egon W Berghout. 1999. The Goal/Question/Metric Method: a practical guide for quality improvement of software development . McGraw-Hill, NY, US
1999
-
[31]
Rini van Solingen. 2014. Agile GQM: Why Goal/Question/Metric is more Rel- evant than Ever and Why It Helps Solving the Agility Challenges of Today’s Organizations. In 2014 Joint Conference of the International Workshop on Software Measurement and the International Conference o...
2014
-
[32]
Luis Von Ahn. 2008. Human computation. In Proceedings of the 2008 IEEE 24th In- ternational Conference on Data Engineering . IEEE Computer Society, Washington, DC, US, 1–2
2008
-
[33]
Luis Von Ahn, Benjamin Maurer, Colin McMillen, David Abraham, and Manuel Blum. 2008. recaptcha: Human-based character recognition via web security measures. Science 321, 5895 (2008), 1465–1468
2008
-
[34]
Sarah West and Rachel Pateman. 2016. Recruiting and Retaining Participants in Citizen Science: What Can Be Learned from the Volunteering Literature? Citizen Science: Theory and Practice 1, 2 (2016), 10
2016
-
[35]
Andrea Wiggins and Kevin Crowston. 2011. From conservation to crowdsourcing: A typology of citizen science. In 44th Hawaii International Conference on System Sciences. IEEE, Washington, DC, US, 1–10
2011
-
[36]
Andrea Wiggins and Kevin Crowston. 2012. Goals and tasks: Two typologies of citizen science projects. In 45th Hawaii International Conference on System Sciences. IEEE, Washington, DC, US, 3426–3435
2012
-
[37]
John Wilson. 2000. Volunteering.Annual review of sociology 26, 1 (2000), 215–240
2000
-
[38]
Poonam Yadav and John Darlington. 2016. Design guidelines for the user-centred collaborative citizen science platforms. Human Computation 3, 11 (2016), 205– 211
2016
-
[2014]
Journal of Internet Services and Applications 5, 1 (2014), 10
Considering human aspects on strategies for designing and managing distributed human computation. Journal of Internet Services and Applications 5, 1 (2014), 10
2014
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.