{"id":"ee7b3be9-bb05-4b4f-8862-8f1e09b5eedf","arxiv_id":"1908.04237","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A smartphone crowdsourcing system collected over 7,000 geo-tagged cassava disease reports from 29 Ugandan volunteers over 76 weeks, with participation shaped by incentive type, but the study was not a controlled experiment.","lead":"The paper reports a 76-week pilot in Uganda where farmers, extension workers, and experts used smartphones to send geo-tagged cassava disease and pest reports, collecting over 7,000 submissions. It describes how different incentives affected participation among four participant groups, though the study was not a controlled experiment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The feasibility claim rests on unvalidated participant labels: the paper admits image analyses are not evaluated, and incentive gaming is documented, so the 7,000 reports cannot yet be treated as surveillance-grade disease data.","rationale":"The reader identified the same load-bearing assumption: participant-supplied labels and images were not evaluated for accuracy, and the surveillance value of the system depends on that accuracy. The paper is honest about this limitation and frames its contribution around the crowdsourcing aspect, which supports the weaker feasibility claim that a deployed mobile system can collect thousands of geo-tagged reports over 76 weeks. The presence of documented incentive gaming in the challenges section makes the accuracy concern concrete rather than hypothetical, but it does not overturn the paper's primary contribution as a deployment and lessons-learned report. A conditional verdict remains appropriate: the system and dataset are plausibly valuable, but the disease-surveillance interpretation should be stated with an explicit validity caveat pending expert validation of the labels. Thus no change to the reader's verdict is needed.","tokens_in":11353,"tokens_out":1944,"duration_ms":23476,"concrete_test":"Select a random sample of reports stratified by contributor category and incentive period, then have two independent NaCRRI pathologists re-label the uploaded images from the images alone (blinded to participant labels) and compare their labels with the participant-supplied labels and comments. Also independently verify a subset of GPS coordinates against known field locations. Report per-category and per-incentive-period label accuracy and inter-rater agreement; if farmer or extension labels agree with expert labels in less than, say, 80% of cases, the central claim should be narrowed from validated surveillance to data-collection feasibility.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that AdSurv enabled farmers, extension workers, and experts to provide near real-time, geo-tagged surveillance data on cassava viral diseases and pests, with over 7,000 reports in 76 weeks. For this to support disease surveillance, the participant-supplied image labels and comments must be accurate enough to reflect true disease and pest incidence. The paper explicitly does not establish this: in Related Work it states, \"Because we presently do not evaluate how accurate these particular analyses are, the work we present here mainly focuses on the crowdsourcing aspect.\" The Limitations section also concedes the pilot was not a controlled experiment. More concretely, the Discussion on challenges documents incentive-driven gaming: after the micropayment was increased to 500 UGX per report, \"an agent would report many reports from within a very small restricted locality,\" with counts jumping from 12 to over 70 per week, and the operators had to cap rewardable reports per village at 35. This shows that the label and location data were not uniformly trustworthy, especially under monetary incentives. Without a validation of participant labels against expert ground truth, the 7,000 reports are strong evidence of participation and data-collection feasibility, but not yet evidence that the resulting national maps are accurate enough for disease surveillance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AdSurv, a mobile crowdsourcing system for cassava disease and pest surveillance, deployed in Uganda over 76 weeks with 29 volunteer participants (farmers, extension workers, crop experts, and partner agents). The authors report collecting more than 7,000 geo-tagged image reports, describe the participation patterns of the different user categories, and discuss the effects of six incentive types (equipment provision, data credit, prompts, subject surveillance, feedback, and micropayments). The paper argues that this ad hoc crowdsourcing approach can supplement annual expert surveys by providing more frequent, spatially distributed, near-real-time surveillance data. The authors acknowledge that the pilot was not designed as a controlled experiment and that the accuracy of the participant-supplied image labels was not evaluated.","tokens_in":11578,"tokens_out":3884,"duration_ms":41557,"significance":"If the feasibility claim holds, the paper provides a valuable case study of mobile crowdsourcing for agricultural disease surveillance in a low-resource setting. Strengths of the work include the long deployment period (76 weeks), the real-world participant pool, and the honest reporting of operational challenges (GPS failures, data credit diversion, and incentive gaming). The paper is also explicit about its own limitations, which aids interpretation. However, the scientific significance is moderated by two factors: the lack of any validation of the participants' image labels and diagnoses against expert ground truth, and the absence of experimental control in the incentive analysis. As a feasibility and deployment report, the paper is informative; as evidence for incentive effects or as a source of surveillance-grade disease data, it remains suggestive rather than conclusive.","major_comments":[{"comment":"The Related Work section states: 'Because we presently do not evaluate how accurate these particular analyses are, the work we present here mainly focuses on the crowdsourcing aspect.' This admission directly undercuts the central claim in the abstract and title that the system provides 'real-time surveillance data on viral disease and pest incidence and severity.' Without validation of participant-supplied labels and images against expert ground truth, the 7,000 reports are strong evidence of participation and data-collection feasibility, but not yet evidence that the resulting maps are surveillance-grade. The authors should either provide a small validation study (e.g., expert review of a random subsample) or revise the title, abstract, and conclusion to explicitly limit claims to 'submitted reports' and discuss data-quality validation as future work that is required before use in disease surveillance.","section":"Related Work"},{"comment":"The causal statements about incentives are not supported by the study design. For example, the Effect of incentives section claims 'The direct monetary rewards of pay-per-report most incentivised the crop experts and the partner agents' and 'Farmers seem to be most motivated by the non-monetary incentives,' yet the pilot was not a controlled experiment (Limitations section), incentives were applied in a 'semi-consecutive, semi-mixed fashion' and 'combined randomly from time to time,' and the participant pool was only 29 people. Multiple incentives changed simultaneously, and there was no control group. The authors themselves note in The pilot section that 'our assertions in this paper are mainly on the conjecture side of the line.' I recommend reframing all incentive-related conclusions as observational hypotheses with clearly identified confounds, rather than as demonstrated findings. In particular, the observed reporting rise from week 55 to week 60 coincided with the micropayment increase, renewal of SMS/call prompts, and the broadcast of two subject surveillance tasks; the isolated effect of the monetary incentive cannot be identified from these data.","section":"Incentives structure / Effect of incentives"},{"comment":"Item 7 of the challenges documents that after the micropayment was increased to 500 UGX per report, 'an agent would report many reports from within a very small restricted locality,' with counts jumping from 12 to over 70 per week, and that a cap of 35 rewardable reports per village had to be imposed. This is not a peripheral operational detail; it demonstrates that the monetary incentive motivated submissions that are unlikely to reflect genuine disease observations, undermining the surveillance value of the data. The paper should discuss the implications of this gaming behavior for data quality, describe any filtering or validation mechanisms applied (before or after the cap was introduced), and outline how future deployments will prevent or detect such behavior. Without such a discussion, the feasibility claim for disease surveillance remains incomplete.","section":"Discussion on challenges"}],"minor_comments":[{"comment":"The phrase 'more than 7,000 reports was collected' should be 'more than 7,000 reports were collected' for subject-verb agreement.","section":"Abstract and Results"},{"comment":"Figures 2 and 3 are referenced in the text but not included in the manuscript text; ensure that all figures are present with captions and readable axis labels, and that the timeseries plots have explicit week numbers.","section":"Figures"},{"comment":"The typo 'voluteered geographical information' should be corrected to 'volunteered geographical information.'","section":"Related Work"},{"comment":"The system name is inconsistently written as both 'AdSurv' and 'Adsurv'; please unify the capitalization.","section":"Throughout"},{"comment":"In the bullet list, 'For the reminder of the pilot' should read 'For the remainder of the pilot.'","section":"Reporting trends by category"},{"comment":"The statement that expert agents 'generally posted higher quality reports especially on the subject surveillance matters probably because of the specialised knowledge they possess' uses 'quality' without any defined quality metric and includes the speculative word 'probably.' This claim should either be removed or supported by a concrete measure (e.g., agreement with expert re-labeling).","section":"Results / Reporting by agent category"}],"recommendation":"major_revision","confidential_remarks":"The skeptic's stress-test concern is valid and lands squarely on the paper's central surveillance claim. However, the paper is not fatally flawed: the authors are transparent about the lack of validation and the non-experimental incentive design. A major revision that (a) tempers the abstract/title claims, (b) reframes incentive observations as hypotheses, and (c) adds a short validation or a more explicit data-quality discussion would bring the paper within scope for publication. I also note that the paper appears to be an extended version of a prior workshop presentation; the authors should ensure the novelty is clearly articulated relative to that prior work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: this is a deployment paper, not a controlled study, and it mostly knows that. The contribution is a 76-week, 29-participant crowdsourced cassava disease surveillance pilot in Uganda that collected over 7,000 geo-tagged reports. That alone is a useful data point: it shows a farmer/extension/expert crowd can sustain reporting far longer than the annual national survey.\n\nThe paper is honest where it counts. It explicitly says the image analyses are not evaluated, so the work focuses on the crowdsourcing aspect. It lists limitations—no controlled experiment, incentives applied ad hoc. The incentive observations (farmers respond to calls, experts to micropayments, etc.) are plausible and grounded in the time series, but they are confounded: multiple incentives changed at once, there is no control group, and category differences (travel patterns, prior relationships) are not disentangled. The paper's own wording about being on the conjecture side of the line is fair.\n\nThe soft spots are real but mostly acknowledged. The largest one is that the 7,000 reports cannot yet be treated as surveillance-grade disease data. The paper admits label accuracy is unverified; and the discussion documents gaming (reports from a single small locality after the micropayment rise, capped at 35 per village). So the national situation map is a participation map until validation happens. That is a serious gap for the stated goal, but it is stated.\n\nThe citation pattern is fine. The system builds on ODK and known crowdsourcing concepts; the paper credits related work. No data or code is released, which limits reproduction. For a pilot case study, most venues would want at least the aggregate dataset or an expert-validation subsample.\n\nWho benefits: practitioners working on mobile data collection in low-resource agriculture, and researchers studying incentives in citizen science. It is not a methods paper. I would cite it as a real-world pilot, and it deserves a serious referee—conditional on the authors being asked to either share data or reframe the claims away from surveillance-grade mapping and toward feasibility of participation.\n\nRecommendation: send it to review; it is a solid empirical case study with honest limitations.","headline":"Honest deployment paper with a real 76-week dataset; the incentive claims are suggestive, not causal, and the surveillance-grade framing outruns the unvalidated labels.","tokens_in":12070,"tokens_out":1858,"would_cite":true,"duration_ms":21455,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a mobile phone-based crowdsourcing system with 29 trained volunteers sustained 76 weeks of near real-time, geo-tagged surveillance of cassava viral diseases and pests in Uganda, yielding more than 7,000 reports.","keywords":["crowdsourcing","cassava disease surveillance","mobile ad hoc surveillance","citizen science","incentive design","Uganda agriculture","viral plant disease","volunteered geographic information"],"falsifier":"Take a random sample of about 200 geo-tagged reports from the collected 7,000 and have cassava experts independently re-classify each submitted image; if expert and participant labels agree on fewer than roughly 70% of cases, the national situation map would be misleading rather than informative. A coarser check is to compare the crowdsourced disease heat map with the next annual national expert survey: district-level conflicts would point to the crowd's label accuracy as the weak link.","tokens_in":11150,"feed_emoji":"📱","tokens_out":6351,"duration_ms":55756,"temperature":0.7,"pith_summary":"The paper argues that mobile-phone crowdsourcing can supplement, and partially replace, expensive annual expert crop surveys in developing countries. It presents a 76-week pilot in Uganda in which 29 farmers, extension workers, agricultural experts, and partner agents used a smartphone application to send geo-tagged images and text reports on cassava viral diseases and pests. The pilot collected more than 7,000 reports and fed a live national situation map. The authors report that different participant groups responded to different incentives: farmers to communication and reputation rewards, experts and partner agents to direct per-report payments. If the claim holds, all-year-round real-time surveillance data can be obtained at low cost where it was previously infeasible.","feed_headline":"7,000 phone reports mapped cassava disease in Uganda","feed_subtitle":"Farmers, extension workers and experts geo-tagged real-time crop disease sightings, a low-cost supplement to Uganda's annual expert surveys.","key_machinery":"The load-bearing mechanism is the AdSurv crowdsourcing chain: a smartphone application that captures an image, a label, a comment, and GPS coordinates for each report; a back-end server that runs summary statistics; and a web dashboard that maps submissions in real time and can render disease-density heat maps. Weekly surveillance tasks broadcast a specific target, such as Cassava Mosaic Disease symptoms, and the responses update a national situation map. This mechanism carries the argument because it converts distributed human visual judgment into structured, spatially referenced surveillance data without requiring experts to travel.","core_discovery":"The central claim is that an ad hoc, phone-based crowdsourcing network can provide near real-time, geo-tagged surveillance of cassava viral disease and pests on a national scale. Over 76 weeks, a crowd of 29 trained volunteers submitted more than 7,000 reports, each containing an image, a text label, a comment, and a GPS reading, and these reports populated a live map and dashboard of crop health across Uganda. Participation reached about a third of what a fully compliant, trouble-free group would have produced, with dropouts caused by stolen or broken devices and by GPS resolution failures on some handsets. On incentives, the authors find that pay-per-report micropayments most raised participation among experts and partner agents, while farmers responded more to follow-up calls, SMS prompts, and reputation broadcasts, and each incentive bundle sustained elevated reporting for roughly ten weeks.","pith_inferences":["Beyond the paper, once label accuracy is measured, reports could be weighted by participant reliability, converting the raw crowd signal into a calibrated surveillance statistic.","The observed drop in participation when airtime was replaced by post-submission micropayments suggests a design rule: keep the marginal cost of reporting at zero for rural participants.","The paper's 35-report-per-village cap is a primitive spatial duplicate control; an automated duplicate-detection step would be the scalable version.","A staggered within-crowd comparison of incentive bundles would turn the reported ten-week decay curves into testable predictions about what keeps each participant type engaged."],"forward_implications":["If the central claim is correct, national crop surveillance can run all year round at roughly the cost of mobile data rather than the cost of annual expert travel.","Incentive design can be tuned by participant category, with direct micropayments for experts and partner agents and communication and reputation rewards for farmers.","Rotating incentive bundles can sustain reporting for about ten weeks per bundle, so a schedule of alternating incentives keeps the crowd active.","The accumulated image-and-label dataset is a resource for training automated disease classification tools.","The architecture transfers to other crops, pests, and countries with similar smallholder farming systems and mobile phone penetration."],"supporting_citations":[{"why":"Supplies the general smartphone-crowdsourcing approach that the AdSurv system builds on.","marker":"[Chatzimilioudis et al. 2012]"},{"why":"Grounds citizen science as a supplement to expert monitoring, the core premise of the paper.","marker":"[Silvertown 2009]"},{"why":"Documents the cassava mosaic disease pandemic, establishing why real-time surveillance is needed.","marker":"[Otim-Nape et al. 2000]"},{"why":"Provides the Knowledge Collection from Volunteer Contributors methodology that motivates the compensation structure.","marker":"[Chklovski and Gil 2005]"},{"why":"Precedent for minimally trained volunteer crowds performing image-classification tasks.","marker":"[Kanefsky, Barlow, and Gulick 2001]"},{"why":"Prior Ugandan mobile agricultural system with incentive design, offering a direct comparison for the incentive findings.","marker":"[Ssekibuule, Quinn, and Leyton-Brown 2013]"},{"why":"Frames volunteered geographic information, the spatial data type central to the surveillance map.","marker":"[Haklay 2013]"}],"fun_headline_variants":["7,000 phone reports map cassava disease in Uganda","Ugandan farmers' phone reports track cassava disease in real time","Crowdsourced mobile surveillance spots cassava pests and viruses","Ad hoc phone network monitors cassava disease nationwide in Uganda","Real-time cassava disease tracking via farmer phone reports"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The surveillance value of the system depends on the untested assumption that the disease labels and judgments supplied by participants are accurate enough to trust on a national disease map, and the paper explicitly does not evaluate this accuracy.","fun_headline_variants_meta":{"raw":{"variants":["7,000 phone reports map cassava disease in Uganda","Ugandan farmers' phone reports track cassava disease in real time","Crowdsourced mobile surveillance spots cassava pests and viruses","Ad hoc phone network monitors cassava disease nationwide in Uganda","Real-time cassava disease tracking via farmer phone reports"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1240,"prompt_tokens":892,"completion_tokens":348,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":265}},"tokens_in":508,"tokens_out":348,"duration_ms":4110,"temperature":1.0,"reasoning_tokens":265,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:15:02.982158+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of about 200 geo-tagged reports from the collected 7,000 and have cassava experts independently re-classify each submitted image; if expert and participant labels agree on fewer than roughly 70% of cases, the national situation map would be misleading rather than informative. A coarser check is to compare the crowdsourced disease heat map with the next annual national expert survey: district-level conflicts would point to the crowd's label accuracy as the weak link.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the general smartphone-crowdsourcing approach that the AdSurv system builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds citizen science as a supplement to expert monitoring, the core premise of the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the cassava mosaic disease pandemic, establishing why real-time surveillance is needed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Knowledge Collection from Volunteer Contributors methodology that motivates the compensation structure."},{"cited_title":"G.; and Gulick, V","cited_arxiv_id":null,"evidence_quote":"Precedent for minimally trained volunteer crowds performing image-classification tasks."},{"cited_title":"A.; and Leyton-Brown, K","cited_arxiv_id":null,"evidence_quote":"Prior Ugandan mobile agricultural system with incentive design, offering a direct comparison for the incentive findings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames volunteered geographic information, the spatial data type central to the surveillance map."}],"review_version":1}