{"id":"54066b5b-846c-4b34-9363-2a05ebd56e91","arxiv_id":"1908.01875","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A framework that uses survey-trained shareability models to correct social media image bias produces Grevy's zebra population estimates that deviate substantially from official counts.","lead":"The paper proposes a machine-learning framework that estimates wildlife population sizes from social media photos by correcting for users' sharing bias, and tests it on Grevy's zebra images from Flickr. A generalist might read it to see whether crowd-sourced social media images can become a scalable conservation monitoring tool, though the reported population estimates are far from official census counts.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The regression model is trained on full SD-card collections, but Flickr inference has only the shared album, so the correction factor k_i is not computable as described.","rationale":"The abstract claims a learnable framework for correcting social media bias in wildlife population estimates. The core correction is the coefficient k_i = 1/f_i, where f_i is the fraction of photographed animals that a user shares. The paper trains a regression model to predict f_i from features of a complete SD-card collection (Section 4). At deployment on Flickr, only the shared collection is observed; the SD card is not. Thus the feature vector used in training cannot be constructed. The paper gives no alternative feature definition, so the pipeline that produced Table 1 is not fully specified and, on the natural reading, is impossible to execute. This is a stronger internal inconsistency than the survey-to-Flickr transfer problem: even if survey labels perfectly matched Flickr user preferences, the model could not be applied to a Flickr album. The reader's verdict of REJECT is therefore supported, and our concern does not change it. We recommend an unchanged verdict, while noting this additional fundamental flaw.","tokens_in":4055,"tokens_out":6710,"duration_ms":102918,"concrete_test":"Obtain the feature list from Section 4 (code is not provided). For each feature, determine whether it can be computed from only the shared Flickr album. If any feature requires knowing the total number of photos or animals in the photographer's SD card (e.g., collection size, number of unshared images, animal counts across the full event), then the model input at inference is undefined. A direct check: instantiate the described feature extractor on a single Flickr album and on a simulated full SD card containing that album plus extra unshared photos; if the resulting feature vectors differ while the album is identical, the model cannot be applied to the album alone. Re-running Table 1 with only album-computable features would show whether the published estimates depend on this unavailable input.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 trains a regression model on GGR SD cards to predict the percentage of animals a survey respondent would share; the inverse gives k_i. The features explicitly 'model the structure of the source collection, in order to account for all the pictures that were taken.' At inference, the model is applied to Flickr albums C_i, which are only the shared subset of the photographer's SD card; the unshared images (and hence the full source collection) are unobserved. Therefore the training-time feature vector is not computable from the available Flickr data. If the authors instead compute features from the shared album alone, then the input distribution differs from training, and the target—the fraction of animals shared out of all photographed animals—has an unobserved denominator, so it is not identifiable from the album. The reported Table 1 estimates therefore rest on an undocumented feature substitution or an invalid application of the model. This is load-bearing because k_i is the sole mechanism for correcting sharing bias; without it the population estimates have no basis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for estimating wildlife population sizes from Flickr image collections by correcting for the bias that users share only a subset of the animal photos they take. The authors collect survey labels from Great Grevy's Rally citizen-science SD-card images, train a regression model to predict the percentage of photographed animals a survey respondent would share, and use the inverse of this percentage as a correction coefficient k_i that multiplies the number of distinct animals seen in a Flickr album. They then apply Jolly-Seber capture-recapture to year-grouped Flickr collections and multiply the result by an ad hoc ``symmetry recapture multiplier'' motivated by the asymmetry of Grevy's zebra stripes. Table 1 reports population estimates for 2011-2017, with large deviations from official censuses (e.g., 67 versus 1,627 for 2017, overall RMSE 1,194).","tokens_in":4229,"tokens_out":4583,"duration_ms":51052,"significance":"If the proposed approach worked, it would be a valuable new direction for low-cost wildlife monitoring: the paper correctly identifies collection-level sharing bias as an important obstacle, and the shareability classification experiments do show that some signal is learnable (accuracy 0.82, F1 0.62 on the combined GGR dataset). The use of collection-structure features is a sensible and nontrivial idea. However, the central population estimates are not validated: the estimate-problem model achieves only R^2 = 0.417, the final numbers deviate strongly from official counts, no uncertainty intervals are provided, and the method rests on a training-inference mismatch in the shareability model. The paper does not provide code or data, and it depends heavily on unpublished prior work for feature definitions. The contribution is therefore currently a plausible research direction rather than a demonstrated method.","major_comments":[{"comment":"The regression model is trained on full GGR SD-card collections to predict \"the percentage of animals that the interviewee would share\", and the feature vector includes \"features modeling the structure of the source collection, in order to account for all the pictures that were taken\". At Flickr inference, only the shared album C_i is available; the full SD card SD_i is unobserved. The paper does not describe how the source-collection features are computed in that setting, and if they are computed from C_i alone, the input distribution differs from training while the target denominator (all photographed animals) is not identifiable from the album. Since k_i = 1 / predicted sharing percentage is the only mechanism correcting the social-media bias, this mismatch undermines the central estimates in Table 1.","section":"Section 4, estimate problem"},{"comment":"The paper does not specify the capture occasions, the definition of a \"capture\" in the Flickr data, or how individual animals are matched across years. Jolly-Seber requires at least two sampling occasions with known numbers of marked and unmarked animals, and no such capture histories are provided. Without this specification, the \"Jolly-Seber\" column in Table 1 cannot be reproduced or interpreted, and the per-year estimates are not connected to the multi-year method described in the text.","section":"Section 4 and Table 1, Jolly-Seber application"},{"comment":"The final population estimate is multiplied by a \"coefficient to take into account the recaptures which cannot be detected ... due to the lack of symmetry in Grevy's zebra stripes on either side\". This multiplier is introduced without estimation, calibration, or sensitivity analysis, and it is a free parameter whose value changes all final estimates. The paper should either estimate it from data or show that the conclusions are robust to a range of plausible values; as written, the final numbers depend on an ad hoc assumption.","section":"Section 4, last paragraph"},{"comment":"The model's estimates deviate from official censuses by a large margin: 67 versus 1,627 for 2017, with an overall RMSE of 1,194, and no uncertainty intervals or error propagation are reported. The estimate-problem model achieves only R^2 = 0.417, and no evidence is given that the model transfers from citizen-science SD cards to Flickr collections. These results do not support the abstract's claim that the social-media-bias correction is \"a learnable and potentially solvable problem\" at the level of population estimates.","section":"Section 6, Table 1"}],"minor_comments":[{"comment":"The meaning of the \"+\" in the official counts for 2012 and 2017 (1,897+ and 1,627+) is not explained; if these are minimum counts, the comparison should state this explicitly.","section":"Table 1"},{"comment":"The survey procedure is under-described: the number of respondents per SD card, the aggregation rule for multiple responses, and inter-annotator agreement are not reported, although the regression target is a personal sharing preference.","section":"Section 4, survey procedure"},{"comment":"The feature set F_i is not listed. A table of feature names would be necessary to understand the estimate-problem model and the ``source collection structure'' features, and to assess the training-inference mismatch.","section":"Sections 4-5"},{"comment":"Calling the framework \"species-independent\" overstates the evidence, because all experiments concern Grevy's zebra; the text should say that the framework is intended to be general and has been tested on one species.","section":"Conclusions"},{"comment":"The feature definitions and the earlier approach are load-bearing for this paper but appear only as an unpublished thesis and report; the authors should summarize the relevant definitions so that the current method is self-contained.","section":"References [7,8]"}],"recommendation":"reject","confidential_remarks":"The manuscript describes an interesting framing, but the central inference step has a training-inference mismatch that cannot be patched with the current data, and the validation against official counts is poor. I recommend rejection, but a substantially revised version that redefines the learning target using only observable shared collections, calibrates the symmetry multiplier, and reports uncertainty would deserve reconsideration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is genuinely worth a look: treat the gap between photographed and shared animals as a learnable per-user correction factor, and use collection-level features rather than single-image features. The paper also shows the shareability sub-problem is learnable to a modest degree (F1 0.62 on the combined data, 0.77 on GGR1 alone), which is a real, reproducible signal. It uses actual citizen-science GGR data and Wildbook identifications, and the authors are honest about limitations. Credit where due: this is a legitimate step beyond the group's prior single-image shareability work [7,8].\n\nThe soft spot is load-bearing, not cosmetic. The regression model in Section 4 is trained on GGR SD cards, where the full set of photos taken is known and the survey labels give the fraction of animals the respondent would share. The features explicitly model the structure of the source collection, accounting for all pictures taken. At inference, the model is applied to Flickr albums C_i, which are only the shared subset; the full SD card is unobserved. So either the training-time features are not computable, or the authors silently substituted features from the shared album, which changes the input distribution and makes the denominator (total animals photographed) unidentifiable. The stress-test note is correct, and the Table 1 results bear it out: the 2017 estimate of 67 vs. official 1627, and RMSE 1194. There is also no uncertainty propagation, the Jolly-Seber step is under-specified, and the symmetry recapture multiplier is introduced ad hoc without estimation. These are not minor quibbles; they break the central claim of producing usable population estimates.\n\nThe paper is not incoherent on every level. The shareability classification is a legitimate contribution, and the framing of the bias-correction problem is useful for the conservation-tech community. But the estimation pipeline, as written, does not work. The transfer from survey respondents to Flickr users is unvalidated, and the train/inference mismatch is a basic methodological error. That said, the problem matters and the authors have real data, so I would not desk-reject. A serious referee could push them to redesign the inference to use observable features (e.g., shared-album statistics alone) or to collect validation data where full SD cards are known for the same users. With major revision, this could become a useful study. As submitted, it is not citeable for its population estimates, but the shareability finding is worth a look in a reading group as a cautionary example.\n\nIf it crosses your desk, send it to review with the train/inference mismatch flagged as the primary issue. It deserves referee time, not because the estimates are convincing, but because the idea is promising and the flaws are fixable in principle.","headline":"Interesting framing of shareability as a learnable bias-correction problem, but the central pipeline trains on full SD-card collections and applies to shared albums only, so the population estimates are unsupported.","tokens_in":4745,"tokens_out":2504,"would_cite":false,"duration_ms":26205,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that social media sharing bias in wildlife images is learnable, and that correcting for it with machine learning can turn Flickr photo collections into wildlife population estimates.","keywords":["population size estimation","social media bias","shareability prediction","machine learning","wildlife conservation","capture-mark-recapture","Grevy's zebra","Flickr"],"falsifier":"Collect ground-truth sharing behavior for a set of Flickr users who post Grevy's zebra images — for example, by obtaining their full camera rolls or having them label their own upload decisions — and compare the actual photographed-to-shared ratio with the $k_i$ the model predicts from their uploaded albums. If the predicted correction factors are systematically off, the framework's population estimates would not survive the test.","tokens_in":3847,"feed_emoji":"🦓","tokens_out":8258,"duration_ms":76949,"temperature":0.7,"pith_summary":"The paper claims that the gap between how many animals a photographer photographs and how many appear in the photos they share online is not random noise but a learnable pattern. It proposes a two-stage framework: a classifier predicts which images in a collection would be shared, and a regression predicts, from the shared collection, how many animals were originally photographed. The inverse of the predicted share percentage acts as a correction factor, and the corrected counts are fed into a capture–recapture estimator to produce population numbers. The authors report that shareability is learnable (0.82 accuracy on combined data) and that the regression beats naive baselines (R2 of 0.417). If the framework holds, it would give conservation biologists a cheap and scalable way to turn the vast, biased resource of social media wildlife images into usable population estimates.","feed_headline":"Photo-sharing bias can be learned, improving wildlife counts","feed_subtitle":"A two-stage model corrects Flickr collections and beats uncorrected baselines for Grevy's zebra estimates.","key_machinery":"The load-bearing object is the collection-correction factor $k_i$, the inverse of the predicted fraction of photographed animals that a photographer would share. It is computed by a regression model that maps features of the shared image collection — individual-level animal counts, image quality and beauty features, and new features describing the structure of the source photo collection — to the share percentage. Multiplying the shared-animal count $n_i$ by $k_i$ recovers the estimated number of animals photographed. The same features feed a binary shareability classifier, and both models are trained on survey labels obtained by asking people whether they would share each citizen-science photo. The final population estimate comes from plugging the corrected per-photographer counts into the Jolly–Seber capture–recapture estimator.","core_discovery":"The central claim is that social media bias in wildlife photography is a solvable, learnable problem: given the images a user shared, a machine-learning model can predict how many animals the user actually photographed. The paper defines two linked problems — the estimate problem, predicting the total photographed animals $N_i$ from the shared collection's features, and the shareability problem, predicting which individual images get shared. Using survey respondents labeling citizen-science photos from two Great Grevy's Rally events, a gradient-boosted tree model achieves 0.82 accuracy on the shareability task and an $R^2$ of 0.417 on the estimate problem, while the mean and mode baselines score negative $R^2$. Applying the learned correction factors to Flickr albums and feeding the corrected counts into the Jolly–Seber estimator yields Grevy's zebra population estimates for 2011–2017, with an RMSE of 1194 against official counts compared to 1612 for the uncorrected estimator. The paper's point is that the remaining deviation is not irreducible: the structure of the original photo collection carries predictive signal about what gets shared.","pith_inferences":["The survey-based labels could be replaced by observing what photographers actually post versus what they capture on their cameras, which would give a direct, non-hypothetical training signal for the shareability model.","A direct transfer test would recruit Great Grevy's Rally participants who also post to Flickr or Instagram, so the model's predicted correction factor could be checked against their real photographed-to-shared ratio.","The sizeable year-to-year deviations in the corrected population estimates suggest the framework is best read as a demonstration that the bias is learnable rather than as a production census tool in its current form."],"forward_implications":["If the framework generalizes, conservation programs could estimate population sizes for data-poor species directly from existing social media images, without mounting new field surveys or citizen-science rallies.","The finding that collection structure helps predict shareability implies that sharing bias cannot be corrected image-by-image; any model of social media wildlife data must account for the whole photo set taken at an event.","Because the corrected per-photographer counts feed a standard capture–recapture estimator, the framework acts as a bias-correction layer that can be inserted into existing population-monitoring pipelines.","The reported accuracy of 0.82 on shareability and R2 of 0.417 on the estimate problem indicate the correction is learnable with modest data, and the authors expect additional species- and platform-specific features to improve the estimates."],"supporting_citations":[{"why":"supplies the capture-recapture estimator that turns per-photographer corrected counts into a population estimate","marker":"[5]"},{"why":"companion to [5]; together they form the Jolly–Seber method used for the final estimates","marker":"[11]"},{"why":"source of the GGR1 and GGR2 citizen-science image collections whose SD-card photos carry the survey labels","marker":"[1]"},{"why":"the gradient-boosting implementation used as the best classifier and regressor in both problems","marker":"[3]"},{"why":"prior work on predicting the likelihood an image is shared; the paper extends this from single images to collections","marker":"[7]"},{"why":"prior Flickr-based animal population estimation that this framework corrects for sharing bias","marker":"[8]"},{"why":"the computer-vision tool used to identify individual animals in photos, providing the counts feeding the regression","marker":"[2]"}],"fun_headline_variants":["Learnable social media bias improves wildlife counts","Correcting photo bias yields better zebra estimates","Learning photo-sharing bias aids wildlife population estimates","Social media bias is learnable for wildlife counts","Model learns photo bias to estimate animal populations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's estimates rest on the assumption that survey respondents' shareability labels on citizen-science SD-card photos predict how real Flickr users decide which wildlife photos to post online.","fun_headline_variants_meta":{"raw":{"variants":["Learnable social media bias improves wildlife counts","Correcting photo bias yields better zebra estimates","Learning photo-sharing bias aids wildlife population estimates","Social media bias is learnable for wildlife counts","Model learns photo bias to estimate animal populations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000396,"raw_usage":{"total_tokens":2034,"prompt_tokens":863,"completion_tokens":1171,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":1103}},"tokens_in":479,"tokens_out":1171,"duration_ms":9423,"temperature":1.0,"reasoning_tokens":1103,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:00:35.429206+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect ground-truth sharing behavior for a set of Flickr users who post Grevy's zebra images — for example, by obtaining their full camera rolls or having them label their own upload decisions — and compare the actual photographed-to-shared ratio with the $k_i$ the model predicts from their uploaded albums. If the predicted correction factors are systematically off, the framework's population estimates would not survive the test.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the capture-recapture estimator that turns per-photographer corrected counts into a population estimate"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"companion to [5]; together they form the Jolly–Seber method used for the final estimates"},{"cited_title":"Berger-Wolf, J","cited_arxiv_id":null,"evidence_quote":"source of the GGR1 and GGR2 citizen-science image collections whose SD-card photos carry the survey labels"},{"cited_title":"Chen and C","cited_arxiv_id":null,"evidence_quote":"the gradient-boosting implementation used as the best classifier and regressor in both problems"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"prior work on predicting the likelihood an image is shared; the paper extends this from single images to collections"},{"cited_title":"Menon, T","cited_arxiv_id":null,"evidence_quote":"prior Flickr-based animal population estimation that this framework corrects for sharing bias"}],"review_version":1}