{"id":"be533ee3-7845-4699-b512-3e8123ac1362","arxiv_id":"1908.11018","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a randomized sample of 637 US medical GoFundMe campaigns, non-white recipients were underrepresented and received fewer and smaller donations, while women performed most campaign organizing.","lead":"Researchers examined 637 medical crowdfunding campaigns on GoFundMe and found that Black and other non-white recipients are underrepresented and tend to raise smaller donations. The study offers evidence that online medical fundraising may reinforce, rather than reduce, existing health and social inequities.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The sampling frame is built from GoFundMe's top-500-per-zip search results, so the headline underrepresentation and outcome gaps could be artifacts of which campaigns the endpoint surfaced rather than true platform disparities.","rationale":"Reading in good faith: the paper asks whether medical crowdfunding exhibits disparities by race, gender, and age, and reports a new hand-coded dataset. The descriptive finding that women do most organizing labor is likely robust to sampling-frame concerns because it concerns campaigner gender within the sampled campaigns, and the gender imbalance is huge (82% female among those fundraising for others). The more fragile parts are the underrepresentation claims and the race-outcome regressions. The selection mechanism is the load-bearing assumption: the randomized sample is only random conditional on the proprietary list of 165,925 campaigns. Randomizing from a biased frame does not remove the bias. The paper itself acknowledges search algorithms prioritize popular/recent content, so the frame can systematically exclude lower-visibility campaigns, which are plausibly over-represented among marginalized groups. The reader's weakest assumption identifies the same point; my concern sharpens it to the 500-per-zip truncation and adds that no diagnostic is provided. I also note secondary concerns—no multiple-comparison correction (p=.030/.022 would not survive a modest correction), no control for campaign length in the Poisson model, and perceived-race coding—but these are less central than the frame. The honest verdict remains CONDITIONAL: the claim is plausible and socially important, but it should be conditioned on demonstrating that the sampling frame is not truncated in a demographically correlated way. I would not reject because the gender-labor result and the direction of effects are consistent with prior work on other platforms (Airbnb/Kickstarter). I would not accept because the key population-level comparisons cannot be verified from the paper as-is.","tokens_in":18799,"tokens_out":6928,"duration_ms":71703,"concrete_test":"Obtain or reconstruct the list of medical campaigns for a set of dense urban zips (e.g., Manhattan 10001, Chicago 60601, Los Angeles 90001) using an independent discovery path, such as date-bounded pagination through GoFundMe's category browse or a saved snapshot from the Internet Archive. For each zip, determine whether the top-500-per-zip endpoint truncates—i.e., whether more than 500 unique medical campaigns exist. If truncation occurs, compare the truncated-out campaigns with included ones on geography, campaign start date, and (if codable) recipient race. If omitted campaigns are disproportionately non-white, urban, or have systematically different donation outcomes, the sampling-frame bias is confirmed and Tables 3, 5, and 6 should be re-estimated on a frame that includes them.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the central claim, the 637-campaign sample must be representative of all US GoFundMe medical campaigns. The frame was constructed by asking GoFundMe's search endpoint for the 500 campaigns 'closest' to each US zip code and deduplicating; this yields a list of 165,925 campaigns. This only approximates the full universe if (a) 'closest' is purely geographic and (b) no zip has more than 500 medical campaigns. The authors themselves note the platform's search prioritizes popular, recent, and geographically proximate campaigns (§3). In dense urban areas, the 500-cap necessarily truncates, and the omitted campaigns are exactly those the platform ranks lower. Because Black and other non-white users are more concentrated in urban areas, the frame can understate their presence even if their true share of campaigns is equal or higher. The same selection process can contaminate Tables 5 and 6 if omitted urban campaigns differ in donation counts or average gift size. No diagnostic is reported: we do not know how many zips hit the cap, whether the frame's urban/rural mix matches expectations, or how the 165,925 list compares with any independent campaign enumeration. Without data/code these checks cannot be done post hoc, so the central descriptive and regression claims hinge on an unvalidated assumption about a proprietary API.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a cross-sectional study of 637 GoFundMe medical campaigns in the United States, drawn by random selection from a frame of 165,925 campaigns constructed in July 2016 by querying GoFundMe's search endpoint for the 500 campaigns 'closest' to each U.S. zip code and deduplicating. The authors hand-code the perceived race, gender, and age of recipients (and gender of campaigners), compare sample demographics to U.S. population estimates from the American Community Survey, and use linear and Poisson regressions to test associations between recipient demographics and two outcomes: average donation amount and number of donations. The main findings are that non-white recipients, especially Black women, are underrepresented relative to the U.S. population; that women perform most campaign-organizing labor; that Black recipients receive about $22 less per donation; that non-white recipients receive fewer donations; and that child recipients receive more donations but of lower average size. The paper presents these results as evidence that medical crowdfunding reproduces and amplifies existing social and health inequities.","tokens_in":19049,"tokens_out":4600,"duration_ms":47322,"significance":"If the results are robust, the paper provides one of the first systematic, large-sample descriptions of demographic inequities in U.S. medical crowdfunding, an important and understudied topic. Strengths include a large sampling frame with random selection within that frame, multi-rater coding with a reported ICC of .819, the use of two complementary outcome measures, and an explicit external benchmark against ACS population data. The authors also candidly acknowledge several limitations. However, the central descriptive and regression claims depend on an unvalidated assumption that the zip-code-based search frame represents all GoFundMe medical campaigns; attrition and the handling of unknown race/gender cases add further risk. These issues are fixable with additional diagnostics or substantially weakened claims, so the manuscript warrants a major revision rather than rejection.","major_comments":[{"comment":"The sampling frame is built from the GoFundMe search endpoint, which returns the 500 campaigns 'closest' to each U.S. zip code; after deduplication this yields 165,925 campaigns. This frame is only a complete enumeration if the search ranking is purely geographic and no zip code contains more than 500 medical campaigns, but the authors themselves note that the endpoint prioritizes popular, recent, and geographically proximate campaigns. Because the 500-campaign cap is more likely to bind in dense urban areas, campaigns in those areas—where non-white and lower-income populations are concentrated—are differentially likely to be excluded. The central underrepresentation claim in Table 3 and the outcome regressions in Tables 5 and 6 assume the sample represents all GoFundMe medical campaigns; without diagnostics (e.g., how many zips hit the cap, comparison of the frame's urban/rural distribution with an independent enumeration, or sensitivity analyses on truncated zips), this assumption is unvalidated. Since the frame was constructed with a proprietary API in 2016, such checks require data or code that the manuscript does not provide.","section":"§3 (Methods), sampling frame"},{"comment":"Of 822 sampled campaigns, 47 were removed from GoFundMe by July 2018 and 3 campaigns that had run fewer than 30 days were excluded. The authors state that removed campaigns were likely shut down by campaigners, but they do not report any comparison of these campaigns' observed characteristics with the retained sample. If campaigns that are removed are more likely to be unsuccessful, to belong to marginalized groups, or to have short durations, the estimates of both representation and outcomes will be biased. The manuscript should either provide a sensitivity analysis treating removed campaigns under best-/worst-case outcome assumptions or explicitly bound the potential impact of attrition on the reported coefficients.","section":"§3 (Methods), attrition"},{"comment":"Race is coded by raters' perception from campaign pages, and cases with unknown race or gender are dropped from the Table 3 comparisons. This is a reasonable design given that donor-perceived race is the relevant quantity, but unknown status may not be missing at random: campaigns with sparse information may differ systematically in outcomes. The paper should report the characteristics of unknown cases and run a sensitivity analysis (e.g., multiple imputation or extreme-case bounds) to show that the $22 average-donation gap for Black recipients and the lower donation counts for non-white recipients in Tables 5 and 6 are not driven by the exclusion of unknown cases. Additionally, comparing perceived race to ACS self-reported race should be stated as an explicit limitation.","section":"Tables 3, 5, and 6; §3 (race coding)"}],"minor_comments":[{"comment":"The phrase 'randomized sample' is imprecise: the sample is randomly drawn from a constructed search-based frame, not from all GoFundMe medical campaigns. Recommend wording such as 'sample randomly drawn from a constructed frame' to avoid implying a true probability sample of campaigns.","section":"Abstract and §3"},{"comment":"The two regression tables report different covariate sets: Table 5 includes 'Unknown Relationship' and 'Log of Number of Residents in State,' while Table 6 omits both and instead includes an 'Unknown' gender category absent from Table 5. Please explain the model specifications or justify the differences.","section":"Tables 5 and 6"},{"comment":"There is a typo in the last paragraph: 'caompanies' should be 'companies'.","section":"Conclusion"},{"comment":"The average donation variable has a minimum of 0, which implies zero-donation campaigns were assigned an average of 0. Please clarify how the average is computed for campaigns with no donations.","section":"Table 1"},{"comment":"The timeline is unclear: the search-based frame was created in July 2016, but campaign outcomes were re-collected in July 2018. State the exact collection and analysis dates in the methods to help readers interpret campaign length and attrition.","section":"§3 (Data collection timeline)"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The manuscript addresses an important and timely topic and is likely to interest the journal's readership. The central problem is that the sampling frame is built from a proprietary search endpoint with no validation against an independent enumeration of campaigns; without such validation or a clearly bounded sensitivity analysis, the headline underrepresentation and outcome-gap estimates could be artifacts of the frame. I recommend requesting a major revision that either provides the diagnostics or considerably softens the causal-sounding language. It would also strengthen the paper to deposit de-identified data and the frame-construction code, if possible, to allow the checks to be performed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi — quick take: this is the first randomized-sample look at racial and gender inequities in US medical crowdfunding, and it’s worth reading even though the sampling frame means the headline numbers should be treated as provisional.\n\nWhat’s new: the authors built a sampling frame of 165k GoFundMe campaigns via zip-code searches, drew a random sample of 637, coded perceived race/gender with good inter-rater agreement, and tested associations with outcomes. The finding that women do the vast majority of organizing labor, especially for others, is a real and underappreciated result. The regressions showing Black recipients get roughly $22 less per donation and non-white recipients get fewer donations are consistent with prior studies on Airbnb and Kickstarter, so the direction is plausible. They also make a sensible methodological choice in rejecting % of goal and total raised as outcome variables, arguing those are confused by goal-setting. Citations look appropriate; they engage the earlier Canadian and transgender-specific work without leaning on self-citation.\n\nSoft spots: the sample is only as good as the platform's search endpoint. Querying the top 500 campaigns closest to each zip code and deduplicating is not a probability sample. The stress-test concern is real: in dense urban areas the 500 cap will truncate, and the omitted campaigns are exactly the less popular ones, which could suppress counts of non-white recipients precisely where they are overrepresented. The authors acknowledge the search algorithm prioritizes popular/recent/proximate campaigns, but they report no diagnostics—how many zips hit the cap, urban/rural mix, or any comparison against an independent list. The perceived race coding is imperfect, though ICC .819 is acceptable. The regressions lack controls for illness severity, medical condition, and SES, so the $22 gap could partly reflect confounding. Multiple comparisons are not corrected; some borderline results (p=.026, p=.030) could be chance. No data or code is provided, so replication is impossible without going back to the API.\n\nThe limitations section is honest and shows they know the data-access constraints. But the central claim—that crowdfunding reproduces and potentially amplifies existing health disparities—is plausible and consistent with a wider literature, so I would not dismiss it.\n\nThis is a paper for health policy, platform governance, and critical tech studies readers. It deserves a serious referee, but the sampling frame needs to be interrogated, and the authors should be pushed to provide robustness checks and share data. I'd bring it to our reading group to argue about the frame.","headline":"First randomized-sample study of US medical crowdfunding inequities—valuable, but the unvalidated sampling frame makes magnitudes provisional.","tokens_in":19536,"tokens_out":4275,"would_cite":true,"duration_ms":39879,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Black recipients get about $22 less per GoFundMe donation","keywords":["medical crowdfunding","GoFundMe","health disparities","race","gender","fundraising outcomes","digital care labor","cross-sectional study"],"falsifier":"A direct falsifier would be platform-wide data from GoFundMe for the same period: if a complete enumeration of U.S. medical campaigns showed no white over-representation and no significant race or gender differences in average donation or donation count after adjusting for goal, length, and engagement, the paper's central claim would fail. An experimental complement would post identical campaigns with randomly assigned recipient names and photos; if Black-coded campaigns draw the same average donations as white-coded ones, the crowd-bias mechanism is not operating.","tokens_in":18623,"feed_emoji":"🩺","tokens_out":12190,"duration_ms":113593,"temperature":0.7,"pith_summary":"The paper sets out to test whether medical crowdfunding—online appeals for donations to cover health costs—reproduces or amplifies social inequities. Using a randomized sample of 637 U.S. GoFundMe campaigns whose recipients' race, gender, and age were hand-coded, it finds that white recipients are over-represented relative to the U.S. population, while Black recipients receive about $22 less per donation and non-white recipients receive fewer donations overall. It also finds a sharp gender split in labor: women organize roughly two-thirds of campaigns for themselves and nearly four-fifths of campaigns on behalf of someone else. Campaign actions that platforms tell users to focus on—photos, videos, updates—show only weak links to outcomes, which the authors read as evidence that identity and social position, not effort, shape who gets help. If the findings hold, crowdfunding is not a neutral emergency fund but a biased marketplace that channels charitable health dollars toward already-advantaged groups.","feed_headline":"Black recipients get about $22 less per GoFundMe donation","feed_subtitle":"A randomized sample of 637 U.S. medical campaigns shows non-white recipients raise less and women do most of the organizing.","key_machinery":"The central object is a hand-coded dataset of 637 randomized U.S. medical campaigns from GoFundMe, built by querying the site's search endpoint for the 500 campaigns nearest every U.S. zip code and then coding perceived race (white, Black, non-black person of color), gender, age, and campaigner–recipient relationship from campaign text, names, and photos. The analytical machinery is a pair of regression models—a linear regression on average donation amount and a Poisson regression on number of donations—with race, gender, age, relationship, and state population as predictors, alongside chi-square goodness-of-fit tests comparing campaign demographics to U.S. population benchmarks. Race coding used three raters from different backgrounds, each assessing every campaign, with an intraclass correlation of .819 that the paper treats as high agreement. What makes this machinery carry the argument is the contrast it draws: demographic variables show significant associations with outcomes, while the engagement behaviors platforms advise (photos, videos, updates, comments, hearts) do not.","core_discovery":"On its own terms, the paper reports that disparities appear at two distinct points in the crowdfunding process. First, in use: compared with U.S. population benchmarks, recipients perceived as white are over-represented (80.75% vs. 73%), recipients perceived as Black under-represented (8.48% vs. 12.7%), and non-black people of color under-represented (10.77% vs. 14.3%); the shortfall is sharpest for Black women, who make up less than 7% of women in the sample. Second, in outcomes: a linear regression on average donation amount gives a statistically significant coefficient of about −$22 for Black recipients relative to white recipients, and a Poisson regression on number of donations shows significantly fewer donations for Black and non-black POC recipients. Women are slightly less likely to receive donations than men, and women provide the overwhelming majority of organizing labor—82% of campaigns run on behalf of others. Children are under-represented among recipients but receive more donations of smaller average size. Campaign engagement variables such as updates, photos, and videos have minimal association with outcomes, leading the paper to conclude that crowd biases, not campaigners' efforts, dominate.","pith_inferences":["The authors' feedback-loop speculation implies a testable prediction: if failed campaigns are visible to potential users, the demographic skew in who starts campaigns should widen over time; longitudinal platform data could check this.","The paper cannot separate visibility from generosity with its data. If page-view counts become available, the key test is whether the race gap in donations comes from fewer views for non-white campaigns or from smaller gifts per view; the paper's crowd-bias reading would be supported only by the latter.","Because race was coded in three broad categories and socioeconomic status was not measured, a natural extension is to test whether the race coefficients survive finer racial categories and class controls.","If these findings generalize, donating through crowdfunding may function less as a remedy for health inequality and more as a mechanism that legitimizes it; this implication goes beyond the paper's explicit policy call for data transparency."],"forward_implications":["Campaign effort advice—updates, photos, videos—finds little support in the data, since these engagement behaviors show only minimal association with donations.","Non-white campaigners face a compounding disadvantage: under-representation at the point of entry plus worse outcomes once online.","Women's near-monopoly on campaign organizing constitutes a new form of unpaid digital care labor, extending feminized care work into online fundraising.","Children's campaigns attract more donations but smaller average gifts, so broad sympathy and viral spread do not translate into large financial commitments.","Medical crowdfunding should be understood as a biased marketplace rather than a neutral safety net, with data transparency a direct policy lever."],"supporting_citations":[{"why":"Prior study of 200 U.S. medical crowdfunding campaigns that supplied the sampling approach and found that few campaigns meet their goals.","marker":"(3)"},{"why":"Establishes GoFundMe's dominance of the donation-based crowdfunding market, justifying its selection as the sampling site.","marker":"(8)"},{"why":"Spatial study of Canadian cancer crowdfunding showing urban and higher-income skew; provides comparison for who uses medical crowdfunding.","marker":"(15)"},{"why":"Canadian study finding older adults, women, and visible minorities had poorer outcomes; the demographic-outcome baseline this paper extends.","marker":"(16)"},{"why":"Study of transgender medical crowdfunding showing predominance of white trans men; motivates intersectional analysis of under-representation.","marker":"(20)"},{"why":"Ethics guidelines for internet research that led the authors to exclude campaigns removed from the platform, shaping the final sample.","marker":"(62)"},{"why":"Supplies the intraclass correlation method used to assess inter-rater reliability of the race coding.","marker":"(68)"},{"why":"Supplies the U.S. population comparison benchmarks used in the chi-square goodness-of-fit tests.","marker":"(74)"}],"fun_headline_variants":["Non-white recipients raise less, women do most organizing in medical crowdfunding","Black recipients receive about $22 less per GoFundMe donation","Crowd biases, not campaign effort, decide medical fundraising success","Study: Non-white and women organizers face worse medical crowdfunding outcomes","Medical crowdfunding: racial and gender disparities persist, study finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that querying GoFundMe for the 500 campaigns closest to each U.S. zip code produced a list in which every qualifying medical campaign had a fair chance of being sampled; if that search favors urban, popular, or recent campaigns, the demographic comparisons and outcome estimates may not represent all U.S. medical crowdfunding.","fun_headline_variants_meta":{"raw":{"variants":["Non-white recipients raise less, women do most organizing in medical crowdfunding","Black recipients receive about $22 less per GoFundMe donation","Crowd biases, not campaign effort, decide medical fundraising success","Study: Non-white and women organizers face worse medical crowdfunding outcomes","Medical crowdfunding: racial and gender disparities persist, study finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000651,"raw_usage":{"total_tokens":3060,"prompt_tokens":1096,"completion_tokens":1964,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":712,"completion_tokens_details":{"reasoning_tokens":1874}},"tokens_in":712,"tokens_out":1964,"duration_ms":16415,"temperature":1.0,"reasoning_tokens":1874,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:29:24.613757+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct falsifier would be platform-wide data from GoFundMe for the same period: if a complete enumeration of U.S. medical campaigns showed no white over-representation and no significant race or gender differences in average donation or donation count after adjusting for goal, length, and engagement, the paper's central claim would fail. An experimental complement would post identical campaigns with randomly assigned recipient names and photos; if Black-coded campaigns draw the same average donations as white-coded ones, the crowd-bias mechanism is not operating.","supporting_citations":[],"review_version":1}