{"id":"031e87a7-5728-499c-b8f7-2fb7fce96e48","arxiv_id":"2506.13842","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"SatHealth combines Ohio satellite images, climate and air quality data, and all-disease prevalence from medical claims into a public dataset, and demonstrates that its environmental embeddings improve health prediction models.","lead":"This paper introduces SatHealth, a public dataset linking satellite images, environmental variables, and medical claims-derived disease prevalence for Ohio from 2016 to 2022. The authors show that adding this environmental information improves AI models for both regional health prediction and personal disease risk.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'significantly improve' claim is not statistically supported: no confidence intervals or tests are reported, and the regional regressions use few MSA/CBSA-level units with hundreds of features, so the Table 4/6 improvements may be overfitting or noise.","rationale":"After reading the paper, I agree the dataset is a substantial resource and the reader's CONDITIONAL verdict is appropriate. The reader's weakest assumption, MarketScan prevalence representativeness, is a real validity threat, and external validation against CDC PLACES/BRFSS or similar sources would materially strengthen the paper. However, I see a more immediately load-bearing weakness: the central claim is expressed in statistical language ('significantly improve') but is never backed by any uncertainty quantification, and the regional modeling experiments that ground the generalizability claim use a very small number of independent spatial units relative to the feature dimensionality. This makes the reported R^2 differences fragile regardless of label representativeness. The concern is not that the authors are dishonest; it is that the evidence as presented cannot discriminate between a real environmental signal and overfitting or noise. If the suggested leave-one-CBSA-out bootstrap shows stable improvements, the concern is resolved and the verdict could move toward ACCEPT. If the intervals include zero, the claim should be weakened to 'can improve in some settings' pending further validation. Therefore I would keep the reader's CONDITIONAL verdict, with an added condition of reporting uncertainty estimates and robustness to spatial-unit splits.","tokens_in":24940,"tokens_out":8702,"duration_ms":90698,"concrete_test":"Re-run the Section 4.2 disease-prevalence regressions with leave-one-CBSA-out cross-validation, then bootstrap the All-vs-DEnv R^2/MAE difference over 1,000 resamples of the MSA/CBSA units and report 95% confidence intervals. If any interval includes zero, or if the leave-one-out ranking of 'All' flips on more than one fold, the 'significantly improve' claim is not supported. Repeat the same procedure on Table 6's spatial interpolation and temporal forecasting splits.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Central claim: environmental information 'significantly improve[s] AI models' performance and temporal-spatial generalizability' (Abstract; Section 4). The empirical support rests on Tables 4-6. No significance tests, confidence intervals, or repeated-seed variation are reported, and the word 'significantly' is never operationalized. More structurally, disease prevalence is available only at MSA/CBSA granularity (Section 3.4; Table 7), giving roughly 15 metropolitan areas in Ohio. Table 4 trains random forests on a combined feature set of ~36 dynamic variables (seasonally expanded to ~144), 9 land-cover fractions, and 300 satellite-image features against about 15 regions x 7 years = ~105 samples. In this regime, R^2 gains of 0.02-0.09 for 'All' over single modalities could reflect overfitting to a few MSAs rather than genuine environmental signal. Table 6's spatial folds are splits of these same regions, so the claimed spatiotemporal generalization is not established. The MSA-only geocoding also excludes non-metropolitan Ohio, and MarketScan covers insured employees and dependents only, so the prevalence labels and population are selective; Section 6's limitations do not address this. A secondary issue: some Table 5 metrics decrease with environment (e.g., RETAIN next-visit Recall@50: 0.456 to 0.445), so the benefit is not uniform. The dataset may still be useful, but the paper's strongest claim is under-supported as reported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript introduces SatHealth, a multimodal public health dataset for Ohio (2016–2022) that combines satellite imagery, climate/air-quality/greenery time series, land cover fractions, Social Deprivation Index (SDI) scores, and disease prevalences estimated from the MarketScan CCAE commercial claims database. The authors propose a deterministic embedding pipeline that fuses these modalities into regional environmental embeddings and then evaluate the dataset in two use cases: regional regression of SDI and disease prevalence, and personalized disease-risk prediction with LSTM, RETAIN, Dipole, and Transformer backbones. A further set of experiments examines spatial interpolation, spatial extrapolation, and temporal forecasting. The central claim is that living-environment information 'significantly improve[s] AI models' performance and temporal-spatial generalizability' (Abstract, Section 4). The paper also describes a web application and a public code repository for data access and reproduction.","tokens_in":25166,"tokens_out":2603,"duration_ms":26753,"significance":"If the central claim were established, SatHealth would be a valuable public resource: it is, to my knowledge, the first US dataset combining regional environmental characteristics with a claims-based all-disease prevalence panel, and the authors provide a web application, downloadable data, and a reproducible pipeline. The two use cases are sensible and the deterministic embedding approach is transparent and easy to reuse. However, the headline claim of significant improvement is currently under-supported: the prevalence targets are derived from an insured commercial population with unknown representativeness, the spatial granularity yields only about fifteen regional units, and the reported R2 differences are not accompanied by any significance testing, confidence intervals, or repeated-seed variability. These issues are load-bearing because the same experimental evidence is used in the abstract, Section 4, and the conclusion.","major_comments":[{"comment":"The disease prevalence targets are estimated only from MarketScan CCAE, which covers employees and dependents in employer-sponsored plans, yet no external validation against population-representative sources (e.g., CDC surveys, Medicare, or state registries) is provided. Since prevalence is the health outcome in both use cases, systematic bias in the enrolled population would propagate into the environment-health correlations and the regression and prediction results. Section 6 lists limitations about Ohio-only coverage, simple embeddings, and coarse residence, but does not acknowledge or mitigate this representativeness concern, which is a material omission.","section":"§3.4, §6"},{"comment":"The claim that environmental information 'significantly' improves predictions is not statistically supported. The regional regressions use about 15 MSA/CBSA-level units and 7 years (on the order of 105 region-year observations) with hundreds of input features, and the reported R2 gains of 0.02–0.09 are not accompanied by confidence intervals, hypothesis tests, or repeated cross-validation with different seeds. Without such evidence, the improvements over single modalities may reflect noise or overfitting rather than genuine environmental signal. The authors should add significance testing or clearly reframe the claim as descriptive rather than inferential.","section":"§4.2, Table 4"},{"comment":"The benefit of environmental information is not uniform, which contradicts the blanket statement of significant improvement. For example, for RETAIN, next-visit Recall@50 decreases from 0.456 to 0.445 when environment is added; for LSTM, 1-year diagnosis Recall@5 decreases from 0.503 to 0.489 and Recall@10 from 0.537 to 0.508. The paper should report per-model, per-metric variability and should temper the general claim accordingly, distinguishing settings where the environmental embedding helps from those where it does not.","section":"§4.3, Table 5"},{"comment":"The spatiotemporal generalization results do not consistently support the claimed generalizability benefit. In the spatial extrapolation scenario, SDoH R2 is near zero or negative for every feature set (e.g., DEnv: -0.186, All: 0.035), and adding temporal information sometimes degrades performance (e.g., DEnv+T SDoH: -0.301 vs. DEnv: -0.186). These patterns should be discussed as evidence that the proposed spatiotemporal enhancement is not universally beneficial, and the claims in the text should be aligned with the full set of results rather than only the 'best or second best' cases.","section":"§4.4, Table 6"}],"minor_comments":[{"comment":"In the sentence referencing Table 3, 'Tabel' should be 'Table'.","section":"§4.1.1"},{"comment":"'Greenary' appears as a typo for 'Greenery' in several places, including Section 3.2 and Table 11.","section":"§3.2, Table 11"},{"comment":"Table 2 lists 'SDI 8 Variables' while Section 3.1 and Appendix B.2.3 describe seven demographic components; the table should clarify whether the eighth variable is the overall SDI score.","section":"Table 2"},{"comment":"The web-application section would benefit from screenshots with readable text, and from stating whether the downloadable embeddings include the exact versions used in the reported experiments.","section":"§5"},{"comment":"The statement that Google Maps images lack timestamps and are treated as static is important, but the assumption that the landscape is stable over 2016–2022 should be stated more explicitly as an assumption in the main text, not only in the appendix.","section":"Appendix B.2.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know up front: SatHealth is a real and useful new dataset, but the paper's headline claim — that environmental information 'significantly improves' AI model performance and spatiotemporal generalizability — is not supported by the reported evidence. The dataset itself is the contribution.\n\nWhat's new: to my knowledge it's the first US public resource linking claims-derived all-ICD prevalence with satellite imagery, environmental variables, and SDoH, at multiple geographic levels. The authors put together 432k Google Maps images, climate, air quality, greenery, land cover, SDI, and MarketScan prevalence for 1,377 codes, and they ship the data, an embedding pipeline, and a web app. That's a useful asset for environmental health informatics. The embedding is simple feature engineering but transparent and reproducible.\n\nThe soft spots are real. The 'significantly' word is doing too much. Tables 4-6 report R^2 differences of 0.02-0.09 with no confidence intervals or significance tests. The regression runs on roughly 15 Ohio MSAs, so with hundreds of features those gains could be noise or overfitting. Table 5 shows improvement in most backbones but not all, and a couple of RETAIN metrics get worse. The stress-test note is fair here.\n\nSecond, the prevalence labels are estimated from MarketScan CCAE, a commercially insured population. The authors don't validate against CDC, Medicare, or state registries. That is a meaningful limitation because prevalence is the target in the main use cases. Section 6 mentions coarse residency and Ohio-only coverage but doesn't address representativeness.\n\nThird, the spatiotemporal generalization results undercut the generalizability claim: the extrapolation rows for SDoH are often negative R^2, and spatial autocorrelation likely inflates the interpolation numbers. So the 'significant generalizability' claim is not established.\n\nNone of that makes the dataset worthless. The data is the product, and it's a good one. The paper should be revised to present the experiments as demonstrations of potential, not as proof of significance, and the authors should add uncertainty quantification and external validation of the prevalence proxy.\n\nBottom line: send it to peer review with a request for revision. It deserves a serious referee.","headline":"A genuinely new and useful dataset for environmental health AI; the 'significantly improve' claim is not statistically supported as reported.","tokens_in":25743,"tokens_out":2981,"would_cite":true,"duration_ms":28753,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SatHealth builds a public Ohio dataset pairing satellite images and environmental variables with claims-derived disease prevalence, and argues that adding regional environment embeddings improves AI models' accuracy and spatiotemporal…","keywords":["satellite imagery","environmental health informatics","social determinants of health","disease prevalence","medical claims","multimodal data","spatiotemporal generalization","regional environment embeddings"],"falsifier":"Using the same Ohio CBSAs and years, compare SatHealth's ICD prevalence estimates with an external population-representative source such as CDC PLACES or Ohio state public health data; then train the paper's regional and patient-level models to predict the external target instead of the MarketScan label. If the environment embeddings stop improving accuracy against the external outcomes, or the urban-rural prevalence gaps disappear, the central claim that living environmental information improves health AI is an artifact of the claims population rather than a general result.","tokens_in":24673,"feed_emoji":"🛰️","tokens_out":6318,"duration_ms":63806,"temperature":0.7,"pith_summary":"SatHealth is a public dataset that brings the living environment into AI health research: for every Ohio county, ZIP-code tabulation area, census tract, and metropolitan area from 2016 to 2022, it joins satellite images, climate and air-quality time series, land cover, a social-deprivation index, and yearly prevalence estimates for 1,377 ICD-coded diseases computed from medical claims. The load-bearing claim is that this environmental information is not decoration: adding it to models improves prediction of regional social-deprivation scores and disease prevalence, and also improves patient-level next-visit and one-year disease-risk prediction across several sequence models. The paper also reports that environment-augmented models generalize better under spatial interpolation, spatial extrapolation, and temporal forecasting, and it releases the embeddings so other models can use them without reprocessing raw imagery. If the claim is right, environment-aware AI can be built on top of existing electronic-health-record backbones rather than treated as a separate data collection burden.","feed_headline":"Ohio environment data boosts health AI, new dataset shows","feed_subtitle":"SatHealth pairs 432k satellite images with claims-based disease rates; adding region embeddings improves risk prediction.","key_machinery":"The regional environment embedding is the object that carries the argument. Dynamic monthly variables (climate, air quality, greenery) are averaged by meteorological season and concatenated into an annual vector; static land-cover fractions and pixel-statistics features derived from satellite images (RGB channels plus nine vegetation and color indices) are added to form a multimodal yearly profile for each region. These embeddings are then fed to random-forest regressors for regional health outcomes and concatenated to patient representations from EHR sequence models, with an optional boosting step that adds neighborhood and historical residual predictors. The same embedding object is what turns raw geospatial data into a plug-in health-AI input.","core_discovery":"The paper's central discovery is a reusable multimodal public resource plus the evidence that it works: SatHealth fuses high-resolution Google Maps satellite imagery (432,918 images, each about 500 meters square), monthly environmental variables (climate, air quality, greenery), static land cover, and 2019 Social Deprivation Index values with disease prevalences estimated from 2,141,777 Ohio patients in the MarketScan commercial claims database. From these features the authors construct regional environment embeddings at four geographic levels, then use them in two tasks. On regional modeling, combining all modalities raises $R^2$ for neoplasm prevalence by 0.086 over the best single-modality features and matches or beats single modalities for diabetes and hypertension targets. On personalized prediction, adding the environmental embedding to LSTM, RETAIN, Dipole, and Transformer backbones improves macro AUROC and recall metrics for most settings; with Dipole, next-visit mAUC rises from 0.600 to 0.722. The authors claim this makes SatHealth the first US dataset to combine regional environmental characteristics with a healthcare database, and they interpret the gains as evidence that living environmental information can significantly improve model performance and temporal-spatial generalizability.","pith_inferences":["If the MarketScan prevalence estimates are not representative of Ohio's general population, the observed urban-rural odds ratios could partly reflect who is insured and which providers bill claims, so the environmental associations should be checked against population-representative surveys before being treated as public-health facts.","The paper's own Limitations section notes that only Ohio is covered, that the embeddings are simple feature-engineered statistics rather than learned representations, and that patient residence is coarse; these bound the generalizability claim but do not undermine the resource itself.","Because the embeddings are computed from public geospatial inputs alone, they could plausibly be reused for tasks the paper does not study, such as hospital-resource planning or environmental-justice screening, but those extensions would need their own outcome validation.","The correlation analyses suggest testable causal hypotheses (green space and soil conditions linked to cardiovascular and metabolic disease), but the dataset's cross-sectional design cannot by itself distinguish environment effects from population sorting."],"forward_implications":["Researchers can add environmental context to any health model with a patient or region location by using SatHealth's precomputed embeddings, avoiding raw satellite and weather processing.","The reported results imply that EHR-based risk prediction is leaving signal on the table when it ignores residence: the largest gains appear on recall, meaning environment helps surface diseases a model would otherwise miss.","Combining dynamic and static modalities seems more useful than any single modality, so future SatHealth-style resources should keep the multimodal design rather than simplify to one data source.","The spatiotemporal-enhancement results suggest that models trained on this dataset should include neighborhood and history features when deployed across regions or years.","The published pipeline and web application allow construction of the same dataset for other US states, which the authors say will be updated toward national coverage."],"supporting_citations":[{"why":"Supplies the MarketScan CCAE claims data from which all disease-prevalence targets are computed.","marker":"[45]"},{"why":"Provides Google Earth Engine, the platform used to collect and align climate, air-quality, greenery, and land-cover variables.","marker":"[29]"},{"why":"Defines the Social Deprivation Index used as the SDoH target in regional public health modeling.","marker":"[34]"},{"why":"Describes MedSat, the England-based precursor whose processing pipeline and variable choices SatHealth adapts.","marker":"[58]"},{"why":"Supplies the Google Static Maps API source for the 432,918 high-resolution satellite images.","marker":"[28]"},{"why":"Documents the protective association between greenness and cardiovascular disease that motivates the disease targets and the interpretation of imagery features.","marker":"[4]"},{"why":"Supports the link between urban green space and heart disease, hypertension, and diabetes used to select regression targets.","marker":"[5]"}],"fun_headline_variants":["SatHealth: satellite data boosts disease prediction AI","SatHealth pairs satellite imagery with claims to improve health AI","Adding environment data to health models lifts accuracy, SatHealth shows","SatHealth: environmental embeddings improve health prediction","Satellite + claims data: new dataset sharpens health AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The health labels are not population-representative: disease prevalence is computed from the MarketScan commercial claims database, which covers insured employees and their dependents, and the paper does not validate these estimates against CDC surveys, Medicare, or state registries; if that insured population differs from Ohio as a whole, the environment-health correlations and prediction gains are systematically biased.","fun_headline_variants_meta":{"raw":{"variants":["SatHealth: satellite data boosts disease prediction AI","SatHealth pairs satellite imagery with claims to improve health AI","Adding environment data to health models lifts accuracy, SatHealth shows","SatHealth: environmental embeddings improve health prediction","Satellite + claims data: new dataset sharpens health AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000562,"raw_usage":{"total_tokens":2707,"prompt_tokens":1023,"completion_tokens":1684,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":1606}},"tokens_in":639,"tokens_out":1684,"duration_ms":12590,"temperature":1.0,"reasoning_tokens":1606,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:56:53.188806+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Using the same Ohio CBSAs and years, compare SatHealth's ICD prevalence estimates with an external population-representative source such as CDC PLACES or Ohio state public health data; then train the paper's regional and patient-level models to predict the external target instead of the MarketScan label. If the environment embeddings stop improving accuracy against the external outcomes, or the urban-rural prevalence gaps disappear, the central claim that living environmental information improves health AI is an artifact of the claims population rather than a general result.","supporting_citations":[{"cited_title":"2024.Real world evidence","cited_arxiv_id":null,"evidence_quote":"Supplies the MarketScan CCAE claims data from which all disease-prevalence targets are computed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Social Deprivation Index used as the SDoH target in regional public health modeling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes MedSat, the England-based precursor whose processing pipeline and variable choices SatHealth adapts."},{"cited_title":"2024.Google Maps Static API","cited_arxiv_id":null,"evidence_quote":"Supplies the Google Static Maps API source for the 432,918 high-resolution satellite images."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the protective association between greenness and cardiovascular disease that motivates the disease targets and the interpretation of imagery features."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the link between urban green space and heart disease, hypertension, and diabetes used to select regression targets."}],"review_version":2}