{"id":"79db2a0d-4d3b-498b-8f6b-dcc08ef73326","arxiv_id":"2507.07421","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An LLM-augmented synthetic data pipeline produces the largest public eviction-focused SDoH dataset (14 categories) and fine-tuned open LLMs that outperform prompt-optimized GPT-4o on the authors' test sets.","lead":"Researchers built a pipeline that uses large language models to rewrite clinical notes and create labeled training data for spotting eviction-related social risks in electronic health records. The resulting dataset and fine-tuned models detect eviction and related hardships like housing instability, with reported F1 scores around 88-90%, at a fraction of the usual annotation cost.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PMC train/test overlap is unverified: the 30% PMC notes added to the fine-tuning set (Sec. 4.3) come from the same PMC-Patients source as the human-rewritten PMC test notes (Sec. 4.2.3), and the paper never demonstrates disjointness, so the headline margin over GPT-4o-APO may be inflated.","rationale":"The central claim is that a cheap synthetic-data pipeline yields models that detect eviction and related SDoH at Macro-F1 88.8% and 90.3%, and outperform GPT-4o-APO. For that claim to hold, the evaluation must be unbiased. The weakest point is the PMC component because the same source corpus (PMC-Patients) is used both as a 30% training augmentation and as the raw material for the human-rewritten test set. The paper reports a large PMC gain from adding real PMC notes (+0.302 Micro-F1), making leakage the most plausible mechanism for inflation. The reported confidence intervals for the key comparison overlap, so even a small test-set contamination could change the headline. The reader identified exactly this unverified disjointness, and I agree that the resource is plausible and likely useful, but the comparison requires this check before accepting the generalization claim. If leakage is absent and the numbers replicate, the central claim survives. I do not see an internal inconsistency that would warrant rejection; the test-set realism concern (rewritten notes may not match naturally occurring eviction documentation) is real but secondary and can be addressed by external validation, while the disjointness check is the more immediate, falsifiable test.","tokens_in":34107,"tokens_out":7026,"duration_ms":78625,"concrete_test":"Use the released code and data manifests to compare the PMC-Patients article IDs (or normalized 13-gram hashes) of the 30% PMC notes in the SFT training split against the 48 PMC test notes (and their source articles). If any overlap exists, remove overlapping test instances and recompute PMC Macro-F1/Micro-F1 and the three-dataset average for all fine-tuned models and for GPT-4o-APO; if the difference between SynthEHR-Eviction and GPT-4o-APO remains within its reported confidence interval, the concern lands. If no overlaps are found, the paper should add a one-paragraph statement of the deduplication procedure to settle the issue.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 states that in the second fine-tuning phase 'we enriched the training set by adding 30% PMC long notes,' and Section 4.2.3 defines the PMC development set as PMC-Patients case reports 'rewritten by humans to include targeted eviction-related scenarios.' Table 4 shows 48 PMC test instances in each multi-class eviction split; the paper does not state that the PMC documents used for the 30% training replacement are disjoint from the PMC documents used to build these test instances. Table 5 reports that adding the 30% PMC training notes lifts Qwen2.5-7B's PMC Micro-F1 by +0.302, so this component is decisive for the PMC numbers. Since the three-way average weights PMC as one-third, and the claimed 'outperforming GPT-4o-APO' margin in eviction is about 0.01 Macro-F1 (0.888 vs 0.878), leakage in even a handful of the 48 PMC test notes could flip the headline. This is not an accusation of misconduct; it is a missing reproducibility check. The central dataset contribution can stand, but the as-stated generalization comparison is not fully supported until disjointness is shown.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SynthEHR-Eviction, a pipeline that combines GPT-based augmentation, DSPy automatic prompt optimization, and human-in-the-loop validation to construct a 14-class dataset of eviction-related social determinants of health (SDoH). The authors fine-tune open LLMs (Qwen2.5, LLaMA-3.1, etc.) and report Macro-F1 scores of 88.8% for eviction subcategories and 90.3% for non-eviction subcategories averaged over synthetic, MIMIC, and PMC test sets, slightly exceeding a GPT-4o-APO baseline while reducing manual annotation effort by over 80%. The paper also presents class-wise analyses, reasoning-annotation ablations, and training-size scaling experiments.","tokens_in":34416,"tokens_out":7752,"duration_ms":76554,"significance":"If the generalization claim survives scrutiny, this is a useful resource: the pipeline is modular, the dataset is publicly released (except MIMIC, gated by DUA), and the use of five-run confidence intervals, reasoning annotations for small models, and systematic ablation of training size are strengths. The central comparison to GPT-4o-APO, however, rests on devsets that are human-rewritten to inject eviction content and on a hybrid training set that may overlap with the PMC test set; until those checks are performed, the real-world claim is not fully established.","major_comments":[{"comment":"Section 4.3 states that the second fine-tuning phase enriches the training set by adding 30% PMC long notes, while Section 4.2.3 constructs the PMC devset from PMC-Patients case reports rewritten by humans to include targeted eviction-related scenarios. The paper does not state whether the PMC documents used for the 30% training replacement are disjoint from the PMC documents used to build the 48 PMC test instances. Table 5 reports that this addition improves Qwen2.5-7B's PMC Micro-F1 by +0.302, so this component is decisive for the PMC results. Since the three-way average weights PMC as one-third and the claimed outperformance margin over GPT-4o-APO in eviction Macro-F1 is about 0.01 (0.888 vs. 0.878), overlap in even a handful of the 48 PMC test notes could change the headline. Please verify disjointness or re-run the evaluation with a guaranteed-disjoint test set.","section":"§4.3, Table 5"},{"comment":"The augmentation pipeline selects over 30,000 MIMIC discharge notes as raw input (Section 4.1.1), and the MIMIC devset consists of MIMIC notes that were rewritten by human experts to include eviction-related content (Section 4.2.3). The paper does not demonstrate that the raw notes used for augmentation are disjoint from the notes selected for rewriting into the MIMIC devset. If the same patient notes or even the same discharge summaries appear in both, the synthetic training data could have been generated from exactly the texts that later appear in the test set, inflating MIMIC F1. Please report document- and patient-level overlap between the augmentation corpus and the MIMIC test notes.","section":"§4.1.1 and §4.2.3"},{"comment":"The MIMIC and PMC devsets are human-rewritten notes rather than naturally occurring eviction documentation. The rewriting instructions in Table 8 ask annotators to preserve the original MIMIC documentation style and incorporate eviction-related circumstances, which produces a specific, possibly stereotyped, distribution of eviction language. The paper does not provide evidence that this distribution matches how eviction is actually documented in clinical notes, for example in the authors' prior VA corpus (reference 27). Without such evidence, the reported MIMIC/PMC F1 scores should be interpreted as performance on human-generated rewrites, not on naturalistic real-world notes. A direct comparison on unmodified eviction-containing notes would strengthen the claim.","section":"§4.2.3"},{"comment":"The description of the five-run procedure for GPT-4o and GPT-4o-mini states that one run used temperature=0 and four additional runs used temperature=0.5. This mixes deterministic and stochastic generation within the same five-run confidence interval, so the 95% intervals and p-values reported in Table 1 are not computed under a single well-defined sampling distribution. For a fair comparison with the fine-tuned models, which vary random seeds under fixed hyperparameters, either use temperature>0 for all runs of the closed models or report the deterministic run separately.","section":"§4.4.1"}],"minor_comments":[{"comment":"There is a typo in the phrase 'black-box eviciton prediction'; it should be 'eviction prediction'.","section":"§2.5"},{"comment":"Several typos appear in the appendix: 'happend' instead of 'happened' in Table 8, 'seires' instead of 'series' in Table 9, and 'Fountuantly' instead of 'Fortunately' in Table 13. A proofread pass is recommended.","section":"Appendix, Table 8/9/13"},{"comment":"The title 'Trainset and Devset' is ambiguous; consider renaming it to 'DSPy Trainset and Devset' to distinguish from the later SFT training set.","section":"§4.2.3"},{"comment":"The caption mentions 'Red solid and Red dashed horizontal lines' but does not clearly map which line corresponds to GPT-4o-APO versus GPT-4o-mini-APO; please clarify the legend and caption.","section":"Figure 2"},{"comment":"The rows for 'Eviction absent' have empty cells for Test-Mimic and Test-PMC; use an explicit dash or 'not applicable' to indicate that the class is intentionally excluded from those sets.","section":"Table 4"},{"comment":"The phrase 'divided into two groups, corresponding to three separate annotators' is confusing; rephrase to describe the two task groups (eviction vs. non-eviction) handled by three annotators (binary, eviction multi-class, non-eviction multi-class).","section":"§4.2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be a useful contribution, but the real-world generalization claim is currently supported only by human-rewritten test notes and a hybrid training set whose overlap with the test set is not verified. I recommend requesting the authors to (i) demonstrate PMC/MIMIC train-test disjointness at document and patient level, (ii) evaluate on naturally occurring eviction notes such as their prior VA dataset, and (iii) report confidence intervals under a consistent sampling protocol. If these checks pass, the manuscript would be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"If you're building SDoH extraction resources, this paper is worth a look for the dataset alone: 8,000 synthetic training notes across 14 fine-grained eviction and housing-related categories, with released code and devsets. But read the evaluation carefully before believing the \"outperforms GPT-4o-APO\" claim. The margin is real on in-distribution synthetic data and on MIMIC; on PMC the fine-tuned models are actually worse, and the reported averages hide that.\n\nWhat's new: this is the first large public benchmark for eviction-related SDoH extraction, filling a genuine gap. The pipeline—LLM augmenter, DSPy-based annotator with chain-of-thought reasoning, human verification, and automated prompt optimization—is not conceptually novel, but the integration is clean and the annotation effort reduction claim (over 80%) is plausibly demonstrated. The 5-run confidence intervals and transparent data tables are a methodological plus.\n\nThe soft spots, in proportion. First, the \"real-world\" test sets are not natural: MIMIC and PMC notes were human-rewritten to inject eviction scenarios (Sec 4.2.3). This tests recognition of rewritten eviction language, not the messy ways eviction appears in actual clinical text. It is a proxy, and the paper should say so more clearly. Second, and more load-bearing, the PMC training/test disjointness is never established. Section 4.3 says 30% of the fine-tuning training set was PMC notes, and the PMC test set comes from the same PMC-Patients source. The paper does not state these are disjoint. Since the 30% PMC mix adds +0.302 Micro-F1 to Qwen2.5-7B on PMC, and the headline margin over GPT-4o-APO is about 0.01 Macro-F1, even a handful of overlapping notes could flip the comparison. This is a missing reproducibility check—not an accusation of misconduct—but it needs to be resolved. Third, the headline comparison is thinner than it looks: Qwen2.5-7B's average Macro-F1 of 0.888 vs GPT-4o-APO's 0.878 has overlapping confidence intervals, and GPT-4o-APO is actually higher on PMC. The abstract overstates the case.\n\nOn the other side, the authors are honest about limitations: they discuss temporal ambiguity, reasoning faithfulness problems, and synthetic-data bias. That candor deserves credit. The dataset itself is a serious resource.\n\nWho this is for: anyone working on clinical NLP for social determinants, especially eviction or housing instability. It is a useful benchmark and a template for synthetic data generation, with the caveat that evaluation against naturally occurring eviction documentation is still needed.\n\nRecommendation: yes, send to peer review. The resource contribution is real, the reporting is transparent, and the flaws are fixable with a disjointness check, a clearer statement about the rewritten test set, and a less triumphalist abstract. A good referee will push on those and the paper will be stronger for it.","headline":"A genuinely useful synthetic dataset and pipeline for eviction-related SDoH extraction, but the headline claims overstate what the evaluation actually supports; worth engaging, with fixes.","tokens_in":34944,"tokens_out":3107,"would_cite":true,"duration_ms":31941,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an LLM-driven pipeline can generate synthetic EHR notes labeled for eviction and related social risks, and that fine-tuned open models trained on this data reach Macro-F1 of 88.8% for eviction and 90.3% for other…","keywords":["eviction","social determinants of health","electronic health records","synthetic data","large language models","prompt optimization","clinical NLP","fine-grained annotation"],"falsifier":"Take a corpus of real clinical notes known to contain eviction documentation (for example, an unmodified EHR sample with verified eviction mentions), run the fine-tuned Qwen2.5-7B model on it, and compare F1 against the reported 88.8%; also check whether any PMC training notes overlap the PMC test set. A large drop or detectable overlap would show the reported scores overstate deployment readiness.","tokens_in":33895,"feed_emoji":"🏠","tokens_out":5757,"duration_ms":58427,"temperature":0.7,"pith_summary":"Eviction is a social determinant of health that appears in clinical notes but is almost never recorded in structured EHR fields, so it escapes detection in routine care. This paper tries to establish that a mostly synthetic training dataset, built by an LLM-driven pipeline with human verification, is enough to make open-weights models detect eviction and related housing risks from free-text notes. On human-validated test sets spanning synthetic, MIMIC, and PMC notes, fine-tuned Qwen2.5-7B and LLaMA-3.1-8B reach Macro-F1 of 88.8% for eviction subcategories and 90.3% for other social determinants, edging past a prompt-optimized GPT-4o baseline (87.8% and 87.3%) while cutting annotation time by more than 80%. If true, the paper offers a low-cost, reproducible path to eviction-risk screening from existing EHR text.","feed_headline":"Synthetic EHR notes train eviction detectors to 88.8% F1","feed_subtitle":"Open models beat GPT-4o-APO on eviction SDoH at under 20% of the annotation cost.","key_machinery":"The central mechanism is a two-stage augmenter–annotator loop. The augmenter is a label-specific LLM prompt optimized against expert feedback; it rewrites real MIMIC social-history sections so they exhibit a target eviction or SDoH class, with human verification filtering outputs until accuracy exceeds 90%. The annotator is a DSPy program trained on a small human-validated set, using chain-of-thought reasoning and BootstrapFewShotWithRandomSearch to produce both a label and a rationale. The fine-tuning recipe then mixes 70% synthetic notes with 30% real PMC notes and includes the reasoning traces, which the paper shows is what lets small open models transfer to long, narrative clinical text.","core_discovery":"The paper's central claim is that eviction, a social determinant of health almost never coded in structured EHR fields, can be detected from clinical free text using models trained on a mostly synthetic dataset produced by its SynthEHR-Eviction pipeline. The pipeline separates generation from verification: label-specific LLM 'augmenters' rewrite real MIMIC social-history sections into eviction-relevant notes under expert feedback, and DSPy-optimized 'annotators' label the notes with 14 fine-grained categories plus chain-of-thought rationales. Fine-tuned open models, especially Qwen2.5-7B and LLaMA-3.1-8B, reach Macro-F1 0.888 for eviction subcategories and 0.903 for other SDoH categories on human-validated test sets spanning synthetic, MIMIC, and PMC notes, slightly exceeding the GPT-4o-APO baseline (0.878 and 0.873) while reducing human annotation time from over 266 hours to under 6.","pith_inferences":["Inference — If the pipeline generalizes as claimed, the largest practical impact may be surveillance: coupling an eviction-risk flag with a generated rationale could let health systems screen whole populations without waiting for structured Z-code documentation.","Inference — The rewritten-test design means real-world F1 is untested; a fair deployment test on unmodified notes with adjudicated eviction mentions would be the first check before clinical use.","Inference — The plateau at 3k–5k examples hints that the synthetic data is information-dense; testing whether the same plateau holds for other SDoH domains would show whether the 80% labor saving transfers.","Inference — Because the training and test PMC notes share a source, part of the reported generalization gain may be source-overlap rather than lexical diversity; constructing a held-out hospital system's notes would separate these effects."],"forward_implications":["Fine-tuned open-weights LLMs trained on the released dataset can be deployed at 3B scale with Macro-F1 near 0.85–0.89 on eviction tasks, making eviction screening feasible without proprietary APIs.","The 3,000–5,000 sample plateau suggests downstream users can reproduce the pipeline's gains without collecting tens of thousands of annotations.","Including reasoning traces in training data lifts small-model accuracy (LLaMA-3.2-3B: +3.8 Macro-F1), so interpretable rationales are a byproduct rather than a cost.","Replacing 30% of synthetic training notes with real PMC notes improves out-of-domain PMC Micro-F1 by up to +0.302, indicating synthetic-only training is insufficient for real-world transfer.","The same augmenter–annotator design is claimed to generalize to other under-coded SDoH such as food insecurity, utility shutoff, and intimate partner violence."],"supporting_citations":[{"why":"Supplies prior LLM-based SDoH extraction work that motivates the approach and frames the comparison with biomedical BERT models.","marker":"[11]"},{"why":"MIMIC-IV is the source of raw social-history notes used for augmentation and of real-world test notes.","marker":"[19]"},{"why":"ICD-10-CM Z59 codes provide the label schema grounding the taxonomy.","marker":"[20]"},{"why":"Establishes the prior VA-based eviction classifier and the presence-by-period label schema reused here.","marker":"[27]"},{"why":"Surveys automated prompt optimization techniques that motivate the APO component.","marker":"[28]"},{"why":"Provides the DSPy framework used to implement the annotator, chain-of-thought reasoning, and prompt optimization.","marker":"[29]"},{"why":"Supplies the PMC-Patients case reports used in the devset and in the 30% hybrid training mix.","marker":"[31]"}],"fun_headline_variants":["Synthetic EHR data cracks eviction detection at 88.8% F1","Open models beat GPT-4o-APO on eviction SDoH from synthetic EHR","SynthEHR-Eviction: 88.8% F1 with 80% less annotation","Eviction from EHR notes: synthetic data slashes annotation by 80%","Open LLMs outdo GPT-4o on eviction SDoH with synthetic EHR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation relies on test notes that human experts rewrote to inject eviction content, and the training set includes PMC notes drawn from the same source as the PMC test notes, so the reported F1 may not reflect performance on naturally occurring eviction documentation.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic EHR data cracks eviction detection at 88.8% F1","Open models beat GPT-4o-APO on eviction SDoH from synthetic EHR","SynthEHR-Eviction: 88.8% F1 with 80% less annotation","Eviction from EHR notes: synthetic data slashes annotation by 80%","Open LLMs outdo GPT-4o on eviction SDoH with synthetic EHR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001799,"raw_usage":{"total_tokens":7121,"prompt_tokens":1018,"completion_tokens":6103,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":634,"completion_tokens_details":{"reasoning_tokens":5988}},"tokens_in":634,"tokens_out":6103,"duration_ms":39929,"temperature":1.0,"reasoning_tokens":5988,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:41:56.383004+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a corpus of real clinical notes known to contain eviction documentation (for example, an unmodified EHR sample with verified eviction mentions), run the fine-tuned Qwen2.5-7B model on it, and compare F1 against the reported 88.8%; also check whether any PMC training notes overlap the PMC test set. A large drop or detectable overlap would show the reported scores overstate deployment readiness.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies prior LLM-based SDoH extraction work that motivates the approach and frames the comparison with biomedical BERT models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"MIMIC-IV is the source of raw social-history notes used for augmentation and of real-world test notes."},{"cited_title":"Improving the collection of social determinants of health (sdoh) data with icd-10-cm z codes (2023)","cited_arxiv_id":null,"evidence_quote":"ICD-10-CM Z59 codes provide the label schema grounding the taxonomy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prior VA-based eviction classifier and the presence-by-period label schema reused here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the PMC-Patients case reports used in the devset and in the 30% hybrid training mix."}],"review_version":1}