{"id":"9c65907c-2ae5-4c1f-9b07-144abd14b206","arxiv_id":"2506.12808","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative survey of MIMIC dataset challenges that is undermined by incorrect citations and unsourced performance tables.","lead":"This paper surveys known problems and progress in research using the MIMIC critical care datasets, such as missing data, interoperability, and reproducibility. It organizes many real challenges, but its reference list and several tables contain inaccurate or unverifiable entries, so the survey cannot be used as a reliable map of the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central claim of being a reliable, comprehensive map of MIMIC-based research fails because its citation base is demonstrably unreliable, making every synthesized claim about open problems and progress unverifiable.","rationale":"The reader's verdict identifies the same load-bearing concern: the survey's credibility depends on its citations existing and supporting the claims made, and that assumption is violated by direct evidence. My independent reading confirms multiple egregious citation mismatches (peptide-design paper for digital health policy, CERN report for clinical notes, placeholder references), unsourced performance tables, and textual corruption. These failures attack the central purpose of a survey: providing a trustworthy map of prior work and open problems. The reviewer is not arguing against the paper's conclusions from a position of scientific disagreement; the issue is internal verifiability. I therefore agree with the REJECT verdict. No adjustment is needed. The one caution is to avoid overgeneralizing: some sections may contain accurate summaries of real work, but the pattern of unverifiable citations is systemic enough to invalidate the survey as a whole unless the authors repair the reference list, source all tables, and remove placeholder entries.","tokens_in":44448,"tokens_out":1719,"duration_ms":19658,"concrete_test":"Select 30 numbered references cited in support of specific factual claims (e.g., refs [1], [10], [13], [23], [34], [50], [62]). Retrieve each from PubMed, arXiv, IEEE, or DOI lookup and record: (a) existence, (b) topical match, (c) support for the specific sentence. If fewer than 20 of 30 pass all three, or any placeholder remains, the survey's literature map is unsupported. Also check Table 6 for a traceable source.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it is a comprehensive survey that uniquely focuses on open problems in MIMIC-based digital health research. For this claim to hold, readers must be able to trust that the paper accurately represents prior work: that cited references exist, that they concern the claimed topic, and that they support the specific statements attributed to them. This condition is visibly violated throughout the manuscript. Section 1.1 cites reference [1] to support claims about informing policy decisions and advancing evidence-informed policy, but reference [1] is a peptide-design paper about binding interface mimicry. The same section cites reference [10] for the statement that MIMIC includes unstructured clinical notes, but reference [10] is listed as a CERN technical report on particle physics. Reference [13], cited for MIMIC-IV statistics and as a source for MIMIC-IV benchmarks, is listed as an ETL/FHIR-OMOP loading paper. Reference [50] is just 'Knight. Imitation.' and reference [62] is 'John Smith and Jane Doe. Title of the study.' — placeholders, not citable works. Table 6 gives AUROC values for 'Best Model' on mortality, sepsis, and phenotyping tasks with no source, and Tables 7–8 present performance ranges that cannot be traced to any listed reference. Additional internal defects compound the problem: repeated and garbled sentences in Section 1.2.2 (including '1A20 000' and 'bassolid undsoliding'), duplicated bullets, and a limitations section that concedes bias. These are not stylistic quibbles; they directly destroy the survey's epistemic value as a synthesis. If citations cannot be trusted, the paper's claims about what progress has been made, which open problems persist, and what future directions are promising are all unsupported. The central claim therefore does not hold.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a narrative survey of research using the MIMIC critical care databases in digital health. It claims to be the first comprehensive survey focused specifically on open problems, organizing prior work on data granularity, cardinality, predictive modeling, data integration, data quality, interoperability, reproducibility, and privacy/ethics, and offering a taxonomy of applications and future directions.","tokens_in":44638,"tokens_out":2497,"duration_ms":24812,"significance":"If the survey were reliable, it would be a useful entry point for the MIMIC research community: it assembles a broad set of themes, distinguishes open problems from progress, and suggests concrete future work such as unified ETL toolkits, federated learning, and standardized benchmarks. The paper also credits specific prior contributions such as MIMIC-Extract, the MIMIC Code Repository, and GRU-D. However, the survey's central value depends on the accuracy of its synthesis, and that accuracy is not supported by the manuscript as submitted.","major_comments":[{"comment":"The survey's core claim of being a reliable map of prior MIMIC research is undermined by citations that do not match the claims they support. Reference [1] is listed as 'F. Johnson et al. Peptide design through binding interface mimicry' yet it is cited for statements about improving patient outcomes and informing policy decisions; reference [2] is a software-architecture paper cited for biomedical innovation and interoperability; and reference [10] is a CERN technical report cited as the source for the claim that MIMIC includes unstructured clinical notes. These are not isolated typos but systematic misattributions that make it impossible for the reader to verify the survey's statements.","section":"Section 1.1, references [1], [2], [10]"},{"comment":"The text contains numerous garbled and duplicated passages that indicate the manuscript is not carefully edited. Section 1.2.2 refers to '1A20 000 ICU stays' and 'a solid bassolid undsoliding the need for a need for any additional supported by significant improvements,' and repeats the identical bullet about predicting mortality and classifying length of stay. Section 2.3 repeats the sentence 'Further standardization frameworks are important; crucially, there are no user friendly tools.' These errors are not merely stylistic: they make specific factual claims unreadable and cast doubt on the care with which the rest of the content was verified.","section":"Section 1.2.2 and Section 2.3"},{"comment":"Table 6 reports AUROC values of 0.88 for mortality, 0.83 for sepsis, and 0.91 for phenotyping under a column 'Best Model' (e.g., XGBoost + LSTM, Temporal Fusion Transformer, BioClinicalBERT + GNN) without any citation, dataset split, or model definition. Tables 7 and 8 give AUROC ranges and MSE/RMSE values for classification and regression tasks with no indication of which studies produced these numbers. Because the paper is a survey, these unsourced performance figures cannot be checked against the literature and therefore fail to support the paper's claims about the state of progress.","section":"Section 4.5.6, Table 6, and Section 5, Tables 7-8"},{"comment":"Several subsections summarize studies that are unrelated to MIMIC or digital health. Section 3.1.5 describes mindfulness-based interventions, Section 3.1.6 describes a global air-quality PM2.5 dataset, and Section 3.1.4 includes 'A Survey on Network Data Analytics.' These passages appear to be abstracted from completely different papers and are presented as MIMIC applications without any connection to the dataset. This materially weakens the survey's claim to comprehensiveness and suggests that the application taxonomy is not grounded in a systematic review of the literature.","section":"Section 3.1, Applications"},{"comment":"Two references are placeholders rather than citable works: reference [50] is listed as 'Knight. Imitation.' and reference [62] as 'John Smith and Jane Doe. Title of the study.' Both are cited in Section 1.2.4 to support substantive claims about imitation learning and prospective clinical validation. A survey whose reference list contains non-existent entries cannot support the assertion that it accurately represents prior work.","section":"References [50] and [62]"}],"minor_comments":[{"comment":"The abstract capitalizes 'Kernel' in a way that is inconsistent with scientific prose, and the phrase 'MIMICs datasets comprises components' in Section 1.1 is grammatically incorrect.","section":"Abstract and Section 1"},{"comment":"Figure 1 is labeled 'The overall framework of our survey' but displays only a generic dataset icon and no actual framework; it should be replaced with a real diagram or removed.","section":"Figure 1"},{"comment":"The phrase 'impleques' appears in the bullet 'The combination of simpleques with traditional classifiers'; this appears to be a corrupted word and should be corrected.","section":"Section 1.2.2"},{"comment":"Tables 2, 3, and 4 largely duplicate the same information in different formats; consolidating them would reduce redundancy and clarify the presentation.","section":"Tables 2-4"},{"comment":"The conclusion of Section 1.3 repeats the same sentence about interoperability improvements twice ('By advancing these taxonomy and classification initiatives' and 'By promoting the evolution of taxonomy'); one copy should be deleted.","section":"Section 1.3"},{"comment":"The BLUE benchmark description is accurate in isolation, but it is not clear how it relates to the MIMIC dataset; the section would benefit from an explicit link to MIMIC if one exists.","section":"Section 3.1.1"}],"recommendation":"reject","confidential_remarks":"The number and severity of citation errors, placeholder references, unsupported performance tables, and off-topic content make the manuscript unsuitable for publication in its current form. Even a thorough revision would need to rebuild the reference base and re-verify every claim, which is beyond the scope of a standard major revision. The paper also appears to have been assembled with low editorial care, including duplicated paragraphs and corrupted sentences."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You asked about the MIMIC survey. My take: the topic is right, the execution is not. The paper collects many genuine themes in MIMIC-based research — missingness, irregular sampling, code mapping, reproducibility, privacy, fairness, federated learning — and its chapter structure (predictive modeling, NLP, temporal mining, multimodal analysis, downstream tasks) roughly matches how the field organizes itself. If the citations were sound, this could serve as a useful orientation for newcomers. That is the extent of the good news.\n\nThe problems are load-bearing and visible on a first read. The reference list contains entries that cannot support the claims attached to them: a peptide-design paper ([1]) is cited for digital health policy; a CERN technical report ([10]) is cited for unstructured clinical notes in MIMIC; an ETL/FHIR-OMOP loading paper ([13]) is cited for MIMIC-IV release statistics. Reference [50] is \"Knight. Imitation.\" and [62] is \"John Smith and Jane Doe. Title of the study.\" — placeholders, not citable works. For a survey, the references are the content. When the attribution layer is this unreliable, claims about what progress has been made and which open problems persist become unverifiable, and the central claim of being a \"comprehensive survey\" fails on its own terms.\n\nCompounding the citation failures: Table 6 lists precise AUROC values for \"best models\" on mortality, sepsis, and phenotyping with no source, and Tables 7–8 present performance ranges that cannot be traced. Section 1.2.2 contains duplicated bullets and garbled phrases (\"1A20 000\", \"bassolid undsoliding\"), and several subsections are simply off-topic — a drug-efficacy subsection on mindfulness-based interventions, a network-data-analytics survey, and a global air-quality dataset in a MIMIC review. These are not stylistic quibbles; they indicate the manuscript was not assembled carefully enough to merit referee time. The limitations section concedes potential bias, which is honest, but the defects go well beyond what that concession covers.\n\nI agree with the reject verdict, and I do not think the stress-test overreached. This is a desk-reject-with-invitation-to-repair situation, not a revise-and-resubmit. The authors should rebuild the reference list from scratch, re-verify every in-text citation against its claim, source or delete every performance table, remove the off-topic sections, and copy-edit the whole text. If they do that, a genuinely comprehensive MIMIC survey would be a valuable addition to the literature. As it stands, I would not cite it and would not bring it to a reading group except as a cautionary example.","headline":"A MIMIC survey on the right topic with a broken citation base—placeholder references, mismatched cites, and unsourced performance tables—so it cannot be trusted as a map of the field.","tokens_in":45319,"tokens_out":5056,"would_cite":false,"duration_ms":46572,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey maps the open problems keeping MIMIC-trained health models from the clinic — data granularity, coding heterogeneity, quality gaps, weak interoperability, and privacy limits — along with the progress and directions that could…","keywords":["MIMIC-III","MIMIC-IV","electronic health records","digital health","critical care","open problems","interoperability","reproducibility"],"falsifier":"Look up the survey's load-bearing citations and compare each to its in-text use: if a substantial share of a sampled list do not exist or do not support the claims they are cited for — starting with the peptide-design paper cited on health policy and the particle-physics report cited on clinical notes — then the survey's map of the field cannot be relied upon; if the mismatches turn out to be isolated, the map stands. A complementary check is to take five published MIMIC mortality-prediction studies and verify that the survey's reported AUROC ranges and benchmark tables match the primary sources.","tokens_in":44174,"feed_emoji":"🏥","tokens_out":16536,"duration_ms":141400,"temperature":0.7,"pith_summary":"What this paper is trying to establish: a structured, up-to-date map of the open problems in research that uses the MIMIC critical-care datasets, so that a reader can see where MIMIC-based machine learning actually falls short and what has been done about it. The paper's contention is that the binding constraints on MIMIC-based digital health are structural rather than algorithmic — data granularity and cardinality that are small next to the feature space, heterogeneous coding schemes that resist standardization, data-quality gaps from missingness and label noise, weak interoperability with other hospital records, reproducibility failures, and unresolved privacy and ethics constraints. Against those problems it catalogues demonstrated progress in dimensionality reduction, missingness-aware temporal modeling, causal inference, and privacy-preserving analytics, and it argues that hybrid modeling, federated learning, standardized preprocessing pipelines, and shared benchmarks are the directions most likely to help. A reader should care because these open problems are precisely what separates high benchmark scores from models that work in real hospitals.","feed_headline":"Survey maps what blocks MIMIC health models from the clinic","feed_subtitle":"One taxonomy of the data problems — granularity, coding, quality, interoperability — plus the field's proposed fixes.","key_machinery":"The central object is the MIMIC dataset series itself — deidentified records of intensive-care admissions spanning structured vitals and labs, unstructured clinical notes, waveforms, and chest images — and the carrying machinery is the survey's taxonomy: two research domains (traditional clinical and computational applications versus novel data-mining approaches) crossed with seven named open-problem categories. That grid does the argument's work by converting scattered complaints about MIMIC into fixed problem families, against which the paper places demonstrated progress (dimensionality reduction, missingness-aware temporal models such as GRU-D, causal-inference methods, differential privacy and federated learning) and future directions (hybrid models, standardized extraction pipelines, shared benchmarks). Named interoperability standards carry the fix the paper believes in most: FHIR, a standard for exchanging health records, and OMOP-CDM, a common data model that aligns different hospital schemas, together with named reproducibility tools such as the MIMIC Code Repository of shared analysis scripts.","core_discovery":"On its own terms, the survey's central claim is a classification: MIMIC research divides into traditional clinical and computational applications on one side and novel data-mining approaches on the other, and the obstacles to progress divide into seven named problem families — data granularity and cardinality limits, the effectiveness of predictive approaches under class imbalance and cohort bias, data integration and preprocessing, data-quality issues, interoperability and accessibility, reproducibility and productivity, and privacy, ethical, and legal constraints. The paper asserts that high-granularity, high-cardinality MIMIC data suffer the curse of dimensionality relative to only tens of thousands of ICU stays; that models reporting AUC above 0.90 on MIMIC frequently fail to replicate on independent data; that missingness-aware architectures such as GRU-D and preprocessing pipelines such as MIMIC-Extract constitute the clearest demonstrated progress; and that the field's promise lies in hybrid modeling, federated learning, FHIR/OMOP harmonization, and community benchmarking. The survey presents itself as the first review focused specifically on these open problems rather than on predictive performance alone.","pith_inferences":["Going beyond the survey: if its problem map is correct, then research effort on MIMIC should shift toward data infrastructure and validation infrastructure rather than incremental model improvements; one testable consequence is that studies adopting standardized preprocessing pipelines should show measurably higher cross-institution reproducibility.","Going beyond the survey: the seven problem families are not independent — heterogeneous coding is a root cause that feeds data quality, interoperability, and reproducibility problems at once — so solutions at the coding and vocabulary layer may cascade benefits across the taxonomy.","Going beyond the survey: because the manuscript's own reference list includes citations that do not match their in-text use (a peptide-design paper cited for digital-health policy claims and a particle-physics report cited for clinical notes), a reader should treat the survey's citations as a starting point rather than a verified map; a systematic citation audit of MIMIC review papers would measur"],"forward_implications":["If the survey's map is right, the next gains in MIMIC-based digital health will come more from standardization than from new network architectures: unified extraction toolkits, automated mapping of local codes to standard vocabularies, and FHIR/OMOP harmonization are the named levers.","Reported AUROC figures on MIMIC should not be read as clinical readiness, since the survey's cited evidence shows performance drops of roughly 5 to 12 percent under external validation; prospective multi-center validation becomes a precondition rather than an optional extra.","Missingness-aware modeling and fairness or calibration metrics should become standard parts of benchmarking suites, because irregular sampling and demographic bias are identified as persistent failure modes even in state-of-the-art models.","Shared, version-controlled analysis code is credited with cutting cohort-size variation in replication studies from over 25 percent to under 5 percent, which makes code sharing a load-bearing part of the field's credibility rather than a courtesy."],"supporting_citations":[{"why":"The MIMIC-III data descriptor: the dataset definition the survey's problem taxonomy is built around, supplying the granularity, cardinality, and missingness concerns.","marker":"[9]"},{"why":"The MIMIC-III first-release description: supplies the scale figures (40,000 ICU stays, minute-level physiology) that anchor the open-problem discussion.","marker":"[12]"},{"why":"The Deep EHR survey of deep learning for electronic health records: the foundational review the paper positions itself against and extends.","marker":"[22]"},{"why":"The MIMIC-Extract preprocessing pipeline: the main cited progress in turning raw MIMIC tables into analysis-ready features.","marker":"[38]"},{"why":"The sepsis treatment-policy study that founded imitation and inverse-reinforcement learning on MIMIC: load-bearing for the decision-support applications section.","marker":"[51]"},{"why":"The GRU-D missingness-aware recurrent architecture: the central cited fix for irregular sampling and missing values.","marker":"[66]"},{"why":"The standardized benchmark suite for MIMIC-III tasks (mortality, decompensation, length of stay, phenotyping): load-bearing for the benchmarking and evaluation sections.","marker":"[72]"},{"why":"The MIMIC Code Repository of shared extraction and analysis scripts: the main cited mechanism for reproducibility and cohort consistency.","marker":"[162]"}],"fun_headline_variants":["Seven bottlenecks keeping MIMIC models out of the clinic","MIMIC's data hurdles: why ICU models still fail at the bedside","Curse of dimensionality: MIMIC survey names 7 barriers","Beyond AUC: MIMIC survey lists real-world blocks to deployment","Seven unsolved data problems keep MIMIC models from the clinic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole review rests on one load-bearing premise: the work it cites exists and supports the statements attributed to it, so that the survey's map of open problems and progress is trustworthy; that premise is already strained in the manuscript itself, which cites a peptide-design paper for claims about health policy and a particle-physics technical report for claims about clinical notes.","fun_headline_variants_meta":{"raw":{"variants":["Seven bottlenecks keeping MIMIC models out of the clinic","MIMIC's data hurdles: why ICU models still fail at the bedside","Curse of dimensionality: MIMIC survey names 7 barriers","Beyond AUC: MIMIC survey lists real-world blocks to deployment","Seven unsolved data problems keep MIMIC models from the clinic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000693,"raw_usage":{"total_tokens":3142,"prompt_tokens":958,"completion_tokens":2184,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":2093}},"tokens_in":574,"tokens_out":2184,"duration_ms":13753,"temperature":1.0,"reasoning_tokens":2093,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:06:05.878586+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Look up the survey's load-bearing citations and compare each to its in-text use: if a substantial share of a sampled list do not exist or do not support the claims they are cited for — starting with the peptide-design paper cited on health policy and the particle-physics report cited on clinical notes — then the survey's map of the field cannot be relied upon; if the mismatches turn out to be isolated, the map stands. A complementary check is to take five published MIMIC mortality-prediction studies and verify that the survey's reported AUROC ranges and benchmark tables match the primary sources.","supporting_citations":[],"review_version":1}