{"id":"61ae2b7d-5296-4030-9e1b-c515738ac5fc","arxiv_id":"2412.04248","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"STARR Tools is Stanford Medicine's self-service platform for compliant secondary-use access to clinical data, described here with its architecture, APIs, and regulatory workflow.","lead":"Stanford Medicine's STARR Tools give researchers self-service access to patient cohorts, charts, and downloadable clinical data, with a compliance framework linking IRB protocols to privacy attestations. The paper is a resource description of this internal system, so its value is as an institutional blueprint rather than a new scientific result.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The compliance claim hinges on a one-time DPA check at cohort-save time; the paper does not show that ChartReviewTool revalidates DPA or IRB status on later access, so continuous compliance is unproven.","rationale":"This read agrees with the reader's broad concern about the compliance framework, but sharpens it from 'no external audit or legal opinion' to a specific, internally testable gap: the only stated gate is a DPA check at cohort-save time, and the paper does not establish that access is revalidated after DPA or IRB status changes. The paper is otherwise internally consistent and gives credible architectural detail (cloud migration, APIs, DevSecOps, eProtocol integration), which supports that a working system exists. The weakness is not a dispute with external consensus; it is a missing control point in the described mechanism. Because the compliance claim is categorical, a single unenforced revocation path would undermine it. The proposed lifecycle test would settle whether the system actually denies access after accreditation is withdrawn.","tokens_in":8817,"tokens_out":5640,"duration_ms":61804,"concrete_test":"Run an authorization lifecycle test in the ChartReviewTool: (1) create a cohort under a valid DPA and grant chart-review access; (2) mark that DPA invalid in eProtocol and set the IRB protocol status to 'expired' via the compliance API integration in a test environment; (3) with the same researcher SSO, attempt to open the cohort, view a chart, and download data. Record whether each action is denied in real time and whether an audit entry is written. If any action succeeds after accreditation is revoked, the compliance claim needs to be scoped down or the enforcement code fixed.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that STARR Tools gives self-service PHI access that is 'handled in compliance with all applicable regulations and rules.' The only enforcement mechanism described in the Chart Review and Regulatory Compliance sections is a one-time Data Privacy Attestation: 'the researcher must have supplied a currently valid DPA when saving the cohort for review.' The paper never states that the system re-checks the DPA or the associated eProtocol IRB status on subsequent logins, chart views, file downloads, or after an annual renewal or a UPO rejection. The compliance verification API is mentioned ('STARR Tools uses the compliance APIs to verify IRB status, researcher's status and other information'), but the trigger conditions and failure behavior are unspecified. If a DPA is revoked, a protocol expires, or a researcher's status changes, previously saved cohorts could remain accessible, so the assurance that self-service access stays compliant is not established. The concern is not that the described process is fraudulent, but that the paper's own text does not demonstrate the continuous enforcement that the word 'compliance' requires.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript describes the STARR ecosystem and its self-service research tools at Stanford Medicine, including the Cohort Discovery Tool, the Chart Review Tool, the data download tool, and supporting APIs for identity, anonymization, and compliance verification. It recounts the technical evolution from STRIDE's EAV model to the in-house Square Table model and the migration from on-premise Oracle to Google Cloud Platform with BigQuery and PostgreSQL. The Regulatory Compliance section presents an institutional governance framework: an IRB-approved protocol with waiver of consent and HIPAA authorization, a Data Privacy Attestation (DPA) initiated from within eProtocol and reviewed by Stanford's University Privacy Office, and compliance API checks. The Results section reports growth in users, projects, and median cohort sizes through November 2024. The abstract and discussion assert that self-service access to detailed clinical data is handled in compliance with all applicable regulations and rules.","tokens_in":8970,"tokens_out":8333,"duration_ms":82814,"significance":"If accurate, this is a useful resource description of a mature production system. The paper's strengths are its concrete architectural detail (cloud migration, dual analytical/transactional data stores, API design, CI/CD security practices) and its transparency about residual data risk, such as the explicit statement that PHI-scrubbed data is not deemed de-identified by the Privacy Office. The usage counts indicate real adoption. However, the central compliance claim is supported only by internal institutional processes; no external audit, certification, or independent legal analysis is presented, and the enforcement timing of the compliance checks is not fully specified. No code or data are shipped, so verification rests on the textual description; this is understandable given PHI constraints, but it means the paper's assertions cannot be independently checked. The descriptive content is valuable, but the headline claim as worded exceeds the evidence presented.","major_comments":[{"comment":"The enforcement of the compliance framework is described as a one-time event. In the Regulatory Compliance section the manuscript states that 'to review charts online, the researcher must have supplied a currently valid DPA when saving the cohort for review.' This places the check at cohort-save time. The paper does not state that the system re-validates the DPA, the eProtocol IRB approval, or the researcher's status on later logins, chart views, data downloads, or after a DPA rejection, a protocol expiration, or a change in researcher status. The compliance verification API is mentioned as verifying IRB status and researcher status, but the trigger conditions and failure behavior are unspecified. If a protocol lapses or a DPA is revoked after the cohort is saved, the current text gives no assurance that previously saved cohorts become inaccessible. Because the title and abstract claim 'compliant' self-service access, this continuous-enforcement gap is load-bearing; please document the re-validation mechanism and its failure modes, or narrow the claim to what the system actually enforces.","section":"Features of Chart ReviewTool; Regulatory Compliance"},{"comment":"The abstract states that data acquired via the self-service tools is 'handled in compliance with all applicable regulations and rules,' and the Discussion calls the system 'HIPAA-compliant self-service access.' The evidence offered is entirely internal: Stanford's IRB protocol, UPO review of DPAs, and RCO oversight. No external audit, certification, or independent legal opinion is presented, and no supporting data such as audit-log analyses, privacy-incident counts, or verification test results are included. For a resource-description paper this may be acceptable if the authors clearly mark the claim as 'designed to comply' or 'governed by Stanford's compliance framework'; as written, the manuscript asserts a legal conclusion it does not substantiate. Please either add the missing evidence if it exists or qualify the assertion in the title, abstract, and discussion.","section":"Abstract; Discussion"}],"minor_comments":[{"comment":"The term 'MasterPersonIndex' should be written as 'Master Person Index' or 'Master Patient Index' with the standard abbreviation MPI defined at first use.","section":"Methods, first paragraph"},{"comment":"The phrase 'that is inactiveuseby the current vendor' is garbled; it should read 'that is in active use by the current vendor.'","section":"APIs section"},{"comment":"The reference labeled [Harris2000] is dated 2008 in the reference list (J Biomed Inform, published online 2008); the citation label in the text should be aligned with the actual publication year, e.g., [Harris2009].","section":"References"},{"comment":"The reference [Lowe2010]] contains a stray double closing bracket and should be corrected to [Lowe2010].","section":"References"},{"comment":"The phrase 'specific of the datatype' should read 'specifics of the data type.'","section":"Regulatory Compliance"},{"comment":"The phrase 'demonstrate compliance with HIPAA export controls' is ambiguous; HIPAA does not regulate export controls. If the intended meaning is HIPAA compliance and export-control compliance, it should be written explicitly as such.","section":"DevSecOps"},{"comment":"Figures 7-9 would be easier to assess if at least one or two sentences in the text gave the actual numbers or described the trend, rather than leaving the reader to read values off the plots.","section":"Results"}],"recommendation":"major_revision","confidential_remarks":"For the editor: this is a system/resource description authored by the team that built and operates the system. The main risk is that the word 'compliant' in the title and abstract converts a descriptive claim into a legal/regulatory assertion that the paper does not substantiate. Because the authors are the operators, independent verification is inherently difficult; the remedy is either to provide supporting evidence or to soften the claim. The continuous-enforcement issue is the most substantive technical concern and should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, detailed case study of how Stanford runs STARR Tools, not a research paper with new methods or evidence. If you work on clinical research informatics governance, it is worth reading; the compliance claim should be read as institutional self-description, not certified fact.\n\nWhat is actually new: the paper documents the current STARR Tools stack—EAV-to-square-table migration, Google Cloud Platform split between BigQuery and CloudSQL, the GWT web app, the compliance API integration with eProtocol/DPA, date jitter via the anonymization API, the OMOP download feature, and usage trends through November 2024. That operational detail is useful to other academic medical centers. The authors are appropriately open about the evolution from STRIDE/STARR, and the self-citations describe the same system, so I do not see a citation problem.\n\nWhere it is soft: the central claim is \"compliant with all applicable regulations and rules,\" but the evidence is entirely self-reported. No external audit, no legal opinion, no privacy-incident statistics. The stress-test concern lands. The chart review section says a researcher must have a currently valid DPA \"when saving the cohort for review,\" but the paper never says the system rechecks DPA or IRB status on later logins, chart views, downloads, or after protocol expiration or UPO rejection. If a DPA is revoked, previously saved cohorts could remain accessible. The compliance API exists and verifies IRB status, but the trigger conditions and failure behavior are unspecified. This is not evidence of misconduct; it is an incompleteness in a paper whose headline word is \"compliant.\" A short addition specifying revalidation triggers and what happens on failure would close most of the gap.\n\nAlso missing: quantitative baselines or shipped code/data. The usage charts are counts only. For a resource description, that is acceptable, but the reader should not expect reproducibility in the ordinary sense.\n\nBottom line: the reader's conditional verdict is fair. For clinical research informatics teams designing self-service access with governance, this is a useful blueprint and deserves a serious referee. I would send it to peer review and ask the authors to temper \"compliance\" to \"processes intended to achieve compliance\" and to document revalidation. I would not treat the compliance claim as established outside Stanford's own assessment.","headline":"A useful operational blueprint for self-service clinical data access; the categorical 'compliant' claim needs tempering and revalidation details, but it deserves peer review as a resource description.","tokens_in":9525,"tokens_out":4572,"would_cite":true,"duration_ms":42435,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Stanford's IRB-approved portal gives researchers self-service clinical data","keywords":["secondary use of clinical data","self-service data access","clinical data repository","cohort discovery tool","chart review","HIPAA compliance","OMOP common data model","data privacy attestation"],"falsifier":"A compliance audit that found a single chart-review session or data download occurring after the associated IRB approval lapsed or without a saved Data Privacy Attestation, while the compliance API still returned valid status, would show the gatekeeping does not actually enforce the framework the paper describes.","tokens_in":8608,"feed_emoji":"🩺","tokens_out":5331,"duration_ms":52587,"temperature":0.7,"pith_summary":"This paper is a resource description of STARR Tools, the self-service layer Stanford Medicine built on its STARR clinical data repository. It claims that researchers can now build patient cohorts, review full charts, and download data, including OMOP-formatted exports, without asking a data team for a custom extract, and that this access is handled in a way the institution treats as compliant with HIPAA and IRB rules. The compliance claim rests on a specific governance mechanism: STARR runs under an IRB-approved protocol with a waiver of consent and HIPAA authorization, and every researcher who opens chart review must first complete a Data Privacy Attestation launched from the eProtocol system and reviewed by Stanford's Privacy Office. If the framework works as described, it is a practical model for letting many researchers use detailed clinical data while keeping privacy oversight inside the research approval pipeline.","feed_headline":"Stanford's IRB-approved portal gives researchers self-service clinical data","feed_subtitle":"Cohort search, chart review, and OMOP downloads run through a privacy-attestation compliance gate.","key_machinery":"The machinery that carries the compliance argument is the integration between Stanford's eProtocol system and the Data Privacy Attestation (DPA), a REDCap survey prefilled with protocol information and requiring the researcher to attest to data-use statements. The survey can only be launched from inside eProtocol; when submitted, it is copied back into the protocol document and reviewed by the University Privacy Office before IRB approval. Around this core sit the SquareTable data model (the surviving STRIDE database used by both query tools), the identity, anonymization, and compliance APIs, and the cloud hosting split between BigQuery for clinical data and PostgreSQL for transactional data.","core_discovery":"The paper's central claim is that the sixteen-year-old STRIDE data warehouse has evolved into a working, policy-controlled self-service platform rather than a broker-mediated data request service. The key move was to shift the burden of compliance from individual data extracts to an institutional pipeline: a single IRB-approved STARR protocol with a waiver of consent and HIPAA authorization supplies the legal basis for transferring clinical data into the repository, and a Data Privacy Attestation embedded in eProtocol defines exactly what protected health information each research project may use. Once that attestation is approved, the researcher can search the SquareTable data model with the Cohort Discovery Tool, review patient charts, and download data in either SquareTable or OMOP format. The paper also describes the supporting infrastructure: a unified Epic Clarity ETL into Google BigQuery, an anonymization API that exchanges MRNs for stable pseudo-identifiers and shifts dates by a study-specific random offset of at most 30 days, a compliance API that checks IRB status in real time, and a cloud security perimeter with audit logging.","pith_inferences":["A natural extension the authors do not develop is that the same eProtocol-to-DPA pattern could be reused by other institutions, but only if their IRB and privacy offices accept self-attestation as a sufficient control; the paper provides no independent legal opinion.","Because compliance depends on a one-time attestation plus review of self-reported answers, the framework's real-world safety would be tested by measuring what happens after approval, for instance whether audit logs or Data Risk Assessment reviews ever catch unauthorized downloads, a statistic the paper does not report.","The 30-day date-shift policy is a deliberately visible trade-off: it preserves seasonal patterns for behavioral research but weakens protection against re-identification by jitter, so the same approach would need re-evaluation if the data included especially rare diagnoses or small demographic cells."],"forward_implications":["Researchers can iterate on cohort definitions and inspect charts without waiting for a data team, shortening study feasibility work.","The compliance pipeline ties every chart-review session to a validated Data Privacy Attestation and IRB protocol, making self-service PHI access auditable at the protocol level.","OMOP-format download lets a self-served cohort move directly into standard observational research workflows and into approved computational environments.","The date-jitter scheme bounded by plus-or-minus 30 days preserves within-patient temporal ordering while obscuring exact dates, which supports behavioral studies that need time-of-year information.","Cloud-based reprocessing allows data-model corrections and new mappings to be applied to the entire historical dataset in hours rather than weeks."],"supporting_citations":[{"why":"Describes the original STRIDE platform whose SquareTable database and query tools STARR Tools evolved from.","marker":"[Lowe2009]"},{"why":"Documents the original self-service cohort-identification and chart-review workflow that STARR Tools preserves and extends.","marker":"[Lowe2010]"},{"why":"Supplies the REDCap system used both as Stanford's research data management platform and as the host of the Data Privacy Attestation survey.","marker":"[Harris2018]"},{"why":"Provides the standard practice the anonymization API follows for preserving temporal relations while masking dates.","marker":"[Hripcsak2016]"},{"why":"Describes the broader Stanford Medicine data science ecosystem that includes STARR and the Nero computational environment.","marker":"[Callahan2023]"},{"why":"Introduces the OMOP common data model at Stanford Medicine, the format now available for self-service data download.","marker":"[Datta2020]"},{"why":"Describes the EAV modeling approach used by the original STRIDE database before its conversion to the SquareTable model.","marker":"[Brandt2002]"},{"why":"Explains the real-time HL7-based research alerting that remains the continuing use of HL7 messages after the switch to Epic Clarity.","marker":"[Weber2010]"}],"fun_headline_variants":["STARR: Self-service clinical data with built-in IRB compliance","Self-serve clinical data, gated by privacy attestation","Stanford's self-service clinical data is IRB-approved and ready","From IRB attestation to OMOP: self-service data pipeline","Stanford researchers self-serve clinical cohorts and charts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the combination of an IRB waiver of consent and HIPAA authorization, the researcher-completed Data Privacy Attestation, and the Privacy Office's review of that attestation actually satisfies all applicable regulations for self-service access to protected health information; if that legal assumption is wrong, or if review of self-reported attestations lets unauthorized use through, the paper's central claim of compliant access collapses.","fun_headline_variants_meta":{"raw":{"variants":["STARR: Self-service clinical data with built-in IRB compliance","Self-serve clinical data, gated by privacy attestation","Stanford's self-service clinical data is IRB-approved and ready","From IRB attestation to OMOP: self-service data pipeline","Stanford researchers self-serve clinical cohorts and charts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000825,"raw_usage":{"total_tokens":3580,"prompt_tokens":890,"completion_tokens":2690,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":2605}},"tokens_in":506,"tokens_out":2690,"duration_ms":21990,"temperature":1.0,"reasoning_tokens":2605,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:35:15.673502+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A compliance audit that found a single chart-review session or data download occurring after the associated IRB approval lapsed or without a saved Data Privacy Attestation, while the compliance API still returned valid status, would show the gatekeeping does not actually enforce the framework the paper describes.","supporting_citations":[],"review_version":1}