REVIEW 2 major objections 7 minor 1 references
Compliant Self Service Access to Secondary Use Clinical Data at Stanford Medicine
T0 review · 2 major / 7 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Stanford's IRB-approved portal gives researchers self-service clinical data
desk verdict A useful operational blueprint for self-service clinical data access; the categorical 'compliant' claim needs tempering and revalidation details, but it deserves peer review as a resource description. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery that carries the compliance argument is the integration between Stanford's eProtocol system and the Data Privacy Attestation (DPA), a REDCap survey prefilled with protocol information and requiring the researcher to attest to data-use statements. The survey can only be launched from inside eProtocol; when submitted, it is copied back into the protocol document and reviewed by the University Privacy Office before IRB approval. Around this core sit the SquareTable data model (the surviving STRIDE database used by both query tools), the identity, anonymization, and compliance APIs, and the cloud hosting split between BigQuery for clinical data and PostgreSQL for transactional data.
What would settle it
A compliance audit that found a single chart-review session or data download occurring after the associated IRB approval lapsed or without a saved Data Privacy Attestation, while the compliance API still returned valid status, would show the gatekeeping does not actually enforce the framework the paper describes.
Extended reading notes
Core claim
The paper's central claim is that the sixteen-year-old STRIDE data warehouse has evolved into a working, policy-controlled self-service platform rather than a broker-mediated data request service. The key move was to shift the burden of compliance from individual data extracts to an institutional pipeline: a single IRB-approved STARR protocol with a waiver of consent and HIPAA authorization supplies the legal basis for transferring clinical data into the repository, and a Data Privacy Attestation embedded in eProtocol defines exactly what protected health information each research project may use. Once that attestation is approved, the researcher can search the SquareTable data model with the Cohort Discovery Tool, review patient charts, and download data in either SquareTable or OMOP format. The paper also describes the supporting infrastructure: a unified Epic Clarity ETL into Google BigQuery, an anonymization API that exchanges MRNs for stable pseudo-identifiers and shifts dates by a study-specific random offset of at most 30 days, a compliance API that checks IRB status in real time, and a cloud security perimeter with audit logging.
Load-bearing premise
The load-bearing premise is that the combination of an IRB waiver of consent and HIPAA authorization, the researcher-completed Data Privacy Attestation, and the Privacy Office's review of that attestation actually satisfies all applicable regulations for self-service access to protected health information; if that legal assumption is wrong, or if review of self-reported attestations lets unauthorized use through, the paper's central claim of compliant access collapses.
Editorial extensions
If this is right
- Researchers can iterate on cohort definitions and inspect charts without waiting for a data team, shortening study feasibility work.
- The compliance pipeline ties every chart-review session to a validated Data Privacy Attestation and IRB protocol, making self-service PHI access auditable at the protocol level.
- OMOP-format download lets a self-served cohort move directly into standard observational research workflows and into approved computational environments.
- The date-jitter scheme bounded by plus-or-minus 30 days preserves within-patient temporal ordering while obscuring exact dates, which supports behavioral studies that need time-of-year information.
- Cloud-based reprocessing allows data-model corrections and new mappings to be applied to the entire historical dataset in hours rather than weeks.
Reading between the lines
- A natural extension the authors do not develop is that the same eProtocol-to-DPA pattern could be reused by other institutions, but only if their IRB and privacy offices accept self-attestation as a sufficient control; the paper provides no independent legal opinion.
- Because compliance depends on a one-time attestation plus review of self-reported answers, the framework's real-world safety would be tested by measuring what happens after approval, for instance whether audit logs or Data Risk Assessment reviews ever catch unauthorized downloads, a statistic the paper does not report.
- The 30-day date-shift policy is a deliberately visible trade-off: it preserves seasonal patterns for behavioral research but weakens protection against re-identification by jitter, so the same approach would need re-evaluation if the data included especially rare diagnoses or small demographic cells.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript describes the STARR ecosystem and its self-service research tools at Stanford Medicine, including the Cohort Discovery Tool, the Chart Review Tool, the data download tool, and supporting APIs for identity, anonymization, and compliance verification. It recounts the technical evolution from STRIDE's EAV model to the in-house Square Table model and the migration from on-premise Oracle to Google Cloud Platform with BigQuery and PostgreSQL. The Regulatory Compliance section presents an institutional governance framework: an IRB-approved protocol with waiver of consent and HIPAA authorization, a Data Privacy Attestation (DPA) initiated from within eProtocol and reviewed by Stanford's University Privacy Office, and compliance API checks. The Results section reports growth in users, projects, and median cohort sizes through November 2024. The abstract and discussion assert that self-service access to detailed clinical data is handled in compliance with all applicable regulations and rules.
Significance. If accurate, this is a useful resource description of a mature production system. The paper's strengths are its concrete architectural detail (cloud migration, dual analytical/transactional data stores, API design, CI/CD security practices) and its transparency about residual data risk, such as the explicit statement that PHI-scrubbed data is not deemed de-identified by the Privacy Office. The usage counts indicate real adoption. However, the central compliance claim is supported only by internal institutional processes; no external audit, certification, or independent legal analysis is presented, and the enforcement timing of the compliance checks is not fully specified. No code or data are shipped, so verification rests on the textual description; this is understandable given PHI constraints, but it means the paper's assertions cannot be independently checked. The descriptive content is valuable, but the headline claim as worded exceeds the evidence presented.
major comments (2)
- [Features of Chart ReviewTool; Regulatory Compliance] The enforcement of the compliance framework is described as a one-time event. In the Regulatory Compliance section the manuscript states that 'to review charts online, the researcher must have supplied a currently valid DPA when saving the cohort for review.' This places the check at cohort-save time. The paper does not state that the system re-validates the DPA, the eProtocol IRB approval, or the researcher's status on later logins, chart views, data downloads, or after a DPA rejection, a protocol expiration, or a change in researcher status. The compliance verification API is mentioned as verifying IRB status and researcher status, but the trigger conditions and failure behavior are unspecified. If a protocol lapses or a DPA is revoked after the cohort is saved, the current text gives no assurance that previously saved cohorts become inaccessible. Because the title and abstract claim 'compliant' self-service access, this continuous-enforcement gap is load-bearing; please document the re-validation mechanism and its failure modes, or narrow the claim to what the system actually enforces.
- [Abstract; Discussion] The abstract states that data acquired via the self-service tools is 'handled in compliance with all applicable regulations and rules,' and the Discussion calls the system 'HIPAA-compliant self-service access.' The evidence offered is entirely internal: Stanford's IRB protocol, UPO review of DPAs, and RCO oversight. No external audit, certification, or independent legal opinion is presented, and no supporting data such as audit-log analyses, privacy-incident counts, or verification test results are included. For a resource-description paper this may be acceptable if the authors clearly mark the claim as 'designed to comply' or 'governed by Stanford's compliance framework'; as written, the manuscript asserts a legal conclusion it does not substantiate. Please either add the missing evidence if it exists or qualify the assertion in the title, abstract, and discussion.
minor comments (7)
- [Methods, first paragraph] The term 'MasterPersonIndex' should be written as 'Master Person Index' or 'Master Patient Index' with the standard abbreviation MPI defined at first use.
- [APIs section] The phrase 'that is inactiveuseby the current vendor' is garbled; it should read 'that is in active use by the current vendor.'
- [References] The reference labeled [Harris2000] is dated 2008 in the reference list (J Biomed Inform, published online 2008); the citation label in the text should be aligned with the actual publication year, e.g., [Harris2009].
- [References] The reference [Lowe2010]] contains a stray double closing bracket and should be corrected to [Lowe2010].
- [Regulatory Compliance] The phrase 'specific of the datatype' should read 'specifics of the data type.'
- [DevSecOps] The phrase 'demonstrate compliance with HIPAA export controls' is ambiguous; HIPAA does not regulate export controls. If the intended meaning is HIPAA compliance and export-control compliance, it should be written explicitly as such.
- [Results] Figures 7-9 would be easier to assess if at least one or two sentences in the text gave the actual numbers or described the trend, rather than leaving the reader to read values off the plots.
Circularity Check
No circularity: the paper is a descriptive system/resource report with no derivation chain that reduces to its own inputs.
full rationale
The manuscript is a research resource description of STARR Tools; it contains no equations, no fitted parameters, and no predictive claims that could be equivalent by construction to its inputs. The central statements are institutional and descriptive: the existence of the Cohort Discovery Tool, Chart Review Tool, data download tool, APIs, and the compliance workflow involving eProtocol and the Data Privacy Attestation. Citations to earlier STRIDE/STARR papers (Lowe 2009/2010, Callahan 2023, Datta 2020) describe the same institutional lineage, but the present paper does not rest a derived conclusion on them; the compliance claim is an assertion about Stanford's IRB/UPO/RCO process, not a result inferred from a self-citation. The reviewer's concern that the paper does not prove continuous re-validation of DPAs after cohort save is a limitation of evidence, not a circularity. Accordingly no circular step is present and the score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Stanford IRB approval with a waiver of informed consent and HIPAA authorization permits transfer of all routine-care clinical data to STARR and its self-service tools.
- domain assumption The Data Privacy Attestation completed by researchers, reviewed by the Privacy Office, and embedded in eProtocol is sufficient to prevent inappropriate uses of PHI.
- domain assumption The nightly Epic Clarity ETL and master patient index produce a single coherent dataset across the two Stanford hospitals.
- domain assumption Stanford's date-shifting policy, a stable random offset between -30 and +30 days, preserves temporal relations while adequately masking identifying dates.
Cite this review
Pith. "Pith review of Compliant Self Service Access to Secondary Use Clinical Data at Stanford Medicine." pith.science (2026). https://pith.science/paper/LIXMSKQS
@misc{pith2026241204248,
author = {Pith},
title = {Pith review of: Compliant Self Service Access to Secondary Use Clinical Data at Stanford Medicine},
year = {2026},
howpublished = {\url{https://pith.science/paper/LIXMSKQS}},
note = {Machine review of arXiv:2412.04248}
}
read the original abstract
STARR (STAnford Research Repository) is a clinical research support ecosystem that supports basic science research, population health research and translational research at Stanford University. STARR consists of raw and analysis ready multi-modal data, and tools for cohort analysis and self service data access. STARR data is accessible on secure shared computing systems for ad hoc analysis. Also present is a suite of services on top of STARR, that allow researchers access to complex purpose built data cuts, common data models and software solutions. This manuscript is a research resource description and describes the evolution of STARR Tools that are used to offer self-service access to detailed clinical data for research purposes to researchers at Stanford Medicine, along with a framework used to ensure that data acquired via the self-service tools is handled in compliance with all applicable regulations and rules.
Reference graph
Works this paper leans on
-
[1]
● [Brandt2002]CABrandt,RMorse,KMatthews,KSun,AMDeshpande,RGadagkar,etal.Metadata-Drivencreationofdatamartsfromaneav-modeledclinical researchdatabase.IntJMedInform.2002Nov12;65(3):225–41.[PubMed][GoogleScholar]● [Callahan2023]ACallahan,EAshley,SDatta,PDesai,TAFerris,JAFries,MHalaas,CPLanglotz,SMackey,JDPosada,MAPfeffer,NHShah.TheStanfordMedicinedatascience...
arXiv 2020
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.