{"id":"b65fa02a-540f-497e-823e-0e4f4c1c7acd","arxiv_id":"2608.08806","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Formalizes the Actionable Clinical Record, a source-grounded representation of patient-specific clinical intent, as the atomic object of a proposed third computational layer of healthcare IT.","lead":"The paper argues that healthcare IT should be organized by what becomes computable, and proposes a third layer that turns a doctor's intended actions, such as a follow-up in two weeks, into structured, machine-readable records. It offers a clear vocabulary for why clinical follow-through fails and what standards and evaluation criteria a 'computable care process' would need.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on an undefined equivalence between a recovered ACR and the clinician's intent, which is asserted but never formally specified or tested.","rationale":"The reader's weakest assumption concerns whether inference can reliably populate ACR attributes from real clinical text. I agree that feasibility is unproven, but the more load-bearing issue is that the paper never defines what it means for a recovered ACR to be equivalent to the clinician's intent. Definition 1 provides a schema, not a correctness criterion. The evaluation framework measures component extraction and temporal normalization, but none of its metrics checks whether executing the ACR would fulfill the original instruction. The paper's own falsifiability claims in §8 are framed as redundancy claims (whether existing standards could do the same), not as equivalence checks. The companion study is controlled and synthetic, so it cannot resolve this. Because the gap is both under-specified and unverified, I keep the reader's CONDITIONAL verdict: the framework is coherent, but the central claim rests on an equivalence relation that is neither formally defined nor demonstrated on authentic data. The reader identified a related but narrower concern; my concern is one level deeper, hence 'partial' agreement.","tokens_in":8231,"tokens_out":1108,"duration_ms":13350,"concrete_test":"Run a small equivalence-check study: take 100 real-world clinical notes, recover ACRs with the proposed hybrid pipeline, then present each recovered ACR to the original clinician and to a second expert panel, asking whether executing the ACR as written would satisfy the clinician's intent. Report agreement rates, and also compare the pipeline's scheduled actions with actions actually performed in the record to see whether loop closure improves. If agreement is high and loop closure is measurably better than baseline, the equivalence assumption survives; if not, the central claim is undercut.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 5's central claim is that recovering ACRs from communication makes patient-specific clinical intent computable. This requires that the recovered ACR—after normalizing time, resolving conditions, and populating actor—captures exactly what the clinician intended, so that executing the ACR fulfills that intent. No such equivalence or verification criterion is defined anywhere in the paper; Definition 1 in §5.1 specifies attributes but not what constitutes a faithful recovery. The evaluation framework in §7 measures actions, arguments, time normalization, and source provenance, but none of these measures establishes that executing the recovered ACR would satisfy the original intent. The companion feasibility study [51] addresses one narrow task, follow-up-instruction extraction, on a controlled benchmark, so it does not test real-world equivalence either. The paper itself flags moving from synthetic corpora to real-world notes as the central empirical risk, but the logical gap precedes that: without a definition or test of 'recovered ACR equals intended action,' the executable-correctness framework cannot measure the property it claims to assess. This is not merely a missing experiment; it is an unspecified relationship at the core of the third layer.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This Perspective paper proposes an organizing framework for healthcare IT based on the unit of information made computable, distinguishing three layers: clinical record (Layer 1), clinical state (Layer 2), and clinical intent (Layer 3, proposed). It introduces the Actionable Clinical Record (ACR) as the atomic object of Layer 3, defined as a nine-attribute tuple (action, target, actor, temporal constraint, condition, dependency, status, provenance, confidence), and describes a readiness ladder from mentioned to executable, a hybrid architecture of learned extraction with deterministic reasoning, and an evaluation framework centered on executable correctness rather than text overlap. The paper is explicitly a conceptual contribution, not an empirical study; the companion feasibility study [51] addresses one narrow subproblem. The authors position the proposal as a research program with falsifiable predictions, including tests of whether the ACR representation is redundant if direct FHIR mapping suffices.","tokens_in":8452,"tokens_out":5889,"duration_ms":55550,"significance":"The paper's main contribution is conceptual: it offers a crisp vocabulary for the problem of computable clinical intent and introduces constructs—the ACR and the executable-correctness evaluation framework—that could be reused and extended by the community. The explicit falsifiability conditions in Section 8 are a notable strength, distinguishing this proposal from unfalsifiable visions. The readiness ladder (mentioned → interpreted → actionable → executable) usefully separates representational status from operational lifecycle. The companion feasibility study, although narrow, provides an existence proof for one subproblem. If the framework gains traction, it could help align clinical NLP research with downstream workflow execution, potentially reducing loop-closure failures that the paper documents. These strengths are real, and the stated scope is appropriately careful.","major_comments":[{"comment":"The central claim that Layer 3 makes patient-specific clinical intent computable requires a precise criterion for when a recovered ACR faithfully represents the clinician's intended action. Definition 1 specifies the tuple's attributes but does not define what constitutes faithful recovery; the evaluation framework in Section 7 measures action detection, argument extraction, time normalization, and provenance, but none of these measures establishes that executing the recovered ACR would satisfy the original intent. Please add an explicit equivalence criterion—for example, an ACR is correct for a source communication if executing it under the relevant workflow semantics produces the effect the communicator intended—and specify how the proposed evaluation benchmarks are annotated to reflect that criterion (e.g., gold-standard ACRs built by clinicians with adjudication). Without such a criterion, the term 'executable correctness' remains ambiguous and the framework cannot be used to falsify the core claim.","section":"Section 5.1 and Section 7"},{"comment":"The phrase 'populated when communicated or explicitly inferred' is underspecified. It does not state what kinds of inference are permissible or what evidence is required for an attribute to be 'explicitly inferred' rather than guessed. For example, for the actor attribute, can the responsible performer be inferred from the speaker role, from institutional conventions, or only when explicitly mentioned? This boundary directly affects the design of the recovery pipeline and the meaning of the confidence map, which is described only as 'calibrated recovery-system uncertainty.' Please provide an explicit taxonomy of inference types (e.g., from co-reference, from semantic role, from institutional knowledge) and state how calibration is achieved or tested. Without this, the operational semantics of the ACR--and the reader's ability to assess the feasibility evidence--are incomplete.","section":"Section 5.1, Definition 1"}],"minor_comments":[{"comment":"The 'Class of computation' for Layer 3 in Table 1 is 'Schedule, coordinate, monitor, close,' while Definition 1 in Section 5.1 states that the ACR's attributes are 'required to interpret, execute, monitor, or audit' the action. Please align the wording across these two places so the markers are consistent.","section":"Section 5.1 vs. Table 1"},{"comment":"Reference [49] lists the venue as 'Pac Symp Biocomput 2026;31:144–157' but the DOI (10.1101/2025.08.14.25332837) points to a 2025 preprint server. Please provide the correct publication DOI or update the reference to note the preprint.","section":"Reference 49"},{"comment":"Section 8 states that 'moving from synthetic corpora to de-identified real-world notes is the central empirical risk,' but Section 7 describes the companion study [51] only as a 'controlled benchmark.' It is unclear whether that benchmark uses synthetic corpora or real-world notes; please clarify, as this affects the interpretation of the risk statement.","section":"Sections 7–8"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-structured perspective, and the authors are appropriately modest about the evidence. The main risk is the undefined equivalence between a recovered ACR and the clinician's intended action; if the authors add a precise criterion and a corresponding annotation protocol, the paper would be a solid contribution to the clinical NLP and workflow communities. The self-citation to the companion study is disclosed and used only as an illustration, which is acceptable. I would not reject the paper, but the requested revisions are load-bearing for the central claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the bottom line: this is a perspectives piece that does what perspectives should do—it gives the field a cleaner way to talk about what becomes computable. The three-layer framing is more than a taxonomy: the prescribed/observed/intended process distinction in Section 4 is genuinely clarifying, and the ACR tuple plus the readiness ladder (mentioned → interpreted → actionable → executable) sharpens the gap between NLP extraction of action items and something a scheduler can act on. The executable-correctness evaluation framework in Section 7 is the most useful practical contribution; measuring action–time binding, provenance, and unsupported-action rates rather than text overlap addresses a real weakness in current shared tasks. The paper is also honestly scoped: it says outright that the companion feasibility study is one narrow subproblem and not evidence for the framework.\n\nThe soft spot is where the stress-test note lands. Section 5 claims the third layer makes clinical intent computable by recovering it from communication as a structured, executable representation. For that to mean anything, the recovered ACR must be faithful to what the clinician intended. The paper never defines that faithfulness—no equivalence relation, no verification criterion, no ground-truth procedure beyond evaluating individual attributes. The evaluation framework measures agreement on actions, arguments, and times, but not whether executing the recovered ACR would satisfy the original intent. That is a genuine gap, and it is a logical one, not just a missing experiment. It is partly a feature of any proposal at this stage; you cannot fully define correctness before building systems. But the paper could go further by proposing an operational definition—for example, a recovered ACR is correct if the clinician, shown the ACR, confirms it captures their intended action, with inter-rater agreement as the early metric. That would make the executable-correctness framework measurable from day one.\n\nTwo smaller notes. Definition 1 is a prose tuple; a bit more formal semantics around attribute types and temporal constraint normalization would help, though for a perspective this is acceptable. And the 'generation' label for a layer that is still hypothetical is a slight overstatement; the paper's own epistemic-status caveat covers it, but the title promises more than the authors themselves claim.\n\nCitation pattern is solid. The one self-citation is a companion preprint, clearly flagged as illustrative. No data or code to check, appropriately so for a perspective. The math is minimal and there are no fitted parameters, so the usual fitting-overfitting concerns do not apply.\n\nWho is this for? Clinical informatics researchers interested in workflow, follow-up failures, and the NLP-to-execution gap. It deserves a serious referee; it is a well-written proposal with a load-bearing gap that a good review can help close. I would accept for peer review and recommend the authors be asked to operationalize the correctness criterion.","headline":"Worth engaging: a clear conceptual framework for computable clinical intent, with a load-bearing gap about what counts as a faithful recovery of intent that a good review can help close.","tokens_in":8948,"tokens_out":2902,"would_cite":true,"duration_ms":27091,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Patient-specific clinical intent can be made computable by recovering intended actions from ordinary clinical communication as structured, executable records.","keywords":["electronic health records","clinical workflow","actionable clinical record","clinical intent","FHIR","computational layers","clinical communication","temporal reasoning"],"falsifier":"Take a set of de-identified outpatient notes with independently labeled intended actions, run the proposed hybrid recovery pipeline, and ask clinicians whether each output ACR is executable and faithful to the original instruction; if a large fraction of actions require correction or the deterministic time reasoning cannot populate temporal constraints, the paper's central claim about recovering executable intent is not supported.","tokens_in":8028,"feed_emoji":"🩺","tokens_out":7310,"duration_ms":67112,"temperature":0.7,"pith_summary":"Healthcare information technology, the paper argues, is best organized by the unit of information a system makes computable rather than by the technology it adopts. On that basis it derives three cumulative computational layers: the clinical record, the clinical state, and a proposed third layer of patient-specific clinical intent. The paper's central claim is that this third layer can be built by recovering intended actions from ordinary clinical communication and representing each action as an Actionable Clinical Record (ACR), a structured object that names the action, target, responsible person, timing, conditions, dependencies, status, provenance, and confidence. If the claim is correct, an instruction such as 'repeat the complete blood count in two weeks' or 'refer to cardiology if symptoms persist' could become a machine-readable task that a scheduler or results system can execute, monitor, and close. That would directly target well-documented loop-closure failures, such as missed test-result follow-up and specialist referrals that never reach a completed visit.","feed_headline":"A doctor's instruction becomes a machine-executable record","feed_subtitle":"The paper defines the Actionable Clinical Record, a structured object that could close missed follow-up and referral loops.","key_machinery":"The central object is the Actionable Clinical Record (ACR), a source-grounded representation of one patient-specific intended clinical action, defined as the tuple above. It carries the argument by giving recovery a precise target: an action is not fully recovered until it can be executed, monitored, and audited, not merely indexed. Supporting machinery includes the five markers that define a computational layer (a new atomic unit, an external driver, an enabling technology, a new class of computation, and a residual limitation), the distinction among prescribed, observed, and intended process, the readiness ladder from mentioned through interpreted and actionable to executable, and an evaluation framework built on executable-correctness measures rather than text overlap. The paper also specifies a hybrid architecture in which learned components extract candidate actions and arguments while deterministic components compute times, due dates, and dependency ordering, with selective prediction routing uncertain cases to clinicians.","core_discovery":"The paper's discovery, stated on its own terms, is that after a first generation made the clinical record computable and a second made the clinical fact computable, a third generation can make patient-specific clinical intent computable by recovering it from communication as a structured, executable representation. The atomic object of that layer is the ACR, a source-grounded tuple $ACR = \\langle \\text{action}, \\text{target}, \\text{actor}, \\text{temporal constraint}, \\text{condition}, \\text{dependency}, \\text{status}, \\text{provenance}, \\text{confidence} \\rangle$, in which action, status, and provenance are required and the remaining attributes are populated when communicated or explicitly inferred. The paper locates the gap that motivates the layer: existing workflow standards such as FHIR represent intended process once it has been structured, but they do not recover it from natural communication, and it is that recovery and conversion into executable form that the third layer supplies. The paper frames the third layer as a prospective hypothesis and analytic lens rather than an established periodization, and it separates representational readiness (mentioned, interpreted, actionable, executable) from the operational lifecycle captured by the status attribute.","pith_inferences":["The ACR schema is presented for follow-up instructions, but the paper names a broader class of communication-derived process objects; one extension is to develop the same tuple for conditional escalations, medication transitions, handoffs, and pending-result responsibility, where the dependency and condition attributes would do most of the work.","If ACR recovery matures, a natural testable extension is a randomized comparison of ACR-driven closed-loop task management against usual care on referral completion and test-result follow-up; the paper itself stops at executable-correctness evaluation, so this outcome study is an inference, not a claim.","The readiness ladder implies a staged adoption path: start with human confirmation of recovered actions, then automate only low-risk, well-calibrated actions; this could make regulatory approval incremental, though the paper does not say so.","Because the ACR separates speaker from responsible actor, it could support accountability and audit use; making that work would require solving role ambiguity, which the paper lists as open research rather than a solved problem."],"forward_implications":["A working third layer would let a follow-up instruction like 'repeat the complete blood count in two weeks' be scheduled automatically, with responsibility assigned and completion tracked, reducing missed-result and referral loop-closure failures.","ACRs are designed to sit upstream of existing workflow infrastructure: they map onto FHIR resources such as CarePlan, ServiceRequest, and Task, so the proposal does not require replacing current systems.","Evaluation of intent recovery would shift from text-overlap scores to executable-correctness measures, including action-time linking error, unsupported-action and omitted-action rates, and calibration of confidence.","The framework predicts that machine-actionable use of workflow resources will grow relative to free-text follow-up, that ambient documentation will extend from notes to tasks and orders, and that reliable systems will expose source-linked, auditable records with selective human review.","Reliability would depend on a hybrid design: language models propose candidate actions, deterministic reasoning resolves times and dependencies, and the system abstains or alerts a clinician on uncertain cases rather than acting automatically."],"supporting_citations":[{"why":"Supplies evidence that test-result follow-up often fails, motivating the loop-closure problem the third layer addresses.","marker":"[3]"},{"why":"Shows only about one-third of specialist referrals reached a completed visit, a concrete loop-closure failure.","marker":"[4]"},{"why":"Defines FHIR workflow resources that represent intended process once structured, the infrastructure ACRs map onto.","marker":"[23]"},{"why":"Supplies a shared-task setting for extracting medical orders from doctor-patient consultations, the nearest existing approach to ACR recovery from conversation.","marker":"[29]"},{"why":"Provides datasets and extractions of physician action items and medical decisions from notes, existing NLP capabilities the ACR seeks to make executable.","marker":"[37,38]"},{"why":"Documents weak general-purpose temporal reasoning in language models, motivating deterministic time resolution in the hybrid architecture.","marker":"[41]"},{"why":"Companion feasibility study demonstrating a hybrid neural-symbolic pipeline for follow-up-instruction extraction on a controlled benchmark.","marker":"[51]"}],"fun_headline_variants":["Clinical intent gets a machine-readable form","From records to executable care intentions","The Actionable Clinical Record makes intent computable","Healthcare's third IT generation: computable intent","Recovering clinical intent from natural communication"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the intended actions expressed in clinical communication contain enough explicit or reliably inferable detail about what should happen, when, under what condition, and by whom for a recovered ACR to be executable and to match what the clinician meant.","fun_headline_variants_meta":{"raw":{"variants":["Clinical intent gets a machine-readable form","From records to executable care intentions","The Actionable Clinical Record makes intent computable","Healthcare's third IT generation: computable intent","Recovering clinical intent from natural communication"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000677,"raw_usage":{"total_tokens":3085,"prompt_tokens":958,"completion_tokens":2127,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":2062}},"tokens_in":574,"tokens_out":2127,"duration_ms":17254,"temperature":1.0,"reasoning_tokens":2062,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:23:20.551457+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of de-identified outpatient notes with independently labeled intended actions, run the proposed hybrid recovery pipeline, and ask clinicians whether each output ACR is executable and faithful to the original instruction; if a large fraction of actions require correction or the deterministic time reasoning cannot populate temporal constraints, the paper's central claim about recovering executable intent is not supported.","supporting_citations":[{"cited_title":"Health Level Seven International; 2023","cited_arxiv_id":null,"evidence_quote":"Defines FHIR workflow resources that represent intended process once structured, the infrastructure ACRs map onto."},{"cited_title":"Overview of the MEDIQA-OE 2025 Shared Task on Medical Order Extraction from Doctor-Patient Consultations","cited_arxiv_id":null,"evidence_quote":"Supplies a shared-task setting for extracting medical orders from doctor-patient consultations, the nearest existing approach to ACR recovery from conversation."}],"review_version":1}