{"id":"6cbabc3f-1c2c-4756-a23d-765ed2f04a2f","arxiv_id":"2507.14859","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper defines a holistic, reusable specification for human digital twins based on six functionality levels and generic requirements.","lead":"This paper proposes a structured specification for a 'human digital twin', a virtual model of a person, organized around stakeholders, user groups, six functionality levels, and a list of generic requirements. It gives researchers and industry a common frame for designing human digital twins that can be reused across applications like healthcare, smart homes, and factory work.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table III's level assignments contradict Section IV's own classifications, so the validation of the six functionality levels is internally inconsistent.","rationale":"The reader's weakest assumption points to the self-generated nature of the application table and the validity of the levels. My concern is a specific, concrete instance of that same weakness: the paper's own Table III and Section IV contradict each other on two of the three detailed applications. This is an internal inconsistency, not just an external validation gap. The paper's validation logic in Section III-E says the table demonstrates that the levels cover all users and functions, but if the level assignments are unstable even for the authors' own examples, the abstraction is not well-defined enough to support the claimed 'holistic specification'. This does not change the overall verdict: the paper is clearly a brainstorming-based guideline proposal, and the requirements list in Table II can still be useful. The inconsistency is addressable by either correcting Table III or revising the level definitions to make classifications unambiguous. Therefore a conditional acceptance remains appropriate. I agree only partially with the reader because my objection is more precise and demonstrable from the text itself, rather than relying on the general risk of self-validation.","tokens_in":10867,"tokens_out":5154,"duration_ms":55709,"concrete_test":"Blindly recode the applications in Table III using only the definitions in §III-D, without peeking at the table's level column. Specifically, apply the definitions to 'smart home: room/device adjustment; environment control' and 'personalized therapy; medication'. If a coder following the definitions assigns Smart Home to Level IV and medication to Level III, as Section IV does, then Table III's Level II entries are misclassified and the table cannot serve as validation. Report the number of table cells whose level would shift under this recoding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the six functionality levels are a valid abstraction rests on Table III as evidence. But Table III assigns 'smart home: room/device adjustment; environment control' to Level II (Personalize), while Section IV-B and Table I explicitly classify Smart Home as Level IV (Control). Similarly, Table III places 'personalized therapy; medication' under Level II (Personalize), whereas Section IV-A classifies the medication decision-support system as Level III (Predict) because there is no immediate feedback loop. These are not edge cases: they are two of the three detailed application profiles meant to validate the framework. The paper even acknowledges the Smart Home level shift in Section IV-B but leaves Table III uncorrected. If the authors cannot consistently map their own generated applications onto the levels, the levels lack operational definitions, and the claimed validation ('we are able to identify sufficiently diverse applications for each level of functionality and user group, validating the model', §III-E) is not reproducible. The requirements list may still be useful, but the foundational abstraction is unsubstantiated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a conceptual specification for a holistic human digital twin (HDT). It identifies stakeholders and user groups, introduces six ordered functionality levels (store, analyze, personalize, predict, control, optimize), and compiles a table of applications for each user group and functionality level. Three applications are elaborated in profiles, and a list of generic requirements is derived. The authors intend the specification to serve as a guideline for implementing reusable, holistic HDTs across domains such as healthcare, smart homes, and factory productivity.","tokens_in":11028,"tokens_out":5748,"duration_ms":57254,"significance":"If the framework were soundly validated, this paper would provide a shared vocabulary and requirements baseline for the growing HDT community, and the functionality-level abstraction could help compare and compose HDT applications. The paper is clearly written, honest about its basis in a thought experiment, and usefully connects to standards (ISO/IEC 30173, GDPR) and prior HDT work. However, the evidence for the validity and completeness of the abstraction is currently insufficient: the classification of applications onto levels is internally inconsistent (see Major Comment 1) and circular (Major Comment 2). These issues undermine the central claim that the six levels are a validated abstraction. With substantial rework, the framework could become a useful contribution; as it stands, the specification is not yet reliable.","major_comments":[{"comment":"Table III assigns 'smart home: room/device adjustment; environment control' to Personalize (Level II) in the 'Me' row, but Section IV-B classifies Smart Home as Level IV (Control), and Table I labels it 'Level IV: Smart Home'. Likewise, Table III lists 'personalized therapy; medication' under Personalize (Level II) for the Psychologist column, while Section IV-A explicitly classifies the medication decision-support system as Level III (Predict) because there is no immediate feedback loop, and Table I labels it 'Level III'. Section IV-B even acknowledges that the scenario title suggests Level II while the device control corresponds to Level IV, yet Table III is not corrected. Since the claim in Section III-E that sufficiently diverse applications per level 'validate the model' rests on Table III, these contradictions mean the levels lack operational definitions and the validation is not reproducible. These are not edge cases; they are two of the three applications selected for detailed validation.","section":"Section III-E, Table III vs. Section IV-A/IV-B, Table I"},{"comment":"The functionality levels are derived from the authors' brainstorming of applications, and the applications in Section III-E are then generated to populate those same levels; the existence of such applications is used to 'validate' the model in Section III-E. This is a self-confirming procedure: any arbitrary level set could be populated by brainstorming fitting examples, so the evidence does not support the claim that the six levels constitute a valid or complete abstraction. To make the validation meaningful, the authors should provide a priori operational definitions with decision rules for assigning an application to a level, or use independent raters, or derive a falsifiable prediction (for example, that model count, interface complexity, or required latency increases monotonically with level). Without such a test, the central claim is unsubstantiated.","section":"Section III-D and III-E"},{"comment":"The derivation is explicitly based on a single structured brainstorming session with four experts (Section III-A), and the authors state that the user list 'can be extended at will' (Section III-C) and that applications are only exemplary. Nevertheless, the title and conclusion claim a 'holistic specification' and a 'comprehensive list of requirements' (Section VI). A vision paper may legitimately propose a framework rather than prove its completeness, but the burden is to justify that a four-expert session can yield comprehensive coverage. Without a systematic stakeholder-analysis method, a literature corpus analysis, or a saturation argument, the 'holistic' and 'comprehensive' terms overstate what is demonstrated. The authors should either temper these claims or provide a reproducible method for establishing completeness.","section":"Sections III-A, III-C, and VI"}],"minor_comments":[{"comment":"The text discusses 'Universal Design paradigm (F3)' and 'economic viability (F4)' as accessibility requirements, but Table II lists only F1 (Participation) and F2 (Human understandable I/O). Please add F3 and F4 to the table or remove the references; otherwise the table is not the 'comprehensive' list it claims to be.","section":"Section V, Table II"},{"comment":"The phrase 'save state' should read 'safe state'.","section":"Table II, item B2"},{"comment":"The word 'paramaterization' should be 'parameterization'.","section":"Table II, item D7"},{"comment":"The terms 'Brown Field' and 'Green Field' should be lowercased and defined on first use; 'brownfield' is the standard spelling.","section":"Sections II and V"}],"recommendation":"major_revision","confidential_remarks":"The paper is a conceptual/vision contribution with a useful requirements list and a plausible functionality abstraction, but the central validation is both internally inconsistent and circular. The required fixes (correcting Table III, operationalizing the level assignment, tempering the comprehensiveness claims) are substantial enough to warrant a major revision. The missing F3/F4 table entries are a symptom of the same need for careful consistency checking in the final deliverable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on human digital twin standards, but treat the middle with caution. What's actually new is a consolidated package: six functionality levels (store, analyze, personalize, predict, control, optimize), a stakeholder/user map, a large application table, and a generic requirements list (Table II). That last table is the most useful part — practical needs like zero obsolescence, model composition, digital sovereignty, and plannable death. The writing is clear and the method is honestly described: four experts, a thought experiment, a brainstorming session.\n\nThe soft spots are real. The validation is circular: applications are generated to fit the six levels, then used as evidence the levels are valid (Section III-E). That's a plausibility check, not validation. The completeness claim also isn't benchmarked against existing capability taxonomies, even though the authors cite DIKW as analogous. And there's an internal contradiction: Table III lists 'personalized therapy; medication' under Level II Personalize, while Section IV-A classifies the same scenario as Level III Predict because there is no feedback loop. That is a load-bearing example, not a typo. For what it's worth, the stress-test's smart-home example misses: Table III actually places smart home under Level IV Control, consistent with Section IV-B.\n\nThese issues don't kill the paper as a position piece. The requirements list and taxonomy are plausible starting points. But 'holistic specification' overstates what a four-person brainstorm can support. The authors could fix this by mapping their levels onto prior taxonomies and checking the application table against their own profiles.\n\nI'd send this to peer review at a venue that takes conceptual papers, with a strong request to resolve the contradiction and temper the validation language. A serious referee can push them to do that. I wouldn't cite it in my own work yet; I'd wait for the revision. The paper is for researchers and engineers starting in HDT who need a checklist of issues to think through — they'll get a good orientation, but they shouldn't mistake the six levels for an empirically grounded standard.","headline":"Useful HDT requirements checklist undermined by self-confirming validation and one internal mapping contradiction.","tokens_in":11542,"tokens_out":5573,"would_cite":false,"duration_ms":51731,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that every human digital twin can be classified by six cumulative functionality levels and that a generic requirement list gives a reusable implementation guideline.","keywords":["human digital twin","digital twin specification","functionality levels","stakeholder analysis","requirements engineering","Industry 5.0","human-system interaction"],"falsifier":"Ask independent HDT practitioners to classify a set of deployed or proposed HDT systems into the six levels, including at least one system that predicts a user's future state without first personalizing any configuration; if such a system cannot be placed on the scale without contradiction, the claim that higher levels include all lower ones is refuted.","tokens_in":10631,"feed_emoji":"🧑‍💻","tokens_out":7151,"duration_ms":76868,"temperature":0.7,"pith_summary":"The paper claims that the scattered field of human digital twins — digital models of a person that mirror physical, biological, psychological, or behavioral data — can be unified by a single specification with three parts: a small set of functionality levels, a map of stakeholders and users, and a generic requirements list. The authors propose six ordered levels — store, analyze, personalize, predict, control, optimize — in which each higher level includes the lower ones, with communication treated as an orthogonal capability. Starting from a thought experiment in which a person carries their HDT as a protected parameter set on a key card, they brainstorm applications for each user group, detail three representative applications, and derive requirements covering safety, reliability, law, structure, connectivity, and accessibility. If the specification holds, HDT implementations can be planned for reuse across domains instead of being rebuilt for each application.","feed_headline":"Six functionality levels can specify any human digital twin","feed_subtitle":"A new specification maps store, analyze, personalize, predict, control, and optimize onto users and applications.","key_machinery":"The load-bearing object is the six-level functionality scale: store, analyze, personalize, predict, control, and optimize. Levels are cumulative, so a higher level presupposes all lower ones, and 'communicate' is deliberately kept orthogonal because every level needs it. The scale is generated by a sentence-completion creativity technique applied to each user group, and the resulting application table doubles as the paper's evidence that the scale covers the space of HDT uses. A second machinery is the derived requirements table, organized into six categories, which turns the stakeholder and user analysis into a reusable implementation checklist.","core_discovery":"The central claim is that what makes a human digital twin is best captured by what it is used for, and that all uses fall into six cumulative levels of functionality: store, analyze, personalize, predict, control, and optimize. The paper derives this scale by completing the sentence 'The USER uses the HDT to FUNCTIONALITY APPLICATION' for each identified user group, then groups the resulting applications into the six levels, producing an appendix table that it presents as validation of the scale. Three selected applications — individualized medication, smart-home control, and factory productivity optimization — are profiled in detail to show that the levels correspond to real and escalating demands for models, data, interfaces, and regulatory compliance. The same holistic analysis yields a generic requirements table intended as a guideline for implementing reusable HDTs.","pith_inferences":["The six-level scale could be used as a maturity model: an HDT implementation is only as advanced as the highest level it reliably supports, which would let buyers and regulators compare systems with a single metric.","The singleton requirement 'one person, one HDT' implies a portable identity and data-portability standard; the key-card mechanism in the thought experiment is one way to realize it, but the requirement itself pushes toward interoperable HDT-to-HDT protocols.","The paper's observation of growing complexity across levels suggests a cost model in which higher-level applications are disproportionately expensive; that claim could be tested by counting models, interfaces, and compliance requirements per application in a larger sample.","The requirements list could be turned into a certification checklist by operationalizing each requirement as a test, for example verifying that an HDT can purge all data by a scheduled date for 'plannable death.'"],"forward_implications":["Developers can decide which functionality level an application targets and reuse models and data from the lower levels instead of starting from scratch.","Researchers and industrial use-case owners can position their HDT prototypes on the same scale, making cross-domain comparison and generalization visible.","The generic requirements table provides a pre-implementation checklist covering safety, digital sovereignty, legal conformity, modularity, interoperability, and accessibility.","Existing systems such as the electronic patient record can be understood as level-0 HDTs that gain higher-level functionality when composed with additional models.","The three detailed application profiles indicate that moving up the levels increases the number of models, interfaces, and legal constraints that an implementation must handle."],"supporting_citations":[{"why":"Documents the absence of an accepted definition of digital twin, motivating the need for a specification.","marker":"[1]"},{"why":"Establishes the human digital twin as a human-centric tool for Industry 5.0, the paper's main motivation.","marker":"[3]"},{"why":"Supplies the authors' earlier model-combination paradigm that the functionality levels are said to complement.","marker":"[5]"},{"why":"Shows an existing HDT architecture as a common interface for operator 4.0 applications, an example of domain-specific design.","marker":"[7]"},{"why":"Provides the vision and architecture for personalized healthcare HDTs, used as an application domain and source of legal and ethical concerns.","marker":"[11]"},{"why":"The electronic patient record is used as a minimal level-0 HDT example from which higher functionality can be composed.","marker":"[15]"},{"why":"The ISO/IEC digital twin standard is cited to ground the interoperability and connectivity requirements.","marker":"[19]"},{"why":"The European Accessibility Act is cited to ground the universal-design accessibility requirement.","marker":"[21]"}],"fun_headline_variants":["Human digital twin defined by six functionality levels","Store to optimize: six levels define human digital twins","What makes a human digital twin? Six use levels","Human digital twins: six cumulative functionality levels","Specifying human digital twins via six functionality levels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The specification assumes that the six functionality levels are the right and complete abstraction for every human digital twin use, and that the authors' own brainstormed application table is enough to show it.","fun_headline_variants_meta":{"raw":{"variants":["Human digital twin defined by six functionality levels","Store to optimize: six levels define human digital twins","What makes a human digital twin? Six use levels","Human digital twins: six cumulative functionality levels","Specifying human digital twins via six functionality levels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1187,"prompt_tokens":899,"completion_tokens":288,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":217}},"tokens_in":515,"tokens_out":288,"duration_ms":3577,"temperature":1.0,"reasoning_tokens":217,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:44:31.095835+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask independent HDT practitioners to classify a set of deployed or proposed HDT systems into the six levels, including at least one system that predicts a user's future state without first personalizing any configuration; if such a system cannot be placed on the scale without contradiction, the claim that higher levels include all lower ones is refuted.","supporting_citations":[{"cited_title":"Preliminary systemic model of (human) digital twin,","cited_arxiv_id":null,"evidence_quote":"Documents the absence of an accepted definition of digital twin, motivating the need for a specification."},{"cited_title":"Human digital twin in the context of industry 5.0,","cited_arxiv_id":null,"evidence_quote":"Establishes the human digital twin as a human-centric tool for Industry 5.0, the paper's main motivation."},{"cited_title":"Perspectives-observer-transparency – a novel paradigm for modelling the human in human-to-anything interaction based on a structured review of the human digital twin,","cited_arxiv_id":null,"evidence_quote":"Supplies the authors' earlier model-combination paradigm that the functionality levels are said to complement."},{"cited_title":"Archi- tecture of a human-digital twin as common interface for operator 4.0 applications,","cited_arxiv_id":null,"evidence_quote":"Shows an existing HDT architecture as a common interface for operator 4.0 applications, an example of domain-specific design."},{"cited_title":"Human digital twin for personalized healthcare: Vision, architecture and future directions,","cited_arxiv_id":null,"evidence_quote":"Provides the vision and architecture for personalized healthcare HDTs, used as an application domain and source of legal and ethical concerns."},{"cited_title":"Personal health records: a scoping review,","cited_arxiv_id":null,"evidence_quote":"The electronic patient record is used as a minimal level-0 HDT example from which higher functionality can be composed."},{"cited_title":"Digital twin - concepts and terminology,","cited_arxiv_id":null,"evidence_quote":"The ISO/IEC digital twin standard is cited to ground the interoperability and connectivity requirements."},{"cited_title":"Directive (EU) 2019/882 of the European Parliament and of the Council of 17 April 2019 on the accessibility requirements for products and services,","cited_arxiv_id":null,"evidence_quote":"The European Accessibility Act is cited to ground the universal-design accessibility requirement."}],"review_version":1}