{"id":"f3d35a60-db00-4ac1-90b2-fbc29116284e","arxiv_id":"2506.20435","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 21 studies shows digital twin composition for systems-of-systems lacks formal methods and relies mostly on semi-formal simulation-based verification.","lead":"This paper is a systematic review of 21 studies on how digital twins are composed and verified in systems-of-systems. It finds that composition is rarely formalized and that semi-formal simulation methods dominate verification, with formal methods underused.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PRISMA flow reports 31 included studies while the text and all percentages use 21; the corpus underlying the central claim is inconsistently defined.","rationale":"The reader's verdict is CONDITIONAL, and I agree that the review has reporting and selection issues. The reader's weakest assumption focuses on the search strategy's requirement for 'verif* OR validat*' in abstracts, which may miss formal-methods work not using those exact terms. That is a valid external-validity threat, acknowledged in Section 5. However, the most load-bearing concern is internal: the PRISMA flow diagram reports 31 included studies, while the text and all derived statistics use 21. This is a direct contradiction in the reported evidence base. Since the review's central claim is an empirical summary of the corpus, an inconsistent corpus definition undermines the claim's traceability and reproducibility regardless of search-strategy adequacy. A reviewer cannot check whether the conclusion follows from the data without knowing which papers were actually analyzed. This is not a fatal flaw—it could be a simple reporting error—but it must be fixed before the review's findings can be accepted. The reader noted the 31-versus-21 discrepancy in the rationale, so there is partial agreement, but did not elevate it to the primary weakness. I therefore keep the verdict CONDITIONAL: the paper should be accepted only after the authors reconcile the flow numbers, itemize the final corpus, and recompute the affected statistics if needed.","tokens_in":10580,"tokens_out":3050,"duration_ms":34544,"concrete_test":"Reconstruct the final included-studies list from the PRISMA flow and from the reference lists in Tables 2–6. Count unique references across these tables and compare with 21 and 31. If exactly 21 unique references appear, the flow diagram's '31' is an error that must be corrected; if 31 unique references appear, recompute all percentages and the formal/semi-formal/informal counts to determine whether the central claim survives. Independently, contact the authors or check the paper's supplementary material for a definitive enumerated list of included studies.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that formal verification is underutilized and simulation-based approaches dominate—is a summary of a 21-paper corpus, yet the PRISMA flow diagram (Figure 2) reports 31 new studies included in the review. Section 3.4 states 'the final selection was 21 papers', and the reported percentages in Section 4.3 (52.4%, 66.7%, 42.9%) correspond exactly to fractions of 21 (11, 14, and 9 papers). The flow diagram's 'Reports excluded (n = 0)' at eligibility, followed by 31 included from databases and 31 from citation searching, is never reconciled with the 21 papers actually analyzed. If the final included set is really 31, then the classifications, percentages, and the conclusion about formal verification underutilization may change; if it is really 21, the flow diagram is incorrect. Either way, the review's data basis is not reproducible, and the central claim rests on an unclearly defined corpus. This internal inconsistency is more load-bearing than the search-term bias the reader emphasizes, because even if the search strategy were perfect, the reported corpus is self-contradictory.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a systematic literature review of digital twin (DT) composition and verification and validation (V&V) approaches in cyber-physical systems-of-systems contexts. The authors follow PRISMA and Kitchenham-style guidelines, define four research questions, and analyze what they state is a final corpus of 21 studies from 2022--2024. They classify DT composition approaches, properties and qualities addressed in V&V, levels of V&V formality, and reported challenges. Their central findings are that composition is discussed but formalization is limited, that semi-formal and simulation-based V&V dominate while formal verification is underutilized, and that there is a lack of standardized DT-specific V&V frameworks.","tokens_in":10751,"tokens_out":3711,"duration_ms":34190,"significance":"If the corpus were consistently and transparently defined, the paper would provide a useful structured map of an emerging area, particularly through the classification tables (Tables 2--6), which are traceable to specific cited studies. The authors also make a genuine effort at transparency by following PRISMA and by candidly listing threats to validity in Section 5, including the search-term limitation and the exclusion of grey literature. However, the review's quantitative claims (e.g., the percentages in Section 4.3) and its central qualitative conclusion about underutilization of formal verification depend on an exact, reproducible corpus, and the manuscript currently contains a material inconsistency between the reported corpus size and the PRISMA flow diagram. That inconsistency must be resolved before the findings can be considered reliable.","major_comments":[{"comment":"The final corpus is inconsistently defined. Section 3.4 states 'the final selection was 21 papers', and the percentages in Section 4.3 (52.4%, 66.7%, 42.9%) correspond exactly to fractions of 21 (11, 14, and 9 papers, respectively). However, the PRISMA flow diagram in Figure 2 reports 'New studies included in review (n = 31)' from databases and 'New studies included in review (n = 31)' from citation searching, and reports 'Reports excluded (n = 0)' at the eligibility stage. These numbers are never reconciled with the 21 papers actually analyzed. If the included set is 31, the classifications and percentages may change; if it is 21, the flow diagram is incorrect. As written, the data basis for the central claim is not reproducible, and this must be corrected.","section":"Section 3.4 and Figure 2"},{"comment":"The stated date window for the final corpus is inconsistent with one cited primary study. Section 3.4 says the final selection is 21 papers 'published between 2022 and 2024', and Section 3.3 says 'No suitable references (except for supporting information) predated 2021.' Reference [42] (Van Den Brand et al., 'Models meet data', 2021) is dated 2021 and is cited as a primary study in Table 2, Table 3, and Table 5. If [42] is part of the 21-paper corpus, the date-window statement is wrong; if it is not, the classification tables include a non-corpus study. The authors should either correct the date window or remove/reclassify [42].","section":"Section 3.4 and reference [42]"},{"comment":"The search strategy embeds a potential selection bias that directly affects the central conclusion about formal verification being underutilized. Requiring 'verif* OR validat*' in the abstract means that formal-methods papers that do not use these terms will not be retrieved, as the authors themselves acknowledge in Section 5. To support the claim that formal verification is underutilized, the review should test the sensitivity of this conclusion, for example by reporting how many of the 21 papers actually use formal methods and whether a supplementary search on terms such as 'model checking', 'theorem proving', and 'formal methods' without the 'verif* OR validat*' constraint would add formal-verification papers to the corpus. Without such a check, the finding may reflect the search string rather than the state of the literature.","section":"Section 3.1 and Section 5"}],"minor_comments":[{"comment":"The sentence 'The resulting corpus consisted of 115 journal articles, 85 papers in conference proceedings, and 11 book sections' appears to describe all screened records, not the final 21 selected studies; please clarify whether these counts refer to the initial screening pool or to the included corpus, and reconcile them with the PRISMA flow numbers.","section":"Section 3.3"},{"comment":"The text states that Figure 1 shows a near doubling of publication volume in 2023--2024, but the figure as presented lacks clear axis labels or a data table; please ensure the figure is legible and that the counts underlying the claim are stated.","section":"Figure 1"},{"comment":"The claim that 'earlier papers (2022--2023) focused more on basic model validation, while recent papers (2023--2024) show increased attention to integration challenges and multi-domain verification' is not supported by a table, figure, or per-year counts; please provide the underlying distribution.","section":"Section 4.3"},{"comment":"The statement that 'a few promising papers were excluded because their full text is unavailable' should specify how many papers were excluded for this reason and whether any of them plausibly used formal verification, since this could affect the main conclusion.","section":"Section 5"},{"comment":"Several references (e.g., [6], [10], [40]) lack complete venue or page information; please complete the bibliography to support reproducibility and reader follow-up.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a competent systematic review with a clear structure, but the inconsistency between the PRISMA flow diagram (31 included studies) and the text/percentages (21 papers) is a load-bearing reporting error that must be fixed before the review's findings can be trusted. The date-window issue with reference [42] and the search-term sensitivity concern are also important but addressable within the manuscript's scope. I see no evidence of misconduct, only a need for careful revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nI read the Khedr and Fitzgerald SLR on digital twin composition with interest. The short version: the paper has a genuinely useful taxonomy, but the central corpus is misreported, and that is a load-bearing flaw, not a cosmetic one.\n\nWhat is new: the seven-way classification of composition approaches (orchestrated, federated, service-based, co-simulation, multi-model, decentralized, hierarchical), the three-level V&V formality scale (formal, semi-formal, informal), and the property/quality tables. These are a real synthesis, traceable to the 21 papers in the tables. The RQs are clear, and the threats-to-validity section is honest about search-term bias and grey literature.\n\nThe soft spots, in order of severity. First, the PRISMA flow diagram reports 31 new studies from the database branch and 31 from citation searching, with zero exclusions at eligibility, yet every percentage and table uses 21 papers (52.4%, 66.7%, 42.9% are exactly fractions of 21). The flow diagram cannot be reconciled with Section 3.4. If the real corpus is 21, the flow diagram is wrong; if it is 31 or 62, the classifications and conclusions may change. Either way, the central claim about formal verification being underutilized rests on an undefined dataset. This is worse than the search-term bias, which the authors already concede. Second, reference [42] is from 2021 while the abstract and Section 3.4 say 2022-2024; Section 3.2 says the timespan was 'restricted to 2021 onwards', so there is an internal contradiction. Third, the 'resulting corpus' sentence listing 115 journal articles, 85 conference papers, and 11 book sections does not match any other number in the process. These are reporting issues, but for an SLR, reporting is the method.\n\nThe authors clearly know the area and the synthesis is useful to anyone planning work on DT composition and V&V. But I would not cite it in its current form. If the authors correct the flow diagram and itemize the final 21 papers, the paper becomes a solid contribution worth a reader's time.\n\nRecommendation: send it to review, but with the explicit expectation that the corpus number is fixed and a supplementary list of the 21 studies is provided. Without that fix, the review's conclusions are not reproducible.\n\nBest,\n[Your name]","headline":"Useful synthesis, but the 31/21 corpus mismatch undermines the results and must be fixed.","tokens_in":11288,"tokens_out":4007,"would_cite":false,"duration_ms":34887,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematic review of 21 studies finds that digital twin composition for systems-of-systems is mostly orchestrated and validated by simulation, with formal verification rarely used.","keywords":["Digital Twin","Cyber-Physical Systems","Systems of Systems","Verification and Validation","Systematic Literature Review","Digital twin composition","Formal verification","Simulation-based validation"],"falsifier":"Re-run the search without the 'verif* OR validat*' abstract requirement, or add a second search stream using terms such as 'model checking', 'theorem proving', 'runtime verification' and 'formal verification' without requiring the 'verif*' stem, and compare how many additional studies on formal verification of composed digital twins are recovered; if the count is substantial, the review's central finding of underutilized formal verification would need to be revised.","tokens_in":10349,"feed_emoji":"🤖","tokens_out":5752,"duration_ms":54005,"temperature":0.7,"pith_summary":"This paper is a systematic literature review of how digital twins are composed for cyber-physical systems-of-systems and how those compositions are verified and validated. Examining 21 studies published between 2022 and 2024, the review finds that while composition is widely discussed—orchestrated integration being the most common pattern—formalization of composition is limited. Verification and validation are dominated by semi-formal methods such as simulation-based testing and co-simulation; formal verification, including model checking and theorem proving, is underutilized. The review also classifies the digital twin properties, quality attributes, and V&V challenges addressed in the literature, highlighting model uncertainty, integration complexity, and the absence of standardized DT-specific V&V frameworks. A sympathetic reader would care because these results indicate that current practice lacks the rigorous guarantees needed to trust digital twins in safety-critical systems-of-systems.","feed_headline":"Digital twin research skips formal verification, review finds","feed_subtitle":"A systematic review of 21 studies finds most digital twin composition relies on simulation, with few formal guarantees.","key_machinery":"The load-bearing mechanism of this paper is its systematic review protocol, conducted following Kitchenham's guidelines and reported under PRISMA, with a Boolean search query applied to five databases and supplemented by manual and forward-citation searching. The query requires the title to contain 'digital twin*' and the abstract to contain both a systems term ('cyber physical', 'system* of systems', or 'complex system*') and a V&V term ('verif*' or 'validat*'). This protocol generated 390 records and, after screening and quality assessment, a final corpus of 21 studies. The analytical core is a structured classification scheme that sorts composition approaches, digital twin properties and qualities, V&V approaches by formality level (formal, semi-formal, informal), and recurring challenges, which then carries the answers to the four research questions.","core_discovery":"On its own terms, the review establishes that the current literature on digital twin composition for systems-of-systems is growing but formally shallow. Among the 21 studies analyzed, seven composition approaches were identified, with orchestrated integration used in 11 papers, followed by federated, service-based, and co-simulation approaches. V&V practice splits across three formality levels—formal, semi-formal, and informal—but the weight of practice falls on semi-formal and informal methods: simulation-based testing, co-simulation, experimental validation, and case studies, while formal model checking and theorem proving appear in only a handful of studies. The review further identifies five V&V scope dimensions (model fidelity validation, behavioural correctness verification, integration validation, performance assessment, and cyber-physical consistency) and a catalogue of technical, methodological, and data-related challenges, with model uncertainty and integration complexity named as key technical obstacles and the lack of standardized DT-specific V&V frameworks as the central methodological gap. The conclusion is that the field needs to move beyond model validation toward integration and cyber-physical consistency, with standardized, scalable V&V and rigorous composition methodologies.","pith_inferences":["Because the corpus is restricted to 2022-2024 and to English, peer-reviewed, access-available publications, the review likely undercounts earlier and grey-literature work on formal digital twin verification, so the reported underutilization may be a lower bound.","The classification of composition approaches against SoS characteristics could be turned into a practical selection guide: federation for distributed autonomy, service-based integration for heterogeneity, orchestration for emergence.","A direct extension of the review would be to build a standardized V&V benchmark from the five scope dimensions it identifies, allowing different composition approaches to be compared on fidelity, interoperability, and cyber-physical consistency.","The review's future direction toward model-based SoS engineering case studies, such as a greenhouse digital twin, offers a concrete testbed for whether runtime verification and co-simulation can be scaled to composed digital twins."],"forward_implications":["If the review's findings hold, practitioners cannot currently rely on composed digital twins to provide formal guarantees in systems-of-systems, so safety-critical deployments must supplement simulation with formal verification.","Orchestrated integration's dominance suggests that centralized coordination is the default pattern, while federated and service-based patterns that preserve constituent autonomy remain less explored.","The absence of standardized DT-specific V&V frameworks means results across studies are difficult to compare, and developing such frameworks is a prerequisite for scalable, trustworthy digital twin deployment.","The growing attention to integration validation and cyber-physical consistency in 2023-2024 studies points to a shift in V&V scope that future research and tooling should target.","Model uncertainty and integration complexity will continue to obstruct rigorous V&V unless addressed by dedicated modelling and analysis techniques."],"supporting_citations":[{"why":"Supplies the PRISMA 2020 reporting framework used to structure and report the systematic review.","marker":"[34]"},{"why":"Supplies the systematic review procedures that define the review protocol.","marker":"[21]"},{"why":"Provides the prior systematic review of digital twin V&V in manufacturing that this review extends and contrasts with.","marker":"[3]"},{"why":"Provides the survey of systems-of-systems and digital twins that motivates the research gap.","marker":"[32]"},{"why":"Identifies the three digital twin integration architectures used as a baseline for discussing composition approaches.","marker":"[6]"},{"why":"Supplies the integration challenges for digital twin systems-of-systems that frame the composition problem.","marker":"[28]"},{"why":"Recent automated digital twin composition work cited as an example of the state of the art in composition approaches.","marker":"[14]"},{"why":"Provides the PRISMA flow diagram tool used to document the screening and selection process.","marker":"[16]"}],"fun_headline_variants":["Digital twin studies lean on simulation, skip formal checks","Formal verification rare in digital twin composition studies","Systematic review: digital twin V&V lacks formality","DT research favors simulation over formal verification","Review finds digital twin verification mostly informal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that requiring the words 'verif*' or 'validat*' together with cyber-physical or systems-of-systems terms in the abstract produces a representative corpus, so studies on formal verification that use different vocabulary could be missed and change the conclusion that formal methods are underutilized.","fun_headline_variants_meta":{"raw":{"variants":["Digital twin studies lean on simulation, skip formal checks","Formal verification rare in digital twin composition studies","Systematic review: digital twin V&V lacks formality","DT research favors simulation over formal verification","Review finds digital twin verification mostly informal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000404,"raw_usage":{"total_tokens":2100,"prompt_tokens":938,"completion_tokens":1162,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":1091}},"tokens_in":554,"tokens_out":1162,"duration_ms":9229,"temperature":1.0,"reasoning_tokens":1091,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:48:06.937866+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the search without the 'verif* OR validat*' abstract requirement, or add a second search stream using terms such as 'model checking', 'theorem proving', 'runtime verification' and 'formal verification' without requiring the 'verif*' stem, and compare how many additional studies on formal verification of composed digital twins are recovered; if the count is substantial, the review's central finding of underutilized formal verification would need to be revised.","supporting_citations":[{"cited_title":"Keele, UK, Keele Univ.33(08 2004)","cited_arxiv_id":null,"evidence_quote":"Supplies the systematic review procedures that define the review protocol."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior systematic review of digital twin V&V in manufacturing that this review extends and contrasts with."},{"cited_title":"In: 2023 18th Annual System of Systems Engineering Conf","cited_arxiv_id":null,"evidence_quote":"Provides the survey of systems-of-systems and digital twins that motivates the research gap."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies the three digital twin integration architectures used as a baseline for discussing composition approaches."},{"cited_title":"In: IEEE/ACM 10th Intl","cited_arxiv_id":null,"evidence_quote":"Supplies the integration challenges for digital twin systems-of-systems that frame the composition problem."},{"cited_title":"Campbell Systematic Reviews18(2), e1230 (2022)","cited_arxiv_id":null,"evidence_quote":"Provides the PRISMA flow diagram tool used to document the screening and selection process."}],"review_version":1}