{"id":"53e449c8-75f8-4266-abba-58b286645a7a","arxiv_id":"2601.14632","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":11,"one_line_summary":"In agent-based simulations for Seoul and Busan, contact tracing becomes ineffective beyond untraced-infector rates of about 4% and 10%, respectively, while omitted contacts have a milder, more gradual effect.","lead":"Using detailed computer simulations of Seoul and Busan, this paper finds that contact tracing loses control once roughly 4% (Seoul) or 10% (Busan) of confirmed cases are not traced, while missing individual contacts does less damage. The result provides city-specific tolerances for information loss in manual contact tracing systems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The city-specific IO thresholds (4% Seoul, 10% Busan) are not robustly defined: no formal threshold-extraction rule is given, and the model parameters driving them are admittedly arbitrary and untested by sensitivity analysis.","rationale":"The reader's weakest_assumption was that the thresholds are conditional on arbitrary or unvalidated parameters and that no sensitivity analysis is provided. I agree that this is a serious limitation, but I would sharpen it: the threshold concept itself is under-specified. The text mentions a threshold but never defines how it is computed, and the only explicitly shown IO rates are 0%, 2%, 4%, and 10%. Thus the 4% vs 10% comparison could be an artifact of both parameter choices and grid resolution. This is not an internal inconsistency, and the qualitative finding—that infector omission matters more than contact omission—is plausible and supported by the model's construction. However, because the paper's headline contribution is quantitative ('approximately 4% vs. 10%'), the lack of a formal threshold rule and the absence of sensitivity analysis are load-bearing. The proposed concrete test would settle whether the threshold values are robust enough for public-health interpretation. The verdict remains CONDITIONAL, consistent with the reader, since the paper should be accepted only after these gaps are addressed or the claims are softened to qualitative trends.","tokens_in":11799,"tokens_out":3548,"duration_ms":39264,"concrete_test":"Re-run the Seoul and Busan simulations on a fine IO grid (0–15% in 0.5% steps) and define the threshold a priori as the smallest IO rate at which the probability of a large outbreak (e.g., cumulative infections > 500) exceeds 50% across the 100 runs. Then repeat this threshold extraction under at least the extreme plausible parameter sets: PA∈{0.1,0.5}, Pt∈{0.2,0.8}, friend/local contact probability∈{1/14,2/7}, and local-community omission∈{0.25,0.75}. If the extracted threshold shifts by more than ±2 percentage points in either city, or if the Seoul/Busan ordering reverses under any of these parameter sets, the paper's threshold claim is not robust.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claim—that CT effectiveness breaks down at an IO rate of roughly 4% in Seoul and 10% in Busan—rests on two linked weaknesses. First, the paper never defines a formal procedure for extracting a 'threshold.' The Results section states that the cumulative number of infections 'increases sharply once the IO rate exceeds 2%' and later identifies a 'threshold of 4%,' while Figure 4 displays only four IO rates (0%, 2%, 4%, 10%). Without a pre-specified criterion (e.g., breakpoint regression, crossing of an outbreak-size threshold, or sustained-transmission probability), the 4% and 10% values may reflect visual inspection of a coarse grid rather than a well-defined system property. Second, even if the thresholds are reproduced under the default parameter set, that parameter set is largely unvalidated: Table 1 lists PA=0.2, Pt=0.5, friend/local contact probability 1/7, homophily h=0.9, and a fixed 50% local-community omission rate, while infectiousness distributions are borrowed from COVID-19 studies. The authors explicitly state in the Discussion that they were 'unable to quantitatively validate the realism or accuracy of the CT processes reproduced by the model.' Because the threshold is an emergent property of these parameters, any of them could shift the 4%/10% values. The qualitative ordering—IO more damaging than CO, and smaller cities more tolerant—may survive, but the precise city-specific thresholds, which are the paper's headline quantitative outputs, are not yet supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a stochastic agent-based model of contact tracing (CT) in two South Korean metropolitan areas, Seoul and Busan, built from synthetic populations and multilayer social networks. It simulates two types of information loss—infector-omission (IO), where confirmed cases' movement trajectories are not traced, and contact-omission (CO), further split into selective (SCO) and uniform (UCO) variants—and reports that IO causes a sharp, threshold-like degradation of CT effectiveness (approximately 4% IO rate in Seoul and 10% in Busan), while CO produces gradual degradation. The paper also analyzes the directed transmission network's diameter and out-degree distribution, finding that information loss increases the diameter and that the IO effect is structurally more disruptive than CO. The authors conclude that accurate trajectory reconstruction is more decisive than complete contact notification and that CT strategies should be tailored to regional population structure.","tokens_in":12216,"tokens_out":3596,"duration_ms":39864,"significance":"If the quantitative thresholds were robust, the paper would provide concrete, actionable targets for CT completeness and a clear argument for prioritizing infector-trajectory reconstruction over exhaustive contact notification. The strengths include a detailed synthetic-population construction from census and survey data, stochastic simulation with 100 runs and reported confidence intervals, and public code/data availability. However, the headline thresholds (4% and 10%) rest on an unspecified threshold-extraction criterion and on a parameter set that the authors themselves acknowledge cannot be quantitatively validated. Because these thresholds are emergent simulation outputs, the paper's central quantitative claims are not yet established at the level of confidence the presentation implies.","major_comments":[{"comment":"The threshold values are not defined by any formal procedure. The text states that the cumulative number of infections 'increases sharply once the IO rate exceeds 2%' and later refers to 'the threshold of 4%,' while Fig. 4 shows only four IO rates (0%, 2%, 4%, 10%). No pre-specified criterion (e.g., breakpoint regression, crossing of an outbreak-size threshold, or sustained-transmission probability) is given. Without a reproducible extraction rule, the 4% Seoul and 10% Busan thresholds are visual interpretations of a coarse grid and cannot be independently verified or compared across scenarios.","section":"Results, 'Impact of infector-omission' and Fig. 4"},{"comment":"The thresholds depend on parameters that are largely arbitrary or borrowed from COVID-19 without sensitivity analysis: PA=0.2, Pt=0.5, friend/local contact probability 1/7, homophily h=0.9, the fixed 50% local-community omission rate, and the infectiousness distributions. The Discussion explicitly states that the authors were 'unable to quantitatively validate the realism or accuracy of the CT processes reproduced by the model.' Since the IO threshold is an emergent property of this parameter set, plausible changes in any of these values could shift the 4%/10% numbers. A sensitivity analysis, even over a subset of the most influential parameters, is essential to support the paper's quantitative claims.","section":"Table 1 and Discussion (Limitations)"},{"comment":"The Busan threshold of 10% is inferred from panels showing only peak time and peak height, with the same unspecified threshold criterion. Moreover, the paper attributes the threshold difference to the 'three-fold difference in population size,' but Seoul and Busan also differ in age structure, commuting patterns, and contact density. No controlled experiment varying population size while holding other factors fixed is performed, so the causal attribution to population size is not supported. The city comparison should either be framed as a case study with demographic covariates or include such a controlled analysis.","section":"Fig. 6 and city comparison"}],"minor_comments":[{"comment":"Typo: 'missingnsuch' should be 'missing such.'","section":"Conclusion"},{"comment":"The relative infectiousness distribution is cited to reference [3] (Ferguson et al.), but that reference does not appear to provide such a distribution. Please verify the citation or supply the correct source.","section":"Table 1"},{"comment":"The figures use a 50% confidence interval. Please justify this choice or show standard 95% intervals for the central quantities, as 50% CIs convey less uncertainty information.","section":"Figures 4 and 7"},{"comment":"The sentence 'The results in potentially infected agents moving freely within society' is grammatically incomplete; likely 'result in.'","section":"Methods, 'Contact Tracing'"},{"comment":"The claim that the model 'can be easily adapted by updating these epidemiological parameters' would be strengthened by a brief note on identifiability or a calibration strategy, especially given the acknowledged lack of validation.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope as a computational social-physics / public-health modeling contribution. The code and data availability statements are commendable and should be credited. The main barrier is the gap between the headline quantitative thresholds and the evidence: the thresholds are not formally extracted, and the parameters are unvalidated. These issues are fixable—with sensitivity analyses and a pre-specified threshold criterion—so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a competent ABM study with a genuinely useful conceptual split — infector omission (IO) vs contact omission (CO) — and a nice diagnostic in the transmission-network diameter. The qualitative finding that IO is much more damaging than CO, and that a smaller city (Busan) tolerates more loss, is believable and coherent. They also ship code and data, which is more than many papers do.\n\nThe soft spot is exactly where the reader's stress-test lands. The 4% and 10% thresholds are the headline outputs, but they are extracted from a coarse grid — Figure 4 shows IO rates 0%, 2%, 4%, 10% — and no formal threshold-extraction rule is given. The text says 'increases sharply once the IO rate exceeds 2%' and later identifies a 'threshold of 4%'; those are not the same thing. Without a breakpoint regression or a pre-specified operational criterion, the values may be visual impressions. That matters because the parameters driving them are admittedly arbitrary: PA=0.2, Pt=0.5, friend contact 1/7, homophily h=0.9, fixed 50% local-community omission, COVID-borrowed shedding. The authors themselves state in the Discussion that they were 'unable to quantitatively validate the realism or accuracy of the CT processes reproduced by the model.' So the precise 4%/10% numbers should not be read as public-health constants.\n\nI don't think this is fatal. The qualitative ordering is robust enough across their runs (100 repetitions, confidence intervals), and the distinction between missing an infector (an entire subtree) and missing a contact (a single edge) is structurally sound. The paper would be strengthened by a defined threshold criterion, a one-parameter-at-a-time sensitivity sweep, and a clear statement of which results depend on the unvalidated parameters. The self-citations to their prior ABM are fine given the shared framework.\n\nWho's this for? Modelers working on contact-tracing policy and public-health planners who want an intuition for why trajectory tracking matters more than perfect contact notification. A serious referee should see it, with major revision. I would not accept the thresholds as stated.","headline":"Useful IO/CO decomposition, but the headline 4%/10% thresholds are not yet supported — needs a defined threshold criterion and sensitivity analysis.","tokens_in":12673,"tokens_out":2479,"would_cite":true,"duration_ms":23566,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Contact tracing collapses once more than about 4% of confirmed cases' movements are never traced in a large city, while smaller Busan holds until about 10%.","keywords":["contact tracing","infector omission","contact omission","agent-based model","transmission network diameter","epidemic containment","Seoul","Busan"],"falsifier":"Use contact-tracing records from a real jurisdiction that include whether each confirmed case's full movement history was obtained. If cumulative infections and transmission-chain depth do not jump once untraced cases cross the 4–10% range, the claimed threshold phenomenon is refuted; if the jump appears, it supports the model.","tokens_in":11703,"feed_emoji":"🦠","tokens_out":4291,"duration_ms":53029,"temperature":0.7,"pith_summary":"The paper asks what happens when contact tracing is imperfect in different ways, and claims the answer depends sharply on which information is missing. Its simulations of a synthetic Seoul and Busan show that untraced confirmed cases — infector omission — behave like a switch: below a city-specific threshold containment holds, and above it the outbreak grows abruptly. Missed contacts, by contrast, only slow or delay control gradually. This asymmetry leads the authors to conclude that reconstructing a confirmed case's full movement history matters more than notifying every identified contact. The finding matters because it gives public-health planners a concrete target: keep infector omission low, and set that target by local population and mobility structure, not by a national average.","feed_headline":"A 4% tracing gap flips containment into outbreak in a city model","feed_subtitle":"Missed cases, not missed contacts, trigger the abrupt failure; smaller Busan holds to 10%.","key_machinery":"The argument is carried by a stochastic agent-based model of epidemic spread on a multilayer contact network—households, school classrooms, workplaces, friendships, and local communities—built from census-based synthetic populations. Its central object is the directed transmission network: a node for each infected agent and an edge from infector to infectee. Two observables do the work: cumulative infections and peak timing for severity, and the network diameter for the depth of transmission chains. Across omission scenarios, the paper tracks how these observables respond to the infector-omission rate versus the contact-omission rate, and reads a small diameter as a signature that tracing is","core_discovery":"On the paper's own terms, the central claim is that infector omission and contact omission are not symmetric failures. Infector omission, where a confirmed case is isolated but their trajectory and all downstream contacts are lost, produces a sharp transition: in the Seoul model cumulative infections jump once the omission rate exceeds about 4%, and the directed transmission network grows a longer diameter, meaning chains of infection run deeper. In the Busan model, with roughly a third of Seoul's population and an older age structure, the same transition appears at about 10%. Contact omission—losing individual notifications rather than whole trajectories—causes only gradual increases in pea","pith_inferences":["The exact 4% and 10% numbers are model outputs, not measured facts; because the paper does not validate its parameters against real tracing logs, they are best read as qualitative evidence of a threshold, not as precise operational limits.","The abrupt IO threshold hints at a more general principle: any intervention that removes the root of a transmission chain (such as early isolation of index cases) will show a critical coverage level, whereas interventions that only prune individual edges will degrade smoothly.","The diameter effect suggests a testable empirical signature: real outbreak datasets with good tracing coverage should show shorter generation-interval chains, and chain length should jump when tracing coverage falls below the local critical level.","For digital contact-tracing apps, partial adoption creates an infector-like omission at the system level, so the model implies an adoption threshold below which app-based tracing silently fails."],"forward_implications":["If the thresholds are right, a tracing system in a dense metropolis must keep untraced confirmed cases below roughly 4% to avoid losing containment, even if contact notification is near complete.","When resources are tight, effort should go to reconstructing the movement trajectories of confirmed cases rather than to chasing every last contact, because contact omission degrades control only gradually.","The city-specific thresholds imply that a one-size-fits-all tracing standard is unsafe; smaller or older cities can tolerate a higher omission rate.","A growing transmission-network diameter could serve as an early warning of hidden chains, visible before cumulative case counts surge.","The same model suggests a quantifiable design target for manual and digital tracing systems: they should be engineered to keep infector omission below the local threshold."],"fun_headline_variants":["Tracing fails at 4% missing cases in Seoul, 10% in Busan","Infector omission, not contact omission, flips Seoul at 4%","Missed infectors, not contacts, push tracing past a 4% cliff","A 4% tracing gap flips containment in Seoul; 10% in Busan","Abrupt breakdown: missing infectors at 4% (Seoul) vs 10% (Busan)"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The quantitative thresholds rest on model parameters—asymptomatic share 0.2, self-testing probability 0.5, contact frequencies, homophily, and COVID-19-borrowed infectiousness—that the paper acknowledges it could not validate, so a different set of values would move the 4% and 10% breakpoints.","fun_headline_variants_meta":{"raw":{"variants":["Tracing fails at 4% missing cases in Seoul, 10% in Busan","Infector omission, not contact omission, flips Seoul at 4%","Missed infectors, not contacts, push tracing past a 4% cliff","A 4% tracing gap flips containment in Seoul; 10% in Busan","Abrupt breakdown: missing infectors at 4% (Seoul) vs 10% (Busan)"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001999,"raw_usage":{"total_tokens":7649,"prompt_tokens":768,"completion_tokens":6881,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":6766}},"tokens_in":512,"tokens_out":6881,"duration_ms":49209,"temperature":1.0,"reasoning_tokens":6766,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T09:06:33.280181+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use contact-tracing records from a real jurisdiction that include whether each confirmed case's full movement history was obtained. If cumulative infections and transmission-chain depth do not jump once untraced cases cross the 4–10% range, the claimed threshold phenomenon is refuted; if the jump appears, it supports the model.","supporting_citations":[],"review_version":1}