{"id":"c0a4c30d-effe-4655-b5b0-7e84d022b9f5","arxiv_id":"2607.17244","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"DynImmune-BERT shows that event-aware continuous-time modeling of longitudinal TCR repertoires improves cancer-status AUC over static and simpler temporal baselines, but external validation is limited by small cohorts.","lead":"A new machine-learning model, DynImmune-BERT, tracks how T-cell receptor clones change over time in a patient's blood and uses those changes to predict cancer status. It combines neural ODEs with event-aware restarts so that irregular sampling times and clone disappearances/reappearances are explicitly modeled.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that the ODE core drives most of the accuracy gain is underdetermined: full DynImmune-BERT differs from its latent-ODE control in event handling, transport losses, and attention simultaneously, and the only ablation supporting the ODE-core attribution is summarized in prose/Fig. 4 without","rationale":"The reader's weakest assumption focused on whether a single shared vector field can transport compositional repertoire states without distortion—an architectural/biophysical concern. My concern is more epistemic: the experimental design does not yet permit attribution of the central accuracy gain to the ODE core, because the full model differs from its closest control in multiple components simultaneously. This matters because the paper's headline contribution is specifically the 'Neural ODE Driven' component, not merely temporal awareness. The internal Table 3 gradient (static → GRU → time-aware Transformer → latent ODE → full) supports the weaker claim that temporal information helps, and I credit the paper for patient-level splits, seed variability reporting, calibration diagnostics, and explicit limitations. However, the claim that the ODE core 'drives most of the accuracy gain' is the piece that would distinguish this architecture from a simpler temporal encoder, and it rests on a figure and a sentence rather than a reported ablation table. The reader's conditional verdict remains appropriate: the paper should be accepted only after the ablation details and code are made available. I therefore keep the verdict unchanged rather than escalate, since the concern is about evidence strength, not a demonstrated internal contradiction.","tokens_in":10734,"tokens_out":5837,"duration_ms":72150,"concrete_test":"Obtain the exact ablation code and run a single additional variant, 'Full minus ODE': replace the ODE integration between observations with a feed-forward/MLP or linear interpolation conditioned on Δt, while keeping event restart, bounded-neighborhood attention, both transport terms, Rtemp, and Rspec identical (and retuning λ's). Compare Lung AUC over ≥10 seeds to the reported 0.982±0.006. If the AUC drop is ≤0.005, the ODE core is not the main driver and the 'Neural ODE Driven' claim should be downgraded. Also publish the seed-level values behind Table 3 and exact definitions of 'Latent ODE control' and the Figure 4 ablation variants.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: (i) temporal structure helps, and (ii) the ODE core drives most of the accuracy gain. Part (ii) is load-bearing because the title and contribution are explicitly 'Neural ODE Driven.' It is not currently supported by a controlled ablation. In Table 3, 'Latent ODE control' (Lung 0.963±0.009, THCA 0.969±0.010) differs from 'Full DynImmune-BERT' (0.982±0.006, 0.984±0.007) in at least three ways simultaneously: presence-gated event restart (Eqs. 12–14), hybrid transport supervision (Eq. 16), and the full bounded-neighborhood attention stack. The prose says 'Ablations indicate that the ODE core drives most of the accuracy gain' and points to Fig. 4, but no numeric ablation table or exact variant definitions are provided. It is therefore impossible to verify whether removing only the ODE integration—while keeping event restart, transport, pseudocounts, and regularizers identical—preserves the gain. Additionally, Eq. 17 augments the loss with λ_W fW + λ_temp Rtemp + λ_spec Rspec. If the temporal controls in Table 3 were trained only with L_CE, part of the improvement could be an artifact of auxiliary losses and regularization rather than temporal modeling. The paper's matched-control claim thus does not yet isolate temporal structure from model capacity, event handling, and loss terms.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DynImmune-BERT, a continuous-time model for longitudinal T-cell receptor repertoires that combines a depth-adaptive centered log-ratio initialization, clone-presence-gated Neural ODE dynamics, bounded-neighborhood self-attention, event-based restart for reappearing clonotypes, a low-rank meta-adapter, and a hybrid transport objective over dominant and rare clone mass. The model is evaluated for cancer status prediction on lung and thyroid cancer cohorts. The evaluation is deliberately separated into literature-reported cross-study comparisons (Table 2), internally matched temporal controls (Table 3), and small external/universal detection cohorts (Table 4), with additional calibration and threshold diagnostics. The main claimed findings are that temporal structure improves over static encoders under matched patient-level splits, and that the ODE core is the largest driver of the accuracy gain.","tokens_in":11184,"tokens_out":4544,"duration_ms":53170,"significance":"If the central claims are supported, the paper would make a useful contribution by showing that irregular sampling time, clone presence events, and sequencing depth can be modeled explicitly rather than treating each repertoire as a static bag of sequences. The internal matched comparisons in Table 3 and the formal patient-level split condition in Eq. 3 are strengths; the paper also honestly reports uncertainty on small external cohorts and provides calibration and threshold diagnostics, which is more careful than many related papers. However, the paper's headline architectural contribution is specifically 'Neural ODE Driven', so the claim that the ODE core is responsible for most of the accuracy gain is load-bearing. That claim is not currently supported by a controlled ablation, and one data provenance issue further weakens the cross-study comparison. The temporal-structure claim itself is better supported.","major_comments":[{"comment":"The claim that 'the ODE core drives most of the accuracy gain' is not supported by a controlled ablation. The 'Latent ODE control' row in Table 3 differs from 'Full DynImmune-BERT' simultaneously in at least three design dimensions: presence-gated event restart (Eqs. 12–14), hybrid transport supervision (Eqs. 15–16), and the full bounded-neighborhood attention stack; it may also use a different loss, since Eq. (17) includes transport and regularizers that may not be present in the controls. Figure 4 is described only in prose and no numeric ablation table or exact variant definitions are given. Therefore, removing the ODE integration alone, while keeping event restart, transport, attention, and loss terms identical, is not shown to preserve the gain. Please provide a numeric ablation table with all one-factor-at-a-time variants and report which loss each variant uses.","section":"§4.3, Table 3, Fig. 4, Eq. (17)"},{"comment":"The data provenance for Table 2 is unclear. The caption says both THCA and lung cancer test samples are from [43], but reference [43] is a lung cancer study ('Spatial heterogeneity of the T cell receptor repertoire reflects the mutational landscape in lung cancer') and does not appear to contain thyroid cancer (THCA) data. Please cite the actual THCA data source, or correct the table if the citation is erroneous. Without this, the THCA numbers in Table 2 and any claims built on them cannot be verified.","section":"§4.2, Table 2"},{"comment":"The external detection check is not sufficiently documented for evaluation. The text mentions a universal setting with 2296 samples from 17 cancer types, but Table 4 reports only five disease subsets with n between 8 and 24, and does not state how these subsets were selected, how healthy controls were chosen, whether the model was trained on the same disease classes, or how thresholds were applied in each subset. The paper appropriately warns that small cohorts limit conclusions, but the missing protocol details prevent the reader from assessing possible selection or threshold effects. Please provide cohort composition, split rules, and threshold definitions for each row.","section":"§4.4, Table 4"}],"minor_comments":[{"comment":"The supports T and U used in the hybrid transport loss are not defined in the text. Please state explicitly how 'top clones' and 'tail clones' are selected, including any hyperparameters such as tail sampling size.","section":"§3.5, Eq. (16)"},{"comment":"The paper states that implicit Runge-Kutta solvers are used for 'stable event handling', but no implementation detail is given and standard torchdiffeq does not by default provide implicit RK event handling. Please specify the solver, event detection mechanism, and how restart discontinuities are handled numerically.","section":"§4.6, §3.6"},{"comment":"The uncertainty notation is inconsistent: THCA is reported as a confidence interval, while the other diseases are reported as standard errors. Please standardize or explicitly label the quantities.","section":"Table 4"},{"comment":"No code repository or data access statement is provided. Given the number of custom components (event restart, low-rank adapter, hybrid transport, neighborhood construction), a code release or detailed pseudocode for the event handling would substantially improve reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the thing to know: this paper gives a credible architecture for longitudinal TCR modeling, and its internal matched comparisons (Table 3) do show that using temporal structure helps. The bigger claim that the Neural ODE core is what drives the gain is not actually isolated by the experiments in the text.\n\nWhat's new is a combination: depth-adaptive CLR states, presence-gated Neural ODE dynamics, bounded-neighborhood attention, low-rank restart for reappearing clones, and a hybrid transport objective. That's a reasonable package for an understudied problem. The evaluation is more careful than most: they separate literature baselines from matched internal controls, report seed variability, and include calibration and threshold diagnostics. Table 3 shows a clean monotone ordering from static controls to the full model, and the gain over the latent ODE control (~0.02 AUC) is consistent across two datasets.\n\nThe soft spot is exactly what the stress-test note flags. The \"Latent ODE control\" differs from the full model in several respects at once: event restart, hybrid transport loss, and the full attention stack. So you can't tell how much of the improvement comes from the ODE integration per se. The ablation pointing to the ODE core is prose plus a figure without the per-variant numbers. That's a real gap, because the title sells the ODE. Also, the hybrid transport loss supervises predicted clone mass with observed clone mass—that's reconstruction, not prediction of unseen biology. The external cohorts are small (8–24 cases) and the Table 2 comparison is cross-study. The paper is upfront about these limitations, which counts in its favor.\n\nWho is this for: anyone building models for longitudinal TCR or other compositional repertoire data. It's a useful architecture baseline and the matched comparison is honest. It deserves a serious referee, but the referee should demand code, exact variant definitions, and a proper one-variable-at-a-time ablation for the ODE core. With that, it could become a solid contribution.","headline":"A credible architecture paper for longitudinal TCR modeling whose temporal-value claim is supported internally, but whose 'ODE core drives the gains' claim is not yet isolated by the experiments.","tokens_in":11648,"tokens_out":2355,"would_cite":true,"duration_ms":39298,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Modeling longitudinal T-cell receptor repertoires as event-aware continuous trajectories improves cancer detection over static bag-of-sequences encoders, with internally matched AUC gains to 0.982 (lung) and 0.984 (thyroid).","keywords":["TCR repertoire","neural ODE","continuous transformer","longitudinal immune monitoring","cancer detection","event-aware modeling","centered log-ratio","optimal transport"],"falsifier":"A strictly matched comparison in which a simpler temporal encoder (e.g., a GRU or time-aware transformer without ODE) is equipped with the same event restarts, bounded-neighborhood attention, and hybrid transport loss, and matches or exceeds the reported AUC on the same patient-level splits, would show the ODE core is not the driver. A large external cohort where the ODE-based model fails to beat a static BertTCR-style baseline would further weaken the general claim.","tokens_in":10644,"feed_emoji":"🧬","tokens_out":2418,"duration_ms":31547,"temperature":0.7,"pith_summary":"This paper tries to establish that the temporal structure of a patient's T-cell receptor repertoire—when clones expand, contract, disappear, and reappear across irregular sampling times—contains signal for cancer detection that static models throw away. It proposes a continuous-time transformer in which each clone's compositional state evolves under a presence-gated neural ODE, interrupted by event-based restarts when clones reappear. In internally controlled comparisons that match patient splits and preprocessing, the full model reaches AUC 0.982 for lung cancer and 0.984 for thyroid cancer, with ablations attributing most of the gain to the ODE core. A sympathetic reader would care because this offers a practical path to using serial immune sequencing for cancer status prediction, while the paper's emphasis on matched controls and calibration diagnostics keeps the claim honest.","feed_headline":"Temporal immune modeling hits 0.984 AUC in matched cancer screens","feed_subtitle":"A continuous-time, event-aware model of T-cell repertoire dynamics beats static encoders when serial samples are available.","key_machinery":"The central mechanism is a presence-gated neural ODE vector field Fθ that transports a clone's stabilized centered-log-ratio state between irregularly spaced observations, interrupted by restart events for reappearing clones. Supporting machinery includes depth-adaptive pseudocounts for compositional stability, bounded-neighborhood self-attention over abundance, sequence-similarity, and tail-sampled neighbors, low-rank meta-adapter initialization for reappearing clones, and a hybrid transport loss that supervises both dominant and rare clone mass via entropic and sliced-Wasserstein terms.","core_discovery":"On its own terms, DynImmune-BERT claims that immune repertoire classification should be treated as a patient-level trajectory problem rather than a single-sample bag-of-sequences problem. The core discovery is that an event-aware continuous transformer, where clone states are integrated by a neural ODE between irregular observations and restarted when clones reappear, outperforms static encoders and simpler temporal encoders under strictly matched protocols. The largest internal comparison shows the full model achieving mean AUC 0.982 on lung cancer and 0.984 on thyroid cancer, with the ODE-driven temporal propagation being the dominant contributor to the improvement over ignoring ordering.","pith_inferences":["The ODE integration may be functioning as an interpolation plus depth-normalization mechanism rather than as a faithful model of biological clone dynamics; a test comparing it to a non-ODE smooth interpolation with the same event restarts and transport loss would isolate what the ODE actually contributes beyond the surrounding design.","The bounded-neighborhood attention could be extended to incorporate epitope or HLA context, which the paper leaves for future work but which would likely sharpen the rare-clone tail supervision that the transport loss already emphasizes.","The reappearance-restart mechanism suggests a direct clinical extension: using the timing and magnitude of clone re-emergence as a biomarker for immune response to therapy, a signal the current binary cancer-status setting only partially captures.","The paper's small external cohorts limit the generalizability claim; a larger prospective multi-disease study with standardized preprocessing would be the natural next test of whether the ODE-driven gain persists beyond the two internal cancer types."],"forward_implications":["If the central claim holds, longitudinal TCR repertoires become a viable input modality for noninvasive cancer detection, with temporal dynamics adding information beyond static diversity and clonality features.","The ODE-driven temporal propagation provides a template for other irregularly sampled molecular measurements, such as B-cell receptor repertoires or serial methylation profiles, where clone presence patterns matter.","The matched-control protocol demonstrates a fair way to evaluate temporal models against static baselines, reducing the risk that cohort or preprocessing differences are mistaken for modeling gains.","The reported calibration and threshold diagnostics suggest that the model's probability outputs can be used for decision-making, not just ranking, if validated prospectively on larger cohorts.","The computational cost estimates indicate the continuous-time approach is practical on modest hardware, keeping overhead small relative to sequencing turnaround."],"fun_headline_variants":["Event-aware immune trajectory model hits 0.984 AUC","Continuous-time TCR model beats static encoders on serial data","T-cell repertoire dynamics as patient trajectory: AUC 0.984","ODE-driven immune model: temporal structure lifts AUC to 0.984","Dynamic immune modeling: 0.984 AUC with neural ODE states"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"A single shared vector field, gated by clone presence, can transport each clone's compositional state across gaps of weeks to months without distorting the signal, and the observed accuracy gain comes from this ODE integration rather than from the event restarts, pseudocounts, or hybrid transport supervision that surround it.","fun_headline_variants_meta":{"raw":{"variants":["Event-aware immune trajectory model hits 0.984 AUC","Continuous-time TCR model beats static encoders on serial data","T-cell repertoire dynamics as patient trajectory: AUC 0.984","ODE-driven immune model: temporal structure lifts AUC to 0.984","Dynamic immune modeling: 0.984 AUC with neural ODE states"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1125,"prompt_tokens":709,"completion_tokens":416,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":337}},"tokens_in":453,"tokens_out":416,"duration_ms":5059,"temperature":1.0,"reasoning_tokens":337,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T04:08:27.554571+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A strictly matched comparison in which a simpler temporal encoder (e.g., a GRU or time-aware transformer without ODE) is equipped with the same event restarts, bounded-neighborhood attention, and hybrid transport loss, and matches or exceeds the reported AUC on the same patient-level splits, would show the ODE core is not the driver. A large external cohort where the ODE-based model fails to beat a static BertTCR-style baseline would further weaken the general claim.","supporting_citations":[],"review_version":2}