{"id":"d46ceae6-042d-4ff1-8ff7-27dfde52c82e","arxiv_id":"2504.17122","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"Implicit neural representations can fit voxel-wise two-tissue compartment parameters from dynamic FDG PET with low TAC error, but the reported metric is the training loss and no ground-truth validation is provided.","lead":"This paper fits the four kinetic parameters of a two-tissue FDG model using a per-patient implicit neural representation, and reports lower curve-reconstruction error and sharper parametric images than a prior deep-learning baseline. The central caveat is that the reported error is the training loss on the same data, so the accuracy of the estimated physiological parameters is not yet demonstrated.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The forward model is underspecified: no plasma input function curve is described, so TAC generation and the reported K1/k2/k3/Vb values in Eq. (1) cannot be reproduced or interpreted; low TAC MSE cannot support the claimed parameter accuracy.","rationale":"The reader's weakest_assumption identified the same region of the paper: the TCKM forward model and the role of the image-derived input function are underspecified, and there is no reference measurement to validate the estimated kinetic parameters. I read the manuscript in good faith and acknowledge the INR approach is plausible, the code is promised, and the qualitative figures suggest lower reconstruction MSE than the baseline. However, the central claim—superior kinetic parameter estimation—requires a correctly specified forward model and evidence that low TAC error implies accurate parameters. Section 2.2 mentions the input function only as a normalization scalar, while Section 2.3 defines the loss via TCKM ODE solution without presenting the input function curve. This is a concrete, checkable technical gap, not a matter of disagreement with current consensus. The lack of any ground-truth kinetic parameter comparison or identifiability analysis further weakens the claim. Since the evaluation is therefore not logically connected to the stated conclusion, the REJECT verdict is appropriate; the method could still become useful after external validation with reference parameters and a properly specified input function.","tokens_in":7318,"tokens_out":2456,"duration_ms":25865,"concrete_test":"Inspect the released code at github.com/tkartikay/PhysNRPET and trace what object is passed to the TCKM ODE solver when computing the predicted TAC in Eq. (1). If no time-varying input function Cp(t) is supplied—e.g., only the maximum-normalized PET data or a scalar normalization factor is used—then the forward model is not the TCKM and the central claim fails. Additionally, generate synthetic TACs from known K1, k2, k3, Vb and a known Cp(t), train the INR with the stated MSE loss, and check whether the recovered parameters match the ground truth; if low MSE coincides with large parameter errors, the claimed kinetic parameter accuracy is not established.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central claim—that the proposed INR estimates accurate TCKM parameters (K1, k2, k3, Vb)—rests entirely on the network's ability to generate a TAC from those parameters via the TCKM ODEs and match the observed TAC (Section 2.3, Eq. 1). However, the TCKM forward model requires a time-varying plasma input function Cp(t). Section 2.2 only states that the dynamic PET data were normalized by the maximum of the image-derived input function; no Cp(t) curve, no compartment-model equation, and no delay or dispersion correction are provided. If the ODE solver does not receive a patient-specific, time-resolved input function, the mapping from the four predicted parameters to a TAC is not the two-tissue compartment model, and the parameters lose their physiological interpretation. This is a more fundamental gap than mere parameter identifiability: the reported loss function cannot be evaluated as described. The paper also provides no comparison against reference kinetic parameters from non-linear least squares fitting or an identifiability analysis, so even a well-specified forward model would not establish that low TAC MSE implies correct K1, k2, k3, and Vb. Together these issues mean the central performance claim is not supported by the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes per-patient implicit neural representations (SIREN with Gaussian Fourier features), optionally conditioned on CT Hounsfield values or on 4096-dimensional features from a 3D CT foundation model, to predict voxel-wise two-tissue compartment model (TCKM) parameters K1, k2, k3, and Vb from dynamic [18F]FDG PET. The parameters are optimized by minimizing the mean-squared error between TACs generated from the predicted parameters and the measured dynamic PET TACs. The authors evaluate multiple INR variants (2D/3D, HiRes/LoRes, with/without CT priors) against the self-supervised spatio-temporal neural network of De Benetti et al. on a 24-patient [18F]FDG dynamic PET/CT dataset, reporting lower voxel MSE, qualitative sharper parametric images, and lower training/inference resource footprints. The stated contributions are the introduction of a physiological INR for tracer kinetic modelling, its extension with CT foundation-model features, and its evaluation against a DNN baseline.","tokens_in":7582,"tokens_out":8098,"duration_ms":77930,"significance":"The idea of fitting a per-patient INR directly to dynamic PET data and reading out TCKM parameters is timely and potentially useful for data-efficient, personalized kinetic modelling. The paper makes its code available, which is a strength, and the qualitative figures suggest the INR can produce visually plausible parametric images. However, the present evaluation cannot support the abstract's claims of superior kinetic parameter estimation. The reported MSE in Table 1 is exactly the training loss minimized during optimization on the same data, there is no ground-truth kinetic parameter comparison, no nonlinear least squares baseline, no identifiability analysis, and the forward TCKM model is underspecified because the time-varying plasma input function is never given. As a result, the central claim that the INR estimates accurate physiological parameters is not established by the evidence presented.","major_comments":[{"comment":"The forward model is underspecified. The text states that the predicted TAC is generated by solving the TCKM ODEs, but it never writes the ODEs, the plasma input function Cp(t), or the delay/dispersion treatment; Section 2.2 only says the dynamic PET data were divided by the maximum of the image-derived input function. As written, Eq. (1) cannot be evaluated by a reader, and the outputs K1, k2, k3, Vb cannot be interpreted as physiological TCKM parameters. This is a core reproducibility and correctness issue.","section":"§2.2–2.3, Eq. (1)"},{"comment":"The reported voxel MSE is the same quantity minimized during training, evaluated on the same patient data used for optimization. Table 1 therefore reports training fit quality, which is expected to be low for an overparameterized per-patient network, and it does not validate the estimated kinetic parameters. The paper needs a comparison against voxel-wise nonlinear least squares fits (or otherwise reference parameter maps) with error and bias metrics, plus an out-of-sample or cross-validation scheme.","section":"Table 1, Eq. (1)"},{"comment":"No identifiability analysis is provided for the four-parameter TCKM under the 62-frame protocol and the scalar-normalized input function described in Section 2.2. Multiple parameter sets can produce nearly identical TACs, so a low training MSE does not imply correct K1, k2, k3, and Vb. A simulation study with known ground-truth parameters should be included to demonstrate that the optimized parameters are recoverable and that low TAC error translates into parameter accuracy.","section":"§2.3, Results"},{"comment":"The abstract claims superior spatial resolution and improved anatomical consistency, but no quantitative metrics for edge preservation, contrast, or anatomical consistency are reported; Figure 4 supports these claims only by visual inspection. In addition, Table 1 shows identical voxel MSE for inr-HiRes-2D, HU-INR-2D-HiRes, and FM-INR-2D-HiRes, so the claimed benefit of the CT and foundation-model priors is not evidenced. The Discussion's statement of \"slightly faster convergence (not shown)\" needs supporting convergence curves or a quantitative comparison.","section":"§3, Figs. 1–4, Table 1"},{"comment":"The manuscript states that the dataset contains 24 patients, but Table 1 reports a single MSE value per variant and Figures 1–4 show one exemplary patient. It is unclear whether the quantitative results are averaged over all patients, over a single slice, or over a single patient. The aggregation and per-patient variability must be specified; otherwise the claim that the results apply to the [18F]FDG dynamic PET/CT dataset is not supported.","section":"§2.1, Table 1"}],"minor_comments":[{"comment":"The Gaussian Fourier feature construction is ambiguous: state whether the entries of B are drawn as N(0, σ²) and then multiplied by 10, and give the value of σ; currently \"a standard deviation chosen to control frequency bandwidth\" is not sufficiently precise for reproduction.","section":"§2.2"},{"comment":"The training and inference time columns are confusing because dashes appear for some entries and both training and inference times are listed for INR variants; clarify which times apply to which model and on which hardware they were measured.","section":"Table 1"},{"comment":"The caption says \"MMSE is higher\" where \"MSE\" is intended; please correct the typo.","section":"Fig. 1 caption"},{"comment":"The phrase \"state-of-the-art DNNs\" overstates the comparison, since only the De Benetti et al. network is evaluated as a baseline; consider saying \"one previously proposed self-supervised DNN.\"","section":"§1"},{"comment":"There are minor typographical errors, including \"de Benetti el al.\" (should be \"de Benetti et al.\"), \"signficantly\" in Fig. 4, and inconsistent notation for Vb/V_B.","section":"Throughout"}],"recommendation":"reject","confidential_remarks":"I agree with the stress-test assessment: the missing input function and the circular evaluation are fundamental, not local. Even though the authors could in principle add NLLS comparisons and a full forward-model specification, the present manuscript does not support its abstract claims. I would be open to considering a substantially revised version with a properly specified model, non-circular validation, and quantitative spatial-resolution metrics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know this paper as the first attempt, as far as I can tell, to fit two-tissue compartment model parameters voxel-wise with per-patient implicit neural representations. The idea is sensible: optimize a SIREN with Fourier features to map spatial coordinates to K1, k2, k3, Vb, and use a differentiable ODE solver to reconstruct TACs. The authors also add CT foundation model features as an optional input. If it worked, it would give parametric images at 1.65 mm without a training cohort. That is genuinely worth trying.\n\nWhat the paper does well: it is straightforward, the architecture choices are standard, the code is linked, and the authors are explicit that hyperparameters were borrowed without ablation. The qualitative figures show sharper edges and lower TAC MSE than the self-supervised DNN baseline, which is a plausible benefit of per-patient fitting.\n\nThe soft spots are not small. The headline claim — that the INR estimates accurate kinetic parameters — is not supported by the evidence. Table 1 reports TAC MSE, which is exactly the loss being minimized during training on the same data. That shows the network can fit the training data, not that K1/k2/k3/Vb are right. There is no comparison against non-linear least squares, no reference kinetic parameters from any external source, and no identifiability analysis. A model with four free parameters per voxel can hit a low TAC error with wrong parameters, especially when the input function is not firmly fixed. And the forward model is underspecified: the paper never writes the TCKM ODEs or provides the plasma input function curve, only saying data were normalized by the maximum of the image-derived input function. Without Cp(t), a reader cannot reproduce the TAC generation, and the physiological meaning of the parameters is hard to assess.\n\nThe baseline comparison is also weaker than it looks: the DNN is pretrained and shares an author with this group, so the comparison is not a controlled test of per-patient versus population training. That is less damning, though.\n\nOverall, the paper is a plausible application note, not a validated method. The idea deserves a serious referee and probably a major revision asking for validation against reference kinetic parameters or simulated ground truth, plus a complete forward model. I would not cite it in its current form, but I would read a revised version.\n\nRecommendation: send to peer review with the expectation of heavy revision.","headline":"A reasonable INR-for-kinetics idea whose evaluation only proves it can fit its own training loss; the parameter accuracy claim needs real validation.","tokens_in":8101,"tokens_out":2804,"would_cite":false,"duration_ms":26577,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that fitting an implicit neural representation to one patient's dynamic PET scan estimates the two-tissue compartment model parameters more accurately and with sharper anatomical boundaries than a self-supervised…","keywords":["Implicit Neural Representations","Tracer Kinetic Modelling","Dynamic PET","Two-tissue compartment model","Parametric imaging","Anatomical priors","[18F]FDG","Time-activity curve"],"falsifier":"Generate synthetic dynamic PET frames from known compartment parameters under the same 62-frame protocol, fit the INR to those frames, and compare the recovered K1, k2, k3, and Vb to the ground truth; if the time-activity-curve error stays low while the parameters drift from the truth, then the identifiability assumption is violated.","tokens_in":7098,"feed_emoji":"","tokens_out":7616,"duration_ms":69493,"temperature":0.7,"pith_summary":"This paper tries to show that a small neural network trained on a single patient's dynamic PET scan can estimate the voxel-wise kinetic parameters of the two-tissue compartment model more accurately than a larger self-supervised deep network. The network maps spatial coordinates to four parameters — tracer influx K1, efflux k2, phosphorylation k3, and blood volume fraction Vb — by solving the kinetic model's differential equations inside the loss and comparing the reconstructed time-activity curve against measured frames. On a 24-patient [18F]FDG dataset, the authors report voxel-level mean-squared error of about 0.009 versus 0.066 for the baseline, with sharper edges in tumour and kidney regions. If this holds, personalised parametric imaging could be obtained per patient without large training corpora, which matters for data-scarce clinical settings.","feed_headline":"Fitting one network per patient cuts PET kinetic error by seven-fold","feed_subtitle":"Single-patient neural fits cut voxel error to 0.009 from 0.066, with sharper tumour edges.","key_machinery":"The load-bearing object is the implicit neural representation: a fully connected network with sinusoidal activations that takes normalised spatial coordinates, encoded by Gaussian Fourier Features, and returns K1, k2, k3, and Vb at that location. The two-tissue compartment model ODEs then synthesise a predicted time-activity curve from these parameters, and the MSE between predicted and measured curves is backpropagated through the ODE solve to update the network. This makes the physics the objective and turns the network into a continuous atlas of kinetic parameters; reconstructing a parametric image at any desired spatial resolution only requires querying the trained network at the relevant coordinates. Optional CT features enter as extra inputs, playing a supporting role rather than changing the final error.","core_discovery":"The central discovery is that a SIREN-based implicit neural representation, conditioned on Gaussian Fourier Features of spatial coordinates and optionally on CT-derived anatomical features, can encode the whole 3D+t dynamic PET signal of one patient as a continuous function whose output parameters obey the two-tissue compartment model. Because the forward model is differentiable, standard gradient descent on the time-activity-curve mean-squared error recovers the parameters. The authors report that all INR variants outperform the baseline in reconstruction error and produce parametric images with visibly higher spatial resolution, especially in the tumour and left kidney, where baseline errors concentrate; adding CT Hounsfield units or 4096-dimensional foundation-model features leaves the final error essentially unchanged while accelerating convergence.","pith_inferences":["Because the paper reports MSE on reconstructed time-activity curves rather than agreement with ground-truth kinetic parameters, the strongest testable extension is to validate against an arterial input function or an independent reference method; low TAC error alone leaves an identifiability gap.","The near-identical MSE across HiRes and LoRes variants and with or without CT features suggests the INR's continuous spatial prior already captures most of the anatomical signal; CT features may mainly help convergence, implying a simpler spatial regulariser could obtain similar maps at lower cost.","If single-patient INRs prove reliable, dynamic PET analysis could shift from population-trained DNNs to per-patient fitting, which would make kinetic parameters portable across scanners and protocols because no cross-patient training distribution is assumed.","The authors note in Section 2.2 that all hyperparameters were taken from prior INR examples without ablation, so the reported error levels may not be the architecture's ceiling; testing whether tuning changes the comparison would be a direct next step."],"forward_implications":["Voxel-wise parametric images at the scanner's native 1.65 mm resolution become feasible from a single patient's dynamic scan without large training databases.","Tumour and highly vascularised regions, where the baseline concentrates most error, are precisely where INR fitting is claimed to improve: sharper boundaries in K1, k2, and k3 maps, aiding lesion delineation.","Per-patient optimisation makes the method personalisable: each patient's anatomy and physiology are encoded in the trained network, so downstream classification or segmentation could consume the INR directly.","Memory and training requirements (about 24 minutes and under 7 GB for the LoRes 3D variant) suggest the method can run on clinical workstation hardware, not only research GPU clusters.","CT priors (Hounsfield units or foundation-model features) do not change the final error but accelerate convergence, so anatomy can be used as a regulariser without sacrificing accuracy."],"supporting_citations":[{"why":"Supplies the self-supervised DNN baseline in 2D and 3D configurations that all INR variants are compared against.","marker":"[5]"},{"why":"Supplies the SIREN architecture with sinusoidal activations and its weight-initialisation scheme, forming the INR backbone.","marker":"[18]"},{"why":"Supplies the Gaussian Fourier Feature encoding that maps normalised spatial coordinates to high-frequency inputs.","marker":"[20]"},{"why":"Supplies the 3D CT foundation model used to extract 4096 per-voxel anatomical features for the CTFM variant.","marker":"[11]"},{"why":"Provides the long-axial-field-of-view dynamic PET imaging protocol and kinetic-modelling context for the 62-frame dataset.","marker":"[14]"}],"fun_headline_variants":["One network per patient: 7x lower PET kinetic error","Implicit neural nets sharpen PET kinetic maps 7-fold","Patient-specific neural fit slashes PET kinetic error","SIREN encodes tumor PET kinetics from one scan","Data-efficient PET kinetics: one INR per patient"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The four kinetic parameters must be identifiable from each voxel's measured time-activity curve alone under the 62-frame protocol; if different parameter sets give nearly identical curves, low reconstruction error will not guarantee correct parameters.","fun_headline_variants_meta":{"raw":{"variants":["One network per patient: 7x lower PET kinetic error","Implicit neural nets sharpen PET kinetic maps 7-fold","Patient-specific neural fit slashes PET kinetic error","SIREN encodes tumor PET kinetics from one scan","Data-efficient PET kinetics: one INR per patient"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1218,"prompt_tokens":896,"completion_tokens":322,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":245}},"tokens_in":512,"tokens_out":322,"duration_ms":3537,"temperature":1.0,"reasoning_tokens":245,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:48:43.468891+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate synthetic dynamic PET frames from known compartment parameters under the same 62-frame protocol, fit the INR to those frames, and compare the recovered K1, k2, k3, and Vb to the ground truth; if the time-activity-curve error stays low while the parameters drift from the truth, then the identifiability assumption is violated.","supporting_citations":[{"cited_title":"In: International Conference on Medical Image Comput- ing and Computer-Assisted Intervention","cited_arxiv_id":null,"evidence_quote":"Supplies the self-supervised DNN baseline in 2D and 3D configurations that all INR variants are compared against."},{"cited_title":"Nature machine intelligence6(3), 354–367 (2024)","cited_arxiv_id":null,"evidence_quote":"Supplies the 3D CT foundation model used to extract 4096 per-voxel anatomical features for the CTFM variant."},{"cited_title":"European journal of nuclear medicine and molecular imaging pp","cited_arxiv_id":null,"evidence_quote":"Provides the long-axial-field-of-view dynamic PET imaging protocol and kinetic-modelling context for the 62-frame dataset."}],"review_version":1}