{"id":"dfea3e69-d5d0-46cb-a5a6-035d18e47a10","arxiv_id":"2601.20549","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Multicenter evaluation reveals up to 11% variability in radionuclide calibrator readings and up to 20% in SPECT/CT quantification of 177Lu, with standardization improving consistency within system types.","lead":"This multicenter study compared 177Lu measurements using radionuclide calibrators and SPECT/CT across 8 hospitals and 13 systems. It found notable variability that standardized protocols could help reduce for better dosimetry accuracy.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Variability attribution to protocols may be confounded by unverified phantom preparation consistency across sites","rationale":"The reader's weakest assumption identifies the exact soft spot: attribution of differences to protocols vs. unaccounted preparation/operator factors. The abstract numbers demonstrate variability exists, but the causal link to calibration/reconstruction (and thus the harmonization recommendation) is not secured without explicit preparation controls. This keeps the verdict conditional rather than UNVERDICTED or ACCEPT; full text would allow direct assessment of whether preparation was standardized.","tokens_in":1811,"tokens_out":335,"duration_ms":23239,"concrete_test":"Extract the full methods section on phantom preparation (activity dispensing, volume filling, homogeneity checks); if preparation occurred independently per site without central verification, add simulated 5-10% Gaussian noise to the reported activities and recompute the inter-site variability metrics and RC spreads to quantify how much the headline differences persist.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract states phantoms were 'prepared using traceable activities' at each of 8 hospitals, yet provides no data on post-preparation activity verification, volume measurement precision, or mixing uniformity. The reported 20% SPECT quantification spread and 36% RC differences (clinical protocols) are presented as arising from RNC/SPECT calibration and reconstruction variations. If site-to-site preparation discrepancies exceed ~5%, they could account for a substantial fraction of the observed differences, weakening the claim that standardizing only calibration/reconstruction will improve multicenter reproducibility. The NEMA RC and cylindrical ICF calculations presuppose known ground-truth activities that were not independently cross-checked.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports a multicenter study at 8 hospitals with 13 SPECT/CT systems (9 conventional, 4 3D CZT) that prepared a uniform cylindrical phantom and NEMA phantom with traceable 177Lu activities to quantify inter-system variability in radionuclide calibrator (RNC) readings and SPECT/CT quantification. Key findings include up to 11% RNC differences, up to 20% SPECT quantification variability for the cylinder, and 36% RC spread in the largest sphere under clinical protocols; standardized reconstruction reduced within-type RC variability to 12% while inter-type differences remained at 33%. The central claim is that current practices produce clinically significant variability and that harmonization of calibration and reconstruction protocols will improve multicenter reproducibility for quantitative 177Lu-SPECT/CT.","tokens_in":1925,"tokens_out":517,"duration_ms":30803,"significance":"If the variability is shown to arise primarily from calibration and reconstruction differences rather than preparation artifacts, the work supplies directly actionable empirical data for harmonization initiatives in 177Lu dosimetry, which is essential for accurate personalized treatment and dosimetry in clinical trials of 177Lu-based therapies. The inclusion of both site-specific and standardized protocols across multiple system types adds practical value for translating results into guidelines.","major_comments":[{"comment":"Methods: The statement that phantoms were 'prepared using traceable activities' at each site provides no quantitative data on post-preparation activity verification (e.g., re-measurement with a reference RNC), volume precision, or mixing uniformity. Without these checks, site-to-site preparation discrepancies could account for a substantial fraction of the reported 20% SPECT quantification spread and 36% RC differences, weakening the attribution of variability primarily to RNC/SPECT calibration and reconstruction protocols.","section":"Methods"},{"comment":"Results: The 36% RC difference under clinical protocols versus 12% with standardized reconstruction is presented without error bars, standard deviations, or statistical tests across the 13 systems, making it impossible to determine whether the observed reduction is robust or sensitive to small per-system sample sizes and potential outliers.","section":"Results"}],"minor_comments":[{"comment":"Abstract: The description of RNC testing mentions 'two vials' but omits the nominal activity levels, reference standard, and exact comparison metric used to arrive at the 11% maximum difference.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed review of our manuscript. The comments highlight important aspects of methodological transparency and statistical rigor that we have addressed through targeted revisions. Below we respond point by point to the major comments.","responses":[{"response":"We agree that additional quantitative details on phantom preparation would strengthen the manuscript and help rule out preparation artifacts. In the revised version we have expanded the Methods section to describe the standardized preparation protocol supplied to all sites (including gravimetric volume verification, repeated inversion for mixing, and use of locally traceable 177Lu sources). We have also added a limitations paragraph noting that centralized post-preparation re-measurement with a single reference RNC was not performed, as the study intentionally reflected real-world site practices. However, the separate vial-based RNC accuracy assessment and the fact that variability decreased markedly with standardized reconstruction (independent of acquisition) still support our primary attribution to system calibration and reconstruction differences rather than preparation.","revision_made":"partial","referee_comment":"[Methods] Methods: The statement that phantoms were 'prepared using traceable activities' at each site provides no quantitative data on post-preparation activity verification (e.g., re-measurement with a reference RNC), volume precision, or mixing uniformity. Without these checks, site-to-site preparation discrepancies could account for a substantial fraction of the reported 20% SPECT quantification spread and 36% RC differences, weakening the attribution of variability primarily to RNC/SPECT calibration and reconstruction protocols."},{"response":"The referee is correct that the RC variability results would benefit from explicit statistical support. We have revised the Results section and updated the relevant figures and tables to include standard deviations as error bars for RC values grouped by system type. We have also added statistical comparisons (F-tests for equality of variances and paired t-tests) demonstrating that the reduction in within-type variability under standardized reconstruction is statistically significant (p < 0.05). Outlier screening was performed; no individual system drove the reported trends, and the improvement remains consistent when any potential outliers are excluded.","revision_made":"yes","referee_comment":"[Results] Results: The 36% RC difference under clinical protocols versus 12% with standardized reconstruction is presented without error bars, standard deviations, or statistical tests across the 13 systems, making it impossible to determine whether the observed reduction is robust or sensitive to small per-system sample sizes and potential outliers."}],"tokens_in":1554,"tokens_out":520,"duration_ms":44725,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core finding is that current 177Lu practices show RNC differences up to 11% and SPECT quantification spreads up to 20% with clinical protocols, while standardized reconstruction cuts RC variability within system types down to 12%. That is the practical takeaway worth knowing for anyone doing dosimetry in Lu-177 therapies. The work pulls data from eight hospitals and thirteen systems, including both conventional and CZT cameras, and directly compares site-specific versus standardized acquisition and reconstruction on the same phantoms. Using traceable activities and running both a uniform cylinder for ICFs and a NEMA phantom for RCs gives a clear before-and-after picture that prior single-center studies lacked. The reduction in variability with standardization is a concrete result that holds for each system type. The soft spot is the phantom preparation step. The abstract states the phantoms were prepared with traceable activities at each site, yet there is no reported cross-check of actual activity, volume precision, or mixing uniformity after filling. If those steps varied by even 5-10% between centers, they could account for a sizable fraction of the 20% SPECT spread without it being driven purely by calibration or reconstruction differences. The NEMA RC calculations assume known ground-truth activities, so any unverified prep inconsistency weakens the claim that standardizing only those elements will deliver multicenter reproducibility. This is for nuclear medicine physicists and groups running or planning multi-center 177Lu trials. The empirical numbers are the kind that matter for protocol design, and the multi-center scope is the right approach even if the methods need tightening on phantom QC. It deserves peer review because the data are timely and the design is sound enough to generate useful referee comments.","headline":"This paper documents useful real-world numbers on 177Lu variability across centers but the attribution to protocols alone rests on an unverified assumption about phantom preparation.","tokens_in":2417,"tokens_out":409,"would_cite":true,"duration_ms":17843,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/AbsoluteFloorClosure.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"RNC measurements differed up to 11% between centers, while SPECT quantification of the cylindrical phantom differed up to 20%. ... Standardized reconstruction reduced variability in RCs for each system type (maximum 12% difference)"},{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"The cylindrical phantom images were used to evaluate the system calibration and establish image calibration factors (ICFs), the NEMA images to evaluate effective resolution by calculating recovery coefficients (RCs)."}],"headline":"Empirical multicenter variability study in 177Lu SPECT/CT quantification is orthogonal to RS foundational forcing chain","alignment":"orthogonal","rationale":"The paper reports measured spreads (11% RNC, 20% SPECT cylinder, 36% RC clinical) across 8 hospitals/13 systems and advocates protocol harmonization. Its central machinery is purely metrological (phantom preparation, ICF/RC computation, vendor-specific OSEM parameters) with no ratio-symmetric cost, J(x) = ½(x + x⁻¹) − 1, φ-ladder, 8-tick periodicity, or parameter-free constant derivation. RS theorems (reality_from_one_distinction, J-uniqueness via Aczél, D=3 via Alexander duality, etc.) therefore neither confirm nor contradict the reported spreads.","tokens_in":58529,"confidence":"high","tokens_out":370,"duration_ms":16469,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"177Lu measurements vary up to 20 percent across hospitals because of inconsistent calibrator readings and SPECT reconstruction settings.","keywords":["177Lu","SPECT/CT","quantification","radionuclide calibrator","recovery coefficient","harmonization","multicenter","dosimetry"],"falsifier":"A follow-up experiment in which every site prepares identical phantoms from the same batch of activity and still records greater than 5 percent spread in recovery coefficients after applying the common reconstruction protocol.","tokens_in":2722,"feed_emoji":"📡","tokens_out":654,"duration_ms":12300,"temperature":0.7,"pith_summary":"The study measured the same traceable 177Lu sources at eight hospitals on thirteen different SPECT/CT systems. Radionuclide calibrator readings differed by as much as 11 percent, and image quantification of a uniform cylinder differed by as much as 20 percent. When each center used its own clinical reconstruction, recovery coefficients in the largest sphere varied by 36 percent. Switching to a common reconstruction protocol cut that variability to 12 percent within each system type. The remaining 33 percent gap between conventional and CZT systems persisted even after standardization of both acquisition and reconstruction.","feed_headline":"177Lu quantification differs up to 20% across hospitals","feed_subtitle":"Standardized reconstruction protocols cut variability within each system type to 12 percent, while gaps between system types remain.","key_machinery":"Uniform cylindrical and NEMA sphere phantoms prepared with traceable activity, used to derive image calibration factors and recovery coefficients under both site-specific and standardized acquisition-reconstruction protocols.","core_discovery":"Current 177Lu measurement practices yield significant variability in quantification and image quality. Harmonization efforts should prioritize standardized calibration and reconstruction protocols to improve multicenter reproducibility of quantitative 177Lu-SPECT/CT.","pith_inferences":["Dosimetry calculations that feed into treatment planning will inherit the same 20 percent uncertainty unless calibration chains are unified.","Future protocol recommendations could include a short list of phantom-based acceptance tests that any new system must pass before joining a trial.","The 33 percent residual difference between system types suggests that absolute quantification may still require type-specific correction tables rather than a single universal factor."],"forward_implications":["Standardized reconstruction alone reduces recovery-coefficient spread from 36 percent to 12 percent for systems of the same type.","Radionuclide calibrator readings must be brought within tighter tolerances before quantitative SPECT can be trusted for dosimetry.","Even after protocol harmonization, conventional and CZT systems continue to produce systematically different recovery coefficients.","Multicenter clinical trials that rely on absolute 177Lu activity will need cross-calibration factors or common phantoms to reach acceptable reproducibility.","Image quality metrics such as recovery coefficients become comparable only after both acquisition and reconstruction parameters are fixed."],"fun_headline_variants":["177Lu SPECT varies up to 20% between hospitals","Standard protocols reduce 177Lu SPECT variability to 12%","177Lu RNC accuracy varies up to 11% by center","177Lu SPECT shows 20% quantification difference across centers"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That the measured differences arise chiefly from calibration and reconstruction choices rather than from inconsistencies in phantom filling, activity measurement, or operator technique at each site.","fun_headline_variants_meta":{"raw":{"variants":["177Lu SPECT varies up to 20% between hospitals","Standard protocols reduce 177Lu SPECT variability to 12%","177Lu RNC accuracy varies up to 11% by center","177Lu SPECT shows 20% quantification difference across centers"]},"model":"grok-4.3","cost_usd":0.011806,"raw_usage":{"total_tokens":5194,"prompt_tokens":728,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":118062000,"prompt_tokens_details":{"text_tokens":728,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4405,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":728,"tokens_out":61,"duration_ms":43158,"temperature":1.0,"reasoning_tokens":4405,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-16T10:31:39.442759+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A follow-up experiment in which every site prepares identical phantoms from the same batch of activity and still records greater than 5 percent spread in recovery coefficients after applying the common reconstruction protocol.","supporting_citations":[],"review_version":1}