{"id":"8cf9d066-45c5-4a75-9f5c-9a08149f0678","arxiv_id":"1906.08427","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Deep learning model achieves 1.15-2.88 mm accuracy for markerless prostate localization from projection X-rays in three patients, matching fiducial-derived positions.","lead":"This paper trains a deep learning model on annotated kV X-ray images to locate the prostate without implanted fiducials. If accurate, it could replace invasive marker placement for daily positioning in prostate radiotherapy.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Small retrospective cohort (n=3) provides insufficient evidence that the DL model generalizes to unseen patients and image qualities","rationale":"The reader's weakest_assumption already isolates the generalization step that the n=3 design cannot secure; the concrete_test directly probes that step without requiring new model training.","tokens_in":1819,"tokens_out":285,"duration_ms":16967,"concrete_test":"Acquire kV projections from 15–20 additional prostate patients (prospective or independent retrospective set) with fiducials; run the trained model, compute per-fraction mean absolute error and 95th-percentile error against fiducial-derived positions; if median error exceeds 3 mm or >15 % of fractions exceed 5 mm, the generalization claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim requires that the model, trained on annotated kV projections, produces clinically usable localization (deviations 1.15–2.88 mm) on new patients without fiducials. The study reports results only from three VMAT patients; no cross-validation, inter-patient error breakdown, or prospective hold-out set is described. Because annotations themselves derive from fiducial-based ground truth on the same limited cohort, the reported consistency does not test whether the learned mapping holds when anatomy, implant status, or kV image statistics differ.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes a deep learning model, pre-trained on thousands of annotated kV projection X-ray images, to perform markerless prostate localization for image-guided radiotherapy. It reports quantitative evaluation on three retrospectively studied VMAT patients, with localization deviations from fiducial-derived annotations of 1.66–2.77 mm (AP) and 1.15–2.88 mm (lateral), concluding that highly accurate markerless localization is achievable and clinically valuable.","tokens_in":1938,"tokens_out":463,"duration_ms":15364,"significance":"If the central claim holds under proper validation, the work would offer a non-invasive alternative to fiducial implantation, eliminating associated procedural risks and costs while enabling daily positioning and real-time tracking from routine kV images. The approach leverages standard OBI hardware and demonstrates consistency with fiducial ground truth on the tested cases.","major_comments":[{"comment":"Evaluation on three patients (abstract and patient-study section): the reported deviations support mm-level accuracy on this small retrospective cohort, but the absence of cross-validation, inter-patient error breakdown, prospective hold-out testing, or details on whether test images come from training patients undermines the generalization claim required for the assertion that the model works on 'new patients without fiducials.' This is load-bearing for the headline result.","section":"patient study / results"},{"comment":"Methods and abstract: no model architecture, training hyperparameters, loss function, data split strategy, or statistical tests (error bars, p-values) are supplied, preventing assessment of whether the 1.15–2.88 mm range reflects robust performance or overfitting to the limited annotated data.","section":"methods / abstract"}],"minor_comments":[{"comment":"The claim of demonstrating the result 'for the first time' would benefit from explicit comparison to prior markerless DL or template-matching approaches in the introduction.","section":"introduction"}],"recommendation":"major_revision","confidential_remarks":"The small n=3 cohort is a common limitation in early DL medical-imaging papers; the manuscript would be strengthened by adding at least one additional validation cohort or leave-one-patient-out analysis rather than expanding scope."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments on our manuscript. We provide point-by-point responses to the major comments and indicate where revisions will be made.","responses":[{"response":"We note that the deep learning model was pre-trained on several thousand annotated kV projection images, which forms the basis for its ability to generalize to new patients. The evaluation on three patients was performed retrospectively to compare against fiducial-derived positions, and the test images are from these patients not involved in the pre-training. We agree that additional analyses such as cross-validation and inter-patient breakdowns would strengthen the paper. We will revise to include an inter-patient error breakdown and clarify the data usage. However, as this is a retrospective study, prospective hold-out testing is not feasible in the current work, and we will discuss this limitation explicitly.","revision_made":"partial","referee_comment":"[patient study / results] Evaluation on three patients (abstract and patient-study section): the reported deviations support mm-level accuracy on this small retrospective cohort, but the absence of cross-validation, inter-patient error breakdown, prospective hold-out testing, or details on whether test images come from training patients undermines the generalization claim required for the assertion that the model works on 'new patients without fiducials.' This is load-bearing for the headline result."},{"response":"We agree with this observation. The current version of the manuscript does not provide these details in the abstract or methods. We will revise the manuscript to include a detailed description of the model architecture, training hyperparameters, loss function, data split strategy, and incorporate statistical tests with error bars and p-values where appropriate to allow proper assessment of the results' robustness.","revision_made":"yes","referee_comment":"[methods / abstract] Methods and abstract: no model architecture, training hyperparameters, loss function, data split strategy, or statistical tests (error bars, p-values) are supplied, preventing assessment of whether the 1.15–2.88 mm range reflects robust performance or overfitting to the limited annotated data."}],"tokens_in":1425,"tokens_out":464,"duration_ms":29856,"standing_objections":["Prospective hold-out testing on new patients (retrospective study design prevents this)"]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this work trains a deep learning model on annotated kV projections and reports 1.15–2.88 mm deviations from fiducial ground truth on three VMAT prostate cases. That is the first quantitative clinical demonstration of the approach they describe, and it shows the basic idea can produce usable numbers without implants or daily CBCT. The setup makes sense: train on existing images, then apply to new projections from the on-board imager. The reported consistency with fiducials on those cases is a reasonable first check. The problem itself matters—avoiding fiducial procedures reduces risk and cost for a common treatment. The paper does that part cleanly. The soft spot is the sample size. Three patients, all drawn from the same limited cohort where the annotations themselves came from fiducials, does not test whether the mapping holds for new anatomies, different image qualities, or patients without implants. No per-patient breakdown, no cross-validation, and no statistical tests are described in the abstract, so the error range cannot be read as a general performance claim. The stress-test note is accurate on this point; the current numbers show consistency within the training distribution but do not yet demonstrate robustness outside it. This is for medical physicists and IGRT researchers who want to explore non-invasive tracking. A reader would get value from the concept and the reported numbers as a starting point, but would treat the accuracy claim as preliminary. It deserves peer review because the clinical need is real and the method is distinct from prior fiducial or CBCT work; referees can push for larger validation and clearer methods without the paper being dismissed outright.","headline":"The paper gets mm-level markerless prostate localization from kV images on three patients, but the tiny retrospective sample leaves generalization untested.","tokens_in":2410,"tokens_out":400,"would_cite":false,"duration_ms":26214,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"DL markerless prostate localization via Faster R-CNN on kV projections has no connection to RS forcing chain","alignment":"orthogonal","rationale":"The paper's core machinery (region-proposal CNN trained on deformed-CT DRRs, validated on n=3 VMAT cases) is a standard empirical computer-vision pipeline. It neither invokes nor parallels any RS element: J-cost functional equation, φ-ladder, 8-tick periodicity, Alexander-duality D=3 forcing, or parameter-free constant derivations. RS modules such as Cost.FunctionalEquation, Foundation.AlexanderDuality, and Foundation.RealityFromDistinction are irrelevant; the domain (clinical IGRT) lies outside the RS structural canon.","tokens_in":44923,"confidence":"high","tokens_out":167,"duration_ms":4269,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Deep learning locates the prostate in routine kV X-ray images with 1-3 mm accuracy without fiducials.","keywords":["prostate cancer","markerless localization","deep learning","kV X-ray","image-guided radiotherapy","target tracking","VMAT","fiducial-free"],"falsifier":"A study comparing the model's predictions to fiducial positions or CBCT in a larger group of unseen patients would confirm or refute the accuracy claims.","tokens_in":2734,"feed_emoji":"🩺","tokens_out":386,"duration_ms":23844,"temperature":0.7,"pith_summary":"This paper demonstrates that a deep learning model trained on thousands of annotated kV projection images can determine prostate position in new images. The approach eliminates the need for implanted fiducials in prostate radiotherapy. If the results hold, it reduces patient risks from invasive procedures while enabling accurate daily setup and tracking. The method was validated on three patients, showing consistency with fiducial-derived positions.","feed_headline":"Deep learning tracks prostate without implanted markers","feed_subtitle":"Routine kV images yield 1-3 mm accuracy matching fiducials in three patients.","key_machinery":"The deep learning model trained to interpret projection kV X-ray images for prostate target identification.","core_discovery":"The authors show that their pre-trained deep learning model identifies the prostate location in projection kV X-ray images. Deviations from annotations were 1.66 mm to 2.77 mm anterior-posterior and 1.15 mm to 2.88 mm lateral. Positions matched those from fiducials, establishing that highly accurate markerless localization is possible.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Deep learning tracks prostate without markers on kV X-rays","Markerless prostate localization using deep learning on X-ray images","Prostate tracked by deep learning model on routine kV projections","Deep learning model localizes prostate in kV images without fiducials"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The annotations accurately represent the true prostate position and the model generalizes to new patients and image qualities.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning tracks prostate without markers on kV X-rays","Markerless prostate localization using deep learning on X-ray images","Prostate tracked by deep learning model on routine kV projections","Deep learning model localizes prostate in kV images without fiducials"]},"model":"grok-4.3","cost_usd":0.005785,"raw_usage":{"total_tokens":2781,"prompt_tokens":719,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":57849500,"prompt_tokens_details":{"text_tokens":719,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1993,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":719,"tokens_out":69,"duration_ms":13552,"temperature":1.0,"reasoning_tokens":1993,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T19:32:41.964266+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A study comparing the model's predictions to fiducial positions or CBCT in a larger group of unseen patients would confirm or refute the accuracy claims.","supporting_citations":[],"review_version":1}