{"id":"554d28fb-fde8-43f3-bf09-1addeb9e0fb2","arxiv_id":"2608.03612","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A shared, source-predicted intensity coordinate for latent breast MRI virtual contrast enhancement improves eight internal-cohort quality metrics over fixed and separate coordinate baselines.","lead":"This paper presents Predictive Enhancement Calibration (PEC), a way to map breast MRI images into the shared intensity range that a pretrained image generator expects, instead of using a fixed or per-image range. On an internal 100-patient cohort, the method improved all eight reported quality metrics for virtual contrast enhancement, with statistically strong evidence for two of them.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MAMA100-informing development plus cohort-only FRD point estimates leave the headline radiomics gain unsupported; an external frozen-pipeline test is needed.","rationale":"The reader's weakest_assumption correctly identifies the central fragility: MAMA100 informed development, yet all headline comparisons use the same cohort, and only 2 of 8 metrics have CIs excluding zero with FRD/AUROC as point estimates. My stress-test converges on the same load-bearing concern. I considered two other candidate concerns: (i) the separate-coordinate baseline confounds sharing with endpoint choice, but the paper's primary comparison is PEC vs fixed-wide, which isolates sharing; (ii) the source-predicted endpoint's poor round-trip FRD (9.186 vs 1.280 oracle) might undermine the conditional gains, but the conditional experiment shows PEC still improves point estimates, and the paper explicitly separates representation-level and synthesis-level errors. Neither is as decisive as the cohort-informed selection risk for the central generalization claim. The paper is internally coherent, discloses its limitations ('external cohorts ... remain future tests'), and provides source code; these are credits. But the strongest claim that PEC is the right interface depends on the reported MAMA100 point estimates being unbiased for future deployment. Since the evaluation cohort informed design and the radiomic metric lacks uncertainty quantification, the claim is conditional at best. The reader's CONDITIONAL verdict is appropriate, so no verdict change is needed. The concrete external-frozen-pipeline test would settle whether the concern lands empirically.","tokens_in":7790,"tokens_out":3457,"duration_ms":35204,"concrete_test":"Hold out a disjoint external cohort (e.g., the MAMA-MIA test split or DUKE/ISPY2 cases not used during any development, sampling selection, or predictor fitting) and evaluate the exact archived pipeline with no re-tuning: PEC versus shared fixed-wide, identical to the final protocol of Table 2. Compute paired bootstrap 95% CIs for all eight metrics, including FRD and both AUROCs via patient-level resampling. If PEC does not improve at least MSE and LPIPS with CIs excluding zero, and FRD's CI for the PEC−fixed-wide difference includes zero, the central claim that PEC is the right interface—and specifically the radiomics gain—fails to generalize. Pre-register this analysis before running it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim (Table 2: PEC improves all eight point estimates, and specifically reduces FRD from 4.838 to 4.429) is the load-bearing evidence for 'PEC is the right interface' for latent breast MRI VCE. But all comparisons are on MAMA100, the fixed internal cohort that the authors state 'informed development' (Sec. 6). Iterative choices—top-eight slice sampling, the Q99.99 endpoint, gamma 2.2, the ROI-loss weight, the FT-Transformer predictor, and model selection—were plausibly guided by metrics computed on the same 100-patient cohort that is then reported as the evaluation set. Only two of eight metrics (MSE, LPIPS) have paired 95% CIs excluding zero; FRD and both AUROCs are cohort-level point estimates with no uncertainty interval. This is especially consequential because the paper's core motivation is radiomic fidelity, yet the only statistically supported gains are global reconstruction metrics, not the radiomics-relevant FRD. The internal evidence is coherent, but without an external, pre-registered evaluation, the point-estimate improvements—particularly the FRD decrease—could reflect selection on the development cohort rather than a genuine advantage of the shared predicted coordinate. This is a correctness/validity risk, not an allegation of misconduct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses a real interface problem for latent breast MRI virtual contrast enhancement: pretrained natural-image autoencoders require bounded inputs, while MRI intensities are non-canonical and patient-specific. The authors propose Predictive Enhancement Calibration (PEC), which uses a shared, case-adaptive intensity interval [ℓ(x), u(y)] for source and target during training, with the target's upper endpoint predicted from source statistics at inference. They integrate PEC with a frozen FLUX.2 latent flow transformer via LoRA reference conditioning and an ROI-weighted flow loss. On a fixed internal 100-patient development cohort (MAMA100), they compare PEC against fixed-wide and separate-coordinate baselines under near-matched training budgets. They report that PEC improves all eight point estimates, with paired bootstrap confidence intervals excluding zero for MSE and LPIPS, while SSIMt, Dice, and HD95 intervals include zero and FRD/AUROCs are cohort-level point estimates. The paper also uses target round trips to separate representation loss from synthesis error and includes a sensitivity control for endpoint-scale error.","tokens_in":8093,"tokens_out":4230,"duration_ms":42481,"significance":"If the result holds, PEC is an appealingly simple and portable calibration layer that could be applied to other latent generative models for medical imaging. The paper is methodical in isolating representation effects before generation, in matching training budgets, and in reporting which confidence intervals exclude zero. The release of code is a further strength. However, the central radiomics motivation is not yet statistically supported: the only metrics with paired CIs excluding zero are global reconstruction metrics, whereas the headline radiomic fidelity metric (FRD) is a cohort-level point estimate. The evaluation cohort is explicitly stated to have informed development, so the risk of selection inflation is real. The idea is sound and worth pursuing, but the evidence as presented is too fragile for acceptance without additional validation.","major_comments":[{"comment":"The development and evaluation cohort are the same: the authors state 'MAMA100 informed development under archived standardization' and all comparative results are reported on this fixed 100-patient cohort. Because iterative choices (top-eight slice sampling, Q99.99 endpoint, gamma=2.2, ROI-loss weight λ=4, FT-Transformer predictor, final-epoch selection) plausibly used MAMA100 metrics, the point-estimate improvements in Table 2 could reflect selection on the evaluation cohort. This is especially consequential for FRD, the metric most aligned with the paper's radiomic-fidelity motivation, but only MSE and LPIPS have paired CIs excluding zero; FRD and both AUROCs are cohort-level point estimates without uncertainty intervals. The internal evidence is coherent but insufficient. Please provide an external frozen-pipeline evaluation, or at minimum a prespecified development/evaluation split","section":"Section 6 / Table 2"},{"comment":"There is a tension between the representation-level and synthesis-level results. Table 1 shows source-predicted PEC round trips have FRD 9.186 and FRD_VAE 9.632, markedly worse than oracle PEC (1.280/2.793) and even fixed-wide (1.663/3.133). Yet Table 2 shows conditional PEC synthesis has FRD 4.429, better than fixed-wide 4.838. If the shared predicted coordinate is the mechanism, one would expect the predicted coordinate's poor representation fidelity to hamper synthesis, not improve it. The sensitivity control (oracle encoding, predicted decoding) helps separate endpoint-scale error from tail clipping, but the conditional result still needs explanation: why does a coordinate that alone degrades radiomic distance produce better synthetic FRD? Please analyze this interaction, e.g., report oracle-coordinate conditional generation as an upper bound, and examine whether the FRD gain is driv","section":"Table 1 vs. Table 2 / Section 5.1"},{"comment":"The endpoint predictor fφ uses 47 hand-selected source statistics and is selected on only 20 held-out Yunnan cases. The reported endpoint MAE 1.371 and Pearson correlation 0.975 come from this small selection set, so the predictive performance may be optimistic. With 47 free statistics and no ablation, it is unclear which features are necessary and whether the predictor is overfitting the small selection set. Please report cross-validated endpoint prediction performance on a larger source-only set (e.g., held-out MAMA training-pool patients), include error bars, and justify the choice of the 47 statistics or provide an ablation.","section":"Section 3.2 / Section 4.1"}],"minor_comments":[{"comment":"The source statistics s(x) are not formally defined; please specify the 47 features and the exact loss terms ('asymmetric and monotonicity terms') used for the predictor.","section":"Eq. (4)"},{"comment":"The term 'MAMA training patients' is ambiguous: clarify whether this means the 1,406-patient training pool, and specify which normalization statistics (e.g., z-score mean/std) are computed on which data.","section":"Section 4.1"},{"comment":"The labels 'PEC gain I' and 'PEC gain II' are not explained in the caption or text; please define how these cases were selected and what the labels indicate.","section":"Fig. 2"},{"comment":"The notation Q99.99(y) is computed over finite pixels in the original field of view, but it is worth clarifying whether background air values are included before any masking and how this interacts with padding.","section":"Section 3.2 / Eq. (1)"},{"comment":"Reference [21] could be cited with more detailed metadata (dataset name, version, access date) so readers can locate the Yunnan cohort used for predictor selection.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent and technically careful, but the evidence for the core radiomic claim is thin. The 'all eight point estimates improve' framing overstates what the statistics support, given that only two of eight metrics have intervals excluding zero and the evaluation cohort informed development. The authors disclose this limitation honestly, so I do not suspect misconduct; the issue is validity risk. A revision that adds an external or properly separated evaluation, uncertainty for FRD/AUROC, and reconciliation of Table 1 vs. Table 2 would substantially strengthen the contribution. The work is likely within scope for a medical-imaging journal if these concerns are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick read of arXiv 2608.03612. The genuinely new piece is PEC: a shared, case-adaptive intensity coordinate for latent breast MRI VCE whose upper endpoint is predicted from the pre-contrast image. The paper states the problem well — bounded natural-image autoencoders clash with MRI's unbounded, scanner-dependent intensity — and shows before generation that endpoint choice changes FRD (the sweep from Q99.5 to Q99.99 drops FRD from 10.598 to 0.090). The target round-trip table is a nice design: it isolates codec/representation loss from the conditional generator, and oracle vs predicted PEC separates ideal capacity from estimation error. The predictor is genuinely out-of-sample (trained on the MAMA training pool plus 80 Yunnan cases, selected on 20 held-out Yunnan, evaluated on MAMA100), and code is posted. That is real work and worth credit.\n\nThe soft spots are real. The paper states flatly in Sec. 6 that MAMA100 informed development, then reports all eight point estimates on that same fixed cohort. Iterative choices — slice cap, Q99.99, gamma 2.2, ROI weight, LoRA rank, predictor architecture — plausibly tracked metrics on the same 100 patients. Only MSE and LPIPS have paired CIs excluding zero; FRD and both AUROCs are cohort-level point estimates with no intervals. Since FRD is the radiomics motivation of the paper, that is the load-bearing number and it is the least supported. Add to that Table 1: predicted-PEC round trips push FRD from 1.280 (oracle) to 9.186, so the conditional gain depends on the generator tolerating tail clipping. And the 'separate coordinates' baseline changes two things at once (separate vs shared and source-max vs predicted endpoint), so it does not isolate sharing alone.\n\nNone of this smells like misconduct. Limitations are disclosed, the sampling change is explained in Sec. 4.2, and the predictor evaluation is genuinely external. It is a correctness/validity risk, not fraud. The fix is straightforward: freeze the pipeline, run an external cohort, report per-metric CIs for FRD and AUROC, and add an ablation that changes only the sharing decision.\n\nWho is this for? People building latent generative VCE in medical imaging. It is a solid submission in need of one more validation round. I would send it to peer review — the problem is well-defined and the mechanism is testable — but I would not take the headline FRD at face value until an external evaluation appears.","headline":"A clearly-posed intensity-calibration fix for latent breast MRI VCE, with one honest limitation: the headline gains all come from the cohort that guided development.","tokens_in":8613,"tokens_out":2335,"would_cite":true,"duration_ms":22273,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A predicted intensity ceiling improves virtual contrast breast MRI on all eight metrics.","keywords":["virtual contrast enhancement","breast MRI","latent flow transformer","intensity calibration","radiomics","predictive enhancement calibration","DCE-MRI","image synthesis"],"falsifier":"Run the near-matched conditional comparison (fixed-wide, shared PEC, separate coordinates) on a held-out cohort not used in development, and test whether PEC's MSE and LPIPS paired differences remain negative with confidence intervals excluding zero; additionally, if replacing the predicted upper endpoint with the oracle endpoint does not substantially close the FRD gap (from 9.186 toward 1.280), the source-predictive interface itself is the weak point.","tokens_in":7627,"feed_emoji":"🩻","tokens_out":7731,"duration_ms":62107,"temperature":0.7,"pith_summary":"Virtual contrast enhancement (VCE) aims to synthesize a post-contrast breast MRI from a pre-contrast scan, avoiding the need for gadolinium. The paper argues that the main barrier to using modern pretrained latent image generators for this task is not the network but the intensity coordinate: MRI has no fixed intensity unit, while a frozen natural-image autoencoder expects bounded inputs. It proposes Predictive Enhancement Calibration (PEC), which places each pre-contrast/peak-contrast pair in a shared, case-adaptive interval during training and, at inference, predicts the unavailable upper endpoint from source-image statistics. On the fixed MAMA100 development cohort, source-only PEC improves all eight evaluation metrics over matched fixed-window and separate-coordinate baselines. If correct, this makes strong pretrained latent generators usable for MRI synthesis with only a small trainable adapter.","feed_headline":"Predicting the intensity ceiling lifts virtual breast MRI contrast","feed_subtitle":"A case-adaptive shared coordinate beats fixed windows on every point estimate in the MAMA100 cohort.","key_machinery":"The key object is the PEC coordinate transform: C(z) = clip(((z − ℓ)/(u − ℓ))^(1/γ), 0, 1) with γ = 2.2, and its inverse D(a) = ℓ + (u − ℓ) a^γ. It is 'shared' because one interval [ℓ, u] is used for both source encoding and output decoding; 'case-adaptive' because ℓ comes from the pre-contrast slice and u from the target's 99.99th percentile; and 'predictive' because at inference u is replaced by an FT-Transformer estimate û(x) built from 47 source-only statistics with asymmetric and monotonicity losses. The gamma exponent allocates more 8-bit code levels to the densely occupied lower intensity range, while the 99.99th-percentile endpoint preserves the sparse enhancement tail. Around this c","core_discovery":"The paper's central claim is that intensity mapping is a learnable interface: the upper endpoint of the encoding window changes radiomic distances before generation, and independent source/target scaling assigns different physical meanings to equal coordinate values. PEC encodes both images into one case-adaptive interval [ℓ(x), u(y)], with ℓ(x) the pre-contrast finite minimum and u(y) the 99.99th percentile of the target; an FT-Transformer predicts û(x) from 47 source statistics, so deployment needs neither target nor tumor mask. Oracle PEC reaches the lowest round-trip Fréchet Radiomic Distance (1.280); source-predicted PEC degrades to 9.186, exposing tail clipping. In conditional generati","pith_inferences":["The same shared-predictive-coordinate idea could transfer to other MRI-to-MRI tasks (e.g., T1-to-T2 synthesis) and other latent generators, since the conflict is between MRI's non-canonical scale and any bounded autoencoder.","A stronger endpoint predictor—perhaps predicting a full quantile curve or using a generative model of the tail—could close the gap between predicted-window FRD (9.186) and oracle PEC (1.280), likely improving the conditional FRD as well.","Because MAMA100 informed development, the honest test of the method is a preregistered or external-cohort comparison; the paper's own sensitivity control (oracle encoding, predicted decoding) offers a cheap way to separate window capacity from predictor error on new data."],"forward_implications":["Intensity calibration should be treated as part of the generative model: a fixed upper endpoint that clips the enhancement tail changes FRD before synthesis, so future VCE systems should report calibration choices.","A pretrained natural-image latent generator can be reused for breast MRI VCE with only a rank-128 LoRA adapter plus a small endpoint predictor, avoiding expensive medical-codec training.","Equivalent-pair coordinates make source and target values comparable at inference; separate source/target scaling is a measurable handicap (MSE 0.9096 vs 0.7493 for PEC).","The endpoint predictor reaches high correlation (0.975) but still leaves tail clipping; improving tail fidelity is the remaining bottleneck for radiomic-level synthesis.","Lesion-level gains (Dice, HD95, tumor-ROI AUROC) are point-estimate improvements with wide confidence intervals; the strongest paired evidence is on global image metrics MSE and LPIPS."],"supporting_citations":[{"why":"FLUX.2 [klein] supplies the pretrained latent flow transformer and autoencoder that PEC adapts.","marker":"[1]"},{"why":"EasyControl provides the reference-token LoRA conditioning principle used to inject the calibrated pre-contrast image.","marker":"[22]"},{"why":"MAMA-MIA is the multicenter dataset from which the training pool and fixed MAMA100 development cohort are drawn.","marker":"[4]"},{"why":"MAMA-SYNTH defines the VCE setting and the peak-lesion endpoint construction used for evaluation.","marker":"[13]"},{"why":"Fréchet Radiomic Distance is the metric that exposes how intensity-window endpoints change radiomic fidelity before generation.","marker":"[8]"},{"why":"FT-Transformer is the architecture of the endpoint predictor that estimates û(x) from source statistics.","marker":"[5]"},{"why":"Latent diffusion motivates the bounded-autoencoder representation that creates the intensity-coordinate conflict.","marker":"[18]"},{"why":"Rectified flow transformers provide the scalable denoiser design underlying the FLUX.2 backbone.","marker":"[3]"},{"why":"The public Yunnan DCE-MRI data are used to fit and select the endpoint predictor.","marker":"[21]"},{"why":"LoRA supplies the parameter-efficient adaptation mechanism for the frozen transformer.","marker":"[6]"}],"fun_headline_variants":["Predicting MRI intensity ceiling lifts virtual contrast","Case-adaptive scaling beats fixed windows in breast MRI","Learn the intensity limit for sharper virtual breast MRI","Fix the tail: adaptive scaling for virtual breast MRI"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The central claim rests on the assumption that iterative design on the fixed MAMA100 cohort did not inflate PEC's advantage over the fixed-wide baseline in the same-cohort evaluation.","fun_headline_variants_meta":{"raw":{"variants":["Predicting MRI intensity ceiling lifts virtual contrast","Case-adaptive scaling beats fixed windows in breast MRI","Learn the intensity limit for sharper virtual breast MRI","Fix the tail: adaptive scaling for virtual breast MRI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00014,"raw_usage":{"total_tokens":992,"prompt_tokens":735,"completion_tokens":257,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":197}},"tokens_in":479,"tokens_out":257,"duration_ms":3965,"temperature":1.0,"reasoning_tokens":197,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:01:10.748701+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the near-matched conditional comparison (fixed-wide, shared PEC, separate coordinates) on a held-out cohort not used in development, and test whether PEC's MSE and LPIPS paired differences remain negative with confidence intervals excluding zero; additionally, if replacing the predicted upper endpoint with the oracle endpoint does not substantially close the FRD gap (from 9.186 toward 1.280), the source-predictive interface itself is the weak point.","supporting_citations":[{"cited_title":"https://github.com/black-forest-labs/ flux2(2026)","cited_arxiv_id":null,"evidence_quote":"FLUX.2 [klein] supplies the pretrained latent flow transformer and autoencoder that PEC adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"MAMA-SYNTH defines the VCE setting and the peak-lesion endpoint construction used for evaluation."},{"cited_title":"In: Advances in Neural Information Processing Systems","cited_arxiv_id":null,"evidence_quote":"FT-Transformer is the architecture of the endpoint predictor that estimates û(x) from source statistics."},{"cited_title":"In: Proceedings of the 41st International Conference on Machine Learning","cited_arxiv_id":null,"evidence_quote":"Rectified flow transformers provide the scalable denoiser design underlying the FLUX.2 backbone."},{"cited_title":"https://doi.org/10.5281/ zenodo.8068383","cited_arxiv_id":null,"evidence_quote":"The public Yunnan DCE-MRI data are used to fit and select the endpoint predictor."}],"review_version":1}