REVIEW 3 major objections 5 minor 1 cited by
A lightweight probe ranks 3D-CT vision–language model candidates almost as well as full fine-tuning.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 06:11 UTC pith:L4GSSI3Y
load-bearing objection Honest pilot with a genuinely useful construction framework; the headline r=0.95 is probably inflated by shared RadBERT labels and should not be treated as established. the 3 major comments →
Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that frozen-token probing is a usable stand-in for 3D-CT encoder and compression selection. Over six (encoder × compression) cells, the AUROC of a disease-finding probe trained on cached embeddings orders the cells almost exactly as full LLM fine-tuning does, with Pearson r=0.95 and Spearman ρ=0.89 on the paired metric, report-generation clinical micro-F1. The paper reads this as an ordinal claim, not an exact estimator: the probe predicts the ranking, and within-encoder near-ties between compression variants are not reliably ordered. The relationship holds across nine of ten read-out heads; only one over-parameterized head collapses and inverts the correlation.
What carries the argument
The correlational identity carrying the argument is the cross-cell rank agreement between disease-probe AUROC and downstream clinical micro-F1, measured over six matched (encoder × compression) cells. The paper's methodological contribution is a pair of validation gates applied to every probe target: scale-sanity, which rejects clinical thresholds that produce near-degenerate class splits, and probe-separability, which drops or flags attributes a strong linear probe cannot decode. These gates ensure the probe labels are well-scaled and decodable, making the predictive-validity claim falsifiable. The probe itself is an attention-pooling read-out head trained on normalized, cached embeddings.
Load-bearing premise
The load-bearing premise is that the six-cell rank correlation—computed on an earlier label revision of the probing pipeline—transfers to the exact gated 15-attribute / 18-finding benchmark and to the broader candidate space; if error bars on those six points include a near-zero correlation, the central claim lacks current support.
What would settle it
Re-run the probe-to-downstream correlation on the exact gated benchmark described in the paper, across an expanded grid of at least 12 cells with multiple seeds per cell. If the 95% confidence interval for Spearman ρ includes values below 0.5, or if adding a new encoder/compression cell breaks the monotonic ordering, the ordinal claim would be effectively refuted.
If this is right
- Encoder and compression choices can be screened in minutes with frozen-token probes, reserving full fine-tuning (about one GPU-day per configuration) for top-ranked finalists.
- The probe ranking reliably identifies the best encoder, though it may not break statistically indistinguishable near-ties between an encoder's own compression variants.
- The two validation gates provide a reusable protocol for constructing probing benchmarks whose labels are well-scaled, decodable, and leakage-free.
- If the ordinal relationship holds beyond the six cells, the approach could cut the compute cost of model-selection sweeps for 3D-CT VLMs by orders of magnitude.
Where Pith is reading between the lines
- The six-cell correlation is a small sample; a re-run on the full gated benchmark with more cells and seed error bars could widen the credible interval to include a much lower rank correlation, so the headline number should be treated as provisional.
- Because both probe and downstream labels derive from the same automated-label-extraction process, shared label noise may inflate the correlation; an independent label source would test this.
- The same cost-geometry argument—cheap probing of frozen features to rank expensive generative fine-tuning—likely transfers to other medical imaging modalities or generation tasks, but external validity is untested.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a cheap-probing methodology for selecting a frozen 3D-CT image encoder and a token-compression scheme for a vision--language model. It constructs a benchmark of (encoder × compression) cells with image-grounded clinical attributes, two validation gates (scale-sanity and probe-separability), and an audit rubric. Ten read-out heads are compared on identical frozen tokens, and the central empirical claim is that disease-probe AUROC predicts report-generation clinical micro-F1 at r = 0.95 / ρ = 0.89 over six matched cells, robustly across read-out choice. The authors frame the claim as ordinal and preliminary, and disclose several limitations: the correlation was computed on an earlier label revision, not the gated benchmark; within-encoder compression near-ties are not reliably ordered; and re-running on the gated benchmark with more cells and error bars is ongoing.
Significance. If the central screening claim holds, the paper offers a practically valuable result: encoder/compression choices could be screened in minutes rather than GPU-days. The methodological contributions are real and should be credited: the two validation gates are a sensible anti-degeneracy mechanism, the probe-audit rubric is a useful protocol, the negative control with shuffled labels is appropriate, and the read-out robustness table directly addresses one natural objection. The paper is also unusually honest about the preliminary nature of the evidence. However, the current quantitative anchor is a six-point correlation computed on an earlier label revision and on labels that share a RadBERT extraction pipeline on both sides of the correlation. These issues are load-bearing for the screening claim and need to be resolved before the paper's central promise can be accepted.
major comments (3)
- [Section 5, Figure 3, Table 1] The central correlation r=0.95/ρ=0.89 is computed on an earlier label revision of the probing pipeline, not on the exact gated 15-attribute / 18-finding benchmark that the paper contributes. This is stated in Section 3 ('From construction to validation') and in Section 6. Consequently, the headline quantitative result is not attached to the benchmark that the paper's methodology section builds. A re-run on the gated benchmark is needed before the paper can claim that probe scores predict performance on that benchmark.
- [Section 3 and Section 5] Both sides of the correlation depend on RadBERT-extracted disease findings. Section 3 states that the 18 disease findings are 'extracted with RadBERT,' and Section 5 defines clinical micro-F1 as 'RadBERT-extracted findings, generated vs. reference.' If the same text model extracts the probe labels and the reference labels, common-mode label noise can inflate r=0.95/ρ=0.89 even if the probe has no true relationship to clinically meaningful report quality. The negative control and read-out robustness do not address this confound. The authors should report results with independently obtained labels (e.g., human-annotated findings or a different extraction method) or quantify the extraction error and show that it does not drive the correlation.
- [Section 5, 'Where the ranking mismatches'] The ordinal claim is not supported at the compression-selection level. The paper discloses that for CoLiPri, the probe ranks pool > adp while report-generation ranks adp > pool. Since compression choice is one of the two axes of the claimed screening space, this is not a minor caveat: it means the probe currently cannot be trusted to choose among an encoder's own compression variants. The six-cell correlation is dominated by cross-encoder differences, and the paper provides no error bars or confidence intervals. With n=6, the credible interval for ρ=0.89 is wide even before considering the disclosed near-tie. This should be stated more prominently, and the screening recommendation should be scoped to encoder selection, not compression selection, pending error-bar analysis.
minor comments (5)
- [Section 3] The text says every attribute is 'image-grounded' and 'never from a report,' then states that the 18 disease findings are 'extracted with RadBERT.' Since RadBERT is a radiology text model, this appears internally inconsistent. Please clarify whether the disease-finding labels are derived from CT images directly, from reports, or from a hybrid pipeline.
- [Section 5] The cell enumeration is inconsistent: the text says 'six matched cells (CoLiPri / CT-CLIP / BTB3D under adp, pool, or none),' which would be nine combinations, while Table 1 and Figure 2 refer to six cells and the reported results correspond to three encoders × two compressions. Please correct the enumeration.
- [References] Several reference entries appear incomplete or mismatched: [1] and [2] share the same arXiv identifier, and the titles are truncated. Please provide complete bibliographic information.
- [Section 4] The 'GLocal' head is mentioned without definition. Please specify its architecture or provide a citation so the read-out comparison is reproducible.
- [Throughout] The paper uses informal phrases such as 'stake-a-claim marker' and 'the money figure.' While the honest tone is appreciated, the wording should be tightened for a journal submission.
Circularity Check
No circularity: the probe and the downstream metric are distinct functions of each cell; shared RadBERT labels are a disclosed confound, not a definitional reduction.
full rationale
The paper's central claim is an empirical rank correlation between cheap disease-probe AUROC and expensive report-generation clinical micro-F1 over six matched cells. The two quantities are not defined in terms of one another: the probe is trained on frozen encoder tokens to classify RadBERT-extracted findings (Section 3), while clinical micro-F1 is produced by a full projector+LoRA fine-tune and scored by comparing RadBERT-extracted findings in generated versus reference reports (Section 3, Section 5). The shared RadBERT label pipeline is a real common-method confound that could inflate the correlation, but it is not a reduction by construction: the probe score and the F1 are different functions of the cell, and F1 additionally depends on the LLM's generation behavior. The paper explicitly discloses that the correlations were computed on an earlier label revision, not on the exact gated benchmark, and that the claim is ordinal and preliminary (Section 5, Section 6, and the 'From construction to validation' passage). That is a limitation on validity, not circularity. No load-bearing self-citation is used: the cited encoders and CT-RATE are external benchmarks. No uniqueness theorem, imported ansatz, or renamed known result appears. The regression/VQA pairing shares identical labels but is explicitly not reported, and a same-target cheap-versus-expensive transfer check would not be circular. The shuffled-label negative control and read-out robustness further support that the probe signal is not a pure artifact. Thus the derivation chain is self-contained as an empirical claim, and no circular step meets the evidentiary bar.
Axiom & Free-Parameter Ledger
free parameters (4)
- per-encoder ZCA whitening / LayerNorm fit =
fit once on training features per cell
- tertile cut-points for bucket labels =
fixed on train split; saved to tertile cutpoints.json
- uniform-pool token budget =
216 tokens
- Gate-2 well-posed separability line =
≈0.58 separability vs 0.333 chance
axioms (4)
- domain assumption RadBERT-extracted findings are valid, leakage-free labels for both probe and downstream scoring
- domain assumption Frozen encoder features are computed from identical volumes and preprocessing across cells (input validity)
- ad hoc to paper A modest read-out measures the LLM-usable information; an over-parameterized head measures the information ceiling
- ad hoc to paper The six matched cells and the earlier label revision are representative of the gated benchmark
read the original abstract
Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates. There are several encoders, several ways to compress their tokens, and several token budgets, and the combinations grow fast. Comparing them the usual way means fine-tuning a large language model (LLM) on each combination, and running the whole sweep this way needs far more compute than most groups can spend. We ask whether a cheap probe on the encoder's cached embeddings can stand in for that comparison. We build an image-grounded probing benchmark over (encoder $\times$ compression) cells, with clinical attribute families and two validation gates, scale-sanity and probe-separability, that keep each attribute well-scaled and decodable. These gates are the main methodological contribution. On this benchmark we compare a range of read-out heads, and in a preliminary study we pair each probe with its matched downstream task. The early signal is encouraging: the cheap probe orders the candidates in close agreement with expensive fine-tuning, at about $r\approx0.95$ on the cells measured so far. We read this as an ordinal claim, a ranking predictor rather than an exact estimate, and we are explicit about where it stays preliminary. If it holds up, encoder and compression choices can be screened in minutes with frozen-token probes, with full training spent only on the finalists.
Figures
Forward citations
Cited by 1 Pith paper
-
ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression
ORCA compresses 3D CT tokens into organ-guided connected regions with sinusoidal centroid encoding, outperforming grid average and other compressors at matched budgets.
Reference graph
Works this paper leans on
-
[3]
Ct-agent: A multimodal-llm agent for 3d ct (qa + report gen).arXiv preprint arXiv:2505.16229, 2025
Mao et al. Ct-agent: A multimodal-llm agent for 3d ct (qa + report gen).arXiv preprint arXiv:2505.16229, 2025. -; agentic planning + memory retrieval; QA+RG; region- LoRA(verify)
Pith/arXiv arXiv 2025
-
[4]
Better tokens for better 3d: Advancing vl modeling in 3d medical imaging
Hamamci et al. Better tokens for better 3d: Advancing vl modeling in 3d medical imaging. arXiv preprint arXiv:2510.20639, 2025. NeurIPS; learned freq-aware 3D tokenizer; +40% clinical F1; our base lineage. 7 Cheap Probes Predict Expensive Training A Preprint
arXiv 2025
-
[5]
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes. InICLR Workshop, 2017
2017
-
[6]
Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022
Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022
2022
-
[7]
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. Designing and interpreting probes with control tasks. In EMNLP, 2019
2019
-
[8]
Nguyen, Tal Hassner, Matthias Seeger, and Cedric Archambeau
Cuong V . Nguyen, Tal Hassner, Matthias Seeger, and Cedric Archambeau. LEEP: A new measure to evaluate transferability of learned representations. InInternational Conference on Machine Learning (ICML), 2020
2020
-
[9]
LogME: Practical assessment of pre-trained models for transfer learning
Kaichao You, Yong Liu, Jianmin Wang, and Mingsheng Long. LogME: Practical assessment of pre-trained models for transfer learning. InInternational Conference on Machine Learning (ICML), 2021
2021
-
[10]
An information-theoretic approach to transferability in task transfer learning
Yajie Bao, Yang Li, Shao-Lun Huang, Lin Zhang, Lizhong Zheng, Amir Zamir, and Leonidas Guibas. An information-theoretic approach to transferability in task transfer learning. InIEEE International Conference on Image Processing (ICIP), 2019
2019
-
[11]
Tran, Cuong V
Anh T. Tran, Cuong V . Nguyen, and Tal Hassner. Transferability and hardness of supervised classification tasks. InIEEE/CVF International Conference on Computer Vision (ICCV), 2019
2019
-
[12]
Anonymous. Feature quality and adaptability of medical foundation models: A comparative evaluation for radiographic classification and segmentation.arXiv preprint arXiv:2511.09742,
-
[13]
M3d: Advancing 3d medical image analysis with multi-modal llms.arXiv preprint arXiv:2404.00578, 2024
Bai et al. M3d: Advancing 3d medical image analysis with multi-modal llms.arXiv preprint arXiv:2404.00578, 2024. -; generalist 3D MLLM; spatial pooling perceiver; M3D- Data/Bench
Pith/arXiv arXiv 2024
-
[14]
Radgenome-chest ct: A grounded vision-language dataset.arXiv preprint arXiv:2404.16754, 2024
Zhang et al. Radgenome-chest ct: A grounded vision-language dataset.arXiv preprint arXiv:2404.16754, 2024. Sci Data; region-grounded reports + masks (665K) on CT-RATE
Pith/arXiv arXiv 2024
-
[15]
Medregion-ct: Region-focused multimodal llm for 3d ct rg.arXiv preprint arXiv:2506.23102, 2025
Kyung et al. Medregion-ct: Region-focused multimodal llm for 3d ct rg.arXiv preprint arXiv:2506.23102, 2025. MICCAI; R2 token pooling + mask-driven extractor + patient at- tributes
Pith/arXiv arXiv 2025
-
[16]
Comprehensive language-image pre-training for 3d medical image.arXiv preprint arXiv:2510.15042, 2025
Wald et al. Comprehensive language-image pre-training for 3d medical image.arXiv preprint arXiv:2510.15042, 2025. -; comprehensive LI pretraining; our encoder (3 variants)
arXiv 2025
-
[17]
Hamamci et al. Ct-clip: Contrastive language-image pretraining for 3d chest ct.arXiv preprint arXiv:2403.17834, 2024. -; contrastive 3D CT-text; released CT-RATE; our encoder
arXiv 2024
-
[18]
Cao et al. Boosting vision semantic density with anatomy normality modeling.arXiv preprint arXiv:2508.03742, 2025. -; anatomy normality modeling for RG
Pith/arXiv arXiv 2025
-
[19]
Shui et al. Large-scale and fine-grained vision-language pre-training for enhanced ct.arXiv preprint arXiv:2501.14548, 2025. -; anatomy-level alignment; 69086 patients; our encoder
Pith/arXiv arXiv 2025
-
[20]
RadBERT: Adapting transformer-based language models to radiology
An Yan, Julian McAuley, Xing Lu, Jiang Du, Eric Y Chang, Amilcare Gentili, and Chun-Nan Hsu. RadBERT: Adapting transformer-based language models to radiology. InRadiology: Artificial Intelligence, volume 4, 2022. 8 Cheap Probes Predict Expensive Training A Preprint
2022
-
[21]
TotalSegmentator: Robust segmentation of 104 anatomical structures in CT images.Radiology: Artificial Intelligence, 5 (5), 2023
Jakob Wasserthal, Hanns-Christian Breit, Manfred T Meyer, et al. TotalSegmentator: Robust segmentation of 104 anatomical structures in CT images.Radiology: Artificial Intelligence, 5 (5), 2023. A Benchmark construction and label validation (gallery) The image-grounded labels are validated at construction time, before any probe or question is gen- erated. ...
2023
-
[2025]
2D chest X-ray; linear probing of 8 foundation encoders
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.