Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

A lightweight probe ranks 3D-CT vision–language model candidates almost as well as full fine-tuning.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 06:11 UTC pith:L4GSSI3Y

load-bearing objection Honest pilot with a genuinely useful construction framework; the headline r=0.95 is probably inflated by shared RadBERT labels and should not be treated as established. the 3 major comments →

arxiv 2607.22771 v1 pith:L4GSSI3Y submitted 2026-07-24 cs.CV cs.AI

Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models

classification cs.CV cs.AI
keywords 3D CTvision-language modelsprobingfrozen embeddingsmodel selectionencoder selectiontoken compressionreport generation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that a cheap probe—a small read-out head trained on frozen image-token embeddings—can predict the downstream quality of a 3D CT vision–language model configuration before any expensive fine-tuning is done. The paper builds a probing benchmark over (encoder × compression) cells with clinical attributes, and finds that disease-probe AUROC correlates with report-generation clinical micro-F1 at r=0.95 / ρ=0.89 across six matched cells. The claim is explicitly ordinal: the probe predicts which configurations will rank best, not the exact performance scale. If correct, this lets practitioners screen dozens of encoder/compression choices in minutes and spend GPU-days of fine-tuning only on finalists. The paper is candid that the correlation is preliminary, computed on an earlier label revision and a small cell grid.

Core claim

The central claim is that frozen-token probing is a usable stand-in for 3D-CT encoder and compression selection. Over six (encoder × compression) cells, the AUROC of a disease-finding probe trained on cached embeddings orders the cells almost exactly as full LLM fine-tuning does, with Pearson r=0.95 and Spearman ρ=0.89 on the paired metric, report-generation clinical micro-F1. The paper reads this as an ordinal claim, not an exact estimator: the probe predicts the ranking, and within-encoder near-ties between compression variants are not reliably ordered. The relationship holds across nine of ten read-out heads; only one over-parameterized head collapses and inverts the correlation.

What carries the argument

The correlational identity carrying the argument is the cross-cell rank agreement between disease-probe AUROC and downstream clinical micro-F1, measured over six matched (encoder × compression) cells. The paper's methodological contribution is a pair of validation gates applied to every probe target: scale-sanity, which rejects clinical thresholds that produce near-degenerate class splits, and probe-separability, which drops or flags attributes a strong linear probe cannot decode. These gates ensure the probe labels are well-scaled and decodable, making the predictive-validity claim falsifiable. The probe itself is an attention-pooling read-out head trained on normalized, cached embeddings.

Load-bearing premise

The load-bearing premise is that the six-cell rank correlation—computed on an earlier label revision of the probing pipeline—transfers to the exact gated 15-attribute / 18-finding benchmark and to the broader candidate space; if error bars on those six points include a near-zero correlation, the central claim lacks current support.

What would settle it

Re-run the probe-to-downstream correlation on the exact gated benchmark described in the paper, across an expanded grid of at least 12 cells with multiple seeds per cell. If the 95% confidence interval for Spearman ρ includes values below 0.5, or if adding a new encoder/compression cell breaks the monotonic ordering, the ordinal claim would be effectively refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Encoder and compression choices can be screened in minutes with frozen-token probes, reserving full fine-tuning (about one GPU-day per configuration) for top-ranked finalists.
  • The probe ranking reliably identifies the best encoder, though it may not break statistically indistinguishable near-ties between an encoder's own compression variants.
  • The two validation gates provide a reusable protocol for constructing probing benchmarks whose labels are well-scaled, decodable, and leakage-free.
  • If the ordinal relationship holds beyond the six cells, the approach could cut the compute cost of model-selection sweeps for 3D-CT VLMs by orders of magnitude.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The six-cell correlation is a small sample; a re-run on the full gated benchmark with more cells and seed error bars could widen the credible interval to include a much lower rank correlation, so the headline number should be treated as provisional.
  • Because both probe and downstream labels derive from the same automated-label-extraction process, shared label noise may inflate the correlation; an independent label source would test this.
  • The same cost-geometry argument—cheap probing of frozen features to rank expensive generative fine-tuning—likely transfers to other medical imaging modalities or generation tasks, but external validity is untested.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a cheap-probing methodology for selecting a frozen 3D-CT image encoder and a token-compression scheme for a vision--language model. It constructs a benchmark of (encoder × compression) cells with image-grounded clinical attributes, two validation gates (scale-sanity and probe-separability), and an audit rubric. Ten read-out heads are compared on identical frozen tokens, and the central empirical claim is that disease-probe AUROC predicts report-generation clinical micro-F1 at r = 0.95 / ρ = 0.89 over six matched cells, robustly across read-out choice. The authors frame the claim as ordinal and preliminary, and disclose several limitations: the correlation was computed on an earlier label revision, not the gated benchmark; within-encoder compression near-ties are not reliably ordered; and re-running on the gated benchmark with more cells and error bars is ongoing.

Significance. If the central screening claim holds, the paper offers a practically valuable result: encoder/compression choices could be screened in minutes rather than GPU-days. The methodological contributions are real and should be credited: the two validation gates are a sensible anti-degeneracy mechanism, the probe-audit rubric is a useful protocol, the negative control with shuffled labels is appropriate, and the read-out robustness table directly addresses one natural objection. The paper is also unusually honest about the preliminary nature of the evidence. However, the current quantitative anchor is a six-point correlation computed on an earlier label revision and on labels that share a RadBERT extraction pipeline on both sides of the correlation. These issues are load-bearing for the screening claim and need to be resolved before the paper's central promise can be accepted.

major comments (3)
  1. [Section 5, Figure 3, Table 1] The central correlation r=0.95/ρ=0.89 is computed on an earlier label revision of the probing pipeline, not on the exact gated 15-attribute / 18-finding benchmark that the paper contributes. This is stated in Section 3 ('From construction to validation') and in Section 6. Consequently, the headline quantitative result is not attached to the benchmark that the paper's methodology section builds. A re-run on the gated benchmark is needed before the paper can claim that probe scores predict performance on that benchmark.
  2. [Section 3 and Section 5] Both sides of the correlation depend on RadBERT-extracted disease findings. Section 3 states that the 18 disease findings are 'extracted with RadBERT,' and Section 5 defines clinical micro-F1 as 'RadBERT-extracted findings, generated vs. reference.' If the same text model extracts the probe labels and the reference labels, common-mode label noise can inflate r=0.95/ρ=0.89 even if the probe has no true relationship to clinically meaningful report quality. The negative control and read-out robustness do not address this confound. The authors should report results with independently obtained labels (e.g., human-annotated findings or a different extraction method) or quantify the extraction error and show that it does not drive the correlation.
  3. [Section 5, 'Where the ranking mismatches'] The ordinal claim is not supported at the compression-selection level. The paper discloses that for CoLiPri, the probe ranks pool > adp while report-generation ranks adp > pool. Since compression choice is one of the two axes of the claimed screening space, this is not a minor caveat: it means the probe currently cannot be trusted to choose among an encoder's own compression variants. The six-cell correlation is dominated by cross-encoder differences, and the paper provides no error bars or confidence intervals. With n=6, the credible interval for ρ=0.89 is wide even before considering the disclosed near-tie. This should be stated more prominently, and the screening recommendation should be scoped to encoder selection, not compression selection, pending error-bar analysis.
minor comments (5)
  1. [Section 3] The text says every attribute is 'image-grounded' and 'never from a report,' then states that the 18 disease findings are 'extracted with RadBERT.' Since RadBERT is a radiology text model, this appears internally inconsistent. Please clarify whether the disease-finding labels are derived from CT images directly, from reports, or from a hybrid pipeline.
  2. [Section 5] The cell enumeration is inconsistent: the text says 'six matched cells (CoLiPri / CT-CLIP / BTB3D under adp, pool, or none),' which would be nine combinations, while Table 1 and Figure 2 refer to six cells and the reported results correspond to three encoders × two compressions. Please correct the enumeration.
  3. [References] Several reference entries appear incomplete or mismatched: [1] and [2] share the same arXiv identifier, and the titles are truncated. Please provide complete bibliographic information.
  4. [Section 4] The 'GLocal' head is mentioned without definition. Please specify its architecture or provide a citation so the read-out comparison is reproducible.
  5. [Throughout] The paper uses informal phrases such as 'stake-a-claim marker' and 'the money figure.' While the honest tone is appreciated, the wording should be tightened for a journal submission.

Circularity Check

0 steps flagged

No circularity: the probe and the downstream metric are distinct functions of each cell; shared RadBERT labels are a disclosed confound, not a definitional reduction.

full rationale

The paper's central claim is an empirical rank correlation between cheap disease-probe AUROC and expensive report-generation clinical micro-F1 over six matched cells. The two quantities are not defined in terms of one another: the probe is trained on frozen encoder tokens to classify RadBERT-extracted findings (Section 3), while clinical micro-F1 is produced by a full projector+LoRA fine-tune and scored by comparing RadBERT-extracted findings in generated versus reference reports (Section 3, Section 5). The shared RadBERT label pipeline is a real common-method confound that could inflate the correlation, but it is not a reduction by construction: the probe score and the F1 are different functions of the cell, and F1 additionally depends on the LLM's generation behavior. The paper explicitly discloses that the correlations were computed on an earlier label revision, not on the exact gated benchmark, and that the claim is ordinal and preliminary (Section 5, Section 6, and the 'From construction to validation' passage). That is a limitation on validity, not circularity. No load-bearing self-citation is used: the cited encoders and CT-RATE are external benchmarks. No uniqueness theorem, imported ansatz, or renamed known result appears. The regression/VQA pairing shares identical labels but is explicitly not reported, and a same-target cheap-versus-expensive transfer check would not be circular. The shuffled-label negative control and read-out robustness further support that the probe signal is not a pure artifact. Thus the derivation chain is self-contained as an empirical claim, and no circular step meets the evidentiary bar.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The central claim rests on pre-built encoders, datasets, and labelers (CT-RATE, TotalSegmentator, RadBERT) treated as inputs; the paper adds hand-chosen normalization, train-fitted tertile cuts, the pool token budget, the separability line, and the six-cell/earlier-revision transfer premise. No new entities (particles, forces, conserved quantities) are introduced. The counts here are the honest measure of what the paper contributes vs. what it inherits: the correlation is defined entirely on inherited label machinery plus four tunable choices.

free parameters (4)
  • per-encoder ZCA whitening / LayerNorm fit = fit once on training features per cell
    Input normalization is an explicit variable in §3; the paper reports CT-CLIP mean read-out moving 0.60→0.72 AUROC under ZCA, 'sharpening' the cross-encoder ranking. If the normalization partly induces the cross-encoder ordering, the probe ranking is partly a transform artifact.
  • tertile cut-points for bucket labels = fixed on train split; saved to tertile cutpoints.json
    The probe and VQA labels are defined by these train-fitted thresholds; different cuts would change probe AUROC values and hence the correlation.
  • uniform-pool token budget = 216 tokens
    Hand-chosen fixed budget so pool-column comparisons isolate the encoder; the budget determines which information survives compression on both the probe and the F1 side.
  • Gate-2 well-posed separability line = ≈0.58 separability vs 0.333 chance
    Hand-chosen acceptance line in Fig. 1 determining which attributes ship; two marginal location attributes (≈0.45) are kept as flagged canaries, so the line affects the benchmark contents.
axioms (4)
  • domain assumption RadBERT-extracted findings are valid, leakage-free labels for both probe and downstream scoring
    The 18 disease findings and clinical micro-F1 both come from RadBERT extraction (§3 and §5). If RadBERT label noise is cell-dependent, the probe→F1 correlation partly measures shared label noise.
  • domain assumption Frozen encoder features are computed from identical volumes and preprocessing across cells (input validity)
    Rubric dimension (2) in §3; if preprocessing differed across cells, probe AUROC differences could be artifacts rather than token-quality differences.
  • ad hoc to paper A modest read-out measures the LLM-usable information; an over-parameterized head measures the information ceiling
    §4: 'an over-parameterized read-out measures the representation's information ceiling rather than the usable information the LLM can exploit.' This interpretation of the TransMIL pathology is an assumption about LLM information use, not a proven equivalence.
  • ad hoc to paper The six matched cells and the earlier label revision are representative of the gated benchmark
    §5 'From construction to validation': the correlations were computed on an earlier label revision, not the exact gated benchmark. Transfer of the correlation to the contributed benchmark is assumed until the re-run ships.

pith-pipeline@v1.3.0-alltime-deepseek · 6911 in / 15753 out tokens · 165899 ms · 2026-08-01T06:11:20.007464+00:00 · methodology

0 comments
read the original abstract

Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates. There are several encoders, several ways to compress their tokens, and several token budgets, and the combinations grow fast. Comparing them the usual way means fine-tuning a large language model (LLM) on each combination, and running the whole sweep this way needs far more compute than most groups can spend. We ask whether a cheap probe on the encoder's cached embeddings can stand in for that comparison. We build an image-grounded probing benchmark over (encoder $\times$ compression) cells, with clinical attribute families and two validation gates, scale-sanity and probe-separability, that keep each attribute well-scaled and decodable. These gates are the main methodological contribution. On this benchmark we compare a range of read-out heads, and in a preliminary study we pair each probe with its matched downstream task. The early signal is encouraging: the cheap probe orders the candidates in close agreement with expensive fine-tuning, at about $r\approx0.95$ on the cells measured so far. We read this as an ordinal claim, a ranking predictor rather than an exact estimate, and we are explicit about where it stays preliminary. If it holds up, encoder and compression choices can be screened in minutes with frozen-token probes, with full training spent only on the finalists.

Figures

Figures reproduced from arXiv: 2607.22771 by Renjie Liang.

Figure 1
Figure 1. Figure 1: Separability gate. Tertile 3-class separability of a strong CoLiPri linear probe on each of the 15 regression attributes (avgpack tokens). Thirteen attributes clear the well-posed line; the two marginal location targets (orange) sit near chance and are kept, flagged, as spatial-position canaries. A bucketing a strong probe cannot separate is dropped or coarsened. attributes → VQA, where probe and VQA share… view at source ↗
Figure 2
Figure 2. Figure 2: Ten read-out heads on identical frozen tokens. Mean disease-probe AUROC per head, averaged over the six grid cells (btb3d/colipri/ct clip × adp/pool) for which every head is instan￾tiated. Heads form a broad strong plateau (0.72–0.75); only the over-parameterized TransMIL collapses toward chance and, as Section 5 shows, inverts the probe→downstream correlation. as evidence for the methodology, that cheap f… view at source ↗
Figure 3
Figure 3. Figure 3: The money figure. Cheap disease-probe AUROC (ABMIL, frozen tokens, minutes) vs. report-generation clinical micro-F1 after full LLM training (∼1 GPU-day), over the six matched cells. Pearson r = 0.95, Spearman ρ = 0.89. Marker color is the encoder; text labels give the compression [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: Quality-control for the vert median bone-density attribute: the target vertebrae (red) against neigh￾bours (yellow), with the per-case me￾dian HU, confirm the correct level is measured. 9 [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Scale-sanity gate (Gate 1): per-attribute measurement distributions with the guard cut [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Confound audit: each candidate HU target is regressed against nuisance factors so con [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression

    cs.CV 2026-07 conditional novelty 6.0

    ORCA compresses 3D CT tokens into organ-guided connected regions with sinusoidal centroid encoding, outperforming grid average and other compressors at matched budgets.

Reference graph

Works this paper leans on

20 extracted references · 6 linked inside Pith · cited by 1 Pith paper

  1. [3]

    Ct-agent: A multimodal-llm agent for 3d ct (qa + report gen).arXiv preprint arXiv:2505.16229, 2025

    Mao et al. Ct-agent: A multimodal-llm agent for 3d ct (qa + report gen).arXiv preprint arXiv:2505.16229, 2025. -; agentic planning + memory retrieval; QA+RG; region- LoRA(verify)

  2. [4]

    Better tokens for better 3d: Advancing vl modeling in 3d medical imaging

    Hamamci et al. Better tokens for better 3d: Advancing vl modeling in 3d medical imaging. arXiv preprint arXiv:2510.20639, 2025. NeurIPS; learned freq-aware 3D tokenizer; +40% clinical F1; our base lineage. 7 Cheap Probes Predict Expensive Training A Preprint

  3. [5]

    Understanding intermediate layers using linear classifier probes

    Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes. InICLR Workshop, 2017

  4. [6]

    Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022

    Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022

  5. [7]

    Designing and interpreting probes with control tasks

    John Hewitt and Percy Liang. Designing and interpreting probes with control tasks. In EMNLP, 2019

  6. [8]

    Nguyen, Tal Hassner, Matthias Seeger, and Cedric Archambeau

    Cuong V . Nguyen, Tal Hassner, Matthias Seeger, and Cedric Archambeau. LEEP: A new measure to evaluate transferability of learned representations. InInternational Conference on Machine Learning (ICML), 2020

  7. [9]

    LogME: Practical assessment of pre-trained models for transfer learning

    Kaichao You, Yong Liu, Jianmin Wang, and Mingsheng Long. LogME: Practical assessment of pre-trained models for transfer learning. InInternational Conference on Machine Learning (ICML), 2021

  8. [10]

    An information-theoretic approach to transferability in task transfer learning

    Yajie Bao, Yang Li, Shao-Lun Huang, Lin Zhang, Lizhong Zheng, Amir Zamir, and Leonidas Guibas. An information-theoretic approach to transferability in task transfer learning. InIEEE International Conference on Image Processing (ICIP), 2019

  9. [11]

    Tran, Cuong V

    Anh T. Tran, Cuong V . Nguyen, and Tal Hassner. Transferability and hardness of supervised classification tasks. InIEEE/CVF International Conference on Computer Vision (ICCV), 2019

  10. [12]

    Feature quality and adaptability of medical foundation models: A comparative evaluation for radiographic classification and segmentation.arXiv preprint arXiv:2511.09742,

    Anonymous. Feature quality and adaptability of medical foundation models: A comparative evaluation for radiographic classification and segmentation.arXiv preprint arXiv:2511.09742,

  11. [13]

    M3d: Advancing 3d medical image analysis with multi-modal llms.arXiv preprint arXiv:2404.00578, 2024

    Bai et al. M3d: Advancing 3d medical image analysis with multi-modal llms.arXiv preprint arXiv:2404.00578, 2024. -; generalist 3D MLLM; spatial pooling perceiver; M3D- Data/Bench

  12. [14]

    Radgenome-chest ct: A grounded vision-language dataset.arXiv preprint arXiv:2404.16754, 2024

    Zhang et al. Radgenome-chest ct: A grounded vision-language dataset.arXiv preprint arXiv:2404.16754, 2024. Sci Data; region-grounded reports + masks (665K) on CT-RATE

  13. [15]

    Medregion-ct: Region-focused multimodal llm for 3d ct rg.arXiv preprint arXiv:2506.23102, 2025

    Kyung et al. Medregion-ct: Region-focused multimodal llm for 3d ct rg.arXiv preprint arXiv:2506.23102, 2025. MICCAI; R2 token pooling + mask-driven extractor + patient at- tributes

  14. [16]

    Comprehensive language-image pre-training for 3d medical image.arXiv preprint arXiv:2510.15042, 2025

    Wald et al. Comprehensive language-image pre-training for 3d medical image.arXiv preprint arXiv:2510.15042, 2025. -; comprehensive LI pretraining; our encoder (3 variants)

  15. [17]

    Ct-clip: Contrastive language-image pretraining for 3d chest ct.arXiv preprint arXiv:2403.17834, 2024

    Hamamci et al. Ct-clip: Contrastive language-image pretraining for 3d chest ct.arXiv preprint arXiv:2403.17834, 2024. -; contrastive 3D CT-text; released CT-RATE; our encoder

  16. [18]

    Boosting vision semantic density with anatomy normality modeling.arXiv preprint arXiv:2508.03742, 2025

    Cao et al. Boosting vision semantic density with anatomy normality modeling.arXiv preprint arXiv:2508.03742, 2025. -; anatomy normality modeling for RG

  17. [19]

    Large-scale and fine-grained vision-language pre-training for enhanced ct.arXiv preprint arXiv:2501.14548, 2025

    Shui et al. Large-scale and fine-grained vision-language pre-training for enhanced ct.arXiv preprint arXiv:2501.14548, 2025. -; anatomy-level alignment; 69086 patients; our encoder

  18. [20]

    RadBERT: Adapting transformer-based language models to radiology

    An Yan, Julian McAuley, Xing Lu, Jiang Du, Eric Y Chang, Amilcare Gentili, and Chun-Nan Hsu. RadBERT: Adapting transformer-based language models to radiology. InRadiology: Artificial Intelligence, volume 4, 2022. 8 Cheap Probes Predict Expensive Training A Preprint

  19. [21]

    TotalSegmentator: Robust segmentation of 104 anatomical structures in CT images.Radiology: Artificial Intelligence, 5 (5), 2023

    Jakob Wasserthal, Hanns-Christian Breit, Manfred T Meyer, et al. TotalSegmentator: Robust segmentation of 104 anatomical structures in CT images.Radiology: Artificial Intelligence, 5 (5), 2023. A Benchmark construction and label validation (gallery) The image-grounded labels are validated at construction time, before any probe or question is gen- erated. ...

  20. [2025]

    2D chest X-ray; linear probing of 8 foundation encoders