Pith. sign in

REVIEW 6 major objections 5 minor 64 references

MvKeTR: Chest CT Report Generation with Multi-View Perception and Knowledge Enhancement

T0 review · 6 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read MvKeTR reads chest CTs from three anatomical planes, retrieves similar reports, and outperforms prior generators on nearly every metric.

desk verdict A sensible new architecture for 3D CT report generation with internally consistent ablations, but the SOTA claim rests on a comparison table that mixes cited and re-run baselines and contains a likely erroneous METEOR value. read the letter →

arxiv 2411.18309 v3 pith:2RWF5TUW submitted 2024-11-27 cs.CV cs.AI

classification cs.CVcs.AI
keywords chestCTreportgenerationmulti-viewperceptionknowledgeenhancementKolmogorov-Arnoldnetworksview-awareattentionCT-CLIPretrievalCT-ViTradiology
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to show that automatic chest CT report generation improves when a model reads a volume the way a radiologist does: from the axial, coronal, and sagittal planes at once, and with reference to similar prior cases. The proposed architecture, MvKeTR, combines a Multi-View Perception Aggregator that fuses the three views through view-aware attention with a Cross-Modal Knowledge Enhancer that retrieves the most similar reports via a CT-CLIP model and feeds them in through cross-attention. It replaces MLP layers with Kolmogorov-Arnold Networks to capture fine high-frequency image details with fewer parameters. On the public CTRG-Chest-548K dataset the paper reports the best scores in nearly every automatic metric, including BLEU-1 58.36, BLEU-4 37.86, and ROUGE-L 54.25. If the comparison holds, the result matters because CT reports are time-consuming to write and error-prone, and the design offers a workflow-aligned template for automating them.

What carries the argument

The load-bearing machinery is the pair of perception and knowledge branches on top of a shared sequence-to-sequence backbone. The Multi-View Perception Aggregator treats each anatomical plane as a separate token sequence and modifies ordinary attention by adding a learnable view embedding $E_v$ to the query–key logits, so the model can weight the axial, coronal, and sagittal evidence differently; the three streams are then concatenated and fused by a KAN layer. The Cross-Modal Knowledge Enhancer retrieves the $k=16$ most similar reports with a frozen CT-CLIP model, feeds their embeddings as keys and values in cross-attention with axial CT tokens, and preserves a residual connection so retrieved knowledge cannot override direct visual evidence. Kolmogorov-Arnold Networks, whose activations are learnable spline functions, replace MLPs throughout both modules; the paper argues from prior theoretical results that KANs scale linearly rather than quadratically in parameters and have less spectral bias, helping the model capture high-frequency lesion features.

What would settle it

Re-run all Table II baselines under the exact 80/20 split, 224-cubed volumes, CT-ViT backbone, and beam size used for MvKeTR, and check whether the MvKeTR margins over SL-DG, Dia-LLaMA, and Reg2RG survive; if the cited rows shift when run in this setting, the state-of-the-art claim collapses. In the same experiment, choose top-k on a validation split rather than the test set to see whether the peak at k=16 is real.

Watch

Extended reading notes

Core claim

The paper's central claim is that a chest CT report generator should mirror a clinician's diagnostic routine, and that doing so yields state-of-the-art results. The model predicts a report from a 224-cubed CT volume by running three independent CT-ViT extractors on the axial, coronal, and sagittal reorderings of the volume; a view-aware attention mechanism adds a learnable view embedding to the attention logits, and the three normalized outputs are concatenated and fused through a KAN layer. A second branch retrieves the top-16 most similar reports from the CT-RATE bank using a frozen CT-CLIP encoder, then attends to those report embeddings with the axial visual features, again with residual connections and a KAN layer. The two feature sets are concatenated and fed to a three-layer cross-modal-memory transformer decoder to generate the findings and impression. On the CTRG-Chest-548K test set the paper reports that this design outperforms previous CT report generators on almost every metric, and the ablations attribute the gains to all three components, with multi-view perception contributing the largest single improvement.

Load-bearing premise

The whole comparison rests on the assumption that the numbers in Table II are mutually comparable, yet the rows marked with a double dagger are copied from their original papers while the rest were re-run, and the paper gives no evidence that every row used the same split, preprocessing, and visual extractor.

Editorial extensions

If this is right

  • If the reported numbers are taken at face value, multi-view perception is the largest single contributor: BASE+MVPA improves the average over all six metrics by 17.2%, versus 14.5% for knowledge enhancement alone.
  • Replacing MLP layers with KAN layers accounts for a 9.5% gap in average improvement between the full model and the MLP variant, suggesting the activation choice matters beyond simply adding parameters.
  • The knowledge branch's residual design means retrieved reports act as a prior that can be overridden; the paper argues this preserves the model's ability to report novel findings not present in retrieved cases.
  • View-aware attention degrades gracefully under rotation and artifact perturbations in the paper's controlled experiments, with the largest rotation-induced drop being 3.80% and some artifact conditions slightly improving scores.
  • The modular design is compatible with alternative 3D extractors; Table V shows CT-ViT works best but 3D ViT, CT-Net, and U-Net can be swapped in.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not test retrieval banks drawn from CTRG-Chest-548K itself; a natural follow-up is to isolate whether the gains come from cross-dataset knowledge transfer or from any similar-case prior.
  • The top-k=16 choice peaks on the test set (Fig. 10); an honest hyperparameter check would select k on a validation split before reporting test numbers.
  • If the multi-view advantage transfers, the same view-token and view-aware attention design could be applied to other volumetric modalities such as MRI, where coronal and sagittal planes also carry complementary diagnostic information.
  • Because METEOR is the one metric where MvKeTR trails (Reg2RG scores 49.71), the 'almost all metrics' claim depends on metric weighting; evaluating with disease-level recall would test whether the qualitative nodule-detection gains generalize.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

6 major / 5 minor

Summary. The paper proposes MvKeTR, a Transformer-based model for chest CT report generation on the CTRG-Chest-548K dataset. The model has three main components: a Multi-View Perception Aggregator (MVPA) that fuses axial, sagittal, and coronal CT features extracted by three CT-ViT encoders using view-aware attention; a Cross-Modal Knowledge Enhancer (CMKE) that retrieves the top-k most similar reports from a CT-CLIP-encoded report bank and integrates them via cross-attention; and a report generator based on R2GenCMN with KAN layers replacing MLPs. The authors report state-of-the-art results across most automatic metrics (BLEU, METEOR, ROUGE-L), provide ablations showing the contribution of each component, include a human evaluation by two radiologists, and make the code publicly available.

Significance. If the empirical claims hold, MvKeTR is a plausible step forward for 3D CT report generation: the multi-view aggregation is well motivated by radiological practice, the retrieval-based knowledge enhancer is a clean way to inject domain knowledge, and the use of KANs is a timely architectural choice. The paper includes useful ablations, a public dataset, a public code link, and a human evaluation, which are all strengths. However, the central claim of surpassing prior state-of-the-art rests on a comparison table whose baselines have mixed provenance and at least one implausible entry, and the reported gains are not accompanied by variance or significance information. The architecture is coherent, but the empirical evidence needs substantial verification before the SOTA claim can be accepted.

major comments (6)
  1. [Section IV-A, IV-C, Table II] The SOTA claim is not established because Table II mixes directly cited results (SL-DG, Dia-LLaMA, Reg2RG, marked with double-dagger) with results the authors re-ran using public code (marked with asterisk). Section IV-A states only that 80% of the data is randomly allocated to training and 20% to testing, and Section IV-C says CT-ViT is used as the visual extractor for both compared methods and MvKeTR. For the cited rows, however, the test split, preprocessing, and evaluation script are those of the original papers, and no evidence is given that they coincide with the authors' split. If the cited baselines were evaluated on a different test set, the claimed improvements could be an artifact of split difficulty. The authors should either re-run all baselines under the same protocol or clearly report the provenance and split for each row and restrict the SOTA claim to comparable settings.
  2. [Table II, Reg2RG row] The METEOR value of 49.71 for Reg2RG is far outside the range of all other methods (19.52-28.36) and is inconsistent with Reg2RG's own BLEU and ROUGE scores in the same row (BLEU-1 49.63, BLEU-4 32.04, ROUGE-L 47.76). This is likely a data-entry error or a metric-implementation mismatch. Because the abstract and Section IV-D claim that MvKeTR surpasses SOTA across 'almost all metrics,' and the only exception is this METEOR value, the correctness of this entry is load-bearing. The authors must verify the value against the original Reg2RG paper or re-run Reg2RG on their split, and correct the table and the surrounding text if needed.
  3. [Section IV-H2, Fig. 10] The hyperparameter top-k is set to 16 because it peaks on the test set, as shown in Fig. 10, which plots BLEU-1 on CTRG-Chest-548K. Selecting a hyperparameter on the test set inflates the reported scores and makes the comparison with baselines unfair if those baselines did not receive the same test-set tuning. The authors should select top-k on a validation split and report the corresponding test performance, or at minimum provide a sensitivity analysis that distinguishes validation-based selection from test-set selection.
  4. [Section IV-B, IV-D, Tables II and III] All quantitative results are reported as single numbers without error bars, confidence intervals, or significance tests. Given that the dataset contains only 1,804 image-report pairs and the split is random, the differences between MvKeTR and strong baselines such as CAMANet could be within run-to-run variance. The authors should report results over multiple seeds with mean and standard deviation, and perform a significance test (e.g., paired bootstrap) for the main comparisons in Table II and the ablation comparisons in Table III.
  5. [Table III, Ours-MLP row] The 'AVG. Δ' values in Table III are not reproducible from the listed metric values. For Ours-MLP, the arithmetic mean of the six per-metric relative improvements over BASE is approximately 16.7%, not the reported 14.1%; for Ours, the arithmetic mean is approximately 22.5%, not 23.6%. The base row also gives inconsistent numbers for R2GenCMN (48.29/28.42 in Table III versus 28.42 in the text). The authors should state the averaging formula and correct the table, because the claimed 9.5% gap between Ours-MLP and Ours is used to support the KAN contribution.
  6. [Section IV-F, Table IV] The human evaluation is based on only 10 randomly selected test cases and two radiologists, with no inter-rater agreement measure and no statistical test for the reported 'significant margins' and '32% improvement.' For a clinical-facing claim, this sample is too small and the scoring procedure is underspecified. The authors should either report per-case scores, inter-rater reliability (e.g., Cohen's kappa), and a paired significance test, or temper the clinical-utility claim accordingly.
minor comments (5)
  1. [Section IV-G] The sentence 'The 14.5% BLEU-4 degradation without CMKE confirms that retrieved knowledge primarily enhances model generalization' appears to misstate the ablation: BASE+CMKE improves BLEU-4 by about 21% over BASE (34.43 vs. 28.42), and 14.5% is the average improvement over all metrics. This should be reworded to avoid confusion.
  2. [Section III-F, Eq. (22)] Equation (22) is called a 'Bayesian posterior estimation,' but the displayed factorization p(Y|X,K) ∝ p(X|Y)p(K|Y) is not a valid Bayesian posterior as written (the right-hand side conditions on Y and omits the prior and evidence terms). This is a narrative interpretation of feature concatenation rather than a derivation, and the framing should be removed or substantially revised.
  3. [Abstract and Title] The abstract contains the typo 'TansfoRmer,' and the paper refers to 'Vit-Transformer' in Section IV-D; these should be corrected for consistency.
  4. [Fig. 8 and Fig. 10] Figures 8 and 10 contain font-encoding artifacts (unprintable glyph sequences) in the submitted PDF, making the axis labels and captions unreadable in places; the figures should be regenerated with proper fonts.
  5. [Notation in Section III-B] The notation for CT patch extraction is inconsistent: the text says 'extracting (28)×(28)×(28) non-overlapping patches' but earlier defines patches Z_v as 8×8×8, and the relationship between the temporal/spatial patch sizes Pt, Ph, Pw and these numbers is not clarified. Please align the notation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity found: MvKeTR's reported results are obtained from an independently trained empirical architecture, not from a derivation that reduces to its inputs.

full rationale

MvKeTR is an empirical deep learning system; none of its reported predictions is derived from its inputs by construction. The training objective (Eq. 2) is standard conditional log-likelihood, and the components (MVPA Eqs. 12-14, CMKE Eqs. 15-21, generator Eqs. 23-27) contain learnable parameters trained on the 80% split and evaluated on the 20% test split. Multi-view aggregation concatenates view-specific attention outputs and fuses them with a KAN layer, and knowledge enhancement performs CT-CLIP retrieval followed by cross-attention; neither operation returns a quantity that was used to define its input. The KAN motivation cites external theorems (Refs. [15], [50], [51]) with no author overlap, so no load-bearing self-citation is present. Equation 22 labels feature concatenation as "Bayesian posterior estimation"; this is a narrative gloss rather than a derivation, and no downstream equation uses the Bayesian interpretation to compute the reported metrics. The Limitations section explicitly acknowledges reliance on CT-RATE quality; this is a stated dependency, not a circular step. The remaining concerns—mixed provenance of cited baselines in Table II, the implausible Reg2RG METEOR value, and top-k=16 selected on the test curve in Fig. 10—are validity or reporting risks, not circular reasoning. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported, and no known result is repackaged as novel. Therefore no circular step is identified.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The contribution is a new architecture, not a new physical entity or conservation law. The main free choices are hyperparameters, most notably top-k tuned on the test set. The assumptions are domain-level claims about retrieval quality, multi-view sufficiency, KAN advantages, and metric validity, none of which are independently verified in the paper.

free parameters (3)
  • top-k = 16
    Number of retrieved reports in CMKE; selected by observing peak BLEU-1 on the test set in Fig. 10 rather than on a validation set.
  • learning rate = 5e-5 for CT-ViTs, 1e-4 for other params, decay 0.8 per epoch
    Hand-chosen optimizer settings; they affect optimization but are not central to the scientific claim.
  • beam size = 3
    Beam search width at decoding, chosen by hand; standard hyperparameter.
assumptions (4)
  • domain assumption KAN theory (parameter efficiency, low spectral bias) from Wang et al. [50] holds and translates into better CT report generation in practice.
    Invoked in Section III-C to motivate replacing MLPs with KANs; the empirical gain is only shown on one small dataset without significance tests.
  • domain assumption CT-CLIP, pretrained on CT-RATE, provides reliable volume-to-report retrieval so the top-k retrieved reports carry clinically relevant priors.
    Section III-E uses CT-CLIP without fine-tuning; if retrieval is poor, CMKE could add noise rather than knowledge.
  • domain assumption Axial, coronal, and sagittal views of a resized 224x224x224 volume contain complementary diagnostic information that can be fused via the described attention.
    Section III-B and III-D assume three orthogonal views are sufficient and that resampling to 224 cubed does not destroy clinically significant findings.
  • domain assumption BLEU, METEOR, and ROUGE-L are valid proxies for clinical report quality.
    Used throughout Tables II and III; these metrics are known to correlate only weakly with clinical correctness.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MvKeTR: Chest CT Report Generation with Multi-View Perception and Knowledge Enhancement." pith.science (2026). https://pith.science/paper/2RWF5TUW

@misc{pith2026241118309,
  author       = {Pith},
  title        = {Pith review of: MvKeTR: Chest CT Report Generation with Multi-View Perception and Knowledge Enhancement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2RWF5TUW}},
  note         = {Machine review of arXiv:2411.18309}
}
read the original abstract

CT report generation (CTRG) aims to automatically generate diagnostic reports for 3D volumes, relieving clinicians' workload and improving patient care. Despite clinical value, existing works fail to effectively incorporate diagnostic information from multiple anatomical views and lack related clinical expertise essential for accurate and reliable diagnosis. To resolve these limitations, we propose a novel Multi-view perception Knowledge-enhanced TansfoRmer (MvKeTR) to mimic the diagnostic workflow of clinicians. Just as radiologists first examine CT scans from multiple planes, a Multi-View Perception Aggregator (MVPA) with view-aware attention is proposed to synthesize diagnostic information from multiple anatomical views effectively. Then, inspired by how radiologists further refer to relevant clinical records to guide diagnostic decision-making, a Cross-Modal Knowledge Enhancer (CMKE) is devised to retrieve the most similar reports based on the query volume to incorporate domain knowledge into the diagnosis procedure. Furthermore, instead of traditional MLPs, we employ Kolmogorov-Arnold Networks (KANs) as the fundamental building blocks of both modules, which exhibit superior parameter efficiency and reduced spectral bias to better capture high-frequency components critical for CT interpretation while mitigating overfitting. Extensive experiments on the public CTRG-Chest-548 K dataset demonstrate that our method outpaces prior state-of-the-art (SOTA) models across almost all metrics. The code is available at https://github.com/xiweideng/MvKeTR.

Figures

Figures reproduced from arXiv: 2411.18309 by the authors.

Figure 1
Figure 1. An example of chest CT volume and its corresponding report, including findings and impression. writing, and subsequent decision-making. After acquiring a patient’s radiology images, physicians examine all anatomical structures and regions of concern, then employ related expert knowledge to write a hand-crafted, clinically coherent report that documents the observations [1]. As evidenced in [PITH_FULL_IMAGE:figures/… view at source ↗
Figure 2
Figure 2. The overview of our proposed MvKeTR, which can be partitioned into four parts: 3D visual extractor, multi-view perception aggregator, cross-modal knowledge enhancer, and report generator. The 3D visual extractor extracts CT patches from the axial, coronal, and sagittal views of the input 3D CT volume. The multi-view perception aggregator aggregates diagnostic information from multiple anatomical views effectively. T… view at source ↗
Figure 3
Figure 3. Illustration of vision feature extraction pipeline through CT-ViT. A. Overview of the Proposed Approach The generation of radiology reports can be considered an image-to-text problem, for which we follow a sequence￾to-sequence paradigm. In doing so, unlike previous ap￾proaches [8]–[10] that typically process a single-view im￾age, our network handles the input 3D CT volume X ∈ R (224)×224×224 by transforming it into … view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: The architecture of Kolmogorov–Arnold Networks (KANs) with a series of KAN layers. Softmax Scale Z × × V K Q H Softmax Scale Z × × V K Q H (a) Attention Softmax Scale Z × × V K Q H View Embedding × + Softmax Scale Z × × V K Q H View Embedding × + (b) View-aware Attenti…
Figure 5
Figure 5. Figure 5: The schematic diagram of Attention and View-aware Attention. “⊗” denotes matrix multiplication. “⊕” represents element-wise addition. 1) View-aware Attention: The vanilla attention(see Fig. 5a) adopted in Transformer [52] is computed through the correla￾tion between th…
Figure 7
Figure 7. Figure 7: Illustration of encoder-decoder architecture in report generator. large-scale, meticulously curated dataset minimizes the risk of retrieving erroneous or biased reports, as the retrieval process inherently prioritizes clinically validated cases through CT￾CLIP’s semant…
Figure 8
Figure 8. Figure 8: The loss curve during training on the CTRG-Chest-548K. Y = {y1, y2, y3, . . . , yT }, and a memory matrix M = {m1, m2, m3, . . . , ml}, the memory responses of FS and Y can be obtained by: rFS = FS · mT √ l d · ml (23) ryt = YT · mT √ l d · ml (24) where l denotes the …
Figure 9
Figure 9. Figure 9: The examples of report from ground-truth, R2GenCMN, and our method. Green/red highlights denotes correct/incorrect content respectively. rating medical expertise from similar cases. The 14.5% BLEU-4 degradation without CMKE confirms that retrieved knowledge primarily e…
Figure 10
Figure 10. Figure 10: Effect of varying top-k on CTRG-Chest-548K. Ours reveals a performance gap of 9.5% (23.6% vs 14.1% in average improvement). This observation owes to the fact that KAN layers are more effective in modeling complicated diagnostic relationships compared to conventional M…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 55 canonical work pages

  1. [1]

    Evidence-based guideline for the written radiology report: Methods, recommendations and implementation challenges,

    S. K. Goergen, F. J. Pool, T. J. Turner, J. E. Grimm, M. N. Appleyard, C. Crock, M. C. Fahey, M. F. Fay, N. J. Ferris, S. M. Liew et al. , “Evidence-based guideline for the written radiology report: Methods, recommendations and implementation challenges,” J. Med. Imaging Radiat. Oncol., vol. 57, no. 1, pp. 1–7, 2013. DENG et al.: MVKETR: CHEST CT REPORT G...

  2. [2]

    The us radiologist workforce: an analysis of temporal and geographic variation by using large national datasets,

    A. B. Rosenkrantz, D. R. Hughes, and R. Duszak Jr, “The us radiologist workforce: an analysis of temporal and geographic variation by using large national datasets,” Radiology, vol. 279, no. 1, pp. 175–184, 2016

  3. [3]

    Radiologist shortage leaves patient care at risk, warns royal college,

    A. Rimmer, “Radiologist shortage leaves patient care at risk, warns royal college,” BMJ-BRIT. MED. J., vol. 359, 2017

  4. [4]

    Understanding and confronting our mistakes: the epidemiology of error in radiology and strategies for error reduction,

    M. A. Bruno, E. A. Walker, and H. H. Abujudeh, “Understanding and confronting our mistakes: the epidemiology of error in radiology and strategies for error reduction,” Radiographics, vol. 35, no. 6, pp. 1668– 1676, 2015

  5. [5]

    Cross-modal memory networks for radiology report generation,

    Z. Chen, Y . Shen, Y . Song, and X. Wan, “Cross-modal memory networks for radiology report generation,” in Proc. 59th Ann. Meeting Assoc. Comput. Linguistics, 11th Int. Joint Conf. Natural Lang. Process , 2021, pp. 5904–5914

  6. [6]

    Improving radiology report generation with d 2-net: When diffusion meets dis- criminator,

    Y . Jin, W. Chen, Y . Tian, Y . Song, C. Yan, and Z. Mao, “Improving radiology report generation with d 2-net: When diffusion meets dis- criminator,” in Proc. IEEE Int. Conf. Acoust. Speech Signal Process. , 2024, pp. 2215–2219

  7. [7]

    Computed tomography and magnetic resonance imaging: past, present and future,

    N. M ¨uller, “Computed tomography and magnetic resonance imaging: past, present and future,” Eur. Respir. J., vol. 19, no. 35 suppl, pp. 3s– 12s, 2002

  8. [8]

    Work like a doctor: Unifying scan localizer and dynamic generator for automated computed tomography report generation,

    Y . Tang, H. Yang, L. Zhang, and Y . Yuan, “Work like a doctor: Unifying scan localizer and dynamic generator for automated computed tomography report generation,” Expert Syst. Appl. , vol. 237, p. 121442, 2024

Show all 64 references
  1. [9]

    Ct2rep: Automated radiology report generation for 3d medical imaging,

    I. E. Hamamci, S. Er, and B. Menze, “Ct2rep: Automated radiology report generation for 3d medical imaging,” in Proc. Int. Conf. Med. Image Comput. Comput.-Assisted Intervention , 2024, pp. 476–486

  2. [10]

    Dia-llama: Towards large language model-driven ct report generation,

    Z. Chen, L. Luo, Y . Bie, and H. Chen, “Dia-llama: Towards large language model-driven ct report generation,” arXiv:2403.16386, 2024

  3. [11]

    Pulmonary nodule detection in ct images: false positive reduction using multi-view convolutional networks,

    A. A. A. Setio, F. Ciompi, G. Litjens, P. Gerke, C. Jacobs, S. J. Van Riel, M. M. W. Wille, M. Naqibullah, C. I. S ´anchez, and B. Van Ginneken, “Pulmonary nodule detection in ct images: false positive reduction using multi-view convolutional networks,”IEEE Trans. Med. Imaging...

  4. [12]

    Recommendations for measuring pulmonary nodules at ct: a statement from the fleischner society,

    A. A. Bankier, H. MacMahon, J. M. Goo, G. D. Rubin, C. M. Schaefer- Prokop, and D. P. Naidich, “Recommendations for measuring pulmonary nodules at ct: a statement from the fleischner society,” Radiology, vol. 285, no. 2, pp. 584–600, 2017

  5. [13]

    Guidelines for management of incidental pulmonary nodules detected on ct images: from the fleischner society 2017,

    H. MacMahon, D. P. Naidich, J. M. Goo, K. S. Lee, A. N. Leung, J. R. Mayo, A. C. Mehta, Y . Ohno, C. A. Powell, M. Prokop et al. , “Guidelines for management of incidental pulmonary nodules detected on ct images: from the fleischner society 2017,” Radiology, vol. 284, no. 1, p...

  6. [14]

    Generating radiology re- ports via memory-driven transformer,

    Z. Chen, Y . Song, T.-H. Chang, and X. Wan, “Generating radiology re- ports via memory-driven transformer,” inProc. Conf. Empirical Methods Natural Lang. Process., Nov. 2020, pp. 1439–1449

  7. [15]

    KAN: Kolmogorov–arnold networks,

    Z. Liu, Y . Wang, S. Vaidya, F. Ruehle, J. Halverson, M. Soljacic, T. Y . Hou, and M. Tegmark, “KAN: Kolmogorov–arnold networks,” in Proc. 13th Int. Conf. Learn. Represent. , 2025

  8. [16]

    A survey on automatic image caption generation,

    S. Bai and S. An, “A survey on automatic image caption generation,” Neurocomputing, vol. 311, pp. 291–304, 2018

  9. [17]

    An overview of image caption generation methods,

    H. Wang, Y . Zhang, and X. Yu, “An overview of image caption generation methods,” Comput. Intell. Neurosci. , vol. 2020, no. 1, p. 3062706, 2020

  10. [18]

    I2t: Image parsing to text description,

    B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu, “I2t: Image parsing to text description,” Proc. IEEE, vol. 98, no. 8, pp. 1485–1508, 2010

  11. [19]

    Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora,

    R. Socher and L. Fei-Fei, “Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2010, pp. 966–973

  12. [20]

    Show and tell: A neural image caption generator,

    O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2015, pp. 3156–3164

  13. [21]

    Self- critical sequence training for image captioning,

    S. J. Rennie, E. Marcheret, Y . Mroueh, J. Ross, and V . Goel, “Self- critical sequence training for image captioning,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2017, pp. 7008–7024

  14. [22]

    Show, attend and tell: Neural image caption generation with visual attention,

    K. Xu, “Show, attend and tell: Neural image caption generation with visual attention,” arXiv:1502.03044, 2015

  15. [23]

    Image caption with global-local attention,

    L. Li, S. Tang, L. Deng, Y . Zhang, and Q. Tian, “Image caption with global-local attention,” in Proc. AAAI Conf. Artif. Intell. , vol. 31, no. 1, 2017

  16. [24]

    Knowing when to look: Adaptive attention via a visual sentinel for image captioning,

    J. Lu, C. Xiong, D. Parikh, and R. Socher, “Knowing when to look: Adaptive attention via a visual sentinel for image captioning,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2017, pp. 375–383

  17. [25]

    Meshed-memory transformer for image captioning,

    M. Cornia, M. Stefanini, L. Baraldi, and R. Cucchiara, “Meshed-memory transformer for image captioning,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 10 578–10 587

  18. [26]

    Cptr: Full transformer network for image captioning,

    W. Liu, S. Chen, L. Guo, X. Zhu, and J. Liu, “Cptr: Full transformer network for image captioning,” arXiv:2101.10804, 2021

  19. [27]

    On the automatic generation of medical imaging reports,

    B. Jing, P. Xie, and E. Xing, “On the automatic generation of medical imaging reports,” arXiv:1711.08195, 2017

  20. [28]

    Multimodal recurrent model with attention for automated radiology report generation,

    Y . Xue, T. Xu, L. Rodney Long, Z. Xue, S. Antani, G. R. Thoma, and X. Huang, “Multimodal recurrent model with attention for automated radiology report generation,” in Proc. 21st Int. Conf. Med. Image Comput. Comput.-Assisted Intervention , 2018, pp. 457–466

  21. [29]

    Natural language generation model for mammography reports simulation,

    A. Hoogi, A. Mishra, F. Gimenez, J. Dong, and D. Rubin, “Natural language generation model for mammography reports simulation,” IEEE J. Biomed. Health Inform. , vol. 24, no. 9, pp. 2711–2717, 2020

  22. [30]

    A self-boosting framework for automated radiographic report generation,

    Z. Wang, L. Zhou, L. Wang, and X. Li, “A self-boosting framework for automated radiographic report generation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2021, pp. 2433–2442

  23. [31]

    Progressive transformer-based generation of radiol- ogy reports,

    F. Nooralahzadeh, N. Perez Gonzalez, T. Frauenfelder, K. Fujimoto, and M. Krauthammer, “Progressive transformer-based generation of radiol- ogy reports,” in Proc. Findings Assoc. Comput. Linguistics: EMNLP 2021, 2021, pp. 2824–2832

  24. [32]

    Tsget: Two-stage global enhanced transformer for automatic radiology report generation,

    X. Yi, Y . Fu, R. Liu, H. Zhang, and R. Hua, “Tsget: Two-stage global enhanced transformer for automatic radiology report generation,” IEEE J. Biomed. Health Inform. , vol. 28, no. 4, pp. 2152–2162, 2024

  25. [33]

    Unsupervised disease tags for automatic radiology report generation,

    X. Yi, Y . Fu, R. Hua, R. Liu, and H. Zhang, “Unsupervised disease tags for automatic radiology report generation,” Biomed Signal Process Control, vol. 89, p. 105742, 2024

  26. [34]

    Lhr-rfl: Linear hybrid-reward based reinforced focal learning for automatic radiology report generation,

    X. Yi, Y . Fu, J. Yu, R. Liu, H. Zhang, and R. Hua, “Lhr-rfl: Linear hybrid-reward based reinforced focal learning for automatic radiology report generation,” IEEE Trans. Med. Imaging , 2024

  27. [35]

    Camanet: Class activation map guided attention network for radiology report generation,

    J. Wang, A. Bhalerao, T. Yin, S. See, and Y . He, “Camanet: Class activation map guided attention network for radiology report generation,” IEEE J. Biomed. Health Inform. , vol. 28, no. 4, pp. 2199–2210, 2024

  28. [36]

    Weakly guided attention model with hierarchical interaction for brain ct report generation,

    X. Zhang, S. Yang, Y . Shi, J. Ji, Y . Liu, Z. Wang, and H. Xu, “Weakly guided attention model with hierarchical interaction for brain ct report generation,” Comput. Biol. Med. , vol. 167, p. 107650, 2023

  29. [37]

    Co-occurrence relation- ship driven hierarchical attention network for brain ct report generation,

    X. Zhang, S. Dou, J. Ji, Y . Liu, and Z. Wang, “Co-occurrence relation- ship driven hierarchical attention network for brain ct report generation,” IEEE Trans. Emerging Top. Comput. Intell. , 2024

  30. [38]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv:2307.09288, 2023

  31. [39]

    Large language model with region-guided referring and grounding for ct report generation,

    Z. Chen, Y . Bie, H. Jin, and H. Chen, “Large language model with region-guided referring and grounding for ct report generation,” IEEE Trans. Med. Imaging, 2025

  32. [40]

    Automatic liver segmentation from ct volumes based on multi-view information fusion and condition random fields,

    Z. Xia, M. Liao, S. Di, Y . Zhao, W. Liang, and N. N. Xiong, “Automatic liver segmentation from ct volumes based on multi-view information fusion and condition random fields,” Optics & Laser Technology , vol. 179, p. 111298, 2024

  33. [41]

    Multi-scale and multi-view network for lung tumor segmentation,

    C. Liu, H. Liu, X. Zhang, J. Guo, and P. Lv, “Multi-scale and multi-view network for lung tumor segmentation,” Comput. Biol. Med., vol. 172, p. 108250, 2024

  34. [42]

    Automatic radiology report generation based on multi-view image fusion and medical concept enrichment,

    J. Yuan, H. Liao, R. Luo, and J. Luo, “Automatic radiology report generation based on multi-view image fusion and medical concept enrichment,” in Proc. 22nd Int. Conf. Med. Image Comput. Comput.- Assisted Intervention, 2019, pp. 721–729

  35. [43]

    Automatic medical image report generation with multi-view and multi-modal attention mechanism,

    S. Yang, J. Niu, J. Wu, and X. Liu, “Automatic medical image report generation with multi-view and multi-modal attention mechanism,” in International Conference on Algorithms and Architectures for Parallel Processing. Springer, 2020, pp. 687–699

  36. [44]

    Radiology report generation with a learned knowledge base and multi-modal alignment,

    S. Yang, X. Wu, S. Ge, Z. Zheng, S. K. Zhou, and L. Xiao, “Radiology report generation with a learned knowledge base and multi-modal alignment,” Med. Image Anal. , vol. 86, p. 102798, 2023

  37. [45]

    “knowledge is power

    K. Kale, P. Bhattacharyya, A. Shetty, M. Gune, K. Shrivastava, R. Lawyer, and S. Biswas, ““knowledge is power”: Constructing knowledge graph of abdominal organs and using them for automatic radiology report generation,” in Proc. 61st Ann. Meeting Assoc. Comput. Linguistics (Vo...

  38. [46]

    Mkcl: medical knowledge with contrastive learning model for radiology report gener- ation,

    X. Hou, Z. Liu, X. Li, X. Li, S. Sang, and Y . Zhang, “Mkcl: medical knowledge with contrastive learning model for radiology report gener- ation,” J. Biomed. Inf. , vol. 146, p. 104496, 2023

  39. [47]

    Kiut: Knowledge-injected u- transformer for radiology report generation,

    Z. Huang, X. Zhang, and S. Zhang, “Kiut: Knowledge-injected u- transformer for radiology report generation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2023, pp. 19 809–19 818

  40. [48]

    Learning repre- sentations by back-propagating errors,

    D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning repre- sentations by back-propagating errors,” nature, vol. 323, no. 6088, pp. 533–536, 1986

  41. [49]

    Multilayer feedforward networks are universal approximators,

    K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural networks, vol. 2, no. 5, pp. 359–366, 1989. 14 IEEE TRANSACTIONS AND JOURNALS TEMPLATE

  42. [50]

    On the expressiveness and spectral bias of KANs,

    Y . Wang, J. W. Siegel, Z. Liu, and T. Y . Hou, “On the expressiveness and spectral bias of KANs,” in Proc. 13th Int. Conf. Learn. Represent. , 2025

  43. [51]

    On the spectral bias of neural networks,

    N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. Hamprecht, Y . Bengio, and A. Courville, “On the spectral bias of neural networks,” in Proc. Int. Conf. Mach. Learn. PMLR, 2019, pp. 5301–5310

  44. [52]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. 31st Int. Conf. Neural Inf. Process. Syst. , 2017, p. 6000–6010

  45. [53]

    A foundation model utilizing chest ct volumes and radiology reports for supervised- level zero-shot detection of abnormalities,

    I. E. Hamamci, S. Er, F. Almas, A. G. Simsek, S. N. Esirgun, I. Dogan, M. F. Dasdelen, B. Wittmann, E. Simsar, M. Simsar et al., “A foundation model utilizing chest ct volumes and radiology reports for supervised- level zero-shot detection of abnormalities,” arXiv:2403.17834, 2024

  46. [54]

    Radbert: adapting transformer-based language models to radiology,

    A. Yan, J. McAuley, X. Lu, J. Du, E. Y . Chang, A. Gentili, and C.-N. Hsu, “Radbert: adapting transformer-based language models to radiology,” Radiol. Artif. Intell. , 2022

  47. [55]

    Bleu: a method for automatic evaluation of machine translation,

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proc. 40th Annu. Meeting Assoc. Comput. Linguistics , Jul. 2002, pp. 311–318

  48. [56]

    Meteor 1.3: Automatic metric for reliable optimization and evaluation of machine translation systems,

    M. Denkowski and A. Lavie, “Meteor 1.3: Automatic metric for reliable optimization and evaluation of machine translation systems,” inProc. 6th Workshop Stat. Mach. Transl., 2011, pp. 85–91

  49. [57]

    ROUGE: A package for automatic evaluation of summaries,

    C.-Y . Lin, “ROUGE: A package for automatic evaluation of summaries,” in Proc. Text Summarization Branches Out , Jul. 2004, pp. 74–81

  50. [58]

    Generatect: Text-conditional generation of 3d chest ct volumes,

    I. E. Hamamci, S. Er, A. Sekuboyina, E. Simsar, A. Tezcan, A. G. Simsek, S. N. Esirgun, F. Almas, I. Do ˘gan, M. F. Dasdelen et al. , “Generatect: Text-conditional generation of 3d chest ct volumes,” in Proc. Eur. Conf. Comput. Vis. Springer, 2024, pp. 126–143

  51. [59]

    Adam: A method for stochastic optimization,

    D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv:1412.6980, 2014

  52. [60]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. Int. Conf. Learn. Represent., 2021

  53. [61]

    Machine-learning-based multiple abnormality prediction with large-scale chest computed tomography volumes,

    R. L. Draelos, D. Dov, M. A. Mazurowski, J. Y . Lo, R. Henao, G. D. Rubin, and L. Carin, “Machine-learning-based multiple abnormality prediction with large-scale chest computed tomography volumes,” Med. Image Anal., vol. 67, p. 101857, 2021

  54. [62]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Proc. Int. Conf. Med. Image Comput. Comput.-Assisted Intervention , 2015, pp. 234–241

  55. [63]

    Computed tomography,

    T. M. Buzug, “Computed tomography,” in Springer handbook of medical technology. Springer, 2011, pp. 311–342

  56. [64]

    Multiplanar and three-dimensional reconstruction techniques in ct: impact on chest diseases,

    J. Remy, M. Remy-Jardin, D. Artaud, and M. Fribourg, “Multiplanar and three-dimensional reconstruction techniques in ct: impact on chest diseases,” Eur. Radiol., pp. 335–351, 1998

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.