Pith. sign in

REVIEW 2 cited by

Activating Associative Disease-Aware Vision Token Memory for LLM-Based X-ray Report Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.03458 v1 pith:A4PZJXAB submitted 2025-01-07 eess.IV cs.AIcs.CV

classification eess.IVcs.AIcs.CV
keywords reportx-raygenerationimageinformationmodelvisualmedical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

X-ray image based medical report generation achieves significant progress in recent years with the help of the large language model, however, these models have not fully exploited the effective information in visual image regions, resulting in reports that are linguistically sound but insufficient in describing key diseases. In this paper, we propose a novel associative memory-enhanced X-ray report generation model that effectively mimics the process of professional doctors writing medical reports. It considers both the mining of global and local visual information and associates historical report information to better complete the writing of the current report. Specifically, given an X-ray image, we first utilize a classification model along with its activation maps to accomplish the mining of visual regions highly associated with diseases and the learning of disease query tokens. Then, we employ a visual Hopfield network to establish memory associations for disease-related tokens, and a report Hopfield network to retrieve report memory information. This process facilitates the generation of high-quality reports based on a large language model and achieves state-of-the-art performance on multiple benchmark datasets, including the IU X-ray, MIMIC-CXR, and Chexpert Plus. The source code of this work is released on \url{https://github.com/Event-AHU/Medical_Image_Analysis}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation

    cs.CV 2025-08 reject novelty 5.0 of 10

    R2GenKG generates X-ray reports with an LLM conditioned on a GPT-4o-built multi-modal knowledge graph, reporting small metric gains on IU-Xray and CheXpert Plus.

  2. XiHeFusion: Harnessing Large Language Models for Science Communication in Nuclear Fusion

    cs.CV 2025-02 reject novelty 4.0 of 10

    XiHeFusion is a Qwen2.5-14B model fine-tuned on 1.2 million fusion knowledge pairs to answer nuclear fusion questions for science communication.

Pith tools