Pith. sign in

Paper Citation Record · LEDGER

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.19217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19217 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:10:46.701581Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0cbf479e-e58c-4054-a24a-01bfb744648e · outbound

This paper cites Computed tomography and magnetic resonance imaging: past, present and future.European Respiratory Journal, 19(35 suppl):3s–12s, 2002.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Computed tomography and magnetic resonance imaging: past, present and future.European Respiratory Journal, 19(35 suppl):3s–12s, 2002

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.485949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:42.598520Z digest=sha256:18f294b3252e5dbb91d4abc4b3fe460b56e14c2dc6635b913022fe85d605ccd3

Observation 9e6857e0-9f0d-4a9e-b64b-987df1fcd7a5 · outbound

This paper cites Should we be concerned about the rapid increase in ct usage?Reviews on environmental health, 25(1):63–68, 2010.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Should we be concerned about the rapid increase in ct usage?Reviews on environmental health, 25(1):63–68, 2010

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.469459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:42.647398Z digest=sha256:7163eaa51ccc426fff261592713cbd5f26dad8044e16a253b3b1ffb2da9005b0

Observation 1786067b-301c-4b76-864b-451275c3d7a2 · outbound

This paper cites Cognitive and system factors contribut- ing to diagnostic errors in radiology.American Journal of Roentgenology, 201(3):611–617, 2013.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Cognitive and system factors contribut- ing to diagnostic errors in radiology.American Journal of Roentgenology, 201(3):611–617, 2013

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.454129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:42.739674Z digest=sha256:83cc7c02f60384d193b2cdc176c0c6cdbe4daba1d36df5f950f36a95304cc408

Observation 2a811af5-e706-4dc4-974a-dacc0cbcc068 · outbound

This paper cites Automated radiology report generation: A review of recent advances.IEEE Reviews in Biomedical Engineer- ing, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Automated radiology report generation: A review of recent advances.IEEE Reviews in Biomedical Engineer- ing, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.433752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:42.826120Z digest=sha256:fae2e9c8accaa2bfe647de4a169a6487c9ab99623012434639713c74a89ca474

Observation 6ef77ae9-f7cf-4339-bb01-206237bc6697 · outbound

This paper cites Comparing diagnostic accuracy of radiolo- gists versus gpt-4v and gemini pro vision using image inputs from diagnosis please cases.Radiology, 312(1), July 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Comparing diagnostic accuracy of radiolo- gists versus gpt-4v and gemini pro vision using image inputs from diagnosis please cases.Radiology, 312(1), July 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.417716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:42.915903Z digest=sha256:b244bc98c2677eecb350777083117eb08264666cf173eed0da0afd9edb7570bb

Observation 9ce878c1-d2f8-4c0f-a405-be2a867a5b76 · outbound

This paper cites Evaluating large language models on medical evidence summarization.NPJ digital medicine, 6(1):158, 2023.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Evaluating large language models on medical evidence summarization.NPJ digital medicine, 6(1):158, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.368002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:43.011449Z digest=sha256:47da467ff4d31f825bb2b7db94e02beaa6aea5dcf5236ea2c7744ba24283a2b1

Observation e1bfba31-07e9-48d4-aa5e-dd1b23790d72 · outbound

This paper cites Embracing large language models for medical applications: opportuni- ties and challenges.Cureus, 15(5), 2023.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Embracing large language models for medical applications: opportuni- ties and challenges.Cureus, 15(5), 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.234294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:43.097684Z digest=sha256:d487b29ef7edeb9562cfe0e853f49489311a2f6f91430283cc19f88cd8691f87

Observation a1cfe936-bb90-4e1f-9b69-6e983c5c4bdd · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.108888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:43.193737Z digest=sha256:e721d52df3c50db7808fc258045a0746f742a9f6e822252e044ddbe4d3f94f18

Observation edb9e078-c00c-4196-aa6c-2063a1b601e6 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.293038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.293038Z digest=sha256:ad0b336d5bd6f861c036cc4e3a8714609fd78d6099c2a707dcd70378ab296e2a

Observation fb21cd16-96b6-45dd-99f6-433cc62987a8 · outbound

This paper cites Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:50.786538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:43.381150Z digest=sha256:775a1e2ba4fdffee58fa7c0dbd6cd226d2643a3652a84fc78e20a8fad56829a6

Observation 428a250c-52c1-46a2-b798-e1b4bd39fbcc · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.479305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.479305Z digest=sha256:9b5686ee7964b17b64a02d6b4a2f709e17877c8883a33071f33cdece4acdfd87

Observation 115ddd63-1894-4bef-9a33-8707ebcfab85 · outbound

This paper cites Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:50.479972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:43.555156Z digest=sha256:e490debe24601e1c002db25048f506bf347553fc18944e58e5a4642e0fc9fca3

Observation 88b74f26-1912-41a0-ab9b-ccc68fbf59d6 · outbound

This paper cites Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37:94327–94427, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37:94327–94427, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:50.210620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:43.645745Z digest=sha256:b8f7548fb3703d6ed92e712066525f68ef202e3f93c9cffa6ee1cb326a8aff9f

Observation 05f1c7c1-6992-4164-925c-2b7b3d5f5976 · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.721022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.721022Z digest=sha256:96c69fffd3e50079472dbd925521884011b5562d66c4a29c410720b67676ce76

Observation a8d1fb6e-ba17-4062-b407-3ce167e99ed7 · outbound

This paper cites Mme-survey: A comprehensive survey on evaluation of multimodal llms,.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Mme-survey: A comprehensive survey on evaluation of multimodal llms,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:49.811271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:43.818300Z digest=sha256:31e3bbff4875f11cadc5fe29f48875d578a4bc7a0009bfc61564a793420d00ed

Observation 69e2268f-8bd6-4168-b4ab-fcec64620eab · outbound

This paper cites Kim and Liem T.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Kim and Liem T

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:49.430981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:43.903843Z digest=sha256:b3f935bc44f5ffd88b5b98a7465af24bd58cb19d69b315a499677f5a0f177860

Observation c25235c2-3a65-424e-9262-fa9a93711142 · outbound

This paper cites Recovery at the edge of error: debunking the myth of the infallible expert.Journal of biomedical in- formatics, 44(3):413–424, 2011.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Recovery at the edge of error: debunking the myth of the infallible expert.Journal of biomedical in- formatics, 44(3):413–424, 2011

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:49.078302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:43.983906Z digest=sha256:9d4881db898570bda065b099a9f766d364cf80031aa555114afed49220e68cf2

Observation e6c3db7a-8e7f-4548-898c-e984a7b7deae · outbound

This paper cites Overview of the mediqa-corr 2024 shared task on medical error detection and correction.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Overview of the mediqa-corr 2024 shared task on medical error detection and correction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.799582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:44.067488Z digest=sha256:11f5f0343689ec1dab6aab9a55722ae3c366cc680031d3d7f91b23e91ae0df6e

Observation 06a40d3c-bd45-40a2-a5e0-0983b2bbd85b · outbound

This paper cites Potential of gpt-4 for detecting errors in radiology reports: implications for report- ing accuracy.Radiology, 311(1):e232714, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Potential of gpt-4 for detecting errors in radiology reports: implications for report- ing accuracy.Radiology, 311(1):e232714, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.626354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:44.141513Z digest=sha256:13e79e57afe7c01c467c528bbafbb967773f4a3a71ce74fb6c1d17bb510789a4

Observation 18e7c4a7-cc8a-4549-9bf3-db27cfb420f7 · outbound

This paper cites Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.228630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.228630Z digest=sha256:db8a093291e3d592d163ad7c109c1fa9b0c14da42fd9895a9948b8faa7f7ec31

Observation 83010895-4fda-422d-8ff1-67939e9412c1 · outbound

This paper cites 3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports 3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.308483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.308483Z digest=sha256:e27d3ac1aea656edd656144f715be1a8a4ebfd2f0a17b6a4ccc830bd37508012

Observation 5353d755-11e3-4d34-a00d-9f1d3d9cb799 · outbound

This paper cites Med3dvlm: An efficient vision-language model for 3d med- ical image analysis, 2025.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Med3dvlm: An efficient vision-language model for 3d med- ical image analysis, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.447620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:44.712221Z digest=sha256:85817f7091f36fd0b90738c89c6fd23ef3a0aca9ee53bc1874036485fc7abbd1

Observation 0382bbba-26ec-4f69-8ff6-65e5914ada57 · outbound

This paper cites Medm-vl: What makes a good medical lvlm?,.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Medm-vl: What makes a good medical lvlm?,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.276586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:44.801787Z digest=sha256:ce57095d52f3959201bf27a72a55e0971b1673069ae8c26e763a74abdd25dddb

Observation 17dd16c3-2a9e-4bc6-b893-7ac6135e1c63 · outbound

This paper cites MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.901214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.901214Z digest=sha256:339d1b2d4f83b2605d2082da60ac8dca8d2aa81f33b849769f10b42e244fcbf3

Observation 33204f4a-b552-4226-9903-d3757868ced9 · outbound

This paper cites ReXErr: Synthesizing Clinically Meaningful Errors in Diagnostic Radiology Reports.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports ReXErr: Synthesizing Clinically Meaningful Errors in Diagnostic Radiology Reports

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.037970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.037970Z digest=sha256:d03265fa641ead4dbea800befcbe446a4003a307b582263fa689391a6acf010a

Observation 44d7a051-144e-4d14-950f-827898976d04 · outbound

This paper cites Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.149757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:45.146921Z digest=sha256:6ff37a4f1820c1183ca0d5974057316fde741d56ccad3c83a3cac9340ff72145

Observation 63cfcc53-e338-4855-aab8-75092ce45f50 · outbound

This paper cites MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.285853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.285853Z digest=sha256:b1675079c91165b7e96220bc12ab9d07ffbb00df181fb398b02b09a559b2e446

Observation bdb16b3b-681a-4534-a31f-e49f2746f9b5 · outbound

This paper cites A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero- shot detection of abnormalities.CoRR, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero- shot detection of abnormalities.CoRR, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.034543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:45.400446Z digest=sha256:38beb589a9bb85c5c73d56dfdcdc43c8a8e8673314a8e19255c016181e0ffc2a

Observation 4f65962b-c11c-4f1c-b76e-b002564a7ca5 · outbound

This paper cites RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.482676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.482676Z digest=sha256:a65cb8e2af769c02fcbb5f096b33b5dc5746a7578b4740181f8d3f0686a3e7a7

Observation 6f7c1a52-eae0-465f-8324-1dfff9514139 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.904119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:45.564153Z digest=sha256:57f80d6a0e33d7f713ef96c771307adf36f87158a380cf8e847a4e6bb6ccfe14

Observation 99978041-19d1-4172-b7a5-b4633f686d55 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.655074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.655074Z digest=sha256:d7e85895298b90a75538ec1f2e3ddbee3c939c793ef6c158c2dadcf9c5ec0256

Observation d873694a-61ff-4e2e-8518-a284e4869ca4 · outbound

This paper cites The Llama 3 Herd of Models.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports The Llama 3 Herd of Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.736107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.736107Z digest=sha256:e7a86ffeb3936b4f1152539bf468e43f8ce2648774e4873b429df07644505108

Observation 21572ae9-4807-4a72-b881-70ca51cc44b3 · outbound

This paper cites Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.868971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.868971Z digest=sha256:62ed9803083395868b54068ff57ed211e4f9f128414664a77fbc95c6eb32c74f

Observation 470dc869-5bc2-4946-ac5e-dde649d80ed8 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.988146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.988146Z digest=sha256:e8deb3cdb46a3050b51391302bb6ab34d8602da90546e80282acce9aa1aa547e

Observation 6d3e0d87-338f-4a8a-abce-dad19db84edf · outbound

This paper cites De- veloping Generalist Foundation Models from a Multi- modal Dataset for 3D Computed Tomography, April 2025.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports De- veloping Generalist Foundation Models from a Multi- modal Dataset for 3D Computed Tomography, April 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:46.089541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:46.089541Z digest=sha256:b1cd9ca89db7891ac8d2b72096abbd932ccc85c64127ec15a493186c601c6c89

Observation 89e9e2f8-049d-44d9-a7e0-021817e7a208 · outbound

This paper cites ROUGE: A Package for Automatic Evalu- ation of Summaries.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports ROUGE: A Package for Automatic Evalu- ation of Summaries

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.764041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:46.201273Z digest=sha256:4a897ac9ff2c38489a2f76a53d8c89931a190b173d2cb2e570b75302dfa3334d

Observation 49635171-5368-4672-9ab1-54707ea48709 · outbound

This paper cites Bleu: a Method for Automatic Evaluation of Ma- chine Translation.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Bleu: a Method for Automatic Evaluation of Ma- chine Translation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.559786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:46.283549Z digest=sha256:32233ce2a874786f20861bed10a74874b6759ab2570e53baa6e0e1e41a1516ba

Observation cbd352a5-33a0-412f-8bd0-aeffdc2de515 · outbound

This paper cites METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.393684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:10:46.443896Z digest=sha256:b50331901ae0c31f46c9ac007a0cb5713fb4629637ab65f388e628d7b670b057

Observation 5fff4125-b51c-4d5c-ba59-8a7cedb4ce4e · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports BERTScore: Evaluating Text Generation with BERT

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:46.566102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:46.566102Z digest=sha256:f4e54a35a5f94e6bf3123365870ad89041758b49041169e47825469bfacfc55f

Observation 6e5a2bf0-69a9-4a05-8e98-d63b4cb7fd0a · outbound

This paper cites GREEN: Generative Radiology Report Evaluation and Error Notation.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports GREEN: Generative Radiology Report Evaluation and Error Notation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:46.701581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:46.701581Z digest=sha256:085070f185aaa6a837a7960dcd3959326301ca3af9d313d58ca07e37790be3bf

Pith citing papers

No inbound Pith citation observations are available.