Pith. sign in

Paper Citation Record · LEDGER

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.19217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19217 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:10:46.701581Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0cbf479e-e58c-4054-a24a-01bfb744648e · outbound

This paper cites Computed tomography and magnetic resonance imaging: past, present and future.European Respiratory Journal, 19(35 suppl):3s–12s, 2002.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Computed tomography and magnetic resonance imaging: past, present and future.European Respiratory Journal, 19(35 suppl):3s–12s, 2002

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.485949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:42.598520Z digest=sha256:f57f3c1096a07dd865df74f9aa978bc1e0e55c4a70dc4fddbe05c99d666ab8d8

Observation 9e6857e0-9f0d-4a9e-b64b-987df1fcd7a5 · outbound

This paper cites Should we be concerned about the rapid increase in ct usage?Reviews on environmental health, 25(1):63–68, 2010.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Should we be concerned about the rapid increase in ct usage?Reviews on environmental health, 25(1):63–68, 2010

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.469459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:42.647398Z digest=sha256:1d8d343416f465a1de704ccf19c37a8008817d4285422e621411f9a82e467604

Observation 1786067b-301c-4b76-864b-451275c3d7a2 · outbound

This paper cites Cognitive and system factors contribut- ing to diagnostic errors in radiology.American Journal of Roentgenology, 201(3):611–617, 2013.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Cognitive and system factors contribut- ing to diagnostic errors in radiology.American Journal of Roentgenology, 201(3):611–617, 2013

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.454129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:42.739674Z digest=sha256:e3469171db9c9916a6a250dbea50590bb6ea991edd71aec2497b67d0ff85c53d

Observation 2a811af5-e706-4dc4-974a-dacc0cbcc068 · outbound

This paper cites Automated radiology report generation: A review of recent advances.IEEE Reviews in Biomedical Engineer- ing, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Automated radiology report generation: A review of recent advances.IEEE Reviews in Biomedical Engineer- ing, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.433752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:42.826120Z digest=sha256:21badf5ceae820413d9fbf1d0160eb75e8e6d54208c3665fd4fece4bc9f5a175

Observation 6ef77ae9-f7cf-4339-bb01-206237bc6697 · outbound

This paper cites Comparing diagnostic accuracy of radiolo- gists versus gpt-4v and gemini pro vision using image inputs from diagnosis please cases.Radiology, 312(1), July 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Comparing diagnostic accuracy of radiolo- gists versus gpt-4v and gemini pro vision using image inputs from diagnosis please cases.Radiology, 312(1), July 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.417716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:42.915903Z digest=sha256:7062af1d2ec9efef47029a36016d7ef001325368bcc618c169930d8ac32a84b1

Observation 9ce878c1-d2f8-4c0f-a405-be2a867a5b76 · outbound

This paper cites Evaluating large language models on medical evidence summarization.NPJ digital medicine, 6(1):158, 2023.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Evaluating large language models on medical evidence summarization.NPJ digital medicine, 6(1):158, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.368002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:43.011449Z digest=sha256:f455af8f65e7e94c0d4a7e1ec0e2290d327ed8f718c2235902d0ef4e5a873624

Observation e1bfba31-07e9-48d4-aa5e-dd1b23790d72 · outbound

This paper cites Embracing large language models for medical applications: opportuni- ties and challenges.Cureus, 15(5), 2023.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Embracing large language models for medical applications: opportuni- ties and challenges.Cureus, 15(5), 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.234294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:43.097684Z digest=sha256:2175d590714680e396d908cc44d93fcb1fcb9d03ff5b173a0b27e8ced10e0e42

Observation a1cfe936-bb90-4e1f-9b69-6e983c5c4bdd · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.108888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:43.193737Z digest=sha256:11a810a6e653c1db350d6badb76197685ff2db8a5fc00dd8b1fbe7977a4d255d

Observation edb9e078-c00c-4196-aa6c-2063a1b601e6 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.293038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.293038Z digest=sha256:ad0b336d5bd6f861c036cc4e3a8714609fd78d6099c2a707dcd70378ab296e2a

Observation fb21cd16-96b6-45dd-99f6-433cc62987a8 · outbound

This paper cites Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:50.786538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:43.381150Z digest=sha256:d2aec1219c46e8d6b8f4588b6f8c8a6b26e60727e37890db0f2e131cc76c5672

Observation 428a250c-52c1-46a2-b798-e1b4bd39fbcc · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.479305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.479305Z digest=sha256:9b5686ee7964b17b64a02d6b4a2f709e17877c8883a33071f33cdece4acdfd87

Observation 115ddd63-1894-4bef-9a33-8707ebcfab85 · outbound

This paper cites Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:50.479972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:43.555156Z digest=sha256:471c365fc27c1ddd02154c13bde3366a13a8f3617985f9672e0109f4e8b4e2c8

Observation 88b74f26-1912-41a0-ab9b-ccc68fbf59d6 · outbound

This paper cites Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37:94327–94427, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37:94327–94427, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:50.210620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:43.645745Z digest=sha256:0c716939404f52f0fe672d4f6e51d01b5dfae266025131700c98897c8d170466

Observation 05f1c7c1-6992-4164-925c-2b7b3d5f5976 · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.721022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.721022Z digest=sha256:96c69fffd3e50079472dbd925521884011b5562d66c4a29c410720b67676ce76

Observation a8d1fb6e-ba17-4062-b407-3ce167e99ed7 · outbound

This paper cites Mme-survey: A comprehensive survey on evaluation of multimodal llms,.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Mme-survey: A comprehensive survey on evaluation of multimodal llms,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:49.811271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:43.818300Z digest=sha256:86ce719fd5dee82550ddac9db5e936bf7879f25274c631e5c387c0ba77ebfe71

Observation 69e2268f-8bd6-4168-b4ab-fcec64620eab · outbound

This paper cites Kim and Liem T.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Kim and Liem T

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:49.430981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:43.903843Z digest=sha256:45714b2bec4705d119a98ca63218f9c4d63884278ea9da54688d6096cd57c44c

Observation c25235c2-3a65-424e-9262-fa9a93711142 · outbound

This paper cites Recovery at the edge of error: debunking the myth of the infallible expert.Journal of biomedical in- formatics, 44(3):413–424, 2011.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Recovery at the edge of error: debunking the myth of the infallible expert.Journal of biomedical in- formatics, 44(3):413–424, 2011

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:49.078302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:43.983906Z digest=sha256:f9bb5a31de244507e99bf7c9dac42fe737cd0a45064e1d9eb3ce9cb5b706074b

Observation e6c3db7a-8e7f-4548-898c-e984a7b7deae · outbound

This paper cites Overview of the mediqa-corr 2024 shared task on medical error detection and correction.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Overview of the mediqa-corr 2024 shared task on medical error detection and correction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.799582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:44.067488Z digest=sha256:7cea634146d8d370507853665682472ab997b850bc0a11f45e5a6de1545086d7

Observation 06a40d3c-bd45-40a2-a5e0-0983b2bbd85b · outbound

This paper cites Potential of gpt-4 for detecting errors in radiology reports: implications for report- ing accuracy.Radiology, 311(1):e232714, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Potential of gpt-4 for detecting errors in radiology reports: implications for report- ing accuracy.Radiology, 311(1):e232714, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.626354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:44.141513Z digest=sha256:037a5b291a6fb0e1cb78cba3619ed25ac0d760348f4c4b0c9429b8d7f3a941ad

Observation 18e7c4a7-cc8a-4549-9bf3-db27cfb420f7 · outbound

This paper cites Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.228630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.228630Z digest=sha256:db8a093291e3d592d163ad7c109c1fa9b0c14da42fd9895a9948b8faa7f7ec31

Observation 83010895-4fda-422d-8ff1-67939e9412c1 · outbound

This paper cites 3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports 3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.308483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.308483Z digest=sha256:e27d3ac1aea656edd656144f715be1a8a4ebfd2f0a17b6a4ccc830bd37508012

Observation 5353d755-11e3-4d34-a00d-9f1d3d9cb799 · outbound

This paper cites Med3dvlm: An efficient vision-language model for 3d med- ical image analysis, 2025.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Med3dvlm: An efficient vision-language model for 3d med- ical image analysis, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.447620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:44.712221Z digest=sha256:ff64f3a441c2b8c149560e0d18675e18a582fa7894d8e9668a2159848ebdd092

Observation 0382bbba-26ec-4f69-8ff6-65e5914ada57 · outbound

This paper cites Medm-vl: What makes a good medical lvlm?,.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Medm-vl: What makes a good medical lvlm?,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.276586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:44.801787Z digest=sha256:1cbf3dd16dfbb92bdd4edd8da7e78230f1517334f5c05390b6c500aea20591a5

Observation 17dd16c3-2a9e-4bc6-b893-7ac6135e1c63 · outbound

This paper cites MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.901214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.901214Z digest=sha256:339d1b2d4f83b2605d2082da60ac8dca8d2aa81f33b849769f10b42e244fcbf3

Observation 33204f4a-b552-4226-9903-d3757868ced9 · outbound

This paper cites ReXErr: Synthesizing Clinically Meaningful Errors in Diagnostic Radiology Reports.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports ReXErr: Synthesizing Clinically Meaningful Errors in Diagnostic Radiology Reports

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.037970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.037970Z digest=sha256:d03265fa641ead4dbea800befcbe446a4003a307b582263fa689391a6acf010a

Observation 44d7a051-144e-4d14-950f-827898976d04 · outbound

This paper cites Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.149757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:45.146921Z digest=sha256:43e7bc101bf81e99d6226acb2be81a929cf1c46d510c061c082bb1a8af3119e1

Observation 63cfcc53-e338-4855-aab8-75092ce45f50 · outbound

This paper cites MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.285853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.285853Z digest=sha256:b1675079c91165b7e96220bc12ab9d07ffbb00df181fb398b02b09a559b2e446

Observation bdb16b3b-681a-4534-a31f-e49f2746f9b5 · outbound

This paper cites A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero- shot detection of abnormalities.CoRR, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero- shot detection of abnormalities.CoRR, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.034543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:45.400446Z digest=sha256:567585826dcdbcad511218c7358bc4f5c3b6c5e478a5dbbd63b4e6119a5715d8

Observation 4f65962b-c11c-4f1c-b76e-b002564a7ca5 · outbound

This paper cites RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.482676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.482676Z digest=sha256:a65cb8e2af769c02fcbb5f096b33b5dc5746a7578b4740181f8d3f0686a3e7a7

Observation 6f7c1a52-eae0-465f-8324-1dfff9514139 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.904119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:45.564153Z digest=sha256:5d0c0995802aff799d0d2c6d1fc538c41d71d372b89bce9eefd0f80fd9170181

Observation 99978041-19d1-4172-b7a5-b4633f686d55 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.655074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.655074Z digest=sha256:d7e85895298b90a75538ec1f2e3ddbee3c939c793ef6c158c2dadcf9c5ec0256

Observation d873694a-61ff-4e2e-8518-a284e4869ca4 · outbound

This paper cites The Llama 3 Herd of Models.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports The Llama 3 Herd of Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.736107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.736107Z digest=sha256:e7a86ffeb3936b4f1152539bf468e43f8ce2648774e4873b429df07644505108

Observation 21572ae9-4807-4a72-b881-70ca51cc44b3 · outbound

This paper cites Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.868971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.868971Z digest=sha256:62ed9803083395868b54068ff57ed211e4f9f128414664a77fbc95c6eb32c74f

Observation 470dc869-5bc2-4946-ac5e-dde649d80ed8 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.988146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.988146Z digest=sha256:e8deb3cdb46a3050b51391302bb6ab34d8602da90546e80282acce9aa1aa547e

Observation 6d3e0d87-338f-4a8a-abce-dad19db84edf · outbound

This paper cites De- veloping Generalist Foundation Models from a Multi- modal Dataset for 3D Computed Tomography, April 2025.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports De- veloping Generalist Foundation Models from a Multi- modal Dataset for 3D Computed Tomography, April 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:46.089541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:46.089541Z digest=sha256:b1cd9ca89db7891ac8d2b72096abbd932ccc85c64127ec15a493186c601c6c89

Observation 89e9e2f8-049d-44d9-a7e0-021817e7a208 · outbound

This paper cites ROUGE: A Package for Automatic Evalu- ation of Summaries.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports ROUGE: A Package for Automatic Evalu- ation of Summaries

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.764041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:46.201273Z digest=sha256:fc737b67740a7780eecf757932414f936b4825c22610fdeb91935a4af056766b

Observation 49635171-5368-4672-9ab1-54707ea48709 · outbound

This paper cites Bleu: a Method for Automatic Evaluation of Ma- chine Translation.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Bleu: a Method for Automatic Evaluation of Ma- chine Translation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.559786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:46.283549Z digest=sha256:997e345440f2892160744aa76163c7d49e5e20440ef6a1e10fe732e721f66051

Observation cbd352a5-33a0-412f-8bd0-aeffdc2de515 · outbound

This paper cites METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.393684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T23:10:46.443896Z digest=sha256:3837562e9e92280cc8235ac6e9b79a295e3b641ed201040b8ce6b34e2b0ef986

Observation 5fff4125-b51c-4d5c-ba59-8a7cedb4ce4e · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports BERTScore: Evaluating Text Generation with BERT

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:46.566102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:46.566102Z digest=sha256:f4e54a35a5f94e6bf3123365870ad89041758b49041169e47825469bfacfc55f

Observation 6e5a2bf0-69a9-4a05-8e98-d63b4cb7fd0a · outbound

This paper cites GREEN: Generative Radiology Report Evaluation and Error Notation.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports GREEN: Generative Radiology Report Evaluation and Error Notation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:46.701581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:46.701581Z digest=sha256:085070f185aaa6a837a7960dcd3959326301ca3af9d313d58ca07e37790be3bf

Pith citing papers

No inbound Pith citation observations are available.