Pith. sign in

Paper Citation Record · LEDGER

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs

As of 6 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2605.05810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05810 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T14:49:53.357083Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact34
  • verified fuzzy10
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2471455d-3f34-45b8-8239-b49ecff50ba2 · outbound

This paper cites Vision-Language Models Do Not Understand Negation.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:09.727332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:26c9c7b47ca30ea85b8a1828276445c88ab59a5ae82b75d44b19dfcdda94674a

Observation 44219f25-ef43-45c8-b1d2-95882c382539 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.638022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:186199309d2b8a4f78a4984ff94cff57dce701790239ea2face428685cbfe9f8

Observation 5f1912dc-96f6-48c8-8016-ee71558519b8 · outbound

This paper cites Qwen2.5-VL Technical Report.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:09.470005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:160a32dcf9506adfed69c41a526ea8e891d725fcd471ebd6aac77d5a31171df0

Observation 4048306c-be35-42aa-bd1a-50e78abf1bcc · outbound

This paper cites GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.736536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:938048e568153bb6c2b8d6f28b58d8b14933d91e5ecf77834e009d4b37db2abd

Observation 69d1c6e1-372d-4f39-b296-3c78d768ec41 · outbound

This paper cites Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.458561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:e15187de32b0ae11bea88be9836ae5c5030dc7ffac04dfd5ba24eecba6608e13

Observation 35d6411b-e66d-4d0f-a9f3-b5bef2cb3633 · outbound

This paper cites CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.562237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:c03d984a66225aaabc04b794ae295a2b7e6dd31f7bfa192c20f548e331f84cc1

Observation 6c511b3d-8290-4743-b502-398a2f03e206 · outbound

This paper cites A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.674350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:f67a3c670c9be6eadb773a38376dec239db26a7faea1dd3603c2b9bb59ec269d

Observation d78b6b7a-e870-4b1c-8d07-bf823778c82c · outbound

This paper cites Kohli, Marc B.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Kohli, Marc B

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.339007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:bee8f1fd9fbcb0649b7c83b8e581404dd738f6516701fd2bdf97ce02a9795ba7

Observation 84981c48-cf8d-4be9-a234-c4371e6ea37e · outbound

This paper cites FaithDial: A Faithful Benchmark for Information-Seeking Dialogue.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs FaithDial: A Faithful Benchmark for Information-Seeking Dialogue

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.706343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:e0119aff46b25ee76bbd290562098b49a80f07cf073e2c687bf4b3756ec68187

Observation 3ab7c346-4322-446d-8c74-814e6ac4f089 · outbound

This paper cites Clip and complementary methods.Nature Reviews Methods Primers, 1(1):20.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Clip and complementary methods.Nature Reviews Methods Primers, 1(1):20

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.312567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:fdd63474100ab3313dd14479fc5d12643fafa7e3ee07bc1dadfd34ac6667fd1e

Observation 7550a738-e98a-4ce9-9ac4-8fbd156446fc · outbound

This paper cites Distribution-aligned decoding for efficient llm task adaptation.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Distribution-aligned decoding for efficient llm task adaptation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.391792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:b48e5d0dde22af0f898a2c397df861d3727f4ecfb5e94039956bebb465689c91

Observation df79bd42-b4bc-4afb-b92a-25a266118d25 · outbound

This paper cites Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.320761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:aa5800278f64f5bea03c862ecf87f3c814bb40df4c649be9282696a459b9baa4

Observation a022d22b-f553-4152-95c7-720b888e518a · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.Scientific Data.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs A dataset of clinically generated visual questions and answers about radiology images.Scientific Data

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.317150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:381484323fd34df5a0c5a86207db85f949969acb324d7d8d99f6550556ae4fff

Observation 92b0e867-6ef1-4a5c-825c-23e013f0b203 · outbound

This paper cites Cxrea- sonbench: A benchmark for evaluating structured diagnostic reasoning in chest x-rays.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Cxrea- sonbench: A benchmark for evaluating structured diagnostic reasoning in chest x-rays

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.744373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:1432699206264af979a5669b5a0703d4633ba9d31f8c51dd682306b83dbeae8f

Observation d62e3273-d962-4c3e-84a7-d0d828451029 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:59:50.699339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:4c3b98643d0b713e1da65897789caf343de573f67f545e2cdebf3e41690bf250

Observation fc633c91-e801-4f69-bcb9-0d5e9ed9645c · outbound

This paper cites LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.532522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:8c6ab2e27663a8b03c2cccc3604a68f8a8c4a6dd9841fd92668b8a2546501700

Observation bc60a18a-ddd8-497b-a0b9-d156d92e05d5 · outbound

This paper cites HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.490126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:fd691dbaa6a4869eca2f8444e532303463705d7403d8ebaafb919876ab1fbb9b

Observation c53b0ae3-d627-4921-a704-895d7261428e · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Evaluating Object Hallucination in Large Vision-Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:09.519197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:833361bdf653ff2dd5d08322d4dcd8539478d79ecdb017961b0cd4462628a6f6

Observation 0202a76e-07b1-40ac-9daa-139cb7ae411e · outbound

This paper cites Holistic Evaluation of Language Models.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Holistic Evaluation of Language Models

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:41:09.399485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:49793967c181164eab4ad5ed63f044202cba269ecb569b817e520c1c54c14319

Observation 6892dd73-c036-41b8-90ff-b455665bcd51 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:48:54.871535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:eb41e77b0625c2e52ee2f30f86243ef61bef6c02ead40367e5251e6cea7a23c7

Observation b3a5dcbc-b58a-4320-95ac-d0a41788426d · outbound

This paper cites SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.615748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:f6056bf060ef04400bcd92a8da6563064a4cabd1dfdcd0272592c2ba6222f19b

Observation 9dabbe39-a8ee-4821-a0c9-482c486bee37 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs MMBench: Is Your Multi-modal Model an All-around Player?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:20:54.147488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:f0aa706df7633cd264f4fe423f511a7d522ae4511e2f9e03371f5b4ad59c459d

Observation 36a07da6-07ec-41da-b34b-6be943e0e5f1 · outbound

This paper cites Med-Flamingo: a Multimodal Medical Few-shot Learner.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Med-Flamingo: a Multimodal Medical Few-shot Learner

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.715829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:5c93276e99743b4cffb08493feae2955822dc332104fec44cbe7c33dd94bb4a4

Observation dee71b23-d184-4802-8ec6-90e91c576b9e · outbound

This paper cites ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.474443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:bfe8595e344a8ed8ee5bb48499d2810d6124df2768a8f323ce27690900f9b37b

Observation 91abc5be-ff90-4a44-8a36-3527ec1df793 · outbound

This paper cites RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.527418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:f8bb5f6dad0c3ea503dcb4ba8cbda42b4328d245df2b26cd3b5bb0b658030866

Observation 8e24b5ac-8be5-4a8d-983c-2fc084f1ce0b · outbound

This paper cites Multimedeval: A benchmark and a toolkit for evaluating medical vision-language models.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Multimedeval: A benchmark and a toolkit for evaluating medical vision-language models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.386088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:3c29bcfa4e023a1f8a71556ffd794d67a48ee3236f3a855436906a3f62a1fb1d

Observation 608480f0-6323-4f3b-9b24-d7cca45bfe38 · outbound

This paper cites MedGemma Technical Report.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs MedGemma Technical Report

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:41:09.405200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:6160908d979f65a64de4bf4a1491edc2fd2c1a39e86c62d76fc33dc15f11383d

Observation 4751bf86-608d-4bd1-b235-33abe50ff23b · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:09.438010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:af13dd74ab4d2d57731fefaa69f08aa30ee5c61af227eed08651064026519ae6

Observation 9b9a227a-1ac3-4b8c-890f-0d793bffa3fb · outbound

This paper cites XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.451674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:062e944beeffd30d7d93fcf430071e3095d8e7c698c2264da4366bb425273537

Observation db1a796a-d77e-45f7-9fe0-900c3e1bdf61 · outbound

This paper cites Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.583258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:2d649ad4b73415f4d6c293a27330b94b5d9588ca628b33ce610e6f1085f99c61

Observation 47752c70-ef9a-488e-afc3-5eb0c072034a · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:09.433527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:24739799d4eb5cf0e752a0e4ece8b19a4aec40ac701d779daba016de57ca3e8d

Observation aa48f6b1-cf68-49b1-a061-8d19c787ca3e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:09.446519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:e19d198413f135e2c24675b95f963c0ab3a7c12336ca00d0b85bc31e29543c42

Observation a1e9628e-5de0-4c35-82db-d6a5ed30e9a3 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:09.522954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:f86bfc94803f8b286dbbab089de8bc14442158d57e23953685be04ce9ff0063d

Observation 5984e1de-d044-4f4e-bdb5-106f189bacd9 · outbound

This paper cites R2GenGPT: Radiology Report Generation with Frozen LLMs.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs R2GenGPT: Radiology Report Generation with Frozen LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.442654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:0ae093ee6b4561126f7fb107c46770013c906ff25b14c3e62157c82a1fe12a5e

Observation b5987a14-3f71-4601-86cc-ed95f5bf0c88 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.335596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:9af3dbe052594f207c492b422391ae6b29524260026f9427a2a1535a319952d8

Observation 5fe832a5-eda4-4556-aa91-8c466c15af78 · outbound

This paper cites MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training in Radiology.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training in Radiology

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.371946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:e26f6154a44691eacad56053132cf826cf0eb9d53d2552b40f6e85736a35ed56

Observation fb04ca38-78fc-48f2-bf67-64300102dd4b · outbound

This paper cites Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data.Nature Communications.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data.Nature Communications

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.303112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:bb8e2b38fbe54864a67b21a083ffe6ea684b576c35c49ddc52c3f468ff0e6975

Observation 42799c1f-05c2-4284-9e77-1bd122a756ab · outbound

This paper cites CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.463291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:392f8a28b4ad274308c632f856bca24be9f251b728be64b344b0ee744d279ed7

Observation b296d591-f873-4b7d-b1aa-5218d466808f · outbound

This paper cites MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.507531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:9ac0be4aa345052d3200d15ca7694cb528906303e81a3723e4e1d5bde8887d0b

Observation 06596d6e-f97f-441e-bace-3626784b30d3 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:09.761994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:041ad6c304d4362c3d889598e18833a894de765a58d9a5eb9d320e4f5e942594

Observation 0127b856-4eb6-4441-a5e5-27a90100a567 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:42.029646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:868441ab6aee75c0ff41f20c93955fc8112b421f962522a1f6fd93d041f80e3f

Observation dbe1d2df-6b2b-4939-980f-5c4ae1b57253 · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:42:22.536059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:6a9f515fe2ffc2417ab61a08d3ab1dfdcb6f31a2fba4249437c0771515ab2a99

Observation d7879569-cc97-4268-a7d4-99fa115bff3d · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:08:22.083800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:9590987dfdfb38840edcc11e6570f4b0ef56fac3d7f957b1290a3841505b7b0e

Observation 8bdb11ae-0eb0-4bda-9118-0f29dfb7a777 · outbound

This paper cites ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.629715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:172f3f43c762cf2134e056b08c1e8aec728dbcb72c8b2ea4ab40db21b8da0235

Observation b0025207-1b1d-4f97-a5d7-74a55f090b00 · outbound

This paper cites ReXGradient-160K: A Large-Scale Publicly Available Dataset of Chest Radiographs with Free-text Reports.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs ReXGradient-160K: A Large-Scale Publicly Available Dataset of Chest Radiographs with Free-text Reports

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:09.604286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:3c642a0d75d5630e13a6c0df8a493ad5cace8e79f9bf1db95f9244ae8ece78d7

Observation 514bb6a4-ba28-487f-97dc-19d867271637 · outbound

This paper cites BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.690348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:b54f6878844a69f2fea4d4feeba61b656ee9f4fc02dc7873f7da871b22ebec62

Observation cd907dbd-47b3-4ad1-aae8-46f4c2f0950f · outbound

This paper cites CorBenchX: Large-Scale Chest X-Ray Error Dataset and Vision-Language Model Benchmark for Report Error Correction.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs CorBenchX: Large-Scale Chest X-Ray Error Dataset and Vision-Language Model Benchmark for Report Error Correction

Reference 48

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T18:41:09.654894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:b16c389f401049a980b2249d4d4b13a55ee1931388d3c4e7a8d6fd25660ef04e

Observation f82ffe7d-259c-488b-881d-11c3b569d149 · outbound

This paper cites Absence of X.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Absence of X

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.327913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:a8cfb29b7a64a796d700e6a4ee6035e593a8b3d773f5e5b5c6445257c08d1378

Observation 91e15e50-9778-4e9b-a2e7-1fd267e482b3 · outbound

This paper cites an unresolved cited work.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:37:16.356538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:52255d9f462673324ea2887ab059bb8747c7a820161763151eb121134c334eb4

Observation 471a74a5-0558-466b-8b08-962d59766c0c · outbound

This paper cites an unresolved cited work.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:37:16.324201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:3f8ae5b90a84748591f381df751db1da1fea82f7ab1c80e600a8fd3911e54fa5

Observation 53e06dae-e3a2-478a-8a66-6dd7cd052cfb · outbound

This paper cites an unresolved cited work.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:37:16.331222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:02825582c21dcb406c575f1774db421f6133dd5cbcaf15fbb3e61f80d61b0e16

Observation 634d41b2-cd91-47bf-9a29-af36e68f75cb · outbound

This paper cites Absence-side scaffold (absence_step) Input format • Question text • A/B/C/D answer options Fixed reasoning steps.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Absence-side scaffold (absence_step) Input format • Question text • A/B/C/D answer options Fixed reasoning steps

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.308682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:b891fccd12c6721b54bc6b527e2d1f36380a8ac9a5dd7c8911cf4cd5e4e37ac4

Observation 3f6d8fa4-8204-4382-81f3-0855b280b9e6 · outbound

This paper cites an unresolved cited work.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:37:16.342175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:3539cc7e6fdc5b2f65efa949b151d615806b14c087443ae0845c1098bd066595

Observation 75d9feae-9296-4ee8-9bec-cbb29e0d4563 · outbound

This paper cites Output Final answer:[A/B/C/D] Presence-side scaffold (polarity_step) Input format • Question text • A/B/C/D answer options Fixed reasoning steps.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Output Final answer:[A/B/C/D] Presence-side scaffold (polarity_step) Input format • Question text • A/B/C/D answer options Fixed reasoning steps

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.345915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:48b820762189b870a56a2197b768bfaa4ebac041261b18d230268755e4e05fe9

Observation 6fd8dae1-3da8-4de9-9a78-8441e66ecfe1 · outbound

This paper cites an unresolved cited work.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:37:16.349466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:dcf0586c65801e292e8652cda032e1ee19798e94f09254109db6e3a2b4f27fb1

Observation feafc1ef-7464-4a9f-bccd-a8e2f4fd3941 · outbound

This paper cites an unresolved cited work.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:37:16.360275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:35a4fdfc845a0c7b267671bccbdabb2d2dac1041c71a3242ccb9a0eb22847e3c

Observation 90703af6-9664-447a-9329-e78c33f4bd62 · outbound

This paper cites Tracheal deviation.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Tracheal deviation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:16.352760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:6b896e847ad8ac270da50da120f5d5b13e3c999e9c31f9cc65cf1d34f5a7052f

Pith citing papers

No inbound Pith citation observations are available.