Pith. sign in

Paper Citation Record · LEDGER

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models

As of 12 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2508.21430.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21430 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:21:57.784836Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38e85c1a-bdd7-4236-85b9-cbcae452f34a · outbound

This paper cites an unresolved cited work.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:01.190651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:56.261839Z digest=sha256:0f65dae5cc0d9d5f90b0744f97bc54b1008d858461c11c557b91a35b12201640

Observation c319738c-c5eb-41f6-af4c-93a91875a08b · outbound

This paper cites an unresolved cited work.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:01.105368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:56.454192Z digest=sha256:a188ea0c19caabda7f501c87ffaf80f72ec33c566d15367be43c77c9fab93ab1

Observation 776a21ff-b22c-4a04-8626-15bd2e718400 · outbound

This paper cites an unresolved cited work.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:00.994829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:56.614861Z digest=sha256:6d7b9509fd111f9d621e86295343266d3241d72bdf34f904cf1c2ed30a691fc8

Observation 67f35a9b-fae1-42e3-9332-c2d97c799206 · outbound

This paper cites an unresolved cited work.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:00.834749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:56.710845Z digest=sha256:bc9876e74e2e2ae32e1b0cbd329d287fceae7484621fa4f5ba53d7e391c5175d

Observation 9d13ea83-753d-468b-896d-04d4127abcf0 · outbound

This paper cites an unresolved cited work.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:00.704757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:56.844756Z digest=sha256:2098755a8393054b3f0116b47d0cb62ea75d08ada62802c16b569f89093b371b

Observation 80d65daf-bf30-4c19-a3bb-6c5f5417cd2a · outbound

This paper cites MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:55.894317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:55.894317Z digest=sha256:9fe22aeb768644640c5b9dc5f8282f77f27c5d10fb31a54def6dad239fa56c10

Observation da040793-6afc-4da9-a976-96a115cf4ac3 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:55.973546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:55.973546Z digest=sha256:cc1d30546e86689f17834886729da31361b53dd2d02b5cea2fa83a6e3572d54b

Observation aa498442-b8c9-46d0-9379-96303932e058 · outbound

This paper cites Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:56.146103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:56.146103Z digest=sha256:7a5feed84bdb8749996cc9f7e7ba75b6e9f91c461b8d07b6349e8d91674721ba

Observation 548f9eee-b684-45d2-80d9-790aefc92624 · outbound

This paper cites Avoid any position biases and ensure that the order in which the responses were presented does not influence your decision.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Avoid any position biases and ensure that the order in which the responses were presented does not influence your decision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:00.585461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:56.957220Z digest=sha256:1c0f8addb0027536f94803c80c8bd295df78425241a7a96557e482030c068d35

Observation ecf46c95-0526-423c-bcef-c21d657f7c6f · outbound

This paper cites immediate PET-CT for a suspected pulmonary nodule,.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models immediate PET-CT for a suspected pulmonary nodule,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:00.434624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:57.040837Z digest=sha256:793ba08ea4c0281741f94d98656763d9dd420f1fca8e7a07bdd43ec691a16d3d

Observation aba0a1e2-a3c7-44fd-9f42-d2956b6cc78f · outbound

This paper cites an unresolved cited work.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:00.224753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:57.212535Z digest=sha256:b58adc2ad10266a88a57f0248353210f74cf1f143f5f4210ac3e78fd044258a5

Observation 2bcceeaf-4caa-4493-97e0-180d924d1cab · outbound

This paper cites an unresolved cited work.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:00.044919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:57.326352Z digest=sha256:0003e011e1c3c642315fcc266df855299a0fea7fae4be2c88efb04069e631e2b

Observation 79d69425-8446-4115-b84c-99b5157d7159 · outbound

This paper cites an unresolved cited work.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:21:59.914761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:57.444933Z digest=sha256:235af9e59389e0bebf239a3d9f5d39d2fa6632c9c260b6e41bcd8c8093ccfd79

Observation d3b7a27a-046c-480a-a762-fbfb440a9ae0 · outbound

This paper cites The most likely cause is... I recommend.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models The most likely cause is... I recommend

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:21:59.734747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:57.596334Z digest=sha256:8cf3d33bfee65972380603b0666cd284e466b23c7657152fae714e4d587cb5de

Observation 6e29a972-f3a3-461c-a627-10b8fdb0f1e4 · outbound

This paper cites Additional Notes - All medical images are de-identified from real clinical cases.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Additional Notes - All medical images are de-identified from real clinical cases

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:21:59.604755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:57.784836Z digest=sha256:235759dd7fa08541ded7123b3fb5abe9a1b2e714d6fcf982820b6800cf93a3fa

Observation 527bb06d-1583-48f6-a375-a6a5466e66eb · outbound

This paper cites Fenglin Liu, Tingting Zhu, Xian Wu, Bang Yang, Chenyu You, Chenyang Wang, Lei Lu, Zhangdai- hong Liu, Yefeng Zheng, Xu Sun, et al.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Fenglin Liu, Tingting Zhu, Xian Wu, Bang Yang, Chenyu You, Chenyang Wang, Lei Lu, Zhangdai- hong Liu, Yefeng Zheng, Xu Sun, et al

Reference 1654

Resolution
verified exact
arxiv_id, observed 2026-08-05T14:21:59.131160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:21:55.651594Z digest=sha256:2b6ec120018d766536c97d983257ed42eead977cead21d7030273ec23d5bf2b5

Observation 33c3338e-64c2-4b7c-a62f-2420b1d25284 · outbound

This paper cites Judge Anything: MLLM as a Judge Across Any Modality.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Judge Anything: MLLM as a Judge Across Any Modality

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:55.725194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:55.725194Z digest=sha256:1eddfb7cb1cb68ed88c545ca2219d72b49b12cc07443c3da7bb8df34ed861405

Observation f89be8b3-0511-4bd6-85c4-d878df93c4c4 · outbound

This paper cites VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:55.825705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:55.825705Z digest=sha256:597f137dc7b76c85aa564701472fc6e53487996cbed3d6c0eda9bcf10305a79e

Observation ce46be49-8ef7-4a6a-b82b-5cca7ba006c3 · outbound

This paper cites GSCo: Towards Generalizable AI in Medicine via Generalist-Specialist Collaboration.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models GSCo: Towards Generalizable AI in Medicine via Generalist-Specialist Collaboration

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:55.534892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:55.534892Z digest=sha256:a59899b9bb9d481b0758443ebac5a8010f648575bf6d5ea801131c3882c107d2

Observation 776d64b9-c31a-422a-b910-27d5a4a555fc · outbound

This paper cites Qwen2.5-VL Technical Report.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:55.432245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:55.432245Z digest=sha256:f09b6867d92799e0139dc518ebe538fd47803488276a67dbb8e3605d4847e4ab

Pith citing papers

No inbound Pith citation observations are available.