Pith. sign in

Paper Citation Record · LEDGER

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 6 inbound Pith citation observations for arXiv:2505.19213.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19213 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:42.605024Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:00:49.780513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:15:45.170778Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 083e3c0e-d94f-4e1d-9da8-01239ac2f2e8 · outbound

This paper cites Qwen2.5-VL Technical Report.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.347787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.347787Z digest=sha256:85a0164efa77eadb59d6b89886239fdc4f289b9ea02f9d9e3b6cff8cfb036f07

Observation 485406a0-4f21-4692-9d72-13746cd7c810 · outbound

This paper cites Curriculum learning.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Curriculum learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.355061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.355061Z digest=sha256:ceabbf157286765c1d8103c7b0630c9e145fec8c3761104221934ed8275f79bf

Observation 4984e00a-9b04-4f94-9a48-9fc79a1f61ca · outbound

This paper cites HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.360661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.360661Z digest=sha256:a064544e57f6a413a357b70f64c58a85a1e2e37dba30e37db55719d5d826c228

Observation f4f7ae01-793e-4ea3-90f5-516f7ba1e176 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.366000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.366000Z digest=sha256:b04bf250212c7d58dd6347714a600af84cca884255f262ba8aa4e1d335144e78

Observation 414891cd-b1a4-4552-9e48-849577895210 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.371834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.371834Z digest=sha256:9ae9ca4ce8c19d8d5c001ba6c9d36d90be913a7797593d94e0e2d2a488b0afd7

Observation 3d957eeb-e9f8-477f-8c4b-6e156613f329 · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.378226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.378226Z digest=sha256:0ec5f5d9a25c6a9252f9fb5f7356b1951004cc3c74bb1af2759afab0360de8e5

Observation 194b85c6-e245-417c-be55-fc959ed2dcad · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.384653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.384653Z digest=sha256:0008ffe3d1086c0fb826349c041f078a1eb6d92309f85db69e590d93c6b80865

Observation dc7f90b0-980f-4fb5-8fb5-3d6f445c2311 · outbound

This paper cites Omnimed- vqa: A new large-scale comprehensive evaluation benchmark for medical lvlm.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Omnimed- vqa: A new large-scale comprehensive evaluation benchmark for medical lvlm

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.390120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.390120Z digest=sha256:8acad8e52ae08f5a88f3ba9344fd6cdac5c263e16431d7e736d5e9a292469b4f

Observation 7294fd7b-2b50-4a7d-8e47-73c13e29ea10 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.395324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.395324Z digest=sha256:daab54684f58cf8fc14e472db82f6aca4f5d814c09363632535ff258b889c087

Observation 71f572d3-4d46-4a38-97da-8872d8fef298 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Gonzalez, Hao Zhang, and Ion Stoica

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.400311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.400311Z digest=sha256:ea5b8dc8694e7fa469cc628839b23ba97a23391540b03cb8bc8ef67dc16178c2

Observation 16a6aac7-759d-4c74-9caf-8ada360f6976 · outbound

This paper cites Med-r1: Reinforce- ment learning for generalizable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Med-r1: Reinforce- ment learning for generalizable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.405019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.405019Z digest=sha256:b391860673c4d1ce9058375ab4328d5de926364f91bbdf6d7305742fc1223b98

Observation 47f3ec62-c8eb-49c3-b21a-2d205f4ca00f · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.409995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.409995Z digest=sha256:e0f1a1c210f0b63103721da765ad3de95eb8b8fca9d40a621d84f7956c0b8682

Observation 924a7d54-5ccd-42f7-a856-9053fd5aeba6 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564, 2023.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Llava-med: Training a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564, 2023

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.415633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.415633Z digest=sha256:053be29e49bb4a902489d8885e45689f22253ac50c49f927c60e551203181bb2

Observation 64b2524f-111e-48c3-88b4-45eb68fc228b · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564, 2023.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Llava-med: Training a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:43.889177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:24:42.421492Z digest=sha256:7fdc2068d09890aaefcaeeafeef99220abaee6d9f66cb6aa02283124b5e68e71

Observation 7cb2d57b-458f-46c5-a3f0-de02b9b13686 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.426638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.426638Z digest=sha256:14a6b6acf4957e13ce55a34c1be3199b26ab4897b62dbd4e2b133ef30357a14e

Observation 11c2e270-d74b-4a94-9f9d-1f2ff245b5f5 · outbound

This paper cites Slake: A semantically- labeled knowledge-enhanced dataset for medical visual question answering.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Slake: A semantically- labeled knowledge-enhanced dataset for medical visual question answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.431160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.431160Z digest=sha256:137686a0d0a66045878f826e9312c71144a2745ef6d5477e34dde905f96bbaa9

Observation 743d4455-6f6c-46ed-9c84-51601b7f11b7 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.435904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.435904Z digest=sha256:c3e88062188431b4c220ea0d17f77d7c4b7ee1e4dab5f3ac3cc583946131b2aa

Observation e165dee1-00b5-4151-89af-8d772649130d · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.441501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.441501Z digest=sha256:3199738c9e0fa7f70ac10fd6ca4cada5cde8c65dedd51e6d2ce3901b9507278f

Observation 30310768-6994-47ff-8bc7-445dbc9c98a9 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.447522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.447522Z digest=sha256:47c4d6663e9967237ae8520ed873b0ec9f57a03560969bc368972d230f1ad3ca

Observation 04e546ed-1663-4d8a-a641-f624cd64382b · outbound

This paper cites Med-flamingo: a multimodal medical few-shot learner.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Med-flamingo: a multimodal medical few-shot learner

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.452658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.452658Z digest=sha256:83d8861bc8a93ce732772f0ef0e8da46a4d0dda084a2fe5328fd8fc46ccb5aae

Observation 16f03c70-230e-4852-847a-cf1b281a323c · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.458642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.458642Z digest=sha256:36217d6a37a54157b866b3b0ad50205931a934a4a2807ef0f94dafca601073ab

Observation c757ce95-577b-468f-a5aa-a7f73106c3ed · outbound

This paper cites Learning transferable visual models from natural language supervision.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Learning transferable visual models from natural language supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.463525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.463525Z digest=sha256:caa72e486f619eb62127559591732b224605ff8f237edbe22d66ee9cd1448f2c

Observation 3229f476-5b6a-4b7f-9d02-e3d08ae3bcaf · outbound

This paper cites Proximal Policy Optimization Algorithms.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.468200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.468200Z digest=sha256:1e8d5f406105e2621f47834482d468cae9089301a083d0a915c334f475c34870

Observation 9de819ef-1b09-46d3-95c6-3a74402fe734 · outbound

This paper cites Quilt-llava: Visual instruction tuning by extracting localized narratives from open- source histopathology videos.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Quilt-llava: Visual instruction tuning by extracting localized narratives from open- source histopathology videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.473462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.473462Z digest=sha256:56ff0dd6e96bca0ab41aadf834cd82fe2467d5dbefc02b47c2d6db3ae6325aa9

Observation 5a9508b1-cedb-452b-9d13-dbb420cc2cdf · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.478426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.478426Z digest=sha256:bd2a4fdf25a1922be05cb8d1ae418d0c94433415953edcf75979117a1e44bec0

Observation 38209900-4f7e-46df-a766-5a5b326affb1 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.483506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.483506Z digest=sha256:dc98c5c329489d151ba70f05e02de03bfb5c7ef110b5a143e3390a8374a35760

Observation 8dfb0a38-c2b7-4582-9303-d6f0c64ab1d6 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.489700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.489700Z digest=sha256:9be37d9900a89693d8ed4b30cf97db0ee5dbeb42e0fa122c70b2149a62502045

Observation 01b412b5-8ee6-4f45-b4e9-76c18a49d7c9 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.496078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.496078Z digest=sha256:82034ed16e011babd95303f9e0ec741be4e635820758e690b81d627079af8604

Observation f66319b4-9117-472e-9c0e-83cf91026761 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.501264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.501264Z digest=sha256:231ffc083f53acc99de16f0a65141140d72a6c11d7971fced88f1980d537809c

Observation 1c994559-1c4f-4a4e-a084-4f9d5dee187f · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.507040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.507040Z digest=sha256:054444ce7bff6f42d4ea72be63bb1c820300e9eac4c920822a8f733667611bda

Observation 19f87a5b-8251-4fa6-865c-5da7a6681925 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.512932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.512932Z digest=sha256:9abddaee99b7ca9798432a4a100ecab12d1d442f24c7963452b3cf0ea6e67700

Observation 2f426e4a-38af-491c-85cf-cda98f68442e · outbound

This paper cites Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.518960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.518960Z digest=sha256:3178f50cd96dc390f5b482c37c84c2e93b62e26962bc335a4adea67d5f05562d

Observation a465c906-8a80-426f-be98-5d5d5a5b2b6a · outbound

This paper cites MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.524211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.524211Z digest=sha256:0b8b2fa4d1a0074dad5169b7073b95b61d5617622ce89fc96dfe4eb828df76eb

Observation 63a3a088-1330-4acf-87fb-41c8a6bf9bc7 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Yi: Open Foundation Models by 01.AI

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.530817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.530817Z digest=sha256:b7ceb9f7530ed65bcc0a3c9c71bfde8346c6cfb5161aa37d55fea6788475b64b

Observation 8aa50367-b97b-4643-ab22-b3351f0a2e56 · outbound

This paper cites FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.536609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.536609Z digest=sha256:bd635a9c44562913e866f4116daa0b16501ee03756c0e4d41e3bfcacc55e99c4

Observation 2c688696-33d5-4283-99ed-7a871a8cb5f4 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.542366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.542366Z digest=sha256:246b9760b479b96c7f35f061cbc05fd9f39920bb6dc6249f5c67cc5175e150a6

Observation 62d3e441-4423-4389-b657-bff9f78b355e · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.550786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.550786Z digest=sha256:83f24014cb777f56b06649292b1355ed62d54abd6608b08f363778e03a3eeb6d

Observation 5cf5abd0-9e3c-46ce-a134-fd05dc177399 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.557957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.557957Z digest=sha256:0f940f005c4771c9edcc7cc4d9aebd66386c17cbf9f55e69e192178f52934df7

Observation 59ac7428-3669-4f1a-8ee1-344b0f87e5e7 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.564168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.564168Z digest=sha256:c837fc9f899a56d5a846e13969fba20318e1efb1a6906f298b22f00ba689a9ab

Observation 18c884f0-3a42-4aea-af1e-14ff3c327dce · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.578202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.578202Z digest=sha256:f879d8d7abcbf0612f480417511539bebb5e5225a2ce5ac7ba8c31f6b547ca33

Observation 4fa16a12-2e15-4336-8e07-ae54486d285e · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.584939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.584939Z digest=sha256:f8074e0b71f912cac2134c88025884a7a94709b1436b964a8b49db073bcb9e14

Observation dc64ed3d-8be0-4113-b7ad-b4f368c8077b · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.590126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.590126Z digest=sha256:a7ed8044c7ed281868dd074216006f73dfa969cc90d32db665fbed1b2199f052

Observation 07d79c18-f00e-4afb-b695-829c38aed203 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.595164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.595164Z digest=sha256:fa2921b3e724ad4e4438dc6cb2050ceff499fc48280377b6509f4d67584aac80

Observation dd59a4c8-da6f-43b8-9998-f4168e7382ee · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.600093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.600093Z digest=sha256:e3d698aa3d1c7c78ed9da4f62a89c2dafdbf93e0632fc452c4fbbd4f9d7f7bb2

Observation 247bc90c-215b-453a-ad87-185f4f606d53 · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.605024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.605024Z digest=sha256:78324b73e4ddbe650f1777b5d5bdd69f91b87fba8d126f833df332b97ab4ef62

Pith citing papers

Observation 67b2389a-d06d-445b-a313-a7fe8f0411be · inbound

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning cites this paper.

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:00:49.780513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:00:49.780513Z digest=sha256:123392d00952d353aeacd11c3de8b012ceefc7776f94fec904fffd3115686f34

Observation 0bd09e0e-b0e5-45bf-b46e-894236e74223 · inbound

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification cites this paper.

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:51.711834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:51.711834Z digest=sha256:0c503e01ca4ff1da37c55bf853d5782215dd03a79a9d25b73c9fcfb999e4996d

Observation b3beaf0a-ec85-472b-8a70-3919146c80c0 · inbound

A global log for medical AI cites this paper.

A global log for medical AI Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-04T11:34:14.723040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:34:14.723040Z digest=sha256:86f511f3a6a1e6f6ed7df4232adf454f5d126e8f1a31e97ddb055841e134133d

Observation 4f11b0e0-75db-4281-8571-257bbcc4e6bb · inbound

Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming cites this paper.

Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:14:45.427393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:13:45.830192Z digest=sha256:b74aa94b3ecb6f18a5bf5f313d50b92d2d23f2a6e4c7fc8e7833fa8401ab5e14

Observation 9c9da16b-edaf-4586-9088-323df898fe22 · inbound

Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming cites this paper.

Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.753056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:48:55.898188Z digest=sha256:0236912f16d9b5fb0f9d68fc29f58ffbf22e3cf94cbdce1133db3bbd6626ffdf

Observation 0f4a817e-1ce2-40c8-a068-9739e97f1dd7 · inbound

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning cites this paper.

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:45.172460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T05:36:43.609602Z digest=sha256:389b5c30ae988016a208966803ba36365ede34f104bd0c9132ca9903db1fbb9d