Pith. sign in

Paper Citation Record · LEDGER

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

As of 9 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 10 inbound Pith citation observations for arXiv:2505.17952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17952 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:40.334058Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:44:27.022437Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.448127Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddff6edb-6561-415f-a0be-7db1e94857ea · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.110924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.110924Z digest=sha256:cc441d57b427ae21725f2d19cac0a2e1a7fce2cbd2f822a24a412164589c0371

Observation 44390b50-5e53-4f62-a2df-4932ae45cb20 · outbound

This paper cites Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.205646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.205646Z digest=sha256:5dc38970b44a0d26e36ec181ffd883b50b78e2f98467bfffd4435bfe276a1d57

Observation 49b20d09-7281-4d9a-ae42-599be00232fb · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.294119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.294119Z digest=sha256:d2dfb037fcf95ee1c39e99c0c1fba81c6b1d3ed9bcbb333064aa06c406d03115

Observation fc8a92c0-c574-4791-b7c8-9f36ba96a05b · outbound

This paper cites Deliberative alignment: Reasoning enables safer language models,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Deliberative alignment: Reasoning enables safer language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.430945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:35.381361Z digest=sha256:5fb124336b2c7a40308bd8d9f1e84f5a1e60099429ae159a50d129984121e728

Observation 270b00cb-7f12-4d70-b0c7-ff752bffd7fe · outbound

This paper cites Capabilities of Gemini Models in Medicine.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Capabilities of Gemini Models in Medicine

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.475924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.475924Z digest=sha256:0cf0584257198c41c5091de637cd0d0c53381708f3eb3ff3f401248195488954

Observation 9196e662-694a-46e4-b70f-c0b95d6fbabb · outbound

This paper cites CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.566459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.566459Z digest=sha256:a6fe5c376230984073bc7c0d73f1a17ba48272228688bf27ce0e546a48e69a4e

Observation 8f56f633-7dac-414b-9fb7-4dba5eda8c96 · outbound

This paper cites Thinking and reasoning in medicine,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Thinking and reasoning in medicine,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.290576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:35.658180Z digest=sha256:800d29213b22a8eb08a0e4f7ee8086e6a2f5f5c298cf714e74e7dc78558b0eb4

Observation dbb93beb-e864-4aac-a750-6ab2d4386646 · outbound

This paper cites Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.780810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.780810Z digest=sha256:47bb10bad97ade1633d5040eeb8f2624eb7fb5feecddfd8f8d9be28b657e5fd1

Observation ccd0b47c-5ec5-4198-b2b7-dfcd303abc82 · outbound

This paper cites Openai o1-preview vs. chatgpt in healthcare: A new frontier in medical ai reasoning,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Openai o1-preview vs. chatgpt in healthcare: A new frontier in medical ai reasoning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.134395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:35.861039Z digest=sha256:ec745eda546bd4354ebed8410073d94f5f068fba81998e1a9baebb5f17792c29

Observation 9fb96894-1c47-4a7a-af57-6915fbf06936 · outbound

This paper cites A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.953046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.953046Z digest=sha256:d8101020723096739df4c9905baf4b34c72553272bc4af7988a413cffbc4e8a0

Observation d1f660c4-a146-46db-9daa-c1e61bf84ea2 · outbound

This paper cites HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.033970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.033970Z digest=sha256:843a68bfc8fc1454efcc9c869ce18b38d5f98274c5860567cbf095e715e82dd4

Observation dfb8cb23-fa89-4910-84ff-5fa8cc3b26fb · outbound

This paper cites Chain- of-thought prompting elicits reasoning in large language models,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Chain- of-thought prompting elicits reasoning in large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.008370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:36.140824Z digest=sha256:0ed902d7096ae882607b91c722e58519b8758bd2748a2567576253abbd792550

Observation e5a910c0-7061-4daf-85c7-5fe14a3228a8 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Scaling Instruction-Finetuned Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.223757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.223757Z digest=sha256:73f3e6ac99d6299bbe82eabe4f9ddf23dddb43d8acf75d9a02131d996af3b52b

Observation e7ea9409-fc35-482b-ae14-effe08d21422 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Star: Bootstrapping reasoning with reasoning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.838945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:36.292069Z digest=sha256:a0293a08fc0f258a4df24b410795b7960afa4c693ba52ba81023d601a8ef2043

Observation d33fe216-0b70-4e32-a9cb-71d7dc42abfc · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.371201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.371201Z digest=sha256:3a56c7ba9ce5b8364087fc3ab1f51f6b0796354c9bcfb6b99a951fcc4751ed3b

Observation 8d8e1ea7-7979-4aa0-bee8-ec6cc2b88527 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.475828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.475828Z digest=sha256:60f8cd4eef57c26481bb626416bcc226691b24350c0a6876ed93cf2ed4b0661a

Observation 957b55df-daf8-471d-a8ba-095d21816c1b · outbound

This paper cites Training language models to follow instructions with human feedback,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Training language models to follow instructions with human feedback,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.540075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.540075Z digest=sha256:4b143fc1594079c0f90471191a0c8bdde21907b8083a8f99f272731644146c20

Observation 8a6bfc6b-3108-4773-bdd1-0cbed13750a6 · outbound

This paper cites To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.635299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.635299Z digest=sha256:ec6134066081a5ade898143237fe5ba88387ca774b1d337fe42d6d9276b6c466

Observation 1fbcc017-d600-47fc-b76d-df04a5fa541f · outbound

This paper cites Direct preference opti- mization: Your language model is secretly a reward model,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Direct preference opti- mization: Your language model is secretly a reward model,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.662548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:36.701546Z digest=sha256:76d3c0dbf9fc14ac5db09bde65ec15da483110957cebdac494c5cab4a41c2cc2

Observation e386da9a-1b0f-43f7-94ec-e59668369e62 · outbound

This paper cites Openbiollm-70b: Advancing open-source biomedical llms with direct preference optimization,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Openbiollm-70b: Advancing open-source biomedical llms with direct preference optimization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.481327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:36.828546Z digest=sha256:709ef0b5641f0afd459d1909aedcb5e232e1ab00aa0427f1bc0249cdbef4db51

Observation 1754908f-7f39-47ad-b414-3c26bd069522 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.929289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.929289Z digest=sha256:3e1ec6df63bb7b1fffda99337d1ae76cd1cc115a5d496cb4e2ceb84ad9ecb0c2

Observation c80f83ec-8b68-4e3d-a607-b2e950a5a54d · outbound

This paper cites ACECODER: Acing Coder RL via Automated Test-Case Synthesis.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL ACECODER: Acing Coder RL via Automated Test-Case Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.015968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.015968Z digest=sha256:bfc5d163be7482b5f77b8920c4b5295673843bfa2ded4ae73f087fd91e3c3410

Observation f99ccc4a-99b6-43ff-8516-31e3ba326ad1 · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.084607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.084607Z digest=sha256:cb4fcbcdec12911543ba1cc4a1e5d7618ee95f4573775cbe4539f5835bd08a86

Observation 29a452ce-3036-4ade-a2c8-bbd105af3e82 · outbound

This paper cites HuatuoGPT, towards Taming Language Model to Be a Doctor.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL HuatuoGPT, towards Taming Language Model to Be a Doctor

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.171031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.171031Z digest=sha256:96677d3c08074fb1b16297069ae86ad1ec31d99da4ef449272da020d05dc5f6d

Observation e8570afb-a017-4443-a685-cdaceee4c015 · outbound

This paper cites BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.248353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.248353Z digest=sha256:58b864b15493272d77e43a9ad8c9f6ae557d1c22dc34bed8639195daf63073ea

Observation 73ddad87-78bb-4a09-a9e0-731211447055 · outbound

This paper cites Continual Pre-training of Language Models.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Continual Pre-training of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.342138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.342138Z digest=sha256:a0e06293dc523e7f9d22024268433747f90bcce2b03b0aa2fc912a62b4b95126

Observation 1473c761-1476-43af-8f1e-91f3ff449fd6 · outbound

This paper cites Ultramedical: Building specialized generalists in biomedicine,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Ultramedical: Building specialized generalists in biomedicine,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.431561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.431561Z digest=sha256:d642538d8039cafea6a85e77854f58ebe913326a3ae5fc473166141804b03964

Observation 4d5da96f-8c4b-4684-96e5-9e805d0d8759 · outbound

This paper cites m1: Unleash the potential of test-time scaling for medical reasoning with large language models,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL m1: Unleash the potential of test-time scaling for medical reasoning with large language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.649916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.649916Z digest=sha256:2c8523625414616eb43b7b30db34a0ca56ce01ef163d0da7f65f3659161bb89f

Observation d6ac4e88-471d-4433-beca-d35c9dca5d08 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.741316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.741316Z digest=sha256:7ca59a1318ff8f81fb621d0b640cf7e312aae054cf195e9ddad73f730eb9c88a

Observation b651a1eb-6835-48f3-8950-f88774b3528c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL HybridFlow: A Flexible and Efficient RLHF Framework

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.853487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.853487Z digest=sha256:7b74ae3b889cc44dd8e30d46ce582fa71b472a7a2c6033e95fe8a53bf549a835

Observation 954f89b7-6c7e-4fb7-885e-eca57af0df40 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.307188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:38.014800Z digest=sha256:b9748b237b6bac2ef9b689e8b5219a9b33aecdefc564b8cee16dc927e79203a4

Observation ea87c37c-55e7-4a5b-a18e-4fc1419b4629 · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.111618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.111618Z digest=sha256:2e25f1b1d028519f7c036d225e13384ed636c735f9ec825edde3d9fc7fc3f0cc

Observation 33dbd54b-e09c-4eb2-8c58-cbe4bddab2b2 · outbound

This paper cites Mmlu- pro: A more robust and challenging multi-task language understanding benchmark,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Mmlu- pro: A more robust and challenging multi-task language understanding benchmark,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.105838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:38.215005Z digest=sha256:b4aaae123aa23fb574a7d3bdc826bec0194c044d0eaf7a6e5b2a732a481343e2

Observation df4ce292-511b-4cf3-b664-3bdd37b2a5fa · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Gpqa: A graduate-level google-proof q&a benchmark,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.322847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.322847Z digest=sha256:3921681a9948f50b3eae5440827c990ed8dbb493ffcc4369699da1439c2a56a5

Observation f4b35547-716f-4666-acfb-e821b41c857e · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.423359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.423359Z digest=sha256:d8c72ace8a227c894c96dfb86c64655183ff90ccda766c54c5fbcf7d4187067c

Observation 2fc5f574-142a-4694-b38a-c57e3ae2a1fa · outbound

This paper cites MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.579027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.579027Z digest=sha256:b202e134ef9ef08d6f5d0082170d0573da5a99c9bbbe0305a7079708440861c5

Observation d514a37b-32bc-4cfd-adef-3c560f2ef65c · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.693288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.693288Z digest=sha256:12097fa8a0a6f5534de8ad5ce2db38a06d6badd2a173752769276db23c9cfc1e

Observation 497ac864-91eb-43b3-bf04-b1d08468b539 · outbound

This paper cites Openbiollms: Advancing open-source large language models for healthcare and life sciences,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Openbiollms: Advancing open-source large language models for healthcare and life sciences,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.905072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:38.796996Z digest=sha256:d00559d1ea99f4d9d82cebe602fc544e90e8887aab0602fbd953577bf67eae2a

Observation a95fa484-fc52-4f60-b28f-9f93522260ab · outbound

This paper cites Towards building multilin- gual language model for medicine,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Towards building multilin- gual language model for medicine,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.710949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:38.908872Z digest=sha256:a371b9d5b37932c4fa915f6039dc253e9e57a87426de898c6884d3178b3c5be9

Observation fee30329-026f-44a0-8b51-306c9d1155a5 · outbound

This paper cites Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.026377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.026377Z digest=sha256:c9db1ea598816846577d5af367c0780cfdac078016d7638e6e6cd79b16456d51

Observation b77641c6-aac8-485d-908b-4405227b7c09 · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL What disease does this patient have? a large-scale open domain question answering dataset from medical exams,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.162856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.162856Z digest=sha256:2f1c763c13565fe84557e6023de5326305da84d4ea9a886c217278a090437726

Observation d4c9f1c4-fb21-43f8-a5c9-812ee0279875 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.514776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:39.278178Z digest=sha256:c22a4fc15ac35b940a0fbd01de07500d00ca816f66074cb531188b11e0b52c0d

Observation e1c5258e-34e5-48bf-8a75-e9254d4a3c30 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL The Llama 3 Herd of Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.409211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.409211Z digest=sha256:48e7048004d6f56869f374a49d9246e541fde9986a7ccbc3b856e6fd891968e4

Observation a0c4a48a-0fd9-4268-b15c-a59d446e6bc3 · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.559067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.559067Z digest=sha256:b7a723cc4e900fb2c0b5b9275c6df66b91deb59010398a50821b34db4bd0b98d

Observation c9a988c5-1a8d-48cb-b5e2-854e6aa9078f · outbound

This paper cites GPT-4o System Card.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL GPT-4o System Card

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.677799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.677799Z digest=sha256:599f8041ccc952c83e9bb50bfb27f0b6af8306f8085aaff235aa160e6c0f36f9

Observation 9c02b31f-638b-4072-b27f-f4f295cdc9f6 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Qwq: Reflect deeply on the boundaries of the unknown,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.352436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:39.815897Z digest=sha256:341ae38359e88ce75beb39958e881863a6caf4ae452c60320f5dd27f62025888

Observation 4906fd2b-fb4d-478b-a2fc-d306805a407f · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL The claude 3 model family: Opus, sonnet, haiku,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.155809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:39.930681Z digest=sha256:289e4869564d525ead0da1987eb0429339c9dce7cae726f7beb3f0fb6f2d4361

Observation 85eec960-6abc-4426-84d3-6e39bbd053cb · outbound

This paper cites DeepSeek-V3 Technical Report.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL DeepSeek-V3 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.068205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:40.068205Z digest=sha256:84dac77a41d41ed69610336b1ce6c4e1d3c8f417ace7915f05c284a34b4d4886

Observation f0042886-c4f9-4cf6-a277-04c03afbcfbe · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.218826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:40.218826Z digest=sha256:622cca6e1e0dde977f3b3eac5e718986b2c04ab9892acc622ac7837c699c5ade

Observation 08e58909-ad10-49a8-9531-4114269efcd3 · outbound

This paper cites Qwen2.5: A party of foundation models,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Qwen2.5: A party of foundation models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:40.962287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:40.334058Z digest=sha256:6452eab4adfc2496972b5ac62eff29ad988ac6a3926c908e9ab9bb820469c85d

Pith citing papers

Observation cd855766-5f9c-42ba-8f38-91dd09040779 · inbound

Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning cites this paper.

Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T10:44:27.022437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:44:27.022437Z digest=sha256:787efdef5fd473be88036f7627e380ae471192794e178670987d04012372da82

Observation 16937fd7-60aa-4f84-b71d-a777a78a7948 · inbound

Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning cites this paper.

Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:18:43.366828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:18:43.366828Z digest=sha256:70b726e28807202535034dc539ee4d76d7c09d10676c0e5fbb560568e69fde1a

Observation 281d21d1-8ff1-4947-b757-d5b19d5a9f99 · inbound

Baichuan-M2: Scaling Medical Capability with Large Verifier System cites this paper.

Baichuan-M2: Scaling Medical Capability with Large Verifier System Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:50:13.102408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:50:13.102408Z digest=sha256:ae537da61a0b0db08cd9fccf5f9edb590ba599521bc8618d859a2f47809f9aea

Observation cb402813-c124-4aaa-ad62-52804fb69d88 · inbound

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification cites this paper.

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:51.561726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:51.561726Z digest=sha256:47b61e22428852025cc19f5b5e9ac211d9b21d7c61a27cf0ac9a6b03b5506e16

Observation 4744b8c1-bbb1-4855-a935-21941b67f116 · inbound

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training cites this paper.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.249504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.249504Z digest=sha256:f580f9e3d17b690d71dd4973788784b1bd08661ca110614e486c87cf9a419d5d

Observation 36baa877-01c2-443a-8b05-923db99d484f · inbound

Medical Reasoning with Large Language Models: A Survey and MR-Bench cites this paper.

Medical Reasoning with Large Language Models: A Survey and MR-Bench Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:25:26.786876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T10:21:39.892271Z digest=sha256:e93eafc68ec659c1ab35f89a38b77344b1d3d997638b691c10d6c88b8a6bbdfa

Observation 5bcc3865-1977-42f6-9ce2-32d32e606e1e · inbound

Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve cites this paper.

Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:46:34.013256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:35:32.915685Z digest=sha256:219e389c413b06443823362fb56df112d35b9ae9fd8c115daa7f64d0cb41facb

Observation 2cf0a86c-9e05-45f4-8702-dc3b3be5d501 · inbound

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics cites this paper.

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:23.869943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T04:32:16.930291Z digest=sha256:e9fa4493e3b91d3b62d4962bb307471da2fe6156031dbde14dd0fb82f795b812

Observation d0d8dc84-b310-48dd-bfd9-68ccac497add · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 186

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.449550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:14aad9009e8c44947b4bbc94e9474bdcf072047596fd9b55a5f08415857486cd

Observation 5b87cf8a-7fa2-4afb-9197-dc5a8f416f8b · inbound

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? cites this paper.

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 155

Resolution
unresolved
no resolver link, observed 2026-07-30T16:01:42.470884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T16:01:42.470884Z digest=sha256:389ffa1273437aa71cf3b4a5cac8275f1486992e44cb0edf8959af18583c0c7c