Pith. sign in

Paper Citation Record · LEDGER

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2506.13102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13102 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:08:56.179861Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T20:21:40.867354Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T20:23:23.797407Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7b0a2c50-eedf-411b-8430-ed1bac389c9e · outbound

This paper cites Language mod- els are few-shot learners,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Language mod- els are few-shot learners,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.124456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.124456Z digest=sha256:999f635cb205d8410b060f595cfaa4b8bfd1edb32ead4481ddae0aa3136dc325

Observation 4753e2e6-d135-4a78-b5b1-c63cd0609a0a · outbound

This paper cites GPT-4 Technical Report.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.141698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.141698Z digest=sha256:94b34b6e9aab2f9940dc3809ce0efeb57a8672486117b5a2e26877b56e6dd687

Observation 4c616e0e-d59c-4497-9610-82db37a225d7 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.147428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.147428Z digest=sha256:770fc817134e91bed20b6b0a701d29779ae05803c5297a18c34a631a44ef9c32

Observation 6e4b2c4d-7cf0-4a35-9bb2-7a1761e7923a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.151764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.151764Z digest=sha256:2423afc5fa806ac61e2ed1836098df6850ac756d37566441673a39efaac00ba1

Observation 06756117-ce0f-416b-9dfd-21e71c5bbe85 · outbound

This paper cites GPT-4o System Card.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs GPT-4o System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.223274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.223274Z digest=sha256:28f703d272af44dc926097139d901a14f0a80dfd013af4a14b526d1f1ea2be78

Observation a593b253-553e-4392-b8e4-642814dda7ea · outbound

This paper cites OpenAI o1 System Card.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs OpenAI o1 System Card

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.227739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.227739Z digest=sha256:a2759b08f82a6d5a577f3ccec9b178022ab736123bded39cf412fe0cb6bd9db7

Observation bbe4fede-64ea-4740-88f7-a8cf9a20e6e5 · outbound

This paper cites The Llama 3 herd of models,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs The Llama 3 herd of models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.800688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.233969Z digest=sha256:b9a693a55ba8f1872ddc7dc887759188482f334dec9940759f8d7921486beb81

Observation c61b68e7-ddc5-4bdb-aaaf-13692b1cc8cc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.238359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.238359Z digest=sha256:d287fe7679b079c62f111129290f95b7706c302d96681946c81197da8fa428e7

Observation 064a8092-116e-445a-bfdf-6a7965ba7994 · outbound

This paper cites Qwen2.5 Technical Report.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Qwen2.5 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.242792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.242792Z digest=sha256:0a351a500dccf8f345e881f502297adb23131fd2ecc6269aaa968dd51b8fa5e9

Observation a3e259a8-4b9a-4b80-b93a-ca1e3fc33bf6 · outbound

This paper cites Flamingo: A visual language model for few-shot learning,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Flamingo: A visual language model for few-shot learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.789598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.366044Z digest=sha256:a1868333bc56d11f7c98a5e85496f8e6b1c7c94953b4605425496b93ef8d926d

Observation 05a1f538-d81f-474c-b8a2-56c9f50484e9 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.710527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.405005Z digest=sha256:b0dd3931be2f705ac252eca8afdf484f34be7c8a0bbe7895ebe74fe23ded1c69

Observation 530c4658-ba4c-41af-bc8e-fce5cd327e10 · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.410289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.410289Z digest=sha256:57423ad819cc978151164b5bed4f1b10914b8fb7bd4e18d3d32db0c7b41a88eb

Observation b95d9010-9af5-43c6-b9e6-b1faae1f807c · outbound

This paper cites Visual instruction tuning,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Visual instruction tuning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.562887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.414433Z digest=sha256:8636c6a7b4b93a1483616542e536b1a1955413830663d6bfbb18acb2ec8615a6

Observation e385f19d-5186-4fbc-9e23-8ac8856a921b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.418177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.418177Z digest=sha256:74445d169ba562e32f1978244ba3c36b482030b0f3e09bb85f61f5f0e8f29902

Observation 22826580-31b0-4246-a375-d883e59e8234 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.462198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.462198Z digest=sha256:5ba29c463c12a11f637ff5835e7a6695dd0ef391e3753f9c8f547393b40437d9

Observation bff8de91-a652-4f31-ac7c-905595c66010 · outbound

This paper cites xGen-MM (BLIP- 3): A family of open large multimodal models,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs xGen-MM (BLIP- 3): A family of open large multimodal models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.535151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.535151Z digest=sha256:b4aa53b034f3e81a74c59840b554426ec19212ed5420533f6fee76ce8a95e5ea

Observation da645a20-c07b-4bfc-8d86-5117e16b7cd0 · outbound

This paper cites s1: Simple test-time scaling.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs s1: Simple test-time scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.538463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.538463Z digest=sha256:679953d1f72fbca46245214e926dc35deabf6d3dec1c6b940165bb9d323fdcad

Observation ec0a4d29-8a6e-4bb1-9bea-0771b8316d95 · outbound

This paper cites Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.550107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.542632Z digest=sha256:c053bed36d7e919e9a55d7ba4611653c519c13ba4af9788b9557cffffdd196fd

Observation 4e9df956-644d-4631-9f00-cffa58cbc8c8 · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for LLM reasoning,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Towards thinking-optimal scaling of test-time compute for LLM reasoning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.546352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.546352Z digest=sha256:4e4cf595eb684953f8065310f1541df1643a3524bf72788d51670fecc5232540

Observation 4bab2202-125b-4014-97e0-44bc5e33e238 · outbound

This paper cites Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.549727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.549727Z digest=sha256:f29e3a326532e89b6e02734e9b51e2cbeea58ca48b75737654db21adb26c98c3

Observation e44cf1a0-81c2-4168-aefa-483f175b4bb3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.562878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.562878Z digest=sha256:bf760891c3bb386e57de7ec01579d93c0924f58dbdf223ac0e67fa39af98a579

Observation 0da8faa2-e808-44a4-beda-1698e698f800 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Direct preference optimization: Your language model is secretly a reward model,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.635279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.635279Z digest=sha256:524d5f2ff6b00918d7572cf2e43b4d5bacb2a8e4e142b1c3ee52509e38d0d288

Observation 76204dec-a21e-4f2b-b428-85aa95f56659 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.690086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.690086Z digest=sha256:458d5d6a33d5f454c699d3be557c3c25db3c3304184c3572f852411b83bfbdfe

Observation 22b1ecc4-a5eb-474e-a1a8-c6eae3be5a34 · outbound

This paper cites Multi- modal understanding and generation for medical images and text via vision-language pre-training,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Multi- modal understanding and generation for medical images and text via vision-language pre-training,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.461742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.694124Z digest=sha256:63f3a2d79811590fa8e828841c5e07c16e96d4adf4a4b0a74d7c31a1b88e532d

Observation 28243b47-b668-4044-bca5-9156feb687cb · outbound

This paper cites an unresolved cited work.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:08:57.448793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.697643Z digest=sha256:68b642052c787a75832fbdd4948fe510126beb949643e8e6fc4e492c2dde2476

Observation bf6b27bb-818f-48c2-8440-a268e68fd01a · outbound

This paper cites Self- supervised multi-modal training from uncurated images and reports enables monitoring AI in radiology,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Self- supervised multi-modal training from uncurated images and reports enables monitoring AI in radiology,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.394904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.701654Z digest=sha256:730f06ebe4612de6c783f3a7f9bf2ac49c1f39215647a2aa5e4d6c17e40321f8

Observation 738fdb9f-5c11-4efc-add8-aa5359ecb97a · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.704689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.704689Z digest=sha256:261c9ed6845ded0cb0a65bb84d25c564e758d992a12efc0eeefb7c8b504add7f

Observation 84343929-6fe0-4463-9fd7-3f7453690286 · outbound

This paper cites A generalist vision–language foundation model for diverse biomedical tasks,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs A generalist vision–language foundation model for diverse biomedical tasks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.381097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.780569Z digest=sha256:7a91ee01e53cbb8eb02fc80ee739b53bcdbb6f1bcaa5c3c8ccc5c7373690c343

Observation f78290e5-db76-4106-868b-4ddcc2383e62 · outbound

This paper cites UltraMedical: Building specialized generalists in biomedicine,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs UltraMedical: Building specialized generalists in biomedicine,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.370481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.804489Z digest=sha256:3f0dcf956148743f901b07b4c54c20822ff80a3ff6f17b7bd700e17b5dfaab77

Observation 9b0666fb-8b0f-4968-b582-caaafe41df4c · outbound

This paper cites Meds3: Towards medical small language models with self-evolved slow think- ing,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Meds3: Towards medical small language models with self-evolved slow think- ing,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.808418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.808418Z digest=sha256:8e7d5d4773fdd2efec31cbc19c23d8984a120778751d5c6f4198c6bccfed5659

Observation f7e4be9b-e182-44f4-ae70-0f754d3eb3d5 · outbound

This paper cites QoQ-Med: Building mul- timodal clinical foundation models with domain-aware GRPO training,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs QoQ-Med: Building mul- timodal clinical foundation models with domain-aware GRPO training,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.811651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.811651Z digest=sha256:a4719714c085c7703c6ac90995804f4980d82831865a51531e54836881f6df6c

Observation 8f69da79-126c-4cd3-8f98-1e119ee4a05e · outbound

This paper cites Med-R1: Reinforcement learning for generalizable medical reasoning in vision-language models,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Med-R1: Reinforcement learning for generalizable medical reasoning in vision-language models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.815623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.815623Z digest=sha256:feb7fea5595a6ad8941ae8133b4852a1ffc6c02570ae5ec6155d651647ee33db

Observation 77892960-1b0e-4769-822a-20f873b956b8 · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.819290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.819290Z digest=sha256:29c4c7b64785c89abb0161da37cee1f9de40419e9a0f90c71190253c826bed35

Observation a9b60ebc-cb9b-46ca-ad37-6b703ba5b9ba · outbound

This paper cites m1: Unleash the potential of test-time scaling for medical reasoning with large language models,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs m1: Unleash the potential of test-time scaling for medical reasoning with large language models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.822853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.822853Z digest=sha256:c9c6837c6f3919fd9bb50272cf880671881d989728c17d934ce0daf3113ff6f6

Observation 46bd650c-b874-468a-ad66-edb98301abcd · outbound

This paper cites O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.826144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.826144Z digest=sha256:a4becd01f25d97244620dc51af652c3463a119e1ac97df20b27286cfbf119368

Observation 3efa8456-6b54-4f43-a834-6df34c35e6b3 · outbound

This paper cites Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.845303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.845303Z digest=sha256:5c801d4de66017231bdfbc0cba067a8abdba4885aa78c1d291ab6f7dc4d8043a

Observation 592a6bd3-963f-47f0-a361-7739fe31f85f · outbound

This paper cites HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.923492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.923492Z digest=sha256:3b2d24984a66b39589c7a27ac370ded7fd982ec811f50c5143a7198fb0337558

Observation 9c0eec9f-606e-4fd9-b6f3-66c7b76437cb · outbound

This paper cites MedGemma: Advanced AI models for medical text and image analysis,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedGemma: Advanced AI models for medical text and image analysis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.331628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.950864Z digest=sha256:b8fe75f780658f814c58c89262347d725da429970fa4ec7023b37acc7947a53f

Observation 8e59b2a0-97c4-4c50-970e-b14c02ad2226 · outbound

This paper cites Qwen2.5-VL Technical Report.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Qwen2.5-VL Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.954295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.954295Z digest=sha256:ddf826eeee34f1b50442a4aa15b2b3bfb8b88340d9c96b7afba16cf2acbae66b

Observation e2a23158-576a-44c3-a1d5-294050a1fbef · outbound

This paper cites Gemma 3 Technical Report.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Gemma 3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.958124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.958124Z digest=sha256:29c7aed19bad8f3d2fd36eed33026e45c32948d3c94a59eff028c3b025b905d0

Observation 1a7385db-4402-4c20-bf65-267dc2d697df · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.962530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.962530Z digest=sha256:2488275447b67cbd76243b8ca28e46d6cb341796b040331760628726998475f7

Observation 754dd459-be40-47fe-961d-457c58b60ac2 · outbound

This paper cites QVQ: To see the world with wisdom,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs QVQ: To see the world with wisdom,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.260355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:55.988044Z digest=sha256:98085366651595bff411f4ecf5135ae8d09351d7195768c1fba8ee7a5a481b1b

Observation 4f5cf823-c308-44ef-a620-093d8701be0c · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Yi: Open Foundation Models by 01.AI

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.037648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.037648Z digest=sha256:744e814d3b601baa5c6694b7cd4cc385397e2dc10fbe0865d695fad6a82a3eda

Observation 6bccdcd1-5787-45f5-802f-6df2a64f15f2 · outbound

This paper cites CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.042381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.042381Z digest=sha256:027a9c30f9c2a955b5ef92239036913d8a24b8a34a641d4186f4a7052960698a

Observation 3d02aa01-2c35-4776-abf1-9888f4d5e465 · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.046556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.046556Z digest=sha256:175452268ea6186b90f36d5d7525e53fd61afcea3c9ad93f6b0d3ae9a65cf251

Observation 7075c5e6-5d92-4af9-8c3a-c80fbf4839fd · outbound

This paper cites What disease does this patient have? A large-scale open domain question answering dataset from medical exams,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs What disease does this patient have? A large-scale open domain question answering dataset from medical exams,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.247426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:56.050738Z digest=sha256:849f09a11045487274c7be8fdda0429ddc4401afd956c611e43368591f02c818

Observation 6a6c7447-317c-477e-a5ff-253fc2ffa83c · outbound

This paper cites Benchmarking large language models on answering and explaining challenging medical questions,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Benchmarking large language models on answering and explaining challenging medical questions,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.167902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:56.055975Z digest=sha256:d4a67d0db27c8a0ff5e9944e7480b5a691ceb1d0d0533e100815b576101fb83c

Observation 79acf541-ad73-46d9-9e9d-714dce4bdc6e · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.122012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.122012Z digest=sha256:dae629dc1f8626af59ab0b049b79bc432f3d19a79813f704398977f5e5119999

Observation 3080d3ec-677f-48eb-839e-1768a50be057 · outbound

This paper cites MedCalc- Bench: Evaluating large language models for medical calculations,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedCalc- Bench: Evaluating large language models for medical calculations,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.157554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:56.152954Z digest=sha256:c388a566f02c439dc12d1f9d22a973b73a9e4d8bfb9ec9590ce3a2ba91f044cc

Observation d43c90ba-8170-4cd3-917a-5021630d60b7 · outbound

This paper cites Susceptibility of Large Language Models to User-Driven Factors in Medical Queries.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Susceptibility of Large Language Models to User-Driven Factors in Medical Queries

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.158706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.158706Z digest=sha256:16041d1afe2de4eaaa5041c304dc2f37006e503747eac27a8643401c558c1f64

Observation b3b38278-e171-4a35-a6a1-0a04189889f8 · outbound

This paper cites Om- niMedVQA: A new large-scale comprehensive evaluation benchmark for medical LVLM,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Om- niMedVQA: A new large-scale comprehensive evaluation benchmark for medical LVLM,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:57.114343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:08:56.163445Z digest=sha256:c150a6999f468411ff2c1e86edff92185fb9a2e11ae62ac1dc37bc2585b0a5b5

Observation e44fe3ff-952c-4e40-b0af-a8b0bbf8a533 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Training language models to follow instructions with human feedback,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.168439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.168439Z digest=sha256:2e5d84c3f9e262ed1ac32057aa2f39dcc1c1c52258ccf73fa7d2aacac986bc46

Observation 6fb26c6c-8d62-420a-8b19-f03bc8da0079 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs KTO: Model Alignment as Prospect Theoretic Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.172268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.172268Z digest=sha256:8908c0826b0c5f9f028623578f68eccad5ddef1bc9c130acbf0aca7466b1e8b2

Observation f1de9616-1e19-45c5-b74f-7bf75913fd1f · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Reasoning Models Don't Always Say What They Think

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.176954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.176954Z digest=sha256:74b23219f6b781c99e11b63abd7ca8071e705be82cdf1207fc0ba8ea3f186f37

Observation fa72c826-4858-41b7-8118-e88fec1060cb · outbound

This paper cites Disentangling Reasoning and Knowledge in Medical Large Language Models.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Disentangling Reasoning and Knowledge in Medical Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.179861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.179861Z digest=sha256:c0d043c81cb567c77cbafff9dd7e9e7e5d9d90e84ebb4fb2e63eef97385f97ab

Pith citing papers

Observation 680b9b8d-f675-48e3-b296-a157c95fb431 · inbound

Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight cites this paper.

Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:23:23.799795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T20:21:40.867354Z digest=sha256:2453c5c496ffd4759307c292130a4cc6bb76cee6269e557b644a5be56bb8253f