Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:08:56.179861Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2506.13102.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:08:56.179861Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-16T20:21:40.867354Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T20:23:23.797407Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7b0a2c50-eedf-411b-8430-ed1bac389c9e · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Language mod- els are few-shot learners,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4753e2e6-d135-4a78-b5b1-c63cd0609a0a · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c616e0e-d59c-4497-9610-82db37a225d7 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e4b2c4d-7cf0-4a35-9bb2-7a1761e7923a · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06756117-ce0f-416b-9dfd-21e71c5bbe85 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs GPT-4o System Card
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a593b253-553e-4392-b8e4-642814dda7ea · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs OpenAI o1 System Card
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe4fede-64ea-4740-88f7-a8cf9a20e6e5 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs The Llama 3 herd of models,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c61b68e7-ddc5-4bdb-aaaf-13692b1cc8cc · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 064a8092-116e-445a-bfdf-6a7965ba7994 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Qwen2.5 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e259a8-4b9a-4b80-b93a-ca1e3fc33bf6 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Flamingo: A visual language model for few-shot learning,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 05a1f538-d81f-474c-b8a2-56c9f50484e9 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 530c4658-ba4c-41af-bc8e-fce5cd327e10 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b95d9010-9af5-43c6-b9e6-b1faae1f807c · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Visual instruction tuning,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e385f19d-5186-4fbc-9e23-8ac8856a921b · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22826580-31b0-4246-a375-d883e59e8234 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bff8de91-a652-4f31-ac7c-905595c66010 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs xGen-MM (BLIP- 3): A family of open large multimodal models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da645a20-c07b-4bfc-8d86-5117e16b7cd0 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs s1: Simple test-time scaling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec0a4d29-8a6e-4bb1-9bea-0771b8316d95 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4e9df956-644d-4631-9f00-cffa58cbc8c8 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Towards thinking-optimal scaling of test-time compute for LLM reasoning,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bab2202-125b-4014-97e0-44bc5e33e238 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44cf1a0-81c2-4168-aefa-483f175b4bb3 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Proximal Policy Optimization Algorithms
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da8faa2-e808-44a4-beda-1698e698f800 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Direct preference optimization: Your language model is secretly a reward model,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76204dec-a21e-4f2b-b428-85aa95f56659 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b1ecc4-a5eb-474e-a1a8-c6eae3be5a34 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Multi- modal understanding and generation for medical images and text via vision-language pre-training,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 28243b47-b668-4044-bca5-9156feb687cb · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bf6b27bb-818f-48c2-8440-a268e68fd01a · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Self- supervised multi-modal training from uncurated images and reports enables monitoring AI in radiology,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 738fdb9f-5c11-4efc-add8-aa5359ecb97a · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84343929-6fe0-4463-9fd7-3f7453690286 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs A generalist vision–language foundation model for diverse biomedical tasks,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f78290e5-db76-4106-868b-4ddcc2383e62 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs UltraMedical: Building specialized generalists in biomedicine,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9b0666fb-8b0f-4968-b582-caaafe41df4c · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Meds3: Towards medical small language models with self-evolved slow think- ing,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e4be9b-e182-44f4-ae70-0f754d3eb3d5 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs QoQ-Med: Building mul- timodal clinical foundation models with domain-aware GRPO training,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f69da79-126c-4cd3-8f98-1e119ee4a05e · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Med-R1: Reinforcement learning for generalizable medical reasoning in vision-language models,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77892960-1b0e-4769-822a-20f873b956b8 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b60ebc-cb9b-46ca-ad37-6b703ba5b9ba · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs m1: Unleash the potential of test-time scaling for medical reasoning with large language models,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46bd650c-b874-468a-ad66-edb98301abcd · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3efa8456-6b54-4f43-a834-6df34c35e6b3 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 592a6bd3-963f-47f0-a361-7739fe31f85f · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c0eec9f-606e-4fd9-b6f3-66c7b76437cb · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedGemma: Advanced AI models for medical text and image analysis,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e59b2a0-97c4-4c50-970e-b14c02ad2226 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Qwen2.5-VL Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a23158-576a-44c3-a1d5-294050a1fbef · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Gemma 3 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7385db-4402-4c20-bf65-267dc2d697df · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 754dd459-be40-47fe-961d-457c58b60ac2 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs QVQ: To see the world with wisdom,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4f5cf823-c308-44ef-a620-093d8701be0c · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Yi: Open Foundation Models by 01.AI
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bccdcd1-5787-45f5-802f-6df2a64f15f2 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d02aa01-2c35-4776-abf1-9888f4d5e465 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs PubMedQA: A Dataset for Biomedical Research Question Answering
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7075c5e6-5d92-4af9-8c3a-c80fbf4839fd · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs What disease does this patient have? A large-scale open domain question answering dataset from medical exams,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6a6c7447-317c-477e-a5ff-253fc2ffa83c · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Benchmarking large language models on answering and explaining challenging medical questions,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 79acf541-ad73-46d9-9e9d-714dce4bdc6e · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3080d3ec-677f-48eb-839e-1768a50be057 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedCalc- Bench: Evaluating large language models for medical calculations,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d43c90ba-8170-4cd3-917a-5021630d60b7 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Susceptibility of Large Language Models to User-Driven Factors in Medical Queries
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b38278-e171-4a35-a6a1-0a04189889f8 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Om- niMedVQA: A new large-scale comprehensive evaluation benchmark for medical LVLM,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e44fe3ff-952c-4e40-b0af-a8b0bbf8a533 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Training language models to follow instructions with human feedback,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb26c6c-8d62-420a-8b19-f03bc8da0079 · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs KTO: Model Alignment as Prospect Theoretic Optimization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1de9616-1e19-45c5-b74f-7bf75913fd1f · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Reasoning Models Don't Always Say What They Think
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa72c826-4858-41b7-8118-e88fec1060cb · outbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs Disentangling Reasoning and Knowledge in Medical Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 680b9b8d-f675-48e3-b296-a157c95fb431 · inbound
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.