Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:22:46.580769Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2510.15859.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:22:46.580769Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T23:18:59.283834Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-29T23:24:01.706149Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a34295d0-34fb-43cc-be62-9b00871745fe · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d06983-6196-491a-a143-7cccab947567 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Language models that think, chat better.arXiv preprint arXiv:2509.20357,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5276c8a0-776e-41c3-b934-af35ee4e937e · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Ace-rl: Adaptive constraint-enhanced reward for long-form generation re- inforcement learning.arXiv preprint arXiv:2509.04903,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63376715-0001-4f25-9e4d-3cea55bbc407 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a388ea47-140f-42eb-bebb-9b2e25ba3a5f · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71bb4fec-3170-44f5-859f-ac3bc0284b3a · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b940e9-440d-40ff-9dc9-27b582546b88 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd3fcd13-dbfe-47c0-a4ad-eea3a1adb75f · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Multichallenge: A realistic multi-turn con- versation evaluation benchmark challenging to frontier llms
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49ab613-422c-446f-ae97-617e899e3be1 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Qa-lign: Aligning llms through constitution- ally decomposed qa.arXiv preprint arXiv:2506.08123,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294d778f-1817-420d-8322-4c8adf78cb60 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Baichuan-M2: Scaling Medical Capability with Large Verifier System
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b9a9b21-0407-4aae-9415-86805e856077 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5406e69-7612-4621-ac78-c534ba31f751 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training m1: Unleash the potential of test-time scaling for medical reasoning with large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793ac8f9-4666-4d31-bcaf-393bc9335441 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 587ece49-1ab1-4654-ac28-0f2667bb4dd4 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Reward Design with Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f751fc81-a66f-4fa7-8f16-a766a6380e8e · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca7aa40-b80c-4670-830c-3d39af3f28e3 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4744b8c1-bbb1-4855-a935-21941b67f116 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6b3a35-f140-4610-ab3a-9d208e1706cc · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Large Language Models: A Survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af21bce-07ba-4364-b4ea-b1fbbe0452e2 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training InFoBench: Evaluating Instruction Following Ability in Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d1f916-1045-44e5-8421-baec885ace3b · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2ac91dd-c48d-493a-82c6-34ae2609fbbf · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training RL's Razor: Why Online Reinforcement Learning Forgets Less
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fb2987-9169-495a-9dff-6ae253a63d3b · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb2586a-26f0-400b-9932-9c66dff9cdbf · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Reasonmed: A 370k multi-agent gen- erated dataset for advancing medical reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa0317e5-b0e1-45ca-ace3-98938490fd10 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Medagents: Large language models as collab- orators for zero-shot medical reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80bf72b1-c7a5-4bf5-98d8-26b69633403a · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Large language models in medicine
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741df60c-8317-4065-983d-3d5ed13632cc · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 678fc548-cf3e-4a30-9635-aa752f0a62c4 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Check- lists are better than reward models for aligning language models.arXiv preprint arXiv:2507.18624,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1870351-44ba-4f14-a312-fa747656f6df · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 802c15a0-7022-4caa-a725-9cef68be0c0c · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Qwen3 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26268332-ed02-44b4-a5b1-ce3c2620eeac · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 836e7d2c-8598-4d25-af18-27d1615279d0 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a10f86-11cb-404d-a60b-1361d6a94a3e · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Group Sequence Policy Optimization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc648d4f-9614-4142-a76f-279483902ed3 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Ask patients with patience: Enabling llms for human-centric medical dialogue with grounded reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2802811a-1746-40bb-a8e7-eff22721f8c9 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e897e7e-7f88-4ebb-a878-fa4b0e154b2e · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training This subplot measures the proportion of rubrics that are satisfied (or penalties avoided) at least once within 40 rollouts for each query
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e83c4e79-9d64-455c-945c-dd0649c27e0b · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training For all other models, the generation parameters were set to align with those specified in the official HealthBench protocol
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4c0393d-1008-4caa-8887-5e2d3577686c · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training A generalist medical language model for disease diagnosis assistance.Nature medicine, 31(3):932–942, 2025e
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 569f998a-5475-4431-8715-ae8ea0b8a4b4 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training gpt-oss-120b & gpt-oss-20b Model Card
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283ec89e-9879-4524-8c41-9972e0ec7745 · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Real-World Doctor Agent with Proactive Consultation through Multi-Agent Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 453fc6cf-e046-4366-a789-5159e2506e8c · outbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training HealthBench: Evaluating Large Language Models Towards Improved Human Health
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372bc0a3-133d-40e8-a4f9-58acde27eaf9 · inbound
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.