Pith. sign in

Paper Citation Record · LEDGER

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training

As of 23 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2510.15859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.15859 v5

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:22:46.580769Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T23:18:59.283834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T23:24:01.706149Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a34295d0-34fb-43cc-be62-9b00871745fe · outbound

This paper cites GPT-4 Technical Report.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:40.110994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:40.110994Z digest=sha256:7d4fd887eed0dbb8aa61354fe340acd86d1cca59a8aebda1c8e1ed90fe15efc2

Observation 51d06983-6196-491a-a143-7cccab947567 · outbound

This paper cites Language models that think, chat better.arXiv preprint arXiv:2509.20357,.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Language models that think, chat better.arXiv preprint arXiv:2509.20357,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:40.930592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:40.930592Z digest=sha256:29d084d3dd605bf5fe118d6f2e6a6c5c93e30ffd4107e33616225761bcf4b78e

Observation 5276c8a0-776e-41c3-b934-af35ee4e937e · outbound

This paper cites Ace-rl: Adaptive constraint-enhanced reward for long-form generation re- inforcement learning.arXiv preprint arXiv:2509.04903,.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Ace-rl: Adaptive constraint-enhanced reward for long-form generation re- inforcement learning.arXiv preprint arXiv:2509.04903,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:41.760730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:41.760730Z digest=sha256:06dda60efe48ea9867a409e4e719cdad463c0bc54a9371b26015fb39118a5b98

Observation 63376715-0001-4f25-9e4d-3cea55bbc407 · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:41.903752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:41.903752Z digest=sha256:31af03db7f4ea726d380fd7db53fd8e63b53a263f782070f2f65a6e23b04673a

Observation a388ea47-140f-42eb-bebb-9b2e25ba3a5f · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.012260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.012260Z digest=sha256:6581333593e756c9f17a38ef84fade7d29b5cf00de21da7267637c5bdde80d64

Observation 71bb4fec-3170-44f5-859f-ac3bc0284b3a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.136987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.136987Z digest=sha256:ddaf814e3c16604703efc330dd31752486b200719ad8df21dcdbb672f97e1d19

Observation 60b940e9-440d-40ff-9dc9-27b582546b88 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.290521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.290521Z digest=sha256:07627db9da95e529ff9bef070573046b4331c3b3c430158a483dce2c4214ee9e

Observation dd3fcd13-dbfe-47c0-a4ad-eea3a1adb75f · outbound

This paper cites Multichallenge: A realistic multi-turn con- versation evaluation benchmark challenging to frontier llms.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Multichallenge: A realistic multi-turn con- versation evaluation benchmark challenging to frontier llms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.370031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.370031Z digest=sha256:2620bcdd06786f5a44fcaa17446692910ec0ecc43b521da6cd1c5e1c2f74615c

Observation c49ab613-422c-446f-ae97-617e899e3be1 · outbound

This paper cites Qa-lign: Aligning llms through constitution- ally decomposed qa.arXiv preprint arXiv:2506.08123,.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Qa-lign: Aligning llms through constitution- ally decomposed qa.arXiv preprint arXiv:2506.08123,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.457913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.457913Z digest=sha256:2bfb4aa85d2ff58c25f021b59506a7ddd923734cb3d93f637dc074ba4e9155b1

Observation 294d778f-1817-420d-8322-4c8adf78cb60 · outbound

This paper cites Baichuan-M2: Scaling Medical Capability with Large Verifier System.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Baichuan-M2: Scaling Medical Capability with Large Verifier System

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.521571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.521571Z digest=sha256:42a3fa6e9b3785400f86e293f16e3c16c7a467bf0e367068583b678ef9d9d965

Observation 2b9a9b21-0407-4aae-9415-86805e856077 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.681512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.681512Z digest=sha256:c553f1f694b585d7bb98702bf8fb8514331fc4539d9b0bd3341e341272eb93c6

Observation a5406e69-7612-4621-ac78-c534ba31f751 · outbound

This paper cites m1: Unleash the potential of test-time scaling for medical reasoning with large language models.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training m1: Unleash the potential of test-time scaling for medical reasoning with large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.737007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.737007Z digest=sha256:6b1d386b2452461e3378c516f8b90122eed5269f4cd13f72326ea9cffd768c52

Observation 793ac8f9-4666-4d31-bcaf-393bc9335441 · outbound

This paper cites Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.829865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.829865Z digest=sha256:2479eb0f6d8bbe4181d2bacb72e4cbf6a805ed7afb7d2a57dbf8e2a309b25efe

Observation 587ece49-1ab1-4654-ac28-0f2667bb4dd4 · outbound

This paper cites Reward Design with Language Models.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Reward Design with Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.928860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.928860Z digest=sha256:333c5544a0340939b71631ff7d8b43784d60982cc9d08b1f8432ca2ead9a620f

Observation f751fc81-a66f-4fa7-8f16-a766a6380e8e · outbound

This paper cites Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.993723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.993723Z digest=sha256:8cd8dc4f2b342c199a7a468b20b375f920f7e58bbea04f1a08cee98d6261b5a3

Observation fca7aa40-b80c-4670-830c-3d39af3f28e3 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.139966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.139966Z digest=sha256:223dc4d50213b3cdc836f494ecc23354d4089ecee0ff6319ba4b1d1a1e8cfacf

Observation 4744b8c1-bbb1-4855-a935-21941b67f116 · outbound

This paper cites Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.249504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.249504Z digest=sha256:a07016ff31a28c86e14c9dc2d82258581fdebbdcf9cb0a0be50cb85ad66cd846

Observation ff6b3a35-f140-4610-ab3a-9d208e1706cc · outbound

This paper cites Large Language Models: A Survey.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Large Language Models: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.509662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.509662Z digest=sha256:c3eb33749fc492dfd6fbdfd34617f6d891c80f24e7d844c50b4574f3e37094c0

Observation 1af21bce-07ba-4364-b4ea-b1fbbe0452e2 · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.631107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.631107Z digest=sha256:c3d689dbfc733f9b8197a114f7cd626207c2a8b1f0892f2df5ed138c154e7779

Observation 38d1f916-1045-44e5-8421-baec885ace3b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.796312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.796312Z digest=sha256:6ba01e7af8386a20140e76a5011bfd2d0e3c9dfe8750e6e689df52cc806b456d

Observation b2ac91dd-c48d-493a-82c6-34ae2609fbbf · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.958435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.958435Z digest=sha256:fec616a8a787d808fe84ff25898046fa481a82b589121f248f3c28a0d172a518

Observation 96fb2987-9169-495a-9dff-6ae253a63d3b · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:44.150906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:44.150906Z digest=sha256:f261d077327db9a259e202ac9fb7468cdc2b7dc6ec4e88a912bfc94738b2eed6

Observation 2eb2586a-26f0-400b-9932-9c66dff9cdbf · outbound

This paper cites Reasonmed: A 370k multi-agent gen- erated dataset for advancing medical reasoning.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Reasonmed: A 370k multi-agent gen- erated dataset for advancing medical reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:44.258716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:44.258716Z digest=sha256:ccb4628bd668a4c8571f03ba4669f340fe45697827233ebb675d2cd4bcc68458

Observation aa0317e5-b0e1-45ca-ace3-98938490fd10 · outbound

This paper cites Medagents: Large language models as collab- orators for zero-shot medical reasoning.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Medagents: Large language models as collab- orators for zero-shot medical reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:44.414968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:44.414968Z digest=sha256:3852ae788e40a41c3b37cf079e323f829267c0116e18b859b1573dbd5f7b77fb

Observation 80bf72b1-c7a5-4bf5-98d8-26b69633403a · outbound

This paper cites Large language models in medicine.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Large language models in medicine

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:44.556266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:44.556266Z digest=sha256:ee610e7e3da0721f06b4fd467803e07b038a2ff1883595527f5026ea78b0e34d

Observation 741df60c-8317-4065-983d-3d5ed13632cc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:44.781689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:44.781689Z digest=sha256:b3df6fb1850b09f3d19510660068242644eb69b6ec1969013b8a233685edc076

Observation 678fc548-cf3e-4a30-9635-aa752f0a62c4 · outbound

This paper cites Check- lists are better than reward models for aligning language models.arXiv preprint arXiv:2507.18624,.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Check- lists are better than reward models for aligning language models.arXiv preprint arXiv:2507.18624,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:44.929913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:44.929913Z digest=sha256:09727e8d3935a39379bfb7eaaae78a2c85408356e1fa72b76980143f5181cfeb

Observation a1870351-44ba-4f14-a312-fa747656f6df · outbound

This paper cites MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:45.109601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:45.109601Z digest=sha256:149c99645d436a87047e4deb8b3b7739d7ddbb1b1c2866371b9822c5d133580f

Observation 802c15a0-7022-4caa-a725-9cef68be0c0c · outbound

This paper cites Qwen3 Technical Report.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Qwen3 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:45.350657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:45.350657Z digest=sha256:c24df1f090e7dd0e23567740e70887f52d103220cf85341cfafdd27d2fb4e632

Observation 26268332-ed02-44b4-a5b1-ce3c2620eeac · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:45.579042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:45.579042Z digest=sha256:df7eef8509837ffab4b2ee529564694db133bc59d18f1d061f6353c21e88e0cc

Observation 836e7d2c-8598-4d25-af18-27d1615279d0 · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:45.779908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:45.779908Z digest=sha256:5cb6353e627875b83c922dc7e18a714706e6a6576f9d04f6c3a4b3f2b90c80cb

Observation 20a10f86-11cb-404d-a60b-1361d6a94a3e · outbound

This paper cites Group Sequence Policy Optimization.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Group Sequence Policy Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:45.967654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:45.967654Z digest=sha256:cee99929351ee15aeab6188430ac5057486d93fa2f43669e2d47b28b395908aa

Observation fc648d4f-9614-4142-a76f-279483902ed3 · outbound

This paper cites Ask patients with patience: Enabling llms for human-centric medical dialogue with grounded reasoning.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Ask patients with patience: Enabling llms for human-centric medical dialogue with grounded reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:46.126866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:46.126866Z digest=sha256:544de8d63388d142545d2a4f7721e87bedb10a97aa9cd19ecec733cffec60764

Observation 2802811a-1746-40bb-a8e7-eff22721f8c9 · outbound

This paper cites an unresolved cited work.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Unresolved cited work

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-04T09:22:46.265245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:46.265245Z digest=sha256:a1341dc59cd3733b02e918918c61669851a00efed9566f951faad6d9c1601551

Observation 2e897e7e-7f88-4ebb-a878-fa4b0e154b2e · outbound

This paper cites This subplot measures the proportion of rubrics that are satisfied (or penalties avoided) at least once within 40 rollouts for each query.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training This subplot measures the proportion of rubrics that are satisfied (or penalties avoided) at least once within 40 rollouts for each query

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:46.416690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:46.416690Z digest=sha256:b6b8e59f690d9c7cf3dbec26cb542842d8b772b663d33ff6908a1b877d359b1a

Observation e83c4e79-9d64-455c-945c-dd0649c27e0b · outbound

This paper cites For all other models, the generation parameters were set to align with those specified in the official HealthBench protocol.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training For all other models, the generation parameters were set to align with those specified in the official HealthBench protocol

Reference 40

Resolution
malformed identifier
no resolver link, observed 2026-08-04T09:22:46.580769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:46.580769Z digest=sha256:73b56954127870663d48199bf48f13a6da1fbe242a27f73510153f4a9d05f75c

Observation f4c0393d-1008-4caa-8887-5e2d3577686c · outbound

This paper cites A generalist medical language model for disease diagnosis assistance.Nature medicine, 31(3):932–942, 2025e.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training A generalist medical language model for disease diagnosis assistance.Nature medicine, 31(3):932–942, 2025e

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.377717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.377717Z digest=sha256:6ff83b514025dfa87d40e97553f10ea820558636462da28933aabcbee22a16b7

Observation 569f998a-5475-4431-8715-ae8ea0b8a4b4 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training gpt-oss-120b & gpt-oss-20b Model Card

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:40.275971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:40.275971Z digest=sha256:b2d0e681e04740747f0c854be87c52dc34d344f2e9b5bb518a922e81211ab7cd

Observation 283ec89e-9879-4524-8c41-9972e0ec7745 · outbound

This paper cites Real-World Doctor Agent with Proactive Consultation through Multi-Agent Reinforcement Learning.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Real-World Doctor Agent with Proactive Consultation through Multi-Agent Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.597472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.597472Z digest=sha256:8ba44632476e5f7df25ec82c16312aba0c911b26e442dc5f711abbc0de05aa7b

Observation 453fc6cf-e046-4366-a789-5159e2506e8c · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:40.462108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:40.462108Z digest=sha256:781fc3a70a4a7c2f92124e16807eb62cb8a4efe9dc642e3e097fdec1b2c47a27

Pith citing papers

Observation 372bc0a3-133d-40e8-a4f9-58acde27eaf9 · inbound

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment cites this paper.

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training

Reference 152

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:24:01.707495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T23:18:59.283834Z digest=sha256:3449e30bebffcbbef0ae8920d6e229262aac4ae0e19a6692b3a25c058dde7adf