Pith. sign in

Paper Citation Record · LEDGER

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon

As of 4 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2606.30445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.30445 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:41:39.266071Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact25
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e6ecc6ef-d751-4055-a081-04adab472a1c · outbound

This paper cites and Ng, A.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon and Ng, A

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:8d7f9084ba41bb32308959a7194ef6e5302013550e4e21d4a4ae5402de50309a

Observation 4197188a-3029-41ef-b840-94aafab0c77d · outbound

This paper cites R., Geist, M., and Bachem, O.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon R., Geist, M., and Bachem, O

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:a73cf9f11ef213bb995405924c915517b83a19a48c35fd87463aeb614a271a4b

Observation 075fac4a-e750-443b-9f34-2f1fac85b87f · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:cc2190491baf12ac757d8386eeebf9c59443ea6617ee45bf7e903446fb049f2c

Observation 65c3ea4b-bb4f-4ec1-bd36-2e919eebf1f5 · outbound

This paper cites Mitigating covariate shift in imitation learning via offline data with partial coverage.Advances in Neural Information Processing Systems, 34: 965–979, 2021.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Mitigating covariate shift in imitation learning via offline data with partial coverage.Advances in Neural Information Processing Systems, 34: 965–979, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:a17b9bb057a0e0b569c5057e8e4d49d0e2c11044dd3dc9050bdac3c521dc8c4c

Observation b32e285a-b95c-423e-862c-f98e37c62af1 · outbound

This paper cites Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.680586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:ecaf7622f4eabb4ccdc7ddaa4ede929635717e6e89e91b64943977c84134febf

Observation 430c45ed-53aa-45f7-9ee2-1b69625823ad · outbound

This paper cites and Jiang, N.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon and Jiang, N

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:87e03ff576167939a7a2070dd8ef6fc01b58c04dbdc5732e93e6d6208b949a38

Observation 49b4608b-25d6-4b4c-90d5-f4e347799661 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:44:21.715939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:23bc1ad885474049bf3b77e3051745d71f74e97dcdada2b2b504136af3c89a87

Observation a2cb1f82-808c-4719-bfc7-532e68c6104b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.694221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:811726661663951b0a6ed2981ad85cafa7b2d4c6e2d3c4888171d6ea6c21f710

Observation 0d477e39-78e9-4623-83d9-a27348479ae4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.689693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:a0a7a80ddebb87fb7fffa8a17dcd40a8ecbafd72ea20e8902bd58a995b1684ab

Observation 8d1f14fc-9eb1-4919-ae0a-bef84651709b · outbound

This paper cites Efficient Imitation under Misspecification.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Efficient Imitation under Misspecification

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:44:21.686151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:837f637a28b5cc4f210abaa729bc65cbcc24e7cf498ac805288e3cf359b5e480

Observation 4ad30175-9af4-4d58-9734-6591a09a9435 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, 2025.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Open r1: A fully open reproduction of deepseek-r1, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:78525bf468f412c7a99d9ae3815dfd68ec148c66ac04b5bfdb7d1b62bf785c9d

Observation e9f2345a-3bb7-4d2d-8768-c967fa4b94c9 · outbound

This paper cites and Rakhlin, A.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon and Rakhlin, A

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:321fea6a12d7dff7df2396cf365311852ce65f8ffeed1d1b4bbaa90e7901a0c6

Observation adbc819d-0c78-460a-aaee-f658c0746e10 · outbound

This paper cites Practical contextual bandits with regression oracles.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Practical contextual bandits with regression oracles

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:9d9e45ecf70e8f6c587a0679a2061b2bc3e1eade91412d2709122fc2137a7f39

Observation dcc258ac-dccc-4fc0-ba6c-3bdc115970d9 · outbound

This paper cites Foundations of Reinforcement Learning and Interactive Decision Making.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Foundations of Reinforcement Learning and Interactive Decision Making

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:21.684171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:051d164bcb2c7a3a5d423aad2d8f7018704af4ee41eace4290f967bef2941c4e

Observation e375f995-898d-4616-8ce7-980ff88982b7 · outbound

This paper cites J., Block, A., and Misra, D.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon J., Block, A., and Misra, D

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:16bb44acc8b34fba3b0f923e1ae9d62badc4755498fc7a102e6ebfd209a48b99

Observation d44558d4-589e-4d64-bb46-56032da01c9d · outbound

This paper cites Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:44:21.676876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:8bfd5ff47aed33afae54aebac83888ea3377030a06859cf9c306a18a6f926c81

Observation 514081ac-b878-47c2-963e-755dfc9be6a5 · outbound

This paper cites Importance-weighted offline learning done right.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Importance-weighted offline learning done right

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:28fbf0ba8f5e0bb7ef944ba6da80f36cda5cc93067ed92a772e3cc11181d64a5

Observation 329e5e25-8a11-4114-b9ac-57280cf47c76 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.660955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:2a502d56d24ceaa1b6702b9895433a73f1a017377437107875399e438c4e2474

Observation 97d1f188-33ff-4a64-82dc-e09b05ef17b1 · outbound

This paper cites The Llama 3 Herd of Models.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon The Llama 3 Herd of Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.691753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:4ae57f86122a2dad8c0d869c4b582c6f97041b5de449adb86a9f6924ad49860b

Observation b872c7ac-445d-4098-9522-e16329086b3d · outbound

This paper cites Minillm: Knowledge distillation of large language models.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Minillm: Knowledge distillation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:bd816e3d0d1d9e7760f7cbad911271b35b9c27cccd79079e62e87b2aa9f063a9

Observation e8833fb0-b182-47b4-9370-24a8a1e22b79 · outbound

This paper cites The false promise of imitating proprietary language models.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon The false promise of imitating proprietary language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:aba3297a3bc095c26c6f24b01315b0d38efba2e2f9e19763e4b7e35624b8cb47

Observation 2bf18e55-98bd-458e-98d3-e2e489b827eb · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Skywork Open Reasoner 1 Technical Report

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:44:21.668673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:a308f849217886196797c74fed9f006db2821e64a88e0edc4850030c5d59d2c3

Observation 9b5002de-bb16-47d0-8f82-447d22aed129 · outbound

This paper cites Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:5c4c923d07d5578557c226f38b6f6ac4ddee0d59ea3b30a51791b86fbc39ecff

Observation 1d988d1d-fc52-4f64-964b-808eac0b58de · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Measuring Mathematical Problem Solving With the MATH Dataset

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.720735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:df07b29ab28504c3892951d5e065f0e3fa8cdb1f37e50df61afae7e568a17804

Observation 26ce6a87-3ecd-41bf-893f-212e74ed801a · outbound

This paper cites J., Rohatgi, D., Zhang, C., Simchowitz, M., Ash, J.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon J., Rohatgi, D., Zhang, C., Simchowitz, M., Ash, J

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:81f3a0e739b0a232209bfd854bf785168d59737df09b1074d754a6e9f1cbdd8c

Observation d2fa29c5-d07f-47bc-a639-8f563056cb2f · outbound

This paper cites D., Sun, W., Krishnamurthy, A., and Foster, D.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon D., Sun, W., Krishnamurthy, A., and Foster, D

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:ce7d06aad2ead67763d4b18aac0ffa33346100aa9eebff031aacdedb2afc4035

Observation ddc7794d-b053-4f93-9638-4d82093666cd · outbound

This paper cites Teach small models to reason by curriculum distillation.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Teach small models to reason by curriculum distillation

Reference 27

Resolution
verified exact
doi, observed 2026-06-30T07:44:20.901838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:360600d90a52b910c65c40ce88ef2881b3196e8e14c75840aa9f675309cad493

Observation 6dc6b239-0651-4de7-8802-b22647781dee · outbound

This paper cites Coverage improvement and fast convergence of on-policy preference learning.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Coverage improvement and fast convergence of on-policy preference learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:21.718311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:3ce5a10ceed651b0ecd1851c3a1e450d73c210b9f96dba895a8281914d7a5e87

Observation 46f206eb-ebbb-4e66-97f4-5c7b02b9da37 · outbound

This paper cites and Rush, A.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon and Rush, A

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:9a837d4f07ad598b7e6af8567f4c9c7a8484d0ceaa9332fa70d9485767ce85f2

Observation 88b622f4-6d1f-4306-b0af-6e1b3356c625 · outbound

This paper cites M., Ma, T., and Liang, P.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon M., Ma, T., and Liang, P

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:2a1553b2e39827f3ab8a134e454e41f1d82827c7421d90992fb3c3ec773d5cdd

Observation 7b9c6d6a-b457-494f-83a3-274ad822b955 · outbound

This paper cites H., Gonzalez, J.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon H., Gonzalez, J

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:4db2a12094905adb72aacebb1e6f0683639137e73242e0efe5d15b65579343df

Observation 1daa683a-9730-441d-aa12-731c07693706 · outbound

This paper cites Small models struggle to learn from strong reasoners.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Small models struggle to learn from strong reasoners

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:44:21.725590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:77066f309e726a9b27b212360330dcc058b95a9e0d58b0f909fdc57d3aac5a68

Observation bfb4fcdf-7b70-486b-af71-298cbce4cab2 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.716774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:c8c2e094665655fd53a2c9da832d7631f0251d0cbbdd23de5195ba5f79656f88

Observation 8c68a5fc-1fa5-47c7-8da4-7aacddbfdaf6 · outbound

This paper cites DeepSeek-V3 Technical Report.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon DeepSeek-V3 Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.728180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:a09714aab3876b46a0715bd25b5870496f67214576b3f77a3d84b297fd54d485

Observation 071bb425-ae7c-47e7-a37b-65d0bc260cf5 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Con- nectionism.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon On-policy distillation.Thinking Machines Lab: Con- nectionism

Reference 35

Resolution
verified exact
doi, observed 2026-06-30T07:44:20.907739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:6c2342138efb7808104c94430cb28f94aa8778d5a3f546272b3dc06731bc4ecf

Observation 9f6dfbeb-c5e1-4b47-b02a-aefa074a40f8 · outbound

This paper cites Y., Roongta, M., Cai, C., Luo, J., Zhang, T., Li, L.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Y., Roongta, M., Cai, C., Luo, J., Zhang, T., Li, L

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:7a2b2b1893dec26a868c13ba0b4f42cfdffb654dd5be98631397b21da73cb862

Observation 535fd6ba-9f55-452c-8b69-ed3c15227e91 · outbound

This paper cites Error bounds for approximate policy iteration.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Error bounds for approximate policy iteration

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:d0ecace92c0e62c69ab15bdd7299dd6b9d852aab54023269778766495435ff0c

Observation f25117d1-1e01-4517-92f0-4b015651280f · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:d60936ba2d9a0f2e1bd9c3688c1107aad770214c106ba759d2953a5e64f8f8c5

Observation e207bce3-a813-4af9-aaa5-593dd7bebf7d · outbound

This paper cites Tinyzero.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Tinyzero

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:5523ce3ea3f44d95ba8cd8fd0448184ce091c69da1916b23d6982c5a7668d9d0

Observation 372fedce-f965-4373-9074-a4318bf0b4b0 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:8ad3e004c45df708c92131a77ed3efc0890a70a4aaf0b1a571ea99e6605a7331

Observation eea2c163-8c6f-43f2-9a5b-d50fa6a4e6ba · outbound

This paper cites D., Ermon, S., and Finn, C.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon D., Ermon, S., and Finn, C

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:be19d9327464f9c422e57fe82c6ed7c3cff8e1cf9dc5ecd68735147f71fd5989

Observation 01f04023-31a1-4d6c-8fd9-cdaa49f6e738 · outbound

This paper cites Toward the fundamental limits of imitation learning.Advances in Neural Information Processing Systems, 33:2914–2924, 2020.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Toward the fundamental limits of imitation learning.Advances in Neural Information Processing Systems, 33:2914–2924, 2020

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:b3ab809a8f6789a739f6c218421acc4185b5451c9203a38c56126018f45e996a

Observation 57687ae9-7134-4ddd-8b39-318901d0021d · outbound

This paper cites On the value of interaction and function approximation in imitation learning.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon On the value of interaction and function approximation in imitation learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:e09a2b130b0798f43288a832b37cc11bfa65b82305a6dffdeca9cc07eda642dd

Observation a7a9aeb8-f620-453b-a73a-2675f5a77010 · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.Advances in Neural Information Processing Systems, 34: 11702–11716, 2021.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Bridging offline reinforcement learning and imitation learning: A tale of pessimism.Advances in Neural Information Processing Systems, 34: 11702–11716, 2021

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:13670f1da6a8ddda69a053e907857e337a4d2a206b84ab881038bf243a17d136

Observation 2a212c74-7aa4-494e-88e0-d435fd51322d · outbound

This paper cites Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:21.722568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:c0d3386d4953af231a3eb991e7a3af5d9c8fd0db380636e4ae4ee0795646dd5c

Observation cbad912d-9437-40b0-b862-9a8d7517e18a · outbound

This paper cites and Bagnell, D.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon and Bagnell, D

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:0d256abaf8d26af0afefe87caea7ae63fa454f63e4ef441230db9a435777b504

Observation 3a4556c6-937b-4a99-b035-2d7504350ce5 · outbound

This paper cites Reinforcement and Imitation Learning via Interactive No-Regret Learning.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.694949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:038a22ed39801b01d5751eaf53ef5fb3c7396f3b057c20c354c808807f475b04

Observation 6e02cead-b9f5-404f-9506-342a0692c9f2 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon A reduction of imitation learning and structured prediction to no-regret online learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:8315afcd6af5d32c4372e9a0a02eff5788cebef1323904514b69e95620cc3fea

Observation 0d4c68e9-8568-45fe-b51e-f274d20b9fb7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:44:21.704969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:8f59f42d3187a46fb6bd28cd5ef94744be081ce2d492778022453b86b4e2d791

Observation 91b6f3bb-e78c-4e41-8d34-2bd1d8aefff8 · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.697599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:f7e5aaa46bc6ee8b9bf79f7861342c1366d7d93613b79d1be91451d32320e59d

Observation 90af2d49-d326-460b-a8e5-3588fe7ae1fd · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon HybridFlow: A Flexible and Efficient RLHF Framework

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.731045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:f44ebfe10988e3e674923efd7eb4da1958d9e47c78804f3785bbe4b3fa5e1315

Observation bc42ebcf-c246-43f2-91fa-658f0a705367 · outbound

This paper cites and Xu, Y.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon and Xu, Y

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:d392115bcf91c91e51852b827c3621d67cdf21a9fe17a855321e18f09ff06712

Observation 500fff68-4926-46d5-afca-510c176a127d · outbound

This paper cites an unresolved cited work.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Unresolved cited work

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:21.736812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:a9f9f81c21ad42a5ebe1cc6e9088d15328b63e8eb95a3c7431523194a1cdef51

Observation dac5e4a0-5f53-42b8-b909-a0faaebf5e78 · outbound

This paper cites and Joachims, T.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon and Joachims, T

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:83af33c732468565ad8f36f8a509f619433b74aacb9f483b9e1f53fc331e73ff

Observation 539c8a90-60c8-45ce-a0f2-35e51b47d52f · outbound

This paper cites A., Wu, S., Jiao, J., and Ramchandran, K.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon A., Wu, S., Jiao, J., and Ramchandran, K

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:7757046b62b4aa3eade4dbbd150c13b7423ceaad916c0538cea2a686cd19a1d7

Observation f0f03878-046d-4363-94fa-03d3f5fac845 · outbound

This paper cites and Schapire, R.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon and Schapire, R

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:090e080a2ca83fbfa15f5ed9435eff5826e38f1172bc9f321954d562f8ada910

Observation e8fc4515-e50b-47dc-a6ef-da31fa0dc109 · outbound

This paper cites an unresolved cited work.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:65f310abe7e41c433197cdb73c68a9cfe1e5a2dda47ba63455d055fa47c8c609

Observation c4614499-a943-4cec-bbe7-863c423253f5 · outbound

This paper cites TRL: Transformers Reinforcement Learning, 2020.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon TRL: Transformers Reinforcement Learning, 2020

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:da77dee398f6546fb2c908760bf3ea13f3acb84112ab63a5dca8dd03598ae3b6

Observation 3fe2005c-ef1b-44da-aa88-7162b875dcb3 · outbound

This paper cites Oracle-efficient pessimism: Offline policy optimization in contextual bandits.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Oracle-efficient pessimism: Offline policy optimization in contextual bandits

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:e06cdf6c02ed5e4ae377fe032c6ea167ad01339bad88ffdf781be4571cf0fb64

Observation 1e5c3aa2-a3d2-4dfd-aaae-a1fe028437dd · outbound

This paper cites MiMo-V2-Flash Technical Report.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon MiMo-V2-Flash Technical Report

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.711026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:87b7ff4ca5fd042c29a9c970238fd4a2115e26c94d33150fe09a9809fb5e16d9

Observation f439881c-abb8-4bfa-a487-04722f306782 · outbound

This paper cites J., Krishnamurthy, A., Rosset, C., Awadallah, A.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon J., Krishnamurthy, A., Rosset, C., Awadallah, A

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:3c30fba4726293d381a5e090839d2d0853b01f57d0339731ddae36f3cd4a12c8

Observation cec4c83d-8ba5-45b8-b569-d375ceb5c8f1 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:21.719763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:b36536057a2cdd0d24d527b01d712fda1c934fb9032c944990dedf51e6736048

Observation 4a6c6593-2f2a-4241-bf63-b04120034149 · outbound

This paper cites Qwen3 Technical Report.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Qwen3 Technical Report

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.702962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:77513ba2634c623694ff77597793f857109be85078c4d3d640ec4039376c77e5

Observation ad3d0918-f783-4a0d-a8e2-8196820c24aa · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.705666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:d494c1b64e9294a60e75cc869c54b5b955cdacf7e8edfd48f0eb5075ee59803b

Observation 3cc4f7cc-812d-45bd-8f87-37922110c0a6 · outbound

This paper cites Black-box on-policy distillation of large language models.arXiv preprint.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Black-box on-policy distillation of large language models.arXiv preprint

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:20.905986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:9855d2d7520bc2eaf8ed7921147d67995915347adb7de89b3187a6e906f936c6

Observation e89c053f-f5e3-4677-8fcb-b908766283f2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.713112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:6a87766afd8fc4ec152abceb179f2d331b0cdb319a6a54e62bb84cd4cadfdafc

Observation b7a0ec17-5518-464c-9918-2a0b2dccfa3a · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.734197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:a052c68c8b96bb07f574d19cceeb8134c5d336e95c1669fc015856c79c842a30

Observation fd06f75f-e44a-4cfd-8d14-2893de821aa8 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:44:20.910562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:0f41c0191023869fc21419299e085fff95f4b4cf1d459e536897fc17ca9bacaf

Observation a719017c-7e4e-4dca-8e43-1d5bfa4bb9c0 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.713644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:dde76a536889e38a08fd21ec33f636111aaaf650c3d7e9f613152f7358d944aa

Observation e5d8dd9d-938e-42d4-840c-c9850ca54a55 · outbound

This paper cites ζ 2 l−1X i=1 1[i̸∈ X(D)] # ≥ 1 2 E.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon ζ 2 l−1X i=1 1[i̸∈ X(D)] # ≥ 1 2 E

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:cf43efa3424baf021233e7ce969dc51a7579fe9de6c350b17528e904db1d0e76

Observation d366ecbd-de05-4adf-9b12-847b66d6646a · outbound

This paper cites Warmup Supervised Fine-Tuning.To enable RL training, we first perform supervised fine-tuning on the OpenR1 dataset [11] to strengthen the model’s reasoning capabilities.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Warmup Supervised Fine-Tuning.To enable RL training, we first perform supervised fine-tuning on the OpenR1 dataset [11] to strengthen the model’s reasoning capabilities

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-30T07:41:39.266071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:9144e543b47eaa171649f6345b2d340eaa0a9946d45ae870f8f448ec4c3cdfef

Pith citing papers

No inbound Pith citation observations are available.