Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.447695Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2505.18952.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.447695Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 59addfc3-77f6-42dc-9b13-99b28a498885 · outbound
Online Knowledge Distillation with Reward Guidance Improved algorithms for linear stochastic bandits
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69b458e-c723-4647-bb7d-58d1371ed75a · outbound
Online Knowledge Distillation with Reward Guidance Gpt-4 technical report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c8fd9c-96b9-4e6a-9be4-f58ce783a0ba · outbound
Online Knowledge Distillation with Reward Guidance On-policy distillation of language models: Learning from self-generated mistakes
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d333c0d-b3e6-439f-b780-8e2532147483 · outbound
Online Knowledge Distillation with Reward Guidance PaLM 2 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2282d970-697f-4073-aa62-40c41c0c9510 · outbound
Online Knowledge Distillation with Reward Guidance Claude 3 family
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d52f406c-99ba-418d-b2c5-263817954980 · outbound
Online Knowledge Distillation with Reward Guidance Gpt-4 is openai’s most advanced system, producing safer and more useful responses, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 710f5ccd-f9e6-472b-989a-e103c6851d94 · outbound
Online Knowledge Distillation with Reward Guidance Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e953fc06-285d-4d15-baab-417c2c7dc0e6 · outbound
Online Knowledge Distillation with Reward Guidance Knowledge distillation of black-box large language models, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72cfc732-41d6-4d40-8a45-21290303f112 · outbound
Online Knowledge Distillation with Reward Guidance Information-theoretic considerations in batch reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 00ef125b-bfbd-47d0-a74b-e3139451a397 · outbound
Online Knowledge Distillation with Reward Guidance Distilling knowledge learned in bert for text generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4208ee3d-3827-491b-b60f-b8609a9c7a79 · outbound
Online Knowledge Distillation with Reward Guidance Gonzalez, Ion Stoica, and Eric P
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34ae95c8-2171-4fa9-8927-572d0ea0e9d7 · outbound
Online Knowledge Distillation with Reward Guidance Deep reinforcement learning from human preferences
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca2004ef-5cdb-4b16-8c46-a47a8af4582d · outbound
Online Knowledge Distillation with Reward Guidance Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b21e71-a03c-41e5-af1b-ba32a9930d78 · outbound
Online Knowledge Distillation with Reward Guidance Training Verifiers to Solve Math Word Problems
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdbadbfd-da79-4e23-a822-dc15546d48e4 · outbound
Online Knowledge Distillation with Reward Guidance Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3317dc0-9d1d-4283-97f0-72bd436f7dcc · outbound
Online Knowledge Distillation with Reward Guidance Ultrafeedback: Boosting language models with scaled ai feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4d41c677-2518-404b-a14c-a747c279317b · outbound
Online Knowledge Distillation with Reward Guidance Stochastic linear optimization under bandit feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce12f250-4f44-48c1-ad61-7200fb8b2439 · outbound
Online Knowledge Distillation with Reward Guidance Openllama: An open reproduction of llama
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad3fb493-2ba3-4271-a22d-85192fab2dc5 · outbound
Online Knowledge Distillation with Reward Guidance Minillm: Knowledge distillation of large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e448cb-6a11-4e17-ae8d-5ecd6eb7f0f3 · outbound
Online Knowledge Distillation with Reward Guidance Direct Language Model Alignment from Online AI Feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3afeb083-cbd3-4ad9-8b15-2882493f894d · outbound
Online Knowledge Distillation with Reward Guidance Measuring massive multitask language understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afd52b7c-3055-4d4d-84ab-0bea56822713 · outbound
Online Knowledge Distillation with Reward Guidance Distilling the Knowledge in a Neural Network
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04deee70-79ff-4e23-b213-0866dca0cf22 · outbound
Online Knowledge Distillation with Reward Guidance Unnatural instructions: Tuning language models with (almost) no human labor
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 095018b2-f406-4c49-a671-cd95e0e668e3 · outbound
Online Knowledge Distillation with Reward Guidance Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ddd20eab-7613-46db-b578-35c42979b48d · outbound
Online Knowledge Distillation with Reward Guidance Adversarial moment-matching distillation of large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d3ce41b-3151-416d-944b-76ebfa0b15a2 · outbound
Online Knowledge Distillation with Reward Guidance Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7de2377-19f5-4805-8bd6-fb1023fece96 · outbound
Online Knowledge Distillation with Reward Guidance Sequence-level knowledge distillation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f367807d-13ac-45ef-b61b-4af49f4a7255 · outbound
Online Knowledge Distillation with Reward Guidance Tinybert: Distilling bert for natural language understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5fbb818f-063b-466e-85ef-e368a5800469 · outbound
Online Knowledge Distillation with Reward Guidance Direct Preference Knowledge Distillation for Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28d439b1-116d-4074-9e5d-17e9966d8356 · outbound
Online Knowledge Distillation with Reward Guidance Distillm: Towards streamlined distillation for large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59d05b33-e740-40c0-bf04-d1503d336845 · outbound
Online Knowledge Distillation with Reward Guidance Autoregressive knowledge distillation through imitation learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a062182a-381a-4113-a7ca-2b9a8f9fb9fa · outbound
Online Knowledge Distillation with Reward Guidance Openorca: An open dataset of gpt augmented flan reasoning traces, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a94db03-1fd9-44d4-b52f-aaccfd86e1d6 · outbound
Online Knowledge Distillation with Reward Guidance TinyGSM: achieving >80% on GSM8k with small language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1527a5-863c-4975-9caf-eb50e87ce8ad · outbound
Online Knowledge Distillation with Reward Guidance Rouge: A package for automatic evaluation of summaries
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d77b095e-aa7a-44bf-bc36-b8993ea832d3 · outbound
Online Knowledge Distillation with Reward Guidance Training language models to follow instructions with human feedback
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c400cc17-38ff-4af9-9e18-eb952fb59898 · outbound
Online Knowledge Distillation with Reward Guidance Orca: Progressive Learning from Complex Explanation Traces of GPT-4
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82c11d5-5432-4c83-9329-afdfa5b23c4e · outbound
Online Knowledge Distillation with Reward Guidance Linearly parameterized bandits
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d098173b-6f47-467a-becf-6decefae2354 · outbound
Online Knowledge Distillation with Reward Guidance Iterative Reasoning Preference Optimization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a182ff-d4d1-48db-b5c8-ccd0febaee66 · outbound
Online Knowledge Distillation with Reward Guidance Hybrid rl: Using both offline and online data can make rl efficient
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf70783c-9ce3-47e1-a898-1a34128b8653 · outbound
Online Knowledge Distillation with Reward Guidance Proximal Policy Optimization Algorithms
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f562761-1290-4e74-8baf-5206bf01bc0a · outbound
Online Knowledge Distillation with Reward Guidance Patient knowledge distillation for bert model compression
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f0d88d0-0d9f-4a1c-8562-0d5b05c465e6 · outbound
Online Knowledge Distillation with Reward Guidance Learning to summarize with human feedback
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8bf6598a-24a4-4fb3-a270-8dd9a7bb1ea4 · outbound
Online Knowledge Distillation with Reward Guidance Of moments and match- ing: A game-theoretic framework for closing the imitation gap
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56170899-96d4-424b-831d-609f23fc7305 · outbound
Online Knowledge Distillation with Reward Guidance Challenging big-bench tasks and whether chain-of-thought can solve them
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2523298-82ef-452f-aff5-ac7a8ab597cb · outbound
Online Knowledge Distillation with Reward Guidance Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e79d0a-3303-4470-a879-5698d3738d7c · outbound
Online Knowledge Distillation with Reward Guidance Hashimoto
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe4ff12-43fe-4232-bcd4-e15dd3add8bc · outbound
Online Knowledge Distillation with Reward Guidance Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b4d144-9a8c-4fbf-88ff-0a0d427ae1d2 · outbound
Online Knowledge Distillation with Reward Guidance Selective knowledge distillation for neural machine translation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f0a7e71-d80b-4079-9c28-82e5807ba4a4 · outbound
Online Knowledge Distillation with Reward Guidance Self-instruct: Aligning language models with self-generated in- structions
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ea241061-40c4-49a8-be59-4fa3a0749da7 · outbound
Online Knowledge Distillation with Reward Guidance Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a7fd4a4a-8696-4e12-878d-91962d7ac8ec · outbound
Online Knowledge Distillation with Reward Guidance f-divergence minimization for sequence-level knowledge distillation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9193a10d-b238-4dfa-b57a-18285898ac5b · outbound
Online Knowledge Distillation with Reward Guidance Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 31f262d2-7e67-4908-bc9b-f3c41959ec52 · outbound
Online Knowledge Distillation with Reward Guidance Qwen2 Technical Report
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76c67b11-d524-4152-8e7a-03e5b443adc0 · outbound
Online Knowledge Distillation with Reward Guidance Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c770c7b2-3bab-4173-b871-f675889b37eb · outbound
Online Knowledge Distillation with Reward Guidance Provable offline preference-based reinforcement learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72392bfe-04c8-4e01-94bd-8013d2fedae8 · outbound
Online Knowledge Distillation with Reward Guidance Online iterative reinforcement learning from human feedback with general preference model
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 10a3c533-330a-43cd-8bf4-351f10f6a8fd · outbound
Online Knowledge Distillation with Reward Guidance Plad: Preference-based large language model distillation with pseudo-preference pairs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a52340ce-2c6a-4184-ad7a-2a7213c3799b · outbound
Online Knowledge Distillation with Reward Guidance TinyLlama: An Open-Source Small Language Model
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84784f89-109a-40bb-bd27-586a98d19ecc · outbound
Online Knowledge Distillation with Reward Guidance Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f336086-d72a-4518-ae58-2840ed8c28f1 · outbound
Online Knowledge Distillation with Reward Guidance Mathematical analysis of machine learning algorithms
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e007a6a-5f47-487e-b2f2-221b56409059 · outbound
Online Knowledge Distillation with Reward Guidance Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d12c72a-b00d-46dc-b77e-05f94e5f29cd · outbound
Online Knowledge Distillation with Reward Guidance Agieval: A human-centric benchmark for evaluating foundation models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1dd11a70-1af2-4288-9ece-00b737f417f7 · outbound
Online Knowledge Distillation with Reward Guidance Fine-Tuning Language Models from Human Preferences
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad94eeee-5986-4c0f-956d-b8a04ca31f6f · outbound
Online Knowledge Distillation with Reward Guidance Measurements of nematic susceptibility with phase sensitive nuclear magnetic resonance in pulsed strain fields
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.