Pith. sign in

Paper Citation Record · LEDGER

Online Knowledge Distillation with Reward Guidance

As of 8 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2505.18952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18952 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.447695Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59addfc3-77f6-42dc-9b13-99b28a498885 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Online Knowledge Distillation with Reward Guidance Improved algorithms for linear stochastic bandits

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.172936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.172936Z digest=sha256:ccd71d16a80da12e4d3ca8cc171ae95f6753530890e07685ec0fd67d741e267b

Observation b69b458e-c723-4647-bb7d-58d1371ed75a · outbound

This paper cites Gpt-4 technical report.

Online Knowledge Distillation with Reward Guidance Gpt-4 technical report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.179006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.179006Z digest=sha256:7acb5416a06ddd6a45c225ae50aca4ac5abb3861ba888e71f9fa8f2b7df15e03

Observation 66c8fd9c-96b9-4e6a-9be4-f58ce783a0ba · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Online Knowledge Distillation with Reward Guidance On-policy distillation of language models: Learning from self-generated mistakes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.184236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.184236Z digest=sha256:5536f7abdbe1f03050c980c6f631e0fe6d4efcdf2820e1ef932281c258acc130

Observation 6d333c0d-b3e6-439f-b780-8e2532147483 · outbound

This paper cites PaLM 2 Technical Report.

Online Knowledge Distillation with Reward Guidance PaLM 2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.189344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.189344Z digest=sha256:db2f8674043f375e66999c86ee06a9dc2a32ad47ddbfa3395b269cb4af2366d6

Observation 2282d970-697f-4073-aa62-40c41c0c9510 · outbound

This paper cites Claude 3 family.

Online Knowledge Distillation with Reward Guidance Claude 3 family

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.139409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.193617Z digest=sha256:68b4ae92f12a3f512f5b338c02e67c09c27b765129e256be9169a4f1539ce4dc

Observation d52f406c-99ba-418d-b2c5-263817954980 · outbound

This paper cites Gpt-4 is openai’s most advanced system, producing safer and more useful responses, 2024.

Online Knowledge Distillation with Reward Guidance Gpt-4 is openai’s most advanced system, producing safer and more useful responses, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.127285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.208501Z digest=sha256:c7a7bddd6962e4badd86997e3497ae8041b53bdf727d2d8c3c70e7a33f4112ae

Observation 710f5ccd-f9e6-472b-989a-e103c6851d94 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Online Knowledge Distillation with Reward Guidance Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.213678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.213678Z digest=sha256:0434825eb0927a22b4f6990edcfa829db38952aebd87e67eca4e9d605621e3fe

Observation e953fc06-285d-4d15-baab-417c2c7dc0e6 · outbound

This paper cites Knowledge distillation of black-box large language models, 2024.

Online Knowledge Distillation with Reward Guidance Knowledge distillation of black-box large language models, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.113522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.219066Z digest=sha256:b16d63c9e9ede529523c3bcd744077ec4ab85a4dde0df27ad2eb859b7d6aeb66

Observation 72cfc732-41d6-4d40-8a45-21290303f112 · outbound

This paper cites Information-theoretic considerations in batch reinforcement learning.

Online Knowledge Distillation with Reward Guidance Information-theoretic considerations in batch reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.100612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.223575Z digest=sha256:371e703de9184b3d5f115d9d4fe242c6bb343423b38bc8bae43fd9514482c3ab

Observation 00ef125b-bfbd-47d0-a74b-e3139451a397 · outbound

This paper cites Distilling knowledge learned in bert for text generation.

Online Knowledge Distillation with Reward Guidance Distilling knowledge learned in bert for text generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.087106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.227911Z digest=sha256:6f5447549026a67f48986e4612d5671f71502214b5347271f06150a1f9c40773

Observation 4208ee3d-3827-491b-b60f-b8609a9c7a79 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Online Knowledge Distillation with Reward Guidance Gonzalez, Ion Stoica, and Eric P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.232858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.232858Z digest=sha256:53e87424ccf87da7ac2591e9e125d26fac48b733456b2949384e65b91b7bd9e9

Observation 34ae95c8-2171-4fa9-8927-572d0ea0e9d7 · outbound

This paper cites Deep reinforcement learning from human preferences.

Online Knowledge Distillation with Reward Guidance Deep reinforcement learning from human preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.236919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.236919Z digest=sha256:ddbceb5d34f9ef49967b68d21784a78d9d0fd2d160230d542d4fa184ad580857

Observation ca2004ef-5cdb-4b16-8c46-a47a8af4582d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Online Knowledge Distillation with Reward Guidance Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.240934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.240934Z digest=sha256:f6446ba31af4be72027cbb278352e48e57e34fb130aca8c00dcc8ea1f2f63ebf

Observation 98b21e71-a03c-41e5-af1b-ba32a9930d78 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Online Knowledge Distillation with Reward Guidance Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.244981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.244981Z digest=sha256:37400e5e117eaebcb87374f613a05da33c60bebc8a3306b4d681d61dcbbd4325

Observation fdbadbfd-da79-4e23-a822-dc15546d48e4 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023.

Online Knowledge Distillation with Reward Guidance Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.249279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.249279Z digest=sha256:bc0f9bc66e594f0061a68050e5bbe097d518f16bfcf36dfc73534d8af8a2a358

Observation f3317dc0-9d1d-4283-97f0-72bd436f7dcc · outbound

This paper cites Ultrafeedback: Boosting language models with scaled ai feedback.

Online Knowledge Distillation with Reward Guidance Ultrafeedback: Boosting language models with scaled ai feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.052465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.253183Z digest=sha256:36c976922ee7c8dae915552bcd01854b4e5fbe87faf299560f93d7673e387c91

Observation 4d41c677-2518-404b-a14c-a747c279317b · outbound

This paper cites Stochastic linear optimization under bandit feedback.

Online Knowledge Distillation with Reward Guidance Stochastic linear optimization under bandit feedback

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.040022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.257789Z digest=sha256:007f504828c7d26e17514e783348d39c55e9846af1409510f1f19df24edfe00e

Observation ce12f250-4f44-48c1-ad61-7200fb8b2439 · outbound

This paper cites Openllama: An open reproduction of llama.

Online Knowledge Distillation with Reward Guidance Openllama: An open reproduction of llama

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.028443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.262377Z digest=sha256:5b6d02a6925ea05272f3ce18c419fd632e45f2e8ac832afa5df694d9aa431911

Observation ad3fb493-2ba3-4271-a22d-85192fab2dc5 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Online Knowledge Distillation with Reward Guidance Minillm: Knowledge distillation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.266807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.266807Z digest=sha256:d14918ec139787c8913fc3c0ababf53a25891bc07824340418911ef942a3c6aa

Observation 49e448cb-6a11-4e17-ae8d-5ecd6eb7f0f3 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Online Knowledge Distillation with Reward Guidance Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.270567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.270567Z digest=sha256:4606f56a18b6aaf5826feb04e362df489293c8949272127c2748f308091a44ff

Observation 3afeb083-cbd3-4ad9-8b15-2882493f894d · outbound

This paper cites Measuring massive multitask language understanding.

Online Knowledge Distillation with Reward Guidance Measuring massive multitask language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.275247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.275247Z digest=sha256:1a540020cf2fb4827895cc9c5190f5916a1afecd3ba77a2d7b365635b74dd78f

Observation afd52b7c-3055-4d4d-84ab-0bea56822713 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Online Knowledge Distillation with Reward Guidance Distilling the Knowledge in a Neural Network

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.279060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.279060Z digest=sha256:6a065403e20060ee92b7f4b35acb8c64b3e9150e7cf05c6eb863b37e2a63ef05

Observation 04deee70-79ff-4e23-b213-0866dca0cf22 · outbound

This paper cites Unnatural instructions: Tuning language models with (almost) no human labor.

Online Knowledge Distillation with Reward Guidance Unnatural instructions: Tuning language models with (almost) no human labor

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.001512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.283639Z digest=sha256:b8c3012cbfd3234da572b040b985bb1f67fcc5f02445ab4128efe5e745472158

Observation 095018b2-f406-4c49-a671-cd95e0e668e3 · outbound

This paper cites Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes.

Online Knowledge Distillation with Reward Guidance Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.983865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.287480Z digest=sha256:8492d0c79160bfa40ba7f48d2ea5c4380b263dd8c9495e9bffc5dd749c58422c

Observation ddd20eab-7613-46db-b578-35c42979b48d · outbound

This paper cites Adversarial moment-matching distillation of large language models.

Online Knowledge Distillation with Reward Guidance Adversarial moment-matching distillation of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.962918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.295167Z digest=sha256:3939728bec25b230c01256c2e2a8821cfd3dd7ffc7b964864e380f73d9b77629

Observation 9d3ce41b-3151-416d-944b-76ebfa0b15a2 · outbound

This paper cites an unresolved cited work.

Online Knowledge Distillation with Reward Guidance Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.291106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.291106Z digest=sha256:b2dd08c63fbcc50b0af1ab45988d10e22e30bb7fa9550a0250081b0482c1518a

Observation b7de2377-19f5-4805-8bd6-fb1023fece96 · outbound

This paper cites Sequence-level knowledge distillation.

Online Knowledge Distillation with Reward Guidance Sequence-level knowledge distillation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.935337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.303395Z digest=sha256:4b6378f1681e09dfd13119ef45fb75ba4e01ac8fd52db9eadc52945e5a6efbea

Observation f367807d-13ac-45ef-b61b-4af49f4a7255 · outbound

This paper cites Tinybert: Distilling bert for natural language understanding.

Online Knowledge Distillation with Reward Guidance Tinybert: Distilling bert for natural language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.948970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.299174Z digest=sha256:2966917c84f868d4e95f51baf73462dc2b4f0c16a70ce1b49ef8178d2e4a2ff9

Observation 5fbb818f-063b-466e-85ef-e368a5800469 · outbound

This paper cites Direct Preference Knowledge Distillation for Large Language Models.

Online Knowledge Distillation with Reward Guidance Direct Preference Knowledge Distillation for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.311523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.311523Z digest=sha256:0ed52010d562f483566228071b405f156360440351dd3f3cbb2a529886f8be23

Observation 28d439b1-116d-4074-9e5d-17e9966d8356 · outbound

This paper cites Distillm: Towards streamlined distillation for large language models.

Online Knowledge Distillation with Reward Guidance Distillm: Towards streamlined distillation for large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.923252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.307693Z digest=sha256:57d776c56bb13f7caa74028a8362ad699c24a13a484a08e45ae26e5b1121bfba

Observation 59d05b33-e740-40c0-bf04-d1503d336845 · outbound

This paper cites Autoregressive knowledge distillation through imitation learning.

Online Knowledge Distillation with Reward Guidance Autoregressive knowledge distillation through imitation learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.899697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.319541Z digest=sha256:79111140d874034f50727c9615081412def07ca6e1df620de0409bd5c7dd155a

Observation a062182a-381a-4113-a7ca-2b9a8f9fb9fa · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces, 2023.

Online Knowledge Distillation with Reward Guidance Openorca: An open dataset of gpt augmented flan reasoning traces, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.911217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.315912Z digest=sha256:026075596f35756ecf66eba653cc413a21eda5af76549fde3f920a0479de3bc8

Observation 2a94db03-1fd9-44d4-b52f-aaccfd86e1d6 · outbound

This paper cites TinyGSM: achieving >80% on GSM8k with small language models.

Online Knowledge Distillation with Reward Guidance TinyGSM: achieving >80% on GSM8k with small language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.328172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.328172Z digest=sha256:22e38d5a992eb519fe2bc494a354f8bfcbe8a23340b55891f0a347d0a96b7bdd

Observation 0b1527a5-863c-4975-9caf-eb50e87ce8ad · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Online Knowledge Distillation with Reward Guidance Rouge: A package for automatic evaluation of summaries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.887728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.324021Z digest=sha256:ce8c9ec13aadf2cd457b2c950f8b85de740d0158ad80b4ab7e920b4586f525a9

Observation d77b095e-aa7a-44bf-bc36-b8993ea832d3 · outbound

This paper cites Training language models to follow instructions with human feedback.

Online Knowledge Distillation with Reward Guidance Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.335177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.335177Z digest=sha256:888333494e7819db426041b998e3e43a9fedc94a9acf7d058b0484c99b49039f

Observation c400cc17-38ff-4af9-9e18-eb952fb59898 · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

Online Knowledge Distillation with Reward Guidance Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.331916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.331916Z digest=sha256:3aa638adc1cdf9a879ace427eddd62c549f387bda8a9157ebba95025d23f7bb9

Observation b82c11d5-5432-4c83-9329-afdfa5b23c4e · outbound

This paper cites Linearly parameterized bandits.

Online Knowledge Distillation with Reward Guidance Linearly parameterized bandits

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.866336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.343111Z digest=sha256:89eb0e418313872a608e93ff2aebba60e8f250e5f4e0d4cd6a4d61ec45acf1d6

Observation d098173b-6f47-467a-becf-6decefae2354 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Online Knowledge Distillation with Reward Guidance Iterative Reasoning Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.338434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.338434Z digest=sha256:d232cf50c8ddaa406f0274472649b073185554a8e2396f5e974beb80ec1ca3a2

Observation 55a182ff-d4d1-48db-b5c8-ccd0febaee66 · outbound

This paper cites Hybrid rl: Using both offline and online data can make rl efficient.

Online Knowledge Distillation with Reward Guidance Hybrid rl: Using both offline and online data can make rl efficient

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.852660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.351084Z digest=sha256:b51d80ae258a9b4e8cc6ad6419073bb8e3b2001348134895dea4d9c062e42ccd

Observation cf70783c-9ce3-47e1-a898-1a34128b8653 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Online Knowledge Distillation with Reward Guidance Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.346880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.346880Z digest=sha256:6de646d756b58cb98c60fbaa21117723e3a6930bae603ded55b174ac1c08dba1

Observation 6f562761-1290-4e74-8baf-5206bf01bc0a · outbound

This paper cites Patient knowledge distillation for bert model compression.

Online Knowledge Distillation with Reward Guidance Patient knowledge distillation for bert model compression

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.827919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.361088Z digest=sha256:974b816c1f354e2929bd8efb489ff87657865a7829c30864c683a5ffb45d20d5

Observation 9f0d88d0-0d9f-4a1c-8562-0d5b05c465e6 · outbound

This paper cites Learning to summarize with human feedback.

Online Knowledge Distillation with Reward Guidance Learning to summarize with human feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.840182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.356124Z digest=sha256:4038cc779441d1a64d520de6520e6cb4a08f02cd42e3900bc03743dcdd21f3f3

Observation 8bf6598a-24a4-4fb3-a270-8dd9a7bb1ea4 · outbound

This paper cites Of moments and match- ing: A game-theoretic framework for closing the imitation gap.

Online Knowledge Distillation with Reward Guidance Of moments and match- ing: A game-theoretic framework for closing the imitation gap

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.803933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.369280Z digest=sha256:3d86f3888a1e01638dbe471e0dff27ade7e2060fc9c5d0cd04cad762c487bd21

Observation 56170899-96d4-424b-831d-609f23fc7305 · outbound

This paper cites Challenging big-bench tasks and whether chain-of-thought can solve them.

Online Knowledge Distillation with Reward Guidance Challenging big-bench tasks and whether chain-of-thought can solve them

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.816354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.365451Z digest=sha256:02d81b5a92d32e04fe899b52a0a1f568a91f53164847d1b32953ec3345ac6611

Observation e2523298-82ef-452f-aff5-ac7a8ab597cb · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Online Knowledge Distillation with Reward Guidance Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.379150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.379150Z digest=sha256:722ae448d77c5c2c6e80456fe4fbb030bec9ca0e4004bb1066133fc06d67a759

Observation 79e79d0a-3303-4470-a879-5698d3738d7c · outbound

This paper cites Hashimoto.

Online Knowledge Distillation with Reward Guidance Hashimoto

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.374653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.374653Z digest=sha256:95535a882082bb5a4433f867efc7eb0b49646c036617830b5a5f961b97755e1f

Observation bbe4ff12-43fe-4232-bcd4-e15dd3add8bc · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

Online Knowledge Distillation with Reward Guidance Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.387131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.387131Z digest=sha256:e8443b8b973dd9a3242b2cc03c6b4a21ad186b8815ed9ce953f07112e990dc19

Observation 30b4d144-9a8c-4fbf-88ff-0a0d427ae1d2 · outbound

This paper cites Selective knowledge distillation for neural machine translation.

Online Knowledge Distillation with Reward Guidance Selective knowledge distillation for neural machine translation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.783300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.382987Z digest=sha256:ab71d0a3de0add84677e9ce025a2f26c56545e6d6a89a05f6652d7ddcaa2abe1

Observation 7f0a7e71-d80b-4079-9c28-82e5807ba4a4 · outbound

This paper cites Self-instruct: Aligning language models with self-generated in- structions.

Online Knowledge Distillation with Reward Guidance Self-instruct: Aligning language models with self-generated in- structions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.751385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.395277Z digest=sha256:faa664d474bbd6ada1375082c8fbbd06636bffabf3cb2ea772317e58c32b3731

Observation ea241061-40c4-49a8-be59-4fa3a0749da7 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Online Knowledge Distillation with Reward Guidance Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.764237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.391263Z digest=sha256:9fc662a47cd1b6af4a2642efbfeb5505a80dd8c0d4ca3ca64ce726af89673380

Observation a7fd4a4a-8696-4e12-878d-91962d7ac8ec · outbound

This paper cites f-divergence minimization for sequence-level knowledge distillation.

Online Knowledge Distillation with Reward Guidance f-divergence minimization for sequence-level knowledge distillation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.726555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.403615Z digest=sha256:04388615eab8a784e4f48851c3401dfd66a484878a17c9640a87daeacb6a2324

Observation 9193a10d-b238-4dfa-b57a-18285898ac5b · outbound

This paper cites Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks.

Online Knowledge Distillation with Reward Guidance Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.739753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.399435Z digest=sha256:4bbe1f81769aa8c7d5c83a34789445fbdb4af2f1e67b73cf380971f93f3dfb8b

Observation 31f262d2-7e67-4908-bc9b-f3c41959ec52 · outbound

This paper cites Qwen2 Technical Report.

Online Knowledge Distillation with Reward Guidance Qwen2 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.412002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.412002Z digest=sha256:5f418a5add02ac6d0f3d48deee7a4732554cbcafe079c6a49c1146226de12ecb

Observation 76c67b11-d524-4152-8e7a-03e5b443adc0 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Online Knowledge Distillation with Reward Guidance Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.408130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.408130Z digest=sha256:8d80cf5314c22af449a7b92a7696f1d1e37084f209117ae57950fdbf9b8375db

Observation c770c7b2-3bab-4173-b871-f675889b37eb · outbound

This paper cites Provable offline preference-based reinforcement learning.

Online Knowledge Distillation with Reward Guidance Provable offline preference-based reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.694198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.420593Z digest=sha256:159bb9cf6ec79e7fdd560965a20b61486be38810a5532d4b10fad8fbbc36dd93

Observation 72392bfe-04c8-4e01-94bd-8013d2fedae8 · outbound

This paper cites Online iterative reinforcement learning from human feedback with general preference model.

Online Knowledge Distillation with Reward Guidance Online iterative reinforcement learning from human feedback with general preference model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.707419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.415999Z digest=sha256:6311dc0994d50ff987f8ed992eeba5b0a4726b62c646f6258c25f4d97fc94c10

Observation 10a3c533-330a-43cd-8bf4-351f10f6a8fd · outbound

This paper cites Plad: Preference-based large language model distillation with pseudo-preference pairs.

Online Knowledge Distillation with Reward Guidance Plad: Preference-based large language model distillation with pseudo-preference pairs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.682195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.428701Z digest=sha256:3a1f5f26da99a922b976936ce1628b7fb9b4e3da17606c54250d9ba493993505

Observation a52340ce-2c6a-4184-ad7a-2a7213c3799b · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Online Knowledge Distillation with Reward Guidance TinyLlama: An Open-Source Small Language Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.424641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.424641Z digest=sha256:371f077b84bc6eaf1bb7e9b99a6744c4aaae8ed6a85e0c45cb71af52cf6beead

Observation 84784f89-109a-40bb-bd27-586a98d19ecc · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Online Knowledge Distillation with Reward Guidance Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.436721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.436721Z digest=sha256:371b11e7758c4308e76adacf9e3ac5d3e08b24bd1a4c22fc819c11caf6e28bdb

Observation 3f336086-d72a-4518-ae58-2840ed8c28f1 · outbound

This paper cites Mathematical analysis of machine learning algorithms.

Online Knowledge Distillation with Reward Guidance Mathematical analysis of machine learning algorithms

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.432511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.432511Z digest=sha256:1289c92c79859ca73e09eafee78919243cf563bbe3952183e61283a9a53ba649

Observation 8e007a6a-5f47-487e-b2f2-221b56409059 · outbound

This paper cites Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023.

Online Knowledge Distillation with Reward Guidance Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.443540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.443540Z digest=sha256:b2ef413b4c026a3d2cfd379857805d83756b82e05de59830885aa3a11bca01e8

Observation 7d12c72a-b00d-46dc-b77e-05f94e5f29cd · outbound

This paper cites Agieval: A human-centric benchmark for evaluating foundation models.

Online Knowledge Distillation with Reward Guidance Agieval: A human-centric benchmark for evaluating foundation models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.654744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:49.440054Z digest=sha256:bed8e63a83241499e8705b1a9261f0d1469a8cf742adfc79486395948ef092e4

Observation 1dd11a70-1af2-4288-9ece-00b737f417f7 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Online Knowledge Distillation with Reward Guidance Fine-Tuning Language Models from Human Preferences

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.447695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.447695Z digest=sha256:cee02f79d9343d6882507070af9d7afc861df5438b87b8ae470b69ed29395e73

Observation ad94eeee-5986-4c0f-956d-b8a04ca31f6f · outbound

This paper cites Measurements of nematic susceptibility with phase sensitive nuclear magnetic resonance in pulsed strain fields.

Online Knowledge Distillation with Reward Guidance Measurements of nematic susceptibility with phase sensitive nuclear magnetic resonance in pulsed strain fields

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.203879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.203879Z digest=sha256:38b87e4cfe3829f45f63d050d5cc2de2785af59514386847034691a2ebb672d8

Pith citing papers

No inbound Pith citation observations are available.