Pith. sign in

Paper Citation Record · LEDGER

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing

As of 5 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2605.05940.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05940 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T14:20:37.141500Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:39:37.096675Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T10:19:47.040092Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact17
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a68abf3-262a-4c50-b57d-f64c0c230539 · outbound

This paper cites GPT-4 Technical Report.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.055846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:1fc0f3f4142bef0dd9471cd790b2b24e163bbd554ba5dcee03f370d36171ede5

Observation 71cce084-14e8-4e71-a4de-8afc5e94be4d · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:11.127436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:c4742f91d61e0df96bcdfd9883643d869e9aaa80eb403da154ff8a109f27d281

Observation 48f882a1-0fdf-4e9a-9830-991c1f4fcdcc · outbound

This paper cites Qwen Technical Report.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Qwen Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.084497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:ff1956b4f97476fcfefb2d12b796a04368a31a15e4f1798b12e7a273ee257f70

Observation ebfcfb0c-5db4-4edf-b438-34c95145772f · outbound

This paper cites Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:11.234190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:f0ef2968c4c99950151fce0afca0560d8c1abed8da0dc843f787268bc45c659a

Observation 389bbd15-3c65-4570-8b76-9ad6092b48a7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Evaluating Large Language Models Trained on Code

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:41:11.032193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:eb6e8e5dc857f4cca2f4aa47cab5087c194be93a920f9b1f92b8442cd801bf88

Observation f967d690-a3f0-4af2-bfbc-5849b5abb4fc · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing MiniLLM: On-Policy Distillation of Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T17:40:27.990750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:bf4be0051aef7e8c749b7b1b710d1a4728a6bf11fdad39661bae40ba589bd750

Observation 8298ea37-677e-49ac-958b-4b81470adfa0 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Measuring Mathematical Problem Solving With the MATH Dataset

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.077759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:6059d14cf836b1fa5b76be8df9376975be50822fa2c3161ba1be8e23082c96b9

Observation b810ca00-7290-4472-a8b8-5246337bf24e · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Distilling the Knowledge in a Neural Network

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.100735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:e212be80326c2412ca7ef83605bb791e025e99ad282b3fad62ef3fd07cb02228

Observation 57e78f15-793d-48d3-b860-f7baa8f580d3 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:11.218058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:3f032e04e32979e33767244888c3a038a7a3cd45621f166fd7ffedaf0a8aa06d

Observation 2ed7d323-8883-457c-bb11-f5564a89be6f · outbound

This paper cites Sequence-Level Knowledge Distillation.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Sequence-Level Knowledge Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:11.201966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:776c37c128d37d58eb17c5d4f12176a04d5c3cfcb730ded2b39e52690283260d

Observation 22128e69-fbca-4a4d-9f98-4c8ee4151840 · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:11.258763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:a56919073cb1ebb2c921c1475813328cfa9c17cb2b6e9f47fef9a6325d1f4c9c

Observation a62b92f5-bb23-42e8-97cb-dfd7691ebed1 · outbound

This paper cites DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:11.069374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:3968de2172efec88dddfc37e5ed2400a9f7ac4c93a00bf88e7cadb3a63be4731

Observation c6133a74-2fe5-44c3-b5f6-4afe2a796849 · outbound

This paper cites From quantity to quality: Boosting llm performance with self-guided data selection for instruction tuning.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing From quantity to quality: Boosting llm performance with self-guided data selection for instruction tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:38.352945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:a745c063466bbebcdc9d4e70b7bbd4d051856c970eca5e96fa90194a06056c98

Observation 04d17247-3b63-42d5-ab12-c97459153abf · outbound

This paper cites DeepSeek-V3 Technical Report.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing DeepSeek-V3 Technical Report

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.177765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:57c494251a13a012e76106a3063771f6a9bbdfc6acbdbd4b4174b7675658ea4e

Observation fa6fe405-048c-4608-b5a5-17128dddb162 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.049615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:4ccb1185a5116be94294fb4532e25570af65c6bc3eb8dcbc3fb802c93fe776dd

Observation 3bb9c390-7dee-40a2-991c-7c9e03ba1aab · outbound

This paper cites Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:11.017629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:924cf1e416f5fe5bd6e0d7b352f42a7cf7032c041f13c7d3d7783b63bd5ff874

Observation 6a589774-f9bc-451f-a4b1-99a1b0f8539f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing LLaMA: Open and Efficient Foundation Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.240651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:5b815dc53663941131d6ae675762978655fe57b05edcdc62e4c8f553ce46fa52

Observation 80367539-ffd1-40c8-9ddd-e2661cd12d81 · outbound

This paper cites CLUE: A Chinese Language Understanding Evaluation Benchmark.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing CLUE: A Chinese Language Understanding Evaluation Benchmark

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:11.208726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:3b5ea56175bf161762caea6e0a480b170e46eeecb46a9e60cd99c64b32bf522d

Observation 82c9ead3-d731-4da1-ba04-ebee15001603 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.023292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:468216f160b750a9294c15a03a6098c414f5c9be142ce662c1c16fece7a89ef2

Observation 6d3c7e77-111a-4434-9631-0d041a9e2854 · outbound

This paper cites Qwen3 Technical Report.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Qwen3 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.037923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:ee88eca82e4c680902851c2ca1bad62a22d6374cf8382b0180ed31714f55f2da

Observation d3d5bb8c-4047-459c-a3ea-0b39b7b2996e · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:07:53.977132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:830eb7d12d514c636d1d74e8d5af159169d04c78d8db5718ea09ef398d078f20

Observation aac6be0d-8f9b-4d5e-8126-adb04a1473a8 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing Instruction-Following Evaluation for Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.167537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:86f722a6c98d77298434b200512646eee24cf2b47f0c4c27ec77502a88cb6ed6

Observation 208c8945-ed74-4fe3-a32a-4f01e1e19a24 · outbound

This paper cites The evaluation metric is the average zero-shot accuracy across eight benchmarks.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing The evaluation metric is the average zero-shot accuracy across eight benchmarks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:11.280183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:7f1bde51c12abdb9011e317c3c806f3ec22b39aa57866bd361c39a1aeddb76a4

Observation 7e83f3f4-347b-44eb-82f8-b081d78f088b · outbound

This paper cites generate-then-filter.

Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing generate-then-filter

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:11.158094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:37.141500Z digest=sha256:84bbb55538ac93f23845e979638f55dd714c71537b10a5a06904f2ce5e054300

Pith citing papers

Observation 410c8209-94a9-4504-8e22-e2a0808e6a8c · inbound

PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation cites this paper.

PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:48:46.307858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T03:39:37.096675Z digest=sha256:de1280739d9fcc28ec4f8882b375530996a4ac99d3320018ba33317e2027221a

Observation e736c71a-3dfa-40d1-9ae1-407af7ebd8f2 · inbound

A Formula-Driven Survey and Research Agenda for On-Policy Distillation cites this paper.

A Formula-Driven Survey and Research Agenda for On-Policy Distillation Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing

Reference 97

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:19:47.041368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T09:02:13.340365Z digest=sha256:e82a7e984d6c2e5a760a54ecc71e2af1bf5a3d908f5eed7c527e409b243b101c