Pith. sign in

Paper Citation Record · LEDGER

Reasoning Compression with Mixed-Policy Distillation

As of 6 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2605.08776.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08776 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:29:02.387691Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact13
  • verified fuzzy21
  • unresolved0
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5239fce8-56e5-45db-9027-10d4d79339f6 · outbound

This paper cites Chi, Quoc V.

Reasoning Compression with Mixed-Policy Distillation Chi, Quoc V

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.099620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:68daa3bc0e191e52dd697afee833dd32c9f57744a288d2931a1f70f82f98513e

Observation 9f5a3dff-e04e-47d2-8389-e468cc4194f6 · outbound

This paper cites Large language models are zero-shot reasoners.

Reasoning Compression with Mixed-Policy Distillation Large language models are zero-shot reasoners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.113459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:9855a3844e4756d0b3f7fc4625370e596d943644cb05e5ec2b8f34394b2f51ef

Observation 906c24a0-c143-4108-91c3-2b3e68c859fb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reasoning Compression with Mixed-Policy Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:56:28.127357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:5270fb1edd6b3d84427413ae298151ee147b7aba7d28d0e480a62cac708451fe

Observation 24f4f96b-4e31-4772-a346-26c320de25b5 · outbound

This paper cites Stop overthinking: A survey on efficient reasoning for large language models.Trans.

Reasoning Compression with Mixed-Policy Distillation Stop overthinking: A survey on efficient reasoning for large language models.Trans

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.108647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:a2022cc568a07ce4ffcba389c4f17f800da57ae21b79dcea0c1b0515a9670464

Observation bb25c0ad-d5a7-45fb-9a78-cf825275f53e · outbound

This paper cites Wait, we don’t need to "wait"! removing thinking tokens improves reasoning efficiency.

Reasoning Compression with Mixed-Policy Distillation Wait, we don’t need to "wait"! removing thinking tokens improves reasoning efficiency

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.090986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:4d059ead9a4ab2a0024b2e6d819d3878d6cde88f28026ee55addb2c92d718348

Observation 36b0bca1-5b46-4e80-bd39-784faaa3023b · outbound

This paper cites The benefits of a concise chain of thought on problem-solving in large language models.

Reasoning Compression with Mixed-Policy Distillation The benefits of a concise chain of thought on problem-solving in large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.095252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:1e8fdca2ce683e6c7b39c98878a39b4c597c4789d0044c15ff18889fff82f950

Observation a1564186-707f-4e3e-9cc6-5b41571fe43f · outbound

This paper cites Chain of Draft: Thinking Faster by Writing Less.

Reasoning Compression with Mixed-Policy Distillation Chain of Draft: Thinking Faster by Writing Less

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:56:28.199335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:ed0a296d05d8ffd5e97e22783f7d19ecc89868e036445b0748172eafe6c91753

Observation 9f5fbbe9-9ba1-4de3-8e0d-9670933f7e90 · outbound

This paper cites C3ot: Generating shorter chain-of-thought without compromising effectiveness.

Reasoning Compression with Mixed-Policy Distillation C3ot: Generating shorter chain-of-thought without compromising effectiveness

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.103993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:f7b337f0bc1adf8ca4491ed7b0cc567d41d23871d85f0867ed73d1634911ae7f

Observation a6d706c0-ba27-4b97-80ac-8f93d0219e6f · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.

Reasoning Compression with Mixed-Policy Distillation Tokenskip: Controllable chain-of-thought compression in llms

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.117779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:5a56b4ee9501c1179fd000c94eaa3a759f42d44f2ef573b1dd4bff0e4ab77afb

Observation e957e801-ff61-474d-a182-30d2420caa32 · outbound

This paper cites Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting.

Reasoning Compression with Mixed-Policy Distillation Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.206267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:a91331399494f85061623a1e2e7d03812da9b22e85e3023283d06f675346b35f

Observation a7e07d58-0a24-4fad-8290-d41b294264da · outbound

This paper cites Recut: Balancing reasoning length and accuracy in llms via stepwise trails and preference optimization.

Reasoning Compression with Mixed-Policy Distillation Recut: Balancing reasoning length and accuracy in llms via stepwise trails and preference optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.086846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:94f03c9606cbd8934a2e541ae84cff1f5095329253d606e5de422a118f96d99e

Observation 4a721cd9-76c9-4ae1-ae28-4d65a3cf2e54 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reasoning Compression with Mixed-Policy Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:56:28.272051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:0611e029bdf9d7e9284f5372bae9c43d2dd526ac4fa55d7aa7eee9c25d4aeba4

Observation 803a5ec3-ef33-4874-93f2-2af4b5f69543 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Reasoning Compression with Mixed-Policy Distillation L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:cc26dcb9715abb1567b7a1516eee466b85cdc4ce564721c999514f63feb9b6dc

Observation 7f5f5274-7069-4df7-a532-eba6699190e2 · outbound

This paper cites Thinkprune: Pruning long chain-of-thought of llms via reinforcement learning.Trans.

Reasoning Compression with Mixed-Policy Distillation Thinkprune: Pruning long chain-of-thought of llms via reinforcement learning.Trans

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.081456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:7b2ae86e04d0586032370fd33d6880d10eafac1e986076d3853cd685517ac122

Observation d2645ce4-997d-4abf-9683-2d3070237991 · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

Reasoning Compression with Mixed-Policy Distillation CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:56:28.173768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:5cfcfd2a5bb09441f42cef885ccaad14cb4c722a2b9632b946eb167d4f1e7787

Observation 0cf3891f-a9f8-4041-9a52-b58291bade57 · outbound

This paper cites Openmathinstruct-2: Accelerating AI for math with massive open-source instruction data.

Reasoning Compression with Mixed-Policy Distillation Openmathinstruct-2: Accelerating AI for math with massive open-source instruction data

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.071867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:72dfbe8a7ffc71a5a1391b9a599353ce5593f51a1e1b35688b214852c13e6b25

Observation 2bb9c946-0fa7-4be5-993d-7d93c839458a · outbound

This paper cites Qwen3 Technical Report.

Reasoning Compression with Mixed-Policy Distillation Qwen3 Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:56:28.182351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:d70ee75285576ad4e4397dd56181cb977d865fcf3e04cfa3ebbdd1a001b2a011

Observation 18f614ed-e3b1-4885-9b78-ed465c708402 · outbound

This paper cites OpenAI o1 System Card.

Reasoning Compression with Mixed-Policy Distillation OpenAI o1 System Card

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:56:28.220558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:9815d909044eb10e90e8491ba50f08057fcaaa5efae1dfc8a2b3585dfc439bc8

Observation d4476742-a5ee-4df9-8521-86ba04693a24 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reasoning Compression with Mixed-Policy Distillation Training Verifiers to Solve Math Word Problems

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:56:28.278238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:a49255b5cacd6e2efd1e0d071cc8b81bec1cd6f52085bb8e3f2fab081f38e456

Observation 8895d4df-acbd-4e1d-86bb-ef10854bbdf5 · outbound

This paper cites Nguyen, Quang Pham, and Nghi D.

Reasoning Compression with Mixed-Policy Distillation Nguyen, Quang Pham, and Nghi D

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.076156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:fb0d3fc88419b5f79f950cc880f67534ab2dd3c2d150b192c866f3fd8dbdf6b3

Observation 9e9a8cf4-2ced-47c9-8f2b-ea9c25b360c4 · outbound

This paper cites Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks.

Reasoning Compression with Mixed-Policy Distillation Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.067394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:1f3502fd4b46aa49c6a06a6513edf0c5d7d7f9238a79c0e93f806e98dc905166

Observation 05e092c1-6179-4af5-88eb-c8fd5ad07e34 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Reasoning Compression with Mixed-Policy Distillation Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:51:30.429014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:19c162c46bdc4d288be8e2366bb6957eb2675bb11eb75daad1789bc319a9273a

Observation 7d6af50d-2d8e-423a-aa47-bba7e3f2f151 · outbound

This paper cites arXiv preprint arXiv:2603.09906 , year=.

Reasoning Compression with Mixed-Policy Distillation arXiv preprint arXiv:2603.09906 , year=

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:56:28.256127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:1ce68f1587a524084be45ab534fc48a10d56f94a4dc910402a54ee5015748be8

Observation f6727f4c-34e4-401c-ab16-6ebfaeb57273 · outbound

This paper cites How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach.

Reasoning Compression with Mixed-Policy Distillation How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.234223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:b40dedeff9b96d97c7e6e46627605d328cf9991e5a66ac0f54a6ce43f792e630

Observation 48faaedd-4bfe-47db-9be2-56981da3159e · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Reasoning Compression with Mixed-Policy Distillation On-policy distillation of language models: Learning from self-generated mistakes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.030470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:a22fa8d8bf09b698735d9dc8b600002e2ba416242507438a0926c875b621472f

Observation 2109f45f-7f28-4ce1-9a4d-4b18c9badbe0 · outbound

This paper cites an unresolved cited work.

Reasoning Compression with Mixed-Policy Distillation Unresolved cited work

Reference 26

Resolution
parse uncertain
raw_fallback, observed 2026-05-14T06:59:04.122087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:6ccae7a970a0cd79c29820ed8aaf37f21541402f5b30662cf8293c6f854f4f7f

Observation 7aadeaee-71b3-4a33-a4a2-a1aa47ce0331 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Reasoning Compression with Mixed-Policy Distillation Minillm: Knowledge distillation of large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.038592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:b574ec3f560fe1bd868bb91ffd3bc4dbeadff2abbff13ce020eb253efb216b31

Observation 6b78cb60-8b28-480e-af82-6aa6984cabf4 · outbound

This paper cites On-policy distillation.

Reasoning Compression with Mixed-Policy Distillation On-policy distillation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.033964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:fb0409adf51eea72817913edb6a037ae5781344b8f43cf998e0da9f995617918

Observation 7d3533df-8ed8-4f1a-ba3d-3842467ce009 · outbound

This paper cites On-Policy Context Distillation for Language Models.

Reasoning Compression with Mixed-Policy Distillation On-Policy Context Distillation for Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:48:41.492604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:c540507a63f7d0f8639d20da318a91b1ce785122464364b8c645ac099484e72f

Observation 906fa986-bfe3-44eb-8b3f-ebee18385ad6 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Reasoning Compression with Mixed-Policy Distillation Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.055042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:bdec91fc79a7992e8f1e67534e579d0acdbc49f56cc637da0db9f2b51d3f7ddc

Observation 1a907381-4167-4cc1-958d-243350e495d5 · outbound

This paper cites American invitational mathematics examination (aime) 2024.

Reasoning Compression with Mixed-Policy Distillation American invitational mathematics examination (aime) 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.058799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:ae464cd726e505878939d35252e348ee0bab190f35007b07b79bcbea6512a3cc

Observation 3d9dc4f0-be9b-4c0e-86b1-00c74d2a1fc1 · outbound

This paper cites American invitational mathematics examination (aime) 2025.

Reasoning Compression with Mixed-Policy Distillation American invitational mathematics examination (aime) 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.052122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:ae77dd2dbc41cded5c1e27a677572105bee48f5db23d43dd95da1cbe04b467e9

Observation 28e6659f-ea38-427a-a1d5-0740c70acad5 · outbound

This paper cites Overconfident errors need stronger correction: Asymmetric confidence penalties for reinforcement learning.arXiv preprint arXiv:2602.21420, 2026a.

Reasoning Compression with Mixed-Policy Distillation Overconfident errors need stronger correction: Asymmetric confidence penalties for reinforcement learning.arXiv preprint arXiv:2602.21420, 2026a

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.161002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:209e3e8769e21334f3c847c6ba123b3c5139c9cc089ef04773b607f3c3e2ac6d

Observation 0dbe3366-048a-4aae-8c47-2d6ca33ac7b8 · outbound

This paper cites Protect our environment from information overload.Nature Human Behaviour, 8(3):402–403.

Reasoning Compression with Mixed-Policy Distillation Protect our environment from information overload.Nature Human Behaviour, 8(3):402–403

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.048275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:832235dca84b9f47ff721ddf2c9eeafd84332c8a6c0881839b1c092991098456

Observation 143c1877-7029-4336-a7d0-ae4624394da3 · outbound

This paper cites Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala.

Reasoning Compression with Mixed-Policy Distillation Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.043625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:f99b87692b341885b58dbfedf7207f5fdd03fcb47b713b50f670f0214344f133

Observation e3d7f1d7-44d6-4e39-a57b-fec434cec536 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Reasoning Compression with Mixed-Policy Distillation HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:56:28.265448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:cf1a1e16e05887af00e3eac2684b3f02ac21d8f208d0c19b218de0b7a8170277

Observation d276b8ec-8e96-4695-867a-2a952046fe64 · outbound

This paper cites PhD thesis, UC Berkeley.

Reasoning Compression with Mixed-Policy Distillation PhD thesis, UC Berkeley

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:59:04.062543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:f5c22d7e8e1e3db7701891c3e50d6720b923526e393b5d8c45bb36980fe706e4

Observation 866e854f-69c8-499a-8fc4-391bd17d07a4 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Reasoning Compression with Mixed-Policy Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 38

Resolution
malformed identifier
local_arxiv, observed 2026-05-12T07:56:28.240279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:8e2c7b8afda76bc28c59843888f6424eccb681a28286ca872374a44a8c683bb6

Pith citing papers

No inbound Pith citation observations are available.