Pith. sign in

Paper Citation Record · LEDGER

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

As of 5 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 6 inbound Pith citation observations for arXiv:2605.01347.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.01347 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T15:18:39.844534Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:31:28.536475Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T21:06:34.834207Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact28
  • verified fuzzy11
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eef348e2-f599-4f49-8978-168331962a8e · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate On-policy distillation of language models: Learning from self-generated mistakes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.198866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:e550bee46dd0ba5eb86217b285c8a7a5fbe84753aca4d0e57fa27d4968fe5f26

Observation 7be22d38-e786-4ce7-99dd-a40a7b03e74f · outbound

This paper cites SMAGDi: Socratic multi agent interaction graph distillation for efficient high accuracy reasoning.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate SMAGDi: Socratic multi agent interaction graph distillation for efficient high accuracy reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:19.166207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:6d07b1f06ba94f6a7512f31e6e6640ff28d3fc6969fe4f2ad92f07ac2871186c

Observation 7efd699f-2ac4-4689-8594-eebdc0807367 · outbound

This paper cites Program Synthesis with Large Language Models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Program Synthesis with Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:19.273752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:f1ce7cb1b38d5c7dd87201751461b7af4f4a8829fca8757d1e5526bbd8e0afb1

Observation 1aaae547-f046-46cc-ae5e-b00e644638b4 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:52:17.902869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:ef5d2cf0bc790d8e2c9083530e5726e79942e4af50a5e67b19df968c8c5b445f

Observation f6ffb5d1-ad7b-4f03-b9a8-614487027435 · outbound

This paper cites MAGDi: Structured distillation of multi-agent interaction graphs improves reasoning in smaller language models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate MAGDi: Structured distillation of multi-agent interaction graphs improves reasoning in smaller language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.201683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:b3c0d1cb1b9ec95bb812401698c7ee1a9a034093356e164e8bfd87c6bc14af86

Observation 36546e9b-1139-4cc4-b9e2-938f1e0799c2 · outbound

This paper cites ReConcile: Round-table con- ference improves reasoning via consensus among diverse LLMs.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate ReConcile: Round-table con- ference improves reasoning via consensus among diverse LLMs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.193083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:6ea011eff6a6f082614960dfe1d61e7a743313ec768c4127a83f08eb817f17de

Observation 1beab1bd-b082-49fe-a436-1411be21ec7f · outbound

This paper cites AMTSS: An Adaptive Multi-Teacher Single-Student Knowledge Distillation Framework For Multilingual Language Inference.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate AMTSS: An Adaptive Multi-Teacher Single-Student Knowledge Distillation Framework For Multilingual Language Inference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:19.781233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:1e9e596adb107f9287209b0030301a51ffdcdfaa31ebd1378ace388b4ba0d4cd

Observation 6b750b7e-e0ce-4b32-aaa1-9dfc0789a0f3 · outbound

This paper cites Improving retrieval-augmented generation through multi-agent reinforcement learning.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Improving retrieval-augmented generation through multi-agent reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.190221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:695c81e4081130add62e91c9ef8c04e70d49484504cdd2b86be4cd8212455615

Observation c154312b-a088-4cc6-972a-556f361a006e · outbound

This paper cites K., Zhu, X., and Li, S.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate K., Zhu, X., and Li, S

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:41:21.555087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:bb7a30d3dd59f8252be816dba472e23d32a16e2810d3bab82a57e42e5561e6aa

Observation 9745eedf-504c-4090-9949-65f8de64a94c · outbound

This paper cites DeepSeek-V4 technical report.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate DeepSeek-V4 technical report

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.204185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:76ed58319a872df6ac6eb924f0db9ef994550b8dc6beff4567b721479cbeb02f

Observation 280a70af-6a0f-45c3-9982-80151f80577d · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:45.301651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:5d0de4a0f3d6a7b14c80813c06932af92c191fee1ff105219f3272552362929a

Observation b34186d4-dd9f-48a7-aff0-eaf6e843043c · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:21.670733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:9c0ab1fddb516754e057ecc454bcf7767ebdb28a23fa9b1acacab0aa33411a6f

Observation fdc7602b-bdfa-4427-8360-9c3d536dc601 · outbound

This paper cites MiniLLM: Knowledge distillation of large language models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate MiniLLM: Knowledge distillation of large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.209815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:32a23b5f84d23eb7854ddffbf667aa71c07ee4be623e93ba3ea5382ea6737a7e

Observation 39037226-ef26-44be-adfb-c67c8ba4ad5a · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate OpenThoughts: Data Recipes for Reasoning Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:57:51.597669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:f21c94dcffc53a0a3daef3f07e0f4102e6225184f22cad7e638e908c505185f5

Observation 0ee95f87-7f24-469e-bd1e-db1b774f7518 · outbound

This paper cites Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:21.665570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:a275d870f0e1563114693ca980f37623730d5c720dadb4f382747c9e984465ef

Observation 72bd42ca-d9d4-4b42-b228-60eb01054e6c · outbound

This paper cites Distilling the Knowledge in a Neural Network.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Distilling the Knowledge in a Neural Network

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:19.795296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:dcb0934ed055ee0bbc1189407c623ffda4e6e27ab568e8ada7ef328ad843a283

Observation f0b10198-68ee-4689-a945-bf7cc93ede90 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Reinforcement Learning via Self-Distillation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:29:18.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:bf81a717d7324f75fdb7a57149e5695b68254f21108ec1c17acde07a2e1eee5d

Observation 3754dbe2-f1f9-4099-abd3-15acf850fcba · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:20.240734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:e47b4a0845da5029fcfdad93610a9a9c4601a29b182e7ed0fa09f055ef948261

Observation 6066dad1-c0ab-4a2e-b0c8-b5f9d322deef · outbound

This paper cites Stable On-Policy Distillation through Adaptive Target Reformulation.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Stable On-Policy Distillation through Adaptive Target Reformulation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:20.224411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:77bdc92eb778a14de4cdaf02093a6032e91fcd2ae2de2977d151e47ba93dd3cb

Observation 12442ef8-7d89-4535-a882-7fa30df4ece8 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Entropy-Aware On-Policy Distillation of Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:02:00.585154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:4a541206760f81798b42cd718de271fa98434e12d1db159c0fa7fdf02b723a55

Observation 65f0a3bc-f6c2-4061-8fe5-f52b27a675f8 · outbound

This paper cites an unresolved cited work.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-26T01:17:17.212724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:7e7c484f9113bec050721323fd3f683caa277e03735946c0f634f95b35206f0c

Observation b36d7507-e2e3-4a42-b496-77e1d1338da2 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:19.052841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:50654abd1e76f80ea791fed69492d2323c101022594292b0f7f40d85ec54e494

Observation 29e3a2a9-3821-4941-8947-568fd2774f1b · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Encouraging divergent thinking in large language models through multi-agent debate

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.181967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:9ff32717711819e52222e9a1ee68bf02f6e3bf1db4dc67b163b4f4aecfde1e47

Observation 4f6b0ee9-1b1e-4774-9a74-766f9b88bcdc · outbound

This paper cites Divergence measures based on the shannon entropy.IEEE Transactions on Information Theory, 37(1):145–151.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Divergence measures based on the shannon entropy.IEEE Transactions on Information Theory, 37(1):145–151

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.206978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:152c7cfdd50213136fc2a5d3a0e4466df6a5133629cb4ee9d6f87d9a60f52cc2

Observation 97a1938b-4618-4f53-adaf-4e8c96672ac3 · outbound

This paper cites Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.195909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:87b2bef76033dfa689f7ac2ee83d26459a9744391043a0a7c89f0018d85e51ed

Observation ef65ccf7-2e1d-49e7-80d7-bdf882c95be4 · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate ToolACE: Winning the Points of LLM Function Calling

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:19.046738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:4e5cd824ebea544da1aa62eae96ae4f81d9171a749c3c73386296a2263e316c6

Observation 677187ca-da3b-429a-a9d4-bf274c1edf3b · outbound

This paper cites Decoupled weight decay regularization.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Decoupled weight decay regularization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.184711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:1c7c7b01ace9f3174a3b4f89e643b45878cb118b628c4c72a1a7a19ec3e4c1de

Observation ce802ad6-1ea6-4a40-b7bd-426062262e7f · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Con- nectionism.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate On-policy distillation.Thinking Machines Lab: Con- nectionism

Reference 28

Resolution
verified exact
doi, observed 2026-05-09T22:08:57.583809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:92c5436ebaedb5b7ea04335582d63c3a33533f126251c7365f8c2613425a467e

Observation 61c22ce0-a9eb-4916-b90d-779314c5c6a1 · outbound

This paper cites Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:20.540725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:06c9b9bc109016a5a4e4ac375ee2884029902deaf0afd07da174c91389fc0153

Observation baaebeef-1d64-40e4-b31f-bffd6e6052f3 · outbound

This paper cites Privileged Information Distillation for Language Models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Privileged Information Distillation for Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:1397c0b267dcec82c84de3459373a5c0b4f8e4707ab52cb7c8174bc24f7a309c

Observation c4348353-cfc9-4e28-b446-83c796480f1e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:20.975823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:e4a9bea8b72c2cb1d20c26616d77e021118368f1e5a669b149e156bf40387c87

Observation 924f58a4-87f7-4381-bfa0-33e598fc3827 · outbound

This paper cites Cycle-instruct: Fully seed- free instruction tuning via dual self-training and cycle consistency.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Cycle-instruct: Fully seed- free instruction tuning via dual self-training and cycle consistency

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:41:19.478364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:7adfa30b7e8fd9aba5e71474304e19db32d89e726ecc484bb4ed2a0c3cd0619b

Observation 6a7ac809-0d5c-4c24-acd0-30eab8b39ff0 · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate A Survey of On-Policy Distillation for Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:47:19.238105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:387844c9c0eb141069a419ed79339c3df4767b8413d14e1240160b2a92a90057

Observation af531e27-ff4d-4dec-b69b-85a21ba8a24c · outbound

This paper cites GKD: A General Knowledge Distillation Framework for Large-scale Pre-trained Language Model.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate GKD: A General Knowledge Distillation Framework for Large-scale Pre-trained Language Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:19.061053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:b1720b7e6bb5af02a78f0c2c9700bdc51b538c222ac964aa01b998cc4f11e415

Observation 3daaa0ac-ade3-4ae6-8fc7-66ed0faa19b2 · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:19.629338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:a360a976bbe9acf6a4dfbe642390e6e569e455562e285273ba1cf618c1f6dbdd

Observation f697bbc9-6970-4099-9497-72d95507cd58 · outbound

This paper cites Towards a Participatory and Social Justice-Oriented Measure of Human-Robot Trust.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Towards a Participatory and Social Justice-Oriented Measure of Human-Robot Trust

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:21.365199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:06edddf1594b939aeb4505dcfd8df41161e498473c9d88743b376a22971ebd59

Observation 32880efb-ae03-4888-8d16-e5b51cdaba33 · outbound

This paper cites Qwen3 Technical Report.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Qwen3 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:20.724373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:3f539516b45ca180bfd59d4f48fbad69900be4ec6ca0848d50400696131f3cf1

Observation b9f40531-4cc6-438f-8d46-816875f81c84 · outbound

This paper cites Self-Distilled RLVR.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Self-Distilled RLVR

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:41:21.254424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:7c382b0c79485e144aa56d545e84468df42d72005af1e5ecf066e872d3e76090

Observation 8288577f-4147-4f79-8902-03196e9bb9bd · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:11:07.754859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:14f437a6600b14e1c1c71e10846905c3af2727a40b54baa4c66a47c8e7bd955e

Observation 0e86687f-99e6-4ac1-ab57-c698271932f2 · outbound

This paper cites On-Policy Context Distillation for Language Models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate On-Policy Context Distillation for Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:48:41.492604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:3e82b5db8d822e3cda73a6ee2afee1c09ec348d60acafdd53fcca22c16dd58c6

Observation 4fc16924-36eb-4e97-83dc-c88954dacf01 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:54:31.187112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:cc0d51da27b610e681c4a1c0ccdfeac0d6a1a710833bf21524a1951fb3435c86

Observation 779565d1-1f1c-430a-ae2f-16347951864e · outbound

This paper cites Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement

Reference 42

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T16:41:20.005103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:4fddc71737f9d2fc7b8e09bcc5c9e8ddbeb40d4f855da2c202b7bff4bf061b9e

Observation a8ba43c8-276c-4018-8787-fc76e122f7a9 · outbound

This paper cites 2020-10-16.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate 2020-10-16

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:17:17.187516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:7b235cc6981755958fc02641a65fc7bf73f92e484086c1b6048041a96abf2d9f

Pith citing papers

Observation 3383638a-2ab5-4588-9be1-96448d6cd801 · inbound

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning cites this paper.

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:11:17.758464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T08:06:30.862911Z digest=sha256:92d573fc03b3cb45d629f3a68e2bdc821a9eb220f46701ee986078fe98d7f608

Observation 7e930519-c915-4227-ae28-ce5401dd800b · inbound

A Formula-Driven Survey and Research Agenda for On-Policy Distillation cites this paper.

A Formula-Driven Survey and Research Agenda for On-Policy Distillation MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:19:47.049206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T09:02:13.340365Z digest=sha256:41072cf0d388f954db603af4bc78ccb6d6477887377b7f2e02de369c80025c0f

Observation 716121cd-cf07-4506-9245-997d95bad734 · inbound

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation cites this paper.

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T21:06:34.835892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T21:05:08.742517Z digest=sha256:f2f4476a0a61ca8f9a590463453ae101649d0e9edfedc3e9fedf89ad3926053d

Observation f7b122f6-3751-40b7-88d4-f1b61f625e2c · inbound

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation cites this paper.

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:11:32.403020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:11:32.403020Z digest=sha256:65155feaa4455135a774de19294fab2afec4f6743b036d356345eed2a40dedc3

Observation f8b2d2c5-8f7d-4aa4-8275-c048f65352d7 · inbound

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation cites this paper.

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:28.536475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:28.536475Z digest=sha256:10281fb83b698507ad5964a9b70651db7df4c9c154ee78d7a76e00f394b67a49

Observation dfd3fc61-235a-4883-9358-1f569245d082 · inbound

MIRROR: Learning from the Other View for Multi-Modal Reasoning cites this paper.

MIRROR: Learning from the Other View for Multi-Modal Reasoning MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T07:15:04.810651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:15:04.810651Z digest=sha256:a94a351193ad72d80df2a5581a7253d004da1e5f66d8ab4e40f29df5e8b3943f