Pith. sign in

Paper Citation Record · LEDGER

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models

As of 13 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.01717.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01717 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:20:08.802604Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e096106b-ccb8-4b44-8d88-e96efa60d70a · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.205227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.662137Z digest=sha256:0b1546192942e3aa1e0b1b2b7fec1d253fae5e37806a902625708f5e15d598b0

Observation d86cc0e2-1e6b-4020-97a5-d3422b44db29 · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.194915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.666455Z digest=sha256:92d8eac5839511f3414611e81497ce53110bf531bbd3206f0fb30754b7083cf1

Observation 4f001faf-74fe-4b1c-9ffa-d3156560e08f · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.183473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.670872Z digest=sha256:2ae4c832bd1ee919780ae51915ea85872f066e937138f378771bd45a363afccc

Observation ec41f3b1-52f3-4a54-972e-7abf67bfe60f · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.172824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.674633Z digest=sha256:244b87ede0d6a76e70b319ea367cf62d7f7511ece414d4b12e572ca620bf361e

Observation 57065abc-120d-4570-bf53-dca3a129b7cd · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.161233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.678892Z digest=sha256:d72c90811a04f1f6278fa5901f2ca00976e364d2d182ebc31d1a60069f4d2357

Observation a104d9a5-d879-4362-a2cf-94fddf81b289 · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.150353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.683125Z digest=sha256:5ade00d856d7b963126dd5b2744b3b0040bdfd72f9e42813c0fb5ae6e7bd45f1

Observation d1be0296-d4e7-4624-8f6a-988cc45148dd · outbound

This paper cites Dream 7B: Diffusion Large Language Models.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Dream 7B: Diffusion Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.688031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.688031Z digest=sha256:4e51f8b38f2b64e64cc6cd4e1ef554d20f9df02e98c7f340ff57322774a02529

Observation 237041a8-3fd2-4a4f-8fcd-d6dd0d2e45be · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.139351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.692135Z digest=sha256:0d8c4db484f836fb99f6b1c82e8c803d87409e5162894c1a17ce2295ec07b138

Observation aca783a0-b0b8-4563-a630-46910f2a96fd · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.695571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.695571Z digest=sha256:253f6d4dc5d7e4c6bb748074e894a00f2a7c480322e0c669545d97a99b8a3e38

Observation 6f174b15-af9e-440a-a6a8-c0062e2d774e · outbound

This paper cites Rojas, J.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Rojas, J

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.699209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.699209Z digest=sha256:cbd3fbc8ee7518f30ed53c3be5302aa6b8b6c2989e1be687f2c3b4087ecca95d

Observation adea863f-3067-4dcf-83c0-6a75b7824776 · outbound

This paper cites SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.702639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.702639Z digest=sha256:08ae4d541d0d627572483c71f55d9ac29ca428b1f48c880a6c96694936ad04b8

Observation 9e802747-b552-46e6-b3ef-2ab327fc7b73 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.707330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.707330Z digest=sha256:d8cf13ec908f47510b1f8a2fdbef0cf5d0935b269bd2dbc0f203a29489723916

Observation 4a4ede1d-cbef-4856-88d1-37037bf52dc4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.711501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.711501Z digest=sha256:9903e1fff887e6929685786ce91a7d2ed7bc826152551f0d7d17d2ebdcd3a33a

Observation 1c04d651-2594-46fd-85ea-11af995948d6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.715123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.715123Z digest=sha256:603b886ab7346ad151c3e99ab82f02599028d23f3e11ba094f7444d2841b4f4b

Observation dc1bb8a5-b3c2-4f92-8171-e11df334ee29 · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.127466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.718494Z digest=sha256:192599fd1689797f775e07c9367ea0f960b7121a9972d059bbd29dff6a311926

Observation e1e6ac47-a1d4-45d5-b481-7419f0e18e4f · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.117792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.721767Z digest=sha256:bc42fc9a68cafb04e376a6dde11952a52c1c705689b249adcbd3d548cf5bc373

Observation d991f920-9f99-4080-aa22-a09209cc8c0d · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.107695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.725274Z digest=sha256:7bf13394717ff97201d617714ed2eefebb8fda05a06c2a82e58223d7cdd492e9

Observation 2a798e5d-cb6e-4a04-9552-425726e7d84f · outbound

This paper cites Lightman, V.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Lightman, V

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:20:09.096121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.728491Z digest=sha256:9e45e0dcbf03fa63343075caa768e2448ab9230575eb968fa78724ab9c0dc625

Observation 24c87849-2400-4ee0-936e-896596840f92 · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.085664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.731558Z digest=sha256:94eda304eeecf49d3c78205d894a612f30525f96948a0c0916e5644ac610bf5f

Observation e4895cd6-cebd-40c9-b749-ece181cdfb8b · outbound

This paper cites Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.735352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.735352Z digest=sha256:a30545f4182ca942ec83b5e203bb3535d11db01f67f76fe7018385846a932d2d

Observation 6d711fa9-e3f3-4e83-a04e-2a99c6136159 · outbound

This paper cites dKV-Cache: The Cache for Diffusion Language Models.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models dKV-Cache: The Cache for Diffusion Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.739938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.739938Z digest=sha256:bc339c798af4b7858bb6a88624e8153f9f1a1abd36cdd24215727a93c6891068

Observation 513cbfe5-4811-442a-b5f1-0ad45ee936e5 · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.075012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.743867Z digest=sha256:3e1ba00862943e43cfc5596788109c1f548d840fa6d79c35e6d42c894c541218

Observation 1e453ca0-cb34-458f-9f89-eb0f03040060 · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.063600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.747038Z digest=sha256:156d2c2ae3ad6d5abc1e4e43f06cebd7b9bb366c515dba7d026b0606a46a3e19

Observation 720df068-b430-4be1-b754-bee7492f274d · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.052684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.750162Z digest=sha256:fb1cd20a779dc0939a7169329357a0b7d39c0a0415f775520706067a41e500f0

Observation 4dc31006-ed98-48a3-9518-22812ce1d8f6 · outbound

This paper cites Arriola, A.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Arriola, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:20:09.039973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.753457Z digest=sha256:a21a9d192dbb0f3c1591f9780ae14fedb44ad7f5cf86196daa5a27458a1aa694

Observation d50e75eb-12d9-417b-9c14-2a49e892f2eb · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.756727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.756727Z digest=sha256:4a4aa0ad671d5d5d3c2f30c48256f883684c83e98869f5612963e5ad0a3c3fc6

Observation 10406f89-3dfb-4774-a6fd-e8f1e0b5ad0f · outbound

This paper cites Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.760164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.760164Z digest=sha256:4f1cb09025c94deb318468402db29a6dd1c790b99085791685ec3c51eb0f2462

Observation af939b89-e9d0-4978-858c-0c7f00239991 · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.029010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.764112Z digest=sha256:deb66b0077447ba23a1134da11d893a8c185409d0f494c6f28795fcc8a6b6243

Observation 4ad35dd3-0b95-467f-9b14-95345db8f117 · outbound

This paper cites DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.768240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.768240Z digest=sha256:f97ccca1d7ebe90b43060597138b09daa754290d4cf624e0cc0c8fef1c655bf5

Observation 9800470a-af6c-4111-adcf-cae611f4a994 · outbound

This paper cites Huang, Z.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Huang, Z

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.772667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.772667Z digest=sha256:284593029ea8b36450cc3bd5a0e199c96069968b9483a1accff201caa400d4a6

Observation 543211e5-3723-4a51-9c8f-0a0dc9d70c27 · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.776663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.776663Z digest=sha256:a28a9eaeff5c4f58c90a21bb0c4b5bd12ec86dd7bad7b76262849068d171939f

Observation 0b3b66b2-8bfd-47d4-815e-adb0f65f1061 · outbound

This paper cites Inpainting-Guided Policy Optimization for Diffusion Large Language Models.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Inpainting-Guided Policy Optimization for Diffusion Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.780638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.780638Z digest=sha256:bc333b587f05b2d8691cf5bc42c358e180814ca1c4c5f883c12d2f17786b50e9

Observation 52a38966-a95b-4c64-8bcc-0026524ed239 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.784761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.784761Z digest=sha256:1b43732927675b324b3dd4d9f2ea4dbc26a00488560089f6b24a2ffcef5531dc

Observation 80e1e70d-e95f-4710-a360-df49cf06192f · outbound

This paper cites Qwen3 Technical Report.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Qwen3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.788734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.788734Z digest=sha256:cc26bef81bd1c9fbd1e2c858a384954451e1c222b73da9277b4ec23d165987d0

Observation 99e55432-bda8-4b63-8efc-135d52b3891e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Evaluating Large Language Models Trained on Code

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T22:20:08.792490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:20:08.792490Z digest=sha256:898cad14a96ed306f64eee751bf4c9b287e77a677117b932fa5fbfea7b42da41

Observation 5a4d79b0-2947-4878-98e7-95b7ba19befe · outbound

This paper cites an unresolved cited work.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:20:09.017340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.796282Z digest=sha256:dacdc5de524d589504fe2cd1ca8d06100e966a095403c44d35b2bc09de9b7f40

Observation 6c0f16d9-8304-4881-917d-e3dc654cfb6d · outbound

This paper cites 24 PREPRINT, 2026.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models 24 PREPRINT, 2026

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:20:08.994240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.802604Z digest=sha256:e310f99e03b08633a4f93fbc42ab066ba2de9ddee5d11b328488d3b6ba0067c1

Observation ebd7d880-1ef7-4521-a9c0-cf257f86a566 · outbound

This paper cites degree in the Interdisciplinary Program in Artificial Intelli- gence at Seoul National University, Seoul, Korea.

Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models degree in the Interdisciplinary Program in Artificial Intelli- gence at Seoul National University, Seoul, Korea

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:20:09.006886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:20:08.799404Z digest=sha256:9270d8aa99d83aeab3e867fe17a5e4d1456731ffdcce14abf9d2353b6d329875

Pith citing papers

No inbound Pith citation observations are available.