Pith. sign in

Paper Citation Record · LEDGER

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

As of 11 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 100 inbound Pith citation observations for arXiv:2503.14476.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.14476 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T23:33:10.824995Z

measured 141 of 141 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 1071 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:35.961914Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact12
  • verified fuzzy27
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

7
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation bf69f2fb-4194-4775-b396-3b04c8c958fa · outbound

This paper cites Learning to reason with llms.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Learning to reason with llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.447246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:e80d3aa545ff7e5a1f71ae853ba1474b693beac5bfb458943067292a68fe5a0f

Observation 82c25e79-80e3-4b0f-b478-56ec6b11807a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.475103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:fa268ee295a7c74c570d424550c8b3d332f16124e3fc1d3adc09a261faf89e1c

Observation a62947e7-0eac-41a5-8f6b-cdaa2a572af1 · outbound

This paper cites GPT-4 Technical Report.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale GPT-4 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.466182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:4324b791988c920574ed7d5c2084c5cec893e8f211f63efd7105b73a40a2fbf4

Observation 7e2e3413-76a2-4acc-b0fe-7b40996b56d7 · outbound

This paper cites Claude 3.5 sonnet.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Claude 3.5 sonnet

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.443465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:9fac7db11eada33af0b23106e884d820ea88f5d477ddced167ccbd7c56645c2f

Observation d2d6d9f1-71d9-4df4-aeee-716c0f95e2b2 · outbound

This paper cites Language models are few-shot learners.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.426165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:4d27946a142c7087438b1b062d9b907f58822accb156c513e7f880b5b038aea4

Observation 781fd30a-94ab-44fe-8414-b59507835446 · outbound

This paper cites Palm: Scaling language modeling with pathways.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Palm: Scaling language modeling with pathways

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.430899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:05af1e406c05a2a3256f00ec39480a2a903594270b57c2bfddd4257fd394c146

Observation f2ddb11d-d609-41ff-9aab-a80de5b2b064 · outbound

This paper cites DeepSeek-V3 Technical Report.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale DeepSeek-V3 Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.482523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:58cef0ba207cf4b0fbdf2d0ffaaba26b9b1346f4dda9ab3f638c98dd94df9ca8

Observation 45bee8c6-9b58-4c89-ba51-2e1c93245934 · outbound

This paper cites Grok 3 beta — the age of reasoning agents.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Grok 3 beta — the age of reasoning agents

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.435347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:1f966385d08af761bf5bcdbd683076e131e1b6d8192dd4437bb870ed9d79df84

Observation 73206345-4efe-4019-9a7e-a589635c6f66 · outbound

This paper cites Gemini 2.0 flash thinking.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Gemini 2.0 flash thinking

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.439254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:0a6351970cfeb6a84276dc7234d799c5b3bf96813e472bb9420e4e2a326ad766

Observation b8e585ba-663c-4daa-91ff-39515fadc3e4 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Qwq-32b: Embracing the power of reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.421829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:50be570c2a566fd3e057e4f01a02e05afa6f4f0b6401463ca26d62e04cb2f420

Observation ed0712cc-0800-4619-b28c-bdfa039ab602 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.448053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:30c13a5fb40c4f06263001357adc1ad13e20758ab45aef8bbe03bfef3a2352cd

Observation 325e40a2-7203-4c23-a302-e01e0e0d2c30 · outbound

This paper cites Qwen2.5 Technical Report.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Qwen2.5 Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.440483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:e503eff9246ccc57afdc510792d751e9abe38e775cffb6193984de2ee1786398

Observation f33de9cf-eec2-4453-83db-1f09bebf594e · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:35:13.444477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:d9471354804c1a84e964d6a41dd64f236202746e01dfbcc05517742c0f07c881

Observation ac6bb475-0f2f-4fbb-b2b8-202823f44509 · outbound

This paper cites Open-reasoner- zero: An open source approach to scaling reinforcement learning on the base model.https://github.com/ Open-Reasoner-Zero/Open-Reasoner-Zero.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Open-reasoner- zero: An open source approach to scaling reinforcement learning on the base model.https://github.com/ Open-Reasoner-Zero/Open-Reasoner-Zero

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.417671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:337f8816c367fe929d6bceaab05aab9bf14865e734c049a1b1f2f00662a949f4

Observation 2ab05f3a-37d7-484c-9821-675e734778c7 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T23:35:13.452225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:3915bac7d202ea93f6a7955f351edf0098d533e65a3010a140bbb93f4e4ef113

Observation 472c55df-44bb-47d7-bd87-d01c7ce34dbe · outbound

This paper cites Process Reinforcement through Implicit Rewards.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Process Reinforcement through Implicit Rewards

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.462057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:b6a0e16188b2ffd7518744a9352a1e4f258e38df3283b8b87716ae0257113148

Observation c4faa081-3ba1-411f-b4ef-6dbde192dc74 · outbound

This paper cites Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:35:13.431013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:e4ee1fb8b0f4da4b21b2dea5a062eeacfb9825d3f4c3ca8e5e7f5c150f1579d2

Observation 531ef94a-12a6-4835-b0ef-ade283579f4a · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T23:35:13.436228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:bdaa7d513630c73181fba59ba23979d6cc52f59cf7b2be4148b43678c9ac2edd

Observation 753f2d48-65e1-43e0-b6fa-b36884c6ae38 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:35:13.471029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:5e2a00b8a3b4211905d2003b80104864a2b1acf7d2e4b33e016de13195af6b80

Observation 8bee5c55-19b4-44cf-9ad3-9094548446d2 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale HybridFlow: A Flexible and Efficient RLHF Framework

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.427523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:cf934d3ec73fe7a37ce3cc1f80c2d21be829a608b897ff824a7ef174130da9d8

Observation fef29dde-0eac-4c94-8e8d-9833432dab40 · outbound

This paper cites Proximal Policy Optimization Algorithms.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Proximal Policy Optimization Algorithms

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.479105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:734ad9e9e8a2912316115905be72279ad295d607d3c317e5bc102a098aaa4935

Observation f96911f5-35ca-4ec7-8304-23d64f07c98c · outbound

This paper cites High-dimensional continuous control using generalized advantage estimation.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale High-dimensional continuous control using generalized advantage estimation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.412096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:85f74dd6bc2c5facc2ac656b583b485682c6870a5161c8ca501c969cfb27897a

Observation 6304fbeb-1beb-4138-8a50-af1a54035d54 · outbound

This paper cites Training language models to follow instructions with human feedback.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Training language models to follow instructions with human feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.406998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:65620607d553c13a451526acbbd40884ac1f8d962233ceed31199bd39067269f

Observation f3c48e26-a5c1-499b-9a92-0ad379b0609b · outbound

This paper cites Concrete problems in ai safety.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Concrete problems in ai safety

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.522375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:d58b13d7677526180c3632488751938255c2e773c84b417869ccc526272c5d8e

Observation 976fcbb8-c2c8-41c0-8565-87a155f023f8 · outbound

This paper cites Reinforcement learning with a corrupted reward channel.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Reinforcement learning with a corrupted reward channel

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.518183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:a4b835dbd580c529b36dbc422c000619ff3c352c2da4c4662e227e35cbb67523

Observation 55f8327b-5fd0-4297-b244-9ec0bff2eb58 · outbound

This paper cites Specification gaming: the flip side of ai ingenuity.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Specification gaming: the flip side of ai ingenuity

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.513859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:b770482fbf7aa37d5f0300d4f53cb7d9fe073a2aa90f1c2a9b222f2038a24857

Observation 2bad6074-0ed0-46ae-bd29-02b1d33a1ce6 · outbound

This paper cites Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.508996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:3fe9fad9e6aed1fcb768224994b8f6847486a001b2bd23da3b5ebe33241707c4

Observation b090dab5-ae1f-4438-9355-0cf49c4b0c52 · outbound

This paper cites Scaling laws for reward model overoptimization.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Scaling laws for reward model overoptimization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.504569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:5609769052015bbed17c8f45d08eeafe2c8ac1ee458695c6e2503a075412157f

Observation 7f0532fb-c2af-498b-95e4-5e53f5f0cb4b · outbound

This paper cites Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.500709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:da8c97702c6f54fd438dbc441d0f15be07b93473f79fa203e167b8731dcef33a

Observation 7f3f4976-6270-472f-a995-54cca691ac76 · outbound

This paper cites Generative language modeling for automated theorem proving.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Generative language modeling for automated theorem proving

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.497137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:2af590a8a5fd79f04f024e16cce89959eb32bf7a9e6e5c21e3193b5faccf7e32

Observation f776b224-262a-4ee5-ad22-abedf1d49b85 · outbound

This paper cites Solving olympiad geometry without human demonstrations.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Solving olympiad geometry without human demonstrations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.493724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:ef10255ac713c21dad25324cba2f09316cdae66894159e6b83170f61a800d9e1

Observation 9671c03c-7442-49fb-a3af-3a94e31bcfb5 · outbound

This paper cites Alphageometry: An olympiad-level ai system for geometry.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Alphageometry: An olympiad-level ai system for geometry

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.489562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:9477b152890ef6497b90674de1e32bb55a73d32baf9da17a3992f8b81a18c3b4

Observation 0c30b8fa-a4e9-4f5f-b6c1-e8f621a071bf · outbound

This paper cites Ai achieves silver-medal standard solving international mathematical olympiad problems.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Ai achieves silver-medal standard solving international mathematical olympiad problems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.485500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:be188c9b59e6fec7a95b9c725dd10d57112b48b50ec42c1955c24e8f7aaf6d81

Observation 8c529730-701f-48a3-8912-a09bac9d6ae4 · outbound

This paper cites Coderl: Mastering code generation through pretrained models and deep reinforcement learning.Advances in Neural Information Processing Systems, 35:21314–21328.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Coderl: Mastering code generation through pretrained models and deep reinforcement learning.Advances in Neural Information Processing Systems, 35:21314–21328

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.481614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:3cb04825c7931f6217dfcd475ccefd686d22564c36e4c8a9d925cd48b6c82b1a

Observation abec4026-a124-4eb8-bfe9-ddb15ca94714 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Reflexion: Language agents with verbal reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.477006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:e38d231d7637dfd90f909c9cf3c1859b1e1b4b0f420387a41e62e26df108a3e5

Observation 36a7154a-69cb-4bb7-9447-d47ad0bde6ba · outbound

This paper cites Teaching large language models to self-debug.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Teaching large language models to self-debug

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.469675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:1b22b61b9c4e93ca4da7ae0a28a9b6253d138345cdbd830939abc84bdb163e33

Observation e4f32f3a-dbf4-4208-8944-6ac7fc227847 · outbound

This paper cites Rlef: Grounding code llms in execution feedback with reinforcement learning.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Rlef: Grounding code llms in execution feedback with reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.464115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:a2ead24a67b0875df3d747e85eec93c51efa428560af349d0181626788db2d03

Observation fc980eca-952f-4e1e-b465-ddf65d8fdb91 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.457464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:cb41468b0cc2d2d78b5eceefd6b910f9c4ecb5f3bd07dbfa60418ea9218c393d

Observation 14fda047-4c20-458a-a442-4e2da1a09b2c · outbound

This paper cites Decoupled weight decay regularization.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Decoupled weight decay regularization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.459927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:00564ecce7ba52b3142251a9bd239ddca6f4b16533153fba9ad78cfe14b9b57f

Observation c51a8b68-f83a-4f56-99fa-7e5af95f863d · outbound

This paper cites First, note that the answer consists of an integer part and a square root term.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale First, note that the answer consists of an integer part and a square root term

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.456109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:0262b024a1e5f234d57cc1e5aa1886d227d848d26927ad3e8d219d77ea9a5e9a

Observation c0e2ce93-1364-4eca-9281-97cb2b63dad3 · outbound

This paper cites Let B be the set of residents who own a set of golf clubs.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale Let B be the set of residents who own a set of golf clubs

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T23:37:16.451910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:2347c8de0bda163b63faa670c9c56b0ecb8feee4fbdee161ce8126e246a59c5f

Pith citing papers

Observation 734e70fe-fa22-492f-b07f-a61db578ec65 · inbound

OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework cites this paper.

OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:28:57.181180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T03:28:57.008431Z digest=sha256:a7157ec29107e8b4fdd0e907eaa2f1891749819b69bdea3503e602d6527a4eb4

Observation 5ba38af3-159d-467c-9c25-3b7f99c5bbf9 · inbound

Learning to Reason at the Frontier of Learnability cites this paper.

Learning to Reason at the Frontier of Learnability DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.203304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:acfe23c0ffc271ffc78e2c290519b755627b48fc23c4c0161d3fc3c716bf63e6

Observation 88fb637f-2325-443c-82ff-053f31c17ad3 · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.501096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:d64c01d072ff06d044d9e64a55108f0dee7fcd95a54c5c08e44e9146f29b7c96

Observation b801f8f0-ec35-4899-9195-21f9995690c3 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:59:03.321024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:e861431e6d6dd7f24498c8c4edc14732c3103ada6b915af92d23c40cfa4acd49

Observation 05d07834-058b-47ae-8bb5-ea8b5f0a98fa · inbound

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks cites this paper.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.788911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:4c17bcdaf57ac20fa85cc406255481f3f59f093327fdcfd115f5b0a7e37eefb0

Observation 9cc0c66f-1531-4f89-8ac8-883264aa5fc7 · inbound

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning cites this paper.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:31:04.843987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:121527dc6cef57b2a8571de15371a8ba27408e8ca29c6c762e9c1d165c44fb63

Observation 9a6a2c45-d65a-4049-aece-f5cc7120eacd · inbound

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning cites this paper.

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:46:56.858678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T18:46:11.571831Z digest=sha256:a6dc18d54775bc3ba93599fae22b24b0273019d6e413bd12a06e978ef533e1c5

Observation 43a80739-3434-4af7-8f5c-87ab5e5b45ba · inbound

ToolRL: Reward is All Tool Learning Needs cites this paper.

ToolRL: Reward is All Tool Learning Needs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T00:26:48.564435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:26:48.291431Z digest=sha256:d28ccc5d4c9080f3671c01a36e30a6a5894245578de985d81e62a0fdc3280893

Observation ce80f86d-eeb6-4241-9a09-21e0f4e91a77 · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.906397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:307daf2c6bfce2a15055e619f89446826ce2be050da2d366015105ee2c7e9cfb

Observation 26fc22d3-9293-4848-b07a-8d6b20e8bb00 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 166

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.824151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:0d3a0b3c70e6093a99e2bc1eabf04e95a98331b21f5d9f4177ca49f5f73cab46

Observation c7cb80a5-9253-417f-a78f-1b3c45744716 · inbound

Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning cites this paper.

Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:51:45.685027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T15:49:44.263123Z digest=sha256:5a4b90c92da4d8e7d101362aff73081da1666c5f9a9b2b0b2c816e73b76baf9d

Observation c9ae3b39-ce40-4366-bf22-bba134d03aec · inbound

Group-in-Group Policy Optimization for LLM Agent Training cites this paper.

Group-in-Group Policy Optimization for LLM Agent Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:15:08.421797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T09:15:08.193357Z digest=sha256:a765a400c2140ed7ffde30aadad8af15cf3c8b4fbfc7008ceb922b2a5e152f5e

Observation b778e018-1636-4f23-891f-dbaaaa36b509 · inbound

The Hallucination Tax of Reinforcement Finetuning cites this paper.

The Hallucination Tax of Reinforcement Finetuning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.961914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.961914Z digest=sha256:5754e22a32fc1c20ce0f1a1a43868adc235355a17f47ef07364726b6f9203e53

Observation fa17cf45-45ea-48df-8151-bfd00c7d1077 · inbound

General-Reasoner: Advancing LLM Reasoning Across All Domains cites this paper.

General-Reasoner: Advancing LLM Reasoning Across All Domains DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:06.180306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:06.180306Z digest=sha256:aca29f535c35a7ceecb56df055a0bafec25357c632aee992619dce922d736f7f

Observation a6fdcaf4-4f08-469b-9515-792345129f24 · inbound

Reward Reasoning Model cites this paper.

Reward Reasoning Model DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.259571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.259571Z digest=sha256:df61fc7a09d1efe6f1c509331c1ecee15550810799c2813b9fab636cf862b37c

Observation a25740f0-4537-4d6f-8f54-d60862085901 · inbound

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models cites this paper.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.407955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.407955Z digest=sha256:5ddf949ff677c56f7db883d8f133b056885bea2384583ace50c85323b4a16398

Observation 33aab401-0b34-4dc2-9521-4f7dd023735e · inbound

An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents cites this paper.

An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:54.687871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:54.687871Z digest=sha256:fee9f52a7ede0c1cf50e56b9dd1f8875636b7d9f7a44c311a9efc6ddc52d37d4

Observation d17c69fb-0447-4821-94ee-8b7e972aff5c · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:02.445271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:02.445271Z digest=sha256:a00e33ad0647adc36be6e9209b471b93361351d32b75bdc4b715375de8db8e0c

Observation 41384598-14da-41dc-ae07-b941ab3cd98d · inbound

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping cites this paper.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.954245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.954245Z digest=sha256:bd6b6721e2be9a0280ba19273a3caef05a7d8b8aa7c82929bc6adadf6f1dd3f5

Observation 474c2c7b-1cc0-44be-99c4-f81dfbc43091 · inbound

NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning cites this paper.

NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:38.902088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:38.902088Z digest=sha256:d919381b5a846589c41c3b07bd00daffc2fa60be47c601de32f12e5f104e77d6

Observation f9423a99-3bc6-4f52-a8cb-ee85a07db547 · inbound

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay cites this paper.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.758735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.758735Z digest=sha256:617f9a74ee44a57b638f476b007b5f5844075f48afe2025e39e1eb9f4f481af2

Observation 510d0643-00d6-4fef-aef8-1ab9ad7221ec · inbound

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning cites this paper.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.487503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.487503Z digest=sha256:42aab4b6578f6ffb8a3fd94aced2fb1641639677ccd3008290f891abab13a497

Observation 4ee74034-adfa-4c42-8c6b-7a8371bd8dbd · inbound

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning cites this paper.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.462641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.462641Z digest=sha256:8458d421478b91ae1cad47608bedd815521977dd4f737367fa42d9dcbf519aae

Observation 6d5116c2-3996-4605-802c-33106f2444d1 · inbound

O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering cites this paper.

O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:52.623494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:52.623494Z digest=sha256:de41086a39f9349423b48e2d5f8a931beff394ee03eafa316d85bfb99efe6366

Observation f137afa9-0a97-4d25-9cc9-51b11a59b0c6 · inbound

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO cites this paper.

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:38.625405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:38.625405Z digest=sha256:97df167b0a6533c45e55f5debf90443164e88c074460acfdf415e7d95456ebf1

Observation 10cbe391-988a-4570-9c01-550f63ae962e · inbound

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation cites this paper.

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:20.133374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:20.133374Z digest=sha256:051438e7431500a702f77ca9126bd676f3b78ab84591efc44de43ee23a56861f

Observation e629553a-ec71-4f4c-aeb8-be8a9c61d419 · inbound

R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning cites this paper.

R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:17.321595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:17.321595Z digest=sha256:c1b58af0967d3077e7ef179abaecfe0c00c01a73403b3cf0bbe866cd60edd2ec

Observation 249c4d65-1e62-49b1-8920-e501a572bab4 · inbound

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO cites this paper.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.192183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:56.192183Z digest=sha256:7299c01bea2de4c18149ad1c2d2caadeaa73b54645015f422cad70faa81f55f6

Observation 62fb29e7-735c-4b34-9cc0-4231cb58ef34 · inbound

Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling cites this paper.

Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:34.445900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:34.445900Z digest=sha256:7b9077520b28ae80e16cd9ce60133389fd36592554bb70247420112e5cb27d7f

Observation fd620298-cab5-4451-a2aa-84df904a8cb4 · inbound

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning cites this paper.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.379000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.379000Z digest=sha256:a0794b3cf4c40f2d412a9ede232c0e55d6405c4830f576b4bd37a2e8f0727afd

Observation 63cbc4b7-953c-499d-986d-d3fc73d92757 · inbound

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning cites this paper.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.685105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.685105Z digest=sha256:3f10f3c256c1efd44922659232a2741355150f511abe302a5ae322369d70ba19

Observation 4b71ca5b-2ab2-46d5-8779-29b6180e8744 · inbound

Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals cites this paper.

Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:39:54.965005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:39:54.965005Z digest=sha256:a666232ef05ff17eb5f4bd1b9f8c5bf143e50d944bf69b41cac4acdebbee6f30

Observation a19cde46-1023-442b-802e-b09dc1152cda · inbound

Formally Solving Answer-Construction Problems in Lean cites this paper.

Formally Solving Answer-Construction Problems in Lean DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:10.188950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:10.188950Z digest=sha256:bf04184c99ac7194445019d40354316b1da1e7ee6383aade4ffb9305f37b8ee9

Observation d9a22cf9-07fd-4372-a117-a775ac9d9bc1 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.705003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.705003Z digest=sha256:66cfa705923466dc8f276161bab51f7c611e184abeb9c195feacb5e3d2c236f0

Observation ad890374-dd81-4fc5-9731-dd33c7d49399 · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:55:40.372850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:c0f6c474e013a18c4f1bbe70a67a42bd5e190326056a6af27f33866536eefd0c

Observation 1e7a4dd1-a25f-4783-b71d-bac8905ba860 · inbound

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization cites this paper.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.323315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.323315Z digest=sha256:cd3455371a0142945b0019bcc3ea2adc7c5c564aeebdeebfcbeefb5720561fac

Observation 69316f9c-050b-4ccf-8937-93ae51d826a2 · inbound

Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs cites this paper.

Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:20:52.124513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T01:16:51.288077Z digest=sha256:8ec573d87605a9ea11c20e9ac3e87b01d04efcdf831c26b008ac14b5062ec2c2

Observation 50d49ff3-dcaf-46aa-922e-d6ae27d2001d · inbound

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking cites this paper.

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:21.123794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:21.123794Z digest=sha256:4948970880cff656c4b7344d547249627dfd7303272aa3d110b8351abe777ada

Observation a53ce49d-145b-4e6a-ac43-cc76e727053d · inbound

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning cites this paper.

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:36.086883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:36.086883Z digest=sha256:4ab891ef721c8ad59253a61ccde0b6433b45ef5e5d6a697153e69de985d8c9f6

Observation df8213ea-708b-44fc-8778-90e3ff9b27be · inbound

SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond cites this paper.

SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:33.391350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:33.391350Z digest=sha256:0a6153796cf03b522f4655439367e263af45edf150fbb034fb7414d02dacb59d

Observation 2ac65f12-04af-4f88-9429-c02619ba8415 · inbound

Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective cites this paper.

Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:12.741388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:12.741388Z digest=sha256:ebf6309e060fef32ca810365765ba429e23965e7a8198cf5af114411776cfa00

Observation 798cd88d-0aea-4720-b162-2f587ba6f6d6 · inbound

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles cites this paper.

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:08.074540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:08.074540Z digest=sha256:b56d0ede7d02462fbb23ea7519a79c862fcf27c13caac5fcb6fc00feece92ce2

Observation b333d74e-723b-4e2f-9f79-a3f8e37df4f0 · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.367525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.367525Z digest=sha256:29d4c4e30f73d1c817d744fd72bb199165913e576f3d2a5a1f1247c0f0f61edc

Observation d5e259d3-2194-40a6-bdbb-67cc08680aa4 · inbound

MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability cites this paper.

MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:16.157460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:16.157460Z digest=sha256:9bf6290e7e38bc3c240e86d8e996a7cda6b7fdfb3e79c19d136102de04f0ba1b

Observation 1e79e282-ae10-462a-b77d-59ed721a99a5 · inbound

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence cites this paper.

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:09.183758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:09.183758Z digest=sha256:5933e6c41a6b5eae26f22e9063055a61c4042cd95890f5ff67dad487fc403ce5

Observation b5f5ff59-2f3b-4638-b652-9df06b60a614 · inbound

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning cites this paper.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.382950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.382950Z digest=sha256:c97c43dab9e2dc7b875abc6473460b581cfb9975c46eab52f4a9d164e75a8d6b

Observation dfbbdad7-3e5b-4f76-a2f8-96a7f26b2e10 · inbound

Reinforcing General Reasoning without Verifiers cites this paper.

Reinforcing General Reasoning without Verifiers DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.837737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.837737Z digest=sha256:e3cc61de2d98c0f49a7a47e45a6fcea0029f1cdd98260851790fcb772ab8967c

Observation 89a3db25-1141-493b-a92f-89c2556b3b23 · inbound

TabReason: A Reinforcement Learning-Enhanced Reasoning LLM for Explainable Tabular Data Prediction cites this paper.

TabReason: A Reinforcement Learning-Enhanced Reasoning LLM for Explainable Tabular Data Prediction DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:16.275086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:28:16.275086Z digest=sha256:296d8e3d9c5b07f5788f39b85eb7269827e537786d1eedad537f18f5acd88613

Observation 3ed4233d-21bb-4557-800f-f04794c7317e · inbound

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training cites this paper.

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.947416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:54.947416Z digest=sha256:7cae56e76767fe147ab7f76f0dd2b91bc85848dcb7ccddfe7a248aba1e8677e1

Observation 9dd26810-09e8-4b80-a6fc-02256ad6fec6 · inbound

Skywork Open Reasoner 1 Technical Report cites this paper.

Skywork Open Reasoner 1 Technical Report DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:26:47.368219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T04:26:47.283983Z digest=sha256:1f7c8c631b2bd9c0edf9dfdabeb53bd3d707c1aa4d67156e5607de859d37d121

Observation f8a98493-4dc1-47da-a516-9671f54832fd · inbound

Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition cites this paper.

Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:18.016091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:18.016091Z digest=sha256:acdd65886aa4dd8898493bf61fb94c9ac39a1d0ca827581348aa689cdd699b56

Observation d53205b8-a3e2-48cf-ad19-68ecbb65810e · inbound

RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning cites this paper.

RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:33.280845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:14:33.280845Z digest=sha256:d6ba496d562e3a0421344e4ae2fa7dae471434f54ef7fc72c31df6a0c55be2c5

Observation 6460a496-1546-47db-89c7-7bfca8d0211a · inbound

EvolveSearch: An Iterative Self-Evolving Search Agent cites this paper.

EvolveSearch: An Iterative Self-Evolving Search Agent DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.475979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.475979Z digest=sha256:89262aff4ea6e783bad6056b4e93f3c52863f33383a949c87b791a13f000bbcc

Observation 820e3921-aa3e-4902-b365-a3e0e076264f · inbound

WebDancer: Towards Autonomous Information Seeking Agency cites this paper.

WebDancer: Towards Autonomous Information Seeking Agency DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:58.025183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:58.025183Z digest=sha256:f29fbbd74f7e2358b8b8cd38edaf17a82579da39d5026f0bc814b31ef96979aa

Observation aa45b29e-9039-478e-b74b-7f210b2826c2 · inbound

Accelerating RLHF Training with Reward Variance Increase cites this paper.

Accelerating RLHF Training with Reward Variance Increase DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.554997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.554997Z digest=sha256:5e58456551607ba506193882cffe730766816710021e6593465b048b5d0364ef

Observation d43a438e-6464-4aa2-9aca-a28846594327 · inbound

Are Reasoning Models More Prone to Hallucination? cites this paper.

Are Reasoning Models More Prone to Hallucination? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:32.967235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:32.967235Z digest=sha256:008c30720e6648b1300abe2d7a7db4964bcfc3df45f73d6c714cb0ec2a2a48b7

Observation 681ebc7e-91ea-4c3c-b865-518350f00f3b · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:52.198652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:6e84927a43898f772fd0fc6dbb587aacd79f689a499058d158af5c5a0edd03cd

Observation 2fa84e31-7281-42fd-bf91-ebeed07c9b66 · inbound

ZeroGUI: Automating Online GUI Learning at Zero Human Cost cites this paper.

ZeroGUI: Automating Online GUI Learning at Zero Human Cost DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:26.439832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:26.439832Z digest=sha256:28ca5ec2143d8b68105ff9deb06bff618f5119109af7ad52d545f5f92dc0da8d

Observation 9c2a96cf-38b2-41b9-9ea6-59fc39a26ca3 · inbound

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning cites this paper.

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:13.610491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:13.610491Z digest=sha256:62dc758f8581f97a7051b17c77c68004cd225e646943200b21ca08b11525778b

Observation 9d8616b4-cc9f-478c-83fe-fa1433f61ea9 · inbound

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM cites this paper.

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:33:12.653152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:33:12.653152Z digest=sha256:d989f0660e9ee2f4f5e52fc9494a97730059cfb15e66af6be6b36825c5dc1e66

Observation 79e61d0b-a2c8-4729-9c60-20162a21aa35 · inbound

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning cites this paper.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 66

Resolution
malformed identifier
local_arxiv, observed 2026-05-15T14:24:21.986322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:82b838f72e7c97b6a59da166ab063d7546fcab5046deb053228041acf0165b6f

Observation 97176b73-2444-4d0e-8d7e-1cff15f52cb7 · inbound

Towards Effective Code-Integrated Reasoning cites this paper.

Towards Effective Code-Integrated Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:25:27.818086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:25:27.818086Z digest=sha256:a9f02e6300fbdfe5a3716e5eceb1e0d7299e10b18ce1c88b4fa8ab60ad3a98bb

Observation 605d6a0f-b46e-45a6-b1fd-25588ab8c0ca · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:12.867048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:12.867048Z digest=sha256:b9c99fa02b45340f138556f2dcf81d504c6f60feae6b87dd659a040c960afbc1

Observation b9d0d6fc-7e1c-441a-8cf2-ed4ba973b7c1 · inbound

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL cites this paper.

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:25.125571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:25.125571Z digest=sha256:8298fc9171591c8ea0714dc6d0d34a1d98fc7df6142e5e645bf861adf6e0be87

Observation a3538b41-65a8-40cd-9408-f3867b7e9582 · inbound

RLAE: Reinforcement Learning-Assisted Ensemble for LLMs cites this paper.

RLAE: Reinforcement Learning-Assisted Ensemble for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:19.084919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:08:19.084919Z digest=sha256:7962db629eeddb9053e91b27709adcef357290e7831b9895419eeb8161e6f085

Observation 9af23f9f-4d9b-47fb-8a1b-8f4bbf3eea20 · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.598747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.598747Z digest=sha256:206a47d19218c02478e62db6556d636ae99b7fa75b57bde7d946321bddb801cd

Observation f2a6f081-263a-452c-8fcf-e02f9268efb1 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.363044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.363044Z digest=sha256:191e66a1a3316b21df3afe2def0b1c480d53ec5329bfec1022085eafa583210d

Observation 8ea6bd0d-5eff-42cd-8422-36b9b3d4ecad · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:29.765282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:29.765282Z digest=sha256:25d13ebc1fdde2f47aeedd04f8dd1070ce61f7987063a02581a7d97d77021533

Observation 4ada27f0-ae0a-47d7-ac19-3c065fd19407 · inbound

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning cites this paper.

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-12T12:12:08.990663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T12:12:08.724844Z digest=sha256:ce0bb11124a5c3b9664c807d0a1d0f4877f1eedeff62028cb80bb93032b022b4

Observation 4d6ed220-c3b6-406f-a283-3abb183608da · inbound

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis cites this paper.

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.704809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:35:47.704809Z digest=sha256:88a4adbee2cf46c38e9d69c412f9b66f7f0f2d7aa75b118c46a45cb823a890b6

Observation 209e468b-62bb-4f5a-b75b-ea1eb0702bd0 · inbound

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts cites this paper.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.907899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.907899Z digest=sha256:f089f0b2abb397ade50593432d7072007598522138307ad976ec0d3c547f1f7e

Observation 24986511-57a8-42cf-965a-fad7904e2c78 · inbound

KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning cites this paper.

KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.474963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.474963Z digest=sha256:f26a1ab32da27952d2616c8b6b1e70883214d4e009ab98f369cbc2c7832c9df3

Observation 378849d3-2eeb-4410-9d09-7f42e5337343 · inbound

Improving LLM-Generated Code Quality with GRPO cites this paper.

Improving LLM-Generated Code Quality with GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.358241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.358241Z digest=sha256:3e1a97e6e0f7094baf4fe907634cac5ce53a0cb36f63e1b7fcedfd9c30f9c14a

Observation d2ab2182-d128-4808-a244-e300fc9dbb07 · inbound

Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening cites this paper.

Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:02.597250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:31:02.597250Z digest=sha256:cd19f86f76b96c933b17962d484f972a0fd2cc7575d84362275dc219405e0b0d

Observation 1d3da582-05a5-4454-8db3-fe030f03c885 · inbound

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs cites this paper.

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:37.613949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:37.613949Z digest=sha256:696dc34ba5a76121741d030e5a4de9c61ce79b576d9fea4cc9a64411d0c34502

Observation cdf9bd70-ff73-41df-b817-ebc7834f7af2 · inbound

Seed-Coder: Let the Code Model Curate Data for Itself cites this paper.

Seed-Coder: Let the Code Model Curate Data for Itself DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:56.907014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:56.907014Z digest=sha256:be405efa6095261b8899266e86d5f77aa40b4024f44b091fa1d75c2af90387e6

Observation ae02c8c2-824d-4188-926e-b4aa1678da7f · inbound

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning cites this paper.

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:07.870856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:07.870856Z digest=sha256:64b1cf01b1b519871a2dd04ea3892357586ca8a0a5dc73779c1c721f25104e42

Observation 98430df0-679c-475b-be78-8a36753bf931 · inbound

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models cites this paper.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.799079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.799079Z digest=sha256:2a4a375e3eafc0dfc37f74c666b1dd45f04406d8518868d0d09ec9b7ecef4a0a

Observation 71e6ada2-e435-4c2c-ac6f-b27145cbf544 · inbound

QiMeng: Fully Automated Hardware and Software Design for Processor Chip cites this paper.

QiMeng: Fully Automated Hardware and Software Design for Processor Chip DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:04.451704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:34:04.451704Z digest=sha256:cb3448be7a3904d95da4146aa0cef6751a2a3650f226a8b7b62f58332e2c6358

Observation 5038f393-df94-4407-8ad9-d8331a54f79a · inbound

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning cites this paper.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T10:53:02.882074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:d566231dcf55ed9f23a5067771f9087992ef089e35974ab1e9cbfac9f32be96c

Observation 381b6018-32b7-4864-a7f0-3657153c080d · inbound

CodeContests+: High-Quality Test Case Generation for Competitive Programming cites this paper.

CodeContests+: High-Quality Test Case Generation for Competitive Programming DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:53.969387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:53.969387Z digest=sha256:19af3ce7b79cd2dd4fb076d9e9b5f50c3198fd8e4eeeaf261780b7268079f275

Observation 2aa1046d-820a-47b7-93e7-aa2b8b71956e · inbound

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library cites this paper.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.052451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.052451Z digest=sha256:a6764c81a5098c3960498426c6c2a85a4abd56da21c139df4fb6220fd06baeda

Observation 6322f138-1f81-47ad-8802-2bdb1a21490b · inbound

How Far Are We from Optimal Reasoning Efficiency? cites this paper.

How Far Are We from Optimal Reasoning Efficiency? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.129855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.129855Z digest=sha256:37a27b200979753f91d94bc5450d6b3577af592aa30434503a670ffd7bbf6d5c

Observation 32bd10b1-7c45-47e1-9654-29027edc5667 · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.224328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.224328Z digest=sha256:283497a70eb9ec6ecb9955b903aa8ee95cdece74fbd968f9439717849984fb7c

Observation 0505c5db-18fa-4fd8-bd05-48cf231bc2a9 · inbound

MiniCPM4: Ultra-Efficient LLMs on End Devices cites this paper.

MiniCPM4: Ultra-Efficient LLMs on End Devices DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:22.420798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:22.420798Z digest=sha256:2846056331318cc46acb750651ffe441bf1284fe537a2f26617732499073c1db

Observation 1605fb19-019b-4012-bcce-2f1da7f47d0e · inbound

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards cites this paper.

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:27.211480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:27.211480Z digest=sha256:9ea46c4ec3ac542710fe646dd0bba389efac37d1d41f94b108c634e6bf7e1a77

Observation 99f5089e-12c0-4f1e-b071-e2dbd9917abd · inbound

Reinforcement Pre-Training cites this paper.

Reinforcement Pre-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.142165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.142165Z digest=sha256:31cf71e4599c07a3dacb36f5e047660e5810c3a859d65af027ff622b8d654aa1

Observation d56dc75e-cda6-408c-84b6-5867ff745e4d · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.543926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.543926Z digest=sha256:c83498e3d02d996c2003248bfdef922e81dd8cc6dbe5f04700dfeb4f85b06329

Observation 56f7a275-2495-411c-a8cd-15faa612de94 · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.745522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.745522Z digest=sha256:176648dda815d0b95ee811aea549142d7a14057ed85dd219ee9c6185411674a2

Observation 42b9d5fd-007a-4d0f-8532-9d229f7347a2 · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.263254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.263254Z digest=sha256:d0acc5cf790c2f936d614f02ebc2ffa3fe0af00f788c92b746d8cfcc72c43100

Observation 8b61a26d-890a-4211-83ab-ee5887694417 · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.861106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.861106Z digest=sha256:4fb057709fcbbdfc96e2c1ac59ff13bd8c3e8b9fc847361f9fee61e25656f636

Observation cfc66114-dd6a-40db-8a4e-b2fb90b34213 · inbound

RePO: Replay-Enhanced Policy Optimization cites this paper.

RePO: Replay-Enhanced Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:58.288675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:58.288675Z digest=sha256:a513627cf3a9e6f31c33cf04c8022e1d651a47a59aa9abad3fb674bcebdb1f77

Observation 2908b548-4c0e-41ef-a3f6-4d729ffb6861 · inbound

CoRT: Code-integrated Reasoning within Thinking cites this paper.

CoRT: Code-integrated Reasoning within Thinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:22.944112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:22.944112Z digest=sha256:31711ed5665e86e3ba2c029eee06f599c4a379caddc21a69bbf064c0adceed28

Observation 02b7edf3-4df0-4b51-9275-cd2534582e1c · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:38.188492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:38.188492Z digest=sha256:38729e6f60801174e0765035089074c3c4cda7f2e670aba443795c90525fcfd8

Observation df998082-8d39-4586-bd26-e31f6501b6fd · inbound

Magistral cites this paper.

Magistral DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:26.415218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:26.415218Z digest=sha256:88e5560a077c725233c2c231430c9e3dc4ab705e91bbcbd3df3929db8ff778f7

Observation 907799d7-17df-4531-9005-ab4ee4715382 · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.797292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.797292Z digest=sha256:0e42419f2c6c2b159c62f4a21192646f301b641442833b2c78b558a1caa85d9f

Observation 30ab0835-abef-46be-821d-04b704d77905 · inbound

Schema-R1: A reasoning training approach for schema linking in Text-to-SQL Task cites this paper.

Schema-R1: A reasoning training approach for schema linking in Text-to-SQL Task DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:26.776323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:26.776323Z digest=sha256:43c567c0fa1c4dc1c840ee1aff9b00e79245045efbfa4fba3aa2cff1bf089c41

Observation 4fe48f52-8d98-45e2-ae22-e99dab76d1b2 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:35.824398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:35.824398Z digest=sha256:85350594da835d23b9192161dd343f6a9fb025323333fa6dd5d654fb08bec316

Observation b93c90ad-17e2-412f-aad4-b3c5f5b4155f · inbound

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy cites this paper.

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:36.411140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:36.411140Z digest=sha256:48a0ae0468d9a2ef10d95d6e829418ffbd8f4b4fea44c53ffb488f4cb4971540

Observation c92f8f91-3a7a-4a5d-94f8-4032ea539520 · inbound

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks cites this paper.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:52:14.083201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:87ebfb07f8ecc3605b1bf92881ff5a8a1a701a9dcd1c15a2c717c097bb8e6a00