Pith. sign in

Paper Citation Record · LEDGER

Agentic Reinforcement Learning with Self-Distilled Reward Shaping

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.03223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03223 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:15:50.932630Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved37
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8675798f-b798-4d3f-8c9c-fdb46028c21f · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.861115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.697103Z digest=sha256:17d93237eabbb75a80fd8f613d00b4f257410a37d64d093e7e4ceb59621b37be

Observation ff4f27e9-3352-42d5-86aa-a439fd126d03 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.846448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.718401Z digest=sha256:854419047087f41d0218738aaba88c08c673eea4c51dbc1e3c809461b8f43ce2

Observation 9adefe45-cc4b-402a-a50d-2d780189b19a · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.724043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.724043Z digest=sha256:fd88517b8432566fd5a983a97799fcd90e7dcc874729e79ad6feaa3c47aebde2

Observation 92a3e15b-408b-41b7-aaf8-4ee8842f5ce9 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.729905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.729905Z digest=sha256:90041bbd275ee7035792e3e297ea5fdf282426ea2ffe2b27d4fa72ed46785976

Observation 13156ec5-b935-4634-b9ce-a745c9ddf324 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.830528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.737402Z digest=sha256:2cb7250c34389e1ffb74b8390ce7c7772e8d402459c8af8801e5857e22eb7d01

Observation b138fd6e-ac23-4e3d-87a5-68837a7b8bbb · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Distilling the Knowledge in a Neural Network

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.742839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.742839Z digest=sha256:30ef7e56db90ca09caa9b5700810fd806dc994e2e238decc54707b4312cab1d2

Observation 71826fd6-80ac-44a6-9ed0-dd1d21f9e9ef · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.749272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.749272Z digest=sha256:3eb7f5e535ea7f3d762bca5d45c774734c367833c9a660c8eef45ab8e599be5d

Observation baa0c435-e836-4b37-91fd-d4d9b6e538de · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Reinforcement Learning via Self-Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.754062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.754062Z digest=sha256:2aaa684f52e996d4cd9938e0f199ac8109e18b00b118a758de61f6b04a2b10b3

Observation fab17f2f-af4e-492e-8d1e-40c658fa80bc · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.758886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.758886Z digest=sha256:ec757eaa69dde2af5245890cd74e25fa172ef9ca5c7d266edf94d664d7980f2a

Observation 7ebed2cd-4208-4aa5-bc04-7de91ad3f08c · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:15:51.804385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.763497Z digest=sha256:fdd45336f23d5735cf1762068751438415eba1fb86cec1ed59d24a0ff471c123

Observation d37c7dd3-f9be-4136-8c78-68c3880e9b1d · outbound

This paper cites Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.767987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.767987Z digest=sha256:65a9f4b6cad948bb62edbff140fc07133945cf025535763bf7e1a45f04aa8ec7

Observation cb700ff0-c4a3-41e1-926a-7527f6a98784 · outbound

This paper cites Naturalquestions:abenchmarkforquestionansweringresearch.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Naturalquestions:abenchmarkforquestionansweringresearch

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:15:51.788550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.773613Z digest=sha256:8cf66157b9033351647cdc0256e7263e228bfaecaeac47fa75a457ee99c749d4

Observation 4bd96b65-717d-43ab-8852-d01b4c418612 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.771787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.778337Z digest=sha256:845586591e0682745c78767731f781cf86e9dee3f144b959a354cd02a1679e2d

Observation e3be0fea-2032-4c84-8e10-cf1b97b8d851 · outbound

This paper cites Self-Distilled Policy Gradient.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled Policy Gradient

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.786116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.786116Z digest=sha256:9913713948d55a5fbcd9024a9fd875896a9ea6dff80d7efcf7c69839c5276f08

Observation 508bf451-91ee-42c2-ab64-7fdc6610c639 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled Agentic Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.790996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.790996Z digest=sha256:2726d2fc5eacc9e58f1706eab0461bcfcca5fe22feffbb14e62faf007653fdd1

Observation 525c4e31-9b8a-4ab2-8e01-d00e5e1783af · outbound

This paper cites Agent Lightning: Train ANY AI Agents with Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.795278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.795278Z digest=sha256:3ac69503f515ed14d1cf50280d96ffa1184db7f758c71cc4bd111c6c87765215

Observation 5b79254f-3fc5-400c-8af9-09f6132a4b10 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.750157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.800345Z digest=sha256:98b3b0cde9edbd60be1242e00be40b78a64ea8fa6dbaabb62e27e46a7468b53f

Observation b605da9c-739d-4e15-9d83-ecdac7487221 · outbound

This paper cites CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:15:51.319505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.804723Z digest=sha256:f43bffb169ad6d5ca9a5f3956dcde6a2b064c5c506d49ed65fcf2c868789d892

Observation 1011db5f-b902-4973-b2a2-3ad1da50d45d · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.730575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.809974Z digest=sha256:e5e84a645851a3ef2aa640c2d441507924d5786363871868ef7465847cdf0531

Observation ecdbda85-cc2f-4523-a3cc-f02fe3097b57 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.713178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.814148Z digest=sha256:3b66e3200862689a14f3aa835272be3d0518fe786b405fb11fb4f1eeb205773b

Observation 5dfa644a-6dee-4f8e-a030-92755311a9dc · outbound

This paper cites Qwen2.5 Technical Report.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Qwen2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.818627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.818627Z digest=sha256:fe58f0b72f989e89090bca42aac2427f17640462c8b1e77ba49da9ff138c0265

Observation fd82cd44-8337-4a08-8b01-2c0467438122 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.693931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.823637Z digest=sha256:c1ff031f540244f3abcfb9270756751d9a9edef15f713c6c1067aad7612bde26

Observation 035ec4cb-5dc6-43fc-bbd5-e08891a2cefe · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.827697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.827697Z digest=sha256:177e90779e1a30f2c439c74f68038eb17b1c072a730df0ae2f5a4a55b7abca12

Observation c6d0de95-b5ef-4323-9231-c65a20275e07 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.837192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.837192Z digest=sha256:2b2a10caacae2cd965c41cdb72be07cc1cb98afe35637bb4f6f64458c997cd13

Observation 81c9f18d-99bd-4bcc-adde-d2766986be77 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.843229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.843229Z digest=sha256:26efe4cf6a4104a20bd09a1d74cb462de875bdf781b2d64aea0bb8e358bd2b37

Observation ff8d01c5-2bd5-42d1-9e20-50c0d9be4d9a · outbound

This paper cites Learning by Distilling Context.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Learning by Distilling Context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.848819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.848819Z digest=sha256:e7a0f77eacb53585db8d7ec2262e28dd2d8aad0c2ac27ef0388b0c50c2c6ec41

Observation f9fb4306-7ed0-4c05-affa-c6e28e54c521 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.853031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.853031Z digest=sha256:5ab8e0ed7967909bd91766f13f9625eecb61acae9d586f3fa1235781b8a5a209

Observation 5cceeed0-e3aa-4b6f-bc35-7df47754fbee · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.636021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.862402Z digest=sha256:09c90eb77fbf7363af972179bd63e330fcd3c5af056fb3783d65c27b666819ca

Observation 38069836-2da3-4d8a-a6ed-bf13e7a964df · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.866749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.866749Z digest=sha256:14a1f36ee8b5a94b8a024cf37d550f4fa0abf3df8c62e752db891e94c76a20b1

Observation 332f1b5f-613f-4323-b26f-5a8389fb2d4c · outbound

This paper cites TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.870908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.870908Z digest=sha256:3c4ac53d9771a4965962e72edb9719a82a523c8d1ade8dc3dd81e74c2f45258c

Observation b747ed9c-5279-4141-b5e1-09f56823bc86 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.619372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.875491Z digest=sha256:ee72a10e987dad707681750471ed4300dbde643a801f167aec997a8533eec4de

Observation 5e49090d-7305-422a-9c4d-dbb8011c04bb · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.886657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.886657Z digest=sha256:9a585db017054c4625e0b62cefc438035e5a61402d414198f191deb072811e6f

Observation f8857167-3609-4823-94d3-4c04bdb123df · outbound

This paper cites KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.891014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.891014Z digest=sha256:0187e79ed0aa16a61a963ed1ceafcab31f6569c3d8a1fd683db720fc24df8466

Observation 580e16ac-39d3-4085-96fb-0f8cca1a2119 · outbound

This paper cites TIP: Token Importance in On-Policy Distillation.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping TIP: Token Importance in On-Policy Distillation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.895596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.895596Z digest=sha256:113b8e5e51e70a5adbdd808d1a4da74258c476dc047563bf5437ca22239ab7b4

Observation 62ec70ad-3164-4a3f-86b3-0c610941cba0 · outbound

This paper cites Qwen3 Technical Report.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Qwen3 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.901475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.901475Z digest=sha256:69d497d3a51bf9c404e394e6279774b3a7363eab67e9c8aab22eeb075dfd40b9

Observation e73c43fc-446a-46e9-bf20-de450cbf756f · outbound

This paper cites Self-Distilled RLVR.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled RLVR

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.907312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.907312Z digest=sha256:8d8d08389a8d6485cddf1a0a49ca46bf65655fb46dc04effdc76d0ad54ecc2a7

Observation fd9b2f95-9564-4cf8-badd-dd4a3e8e1e6f · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.913796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.913796Z digest=sha256:18e84111d1c00a5d1de46d229cbf8ac9af0a6b7daa5f3f8522af2516c160d96b

Observation cce0ed66-ec96-4eae-b737-28ed9bc82c38 · outbound

This paper cites WebShop: Towards scalable real-world web interaction with grounded language agents.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping WebShop: Towards scalable real-world web interaction with grounded language agents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:15:51.590013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.921335Z digest=sha256:40c179c4cc12b2551a641a726816e38c681d59e76d23df0efeeb6eb554b8fef1

Observation b3da775b-c531-4833-81db-c3208aff0302 · outbound

This paper cites On-Policy Context Distillation for Language Models.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping On-Policy Context Distillation for Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.925900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.925900Z digest=sha256:4d3aa95d6e41d76215d841e2bbe52871f25bed7df80625064075725c15edb1da

Observation e5ec1a51-7670-44cb-ae15-62522eb0765f · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 40

Resolution
malformed identifier
no resolver link, observed 2026-08-05T23:15:50.932630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.932630Z digest=sha256:a5e6aec8fd15f7752d6152c3da3ee981e7ddadfbf3ffb720e0be48e87775585b

Observation 24d858c6-4fac-43ac-9b1d-6f58d24bf818 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.832775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.832775Z digest=sha256:b26c6fa1de17efd4779b9c113aaca52f37864d63879cf00112417c0d1d4234e3

Observation 73286c70-c049-4d27-ad0f-0f428c56a001 · outbound

This paper cites Transactions of the Association for Computational Linguistics10 (2022), 539–554.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Transactions of the Association for Computational Linguistics10 (2022), 539–554

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:15:51.654412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:15:50.857722Z digest=sha256:af48f0dc710da829f168ab92ddbe4dc6e6d056f226b6e8334506645caf1fe614

Observation d9033018-4208-4139-a8c5-f5fe7d1a2488 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.881548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.881548Z digest=sha256:59f4415ad8198d7376c5e984cd91c7067ce9f383acaab4db43c409a3d553dfee

Pith citing papers

No inbound Pith citation observations are available.