Pith. sign in

Paper Citation Record · LEDGER

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2608.05987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05987 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:53:26.173725Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82d516c7-7361-4464-8e74-9c0c1715ae86 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:24.992869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:24.992869Z digest=sha256:897cdb2d180fbfff03795a5b05aae1df84ae05383b993656b558d2a644e67dfc

Observation 9b876981-78b5-410b-af7a-3a85d061131b · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:24.998800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:24.998800Z digest=sha256:64386a532dcb276e004b17d9c98aa28d67f97663a3e695355071bf60c91d9dea

Observation 942e9632-ec05-48bf-9c02-77f51e92b41e · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.004243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.004243Z digest=sha256:f63064acd2c4100de2c887bb92d2f1b73b845c41b519e94dbbf71a7ad9c8a2ae

Observation 39cb9bad-9b74-4640-8722-3fa2f343bdf6 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.009238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.009238Z digest=sha256:5825dfe4d9bfd2cd934da817833872ac7e65b5a95a1712285d4fc4d0ad5b31e3

Observation 667a98ec-614e-47d4-8999-26fb44de5b99 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.034126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.034126Z digest=sha256:a72201d032bb5ecd8823e4e337798e1138e0358c8ab89f70408ea245bb22fcd3

Observation 7d3d36e5-cd90-433b-bacb-6f821ec383a5 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:29.145173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.085810Z digest=sha256:4eef8af109a1aab7249739f4ae79916c1d903414d2756d2e056ce8dd7d294222

Observation 38f9310c-887e-4e49-82dc-a9f20ac46ec1 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.115986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.115986Z digest=sha256:034e1c0388d279306fcb39f2c28391014cbca7a97380cd8e7b6a0a964db5d141

Observation e7c5d4b0-11bd-4d90-8efc-386bb35b325f · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.149185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.149185Z digest=sha256:fa085038e2f10f6b2f32fc8ec34e3075a3a9ed0f8fca17a59dbd587e179f9ebc

Observation b38d946d-e347-49ed-9f92-f7c6a045cff9 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.176942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.176942Z digest=sha256:8f0b9f77a81540645e5d351388cb6b3813248397846b582d0f70f71008010c46

Observation f6aa2346-0797-4f1c-b7eb-0d3b0690e0f3 · outbound

This paper cites The eleventh international conference on learning representations , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning The eleventh international conference on learning representations , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.189302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.189302Z digest=sha256:dab1cf2eb97ce8738239e903b54103d7ed77c303f4a32a267f0b0c93d3f096f9

Observation 46fbc3ad-f81c-4ce0-a59e-7e629cc2bcbf · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.218055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.218055Z digest=sha256:7559179641d8e0a375152a436b1a140fd8e7180b8c50b0d0093a3ff1b0acb4d4

Observation 93c6b923-0541-4cbf-8dfc-cf2d80c5f27a · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.273982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.273982Z digest=sha256:b93ec4de539f5fb8258b46116120f1ad69c63f3eb4f6d8c001f2d4a1c8dc6757

Observation 80ccb4e7-5879-4ca6-8af8-86c1b2414914 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Transactions of the Association for Computational Linguistics , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.325014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.325014Z digest=sha256:aab3ac37afa3bc91f6999f96a3d853465b0b4edf44501bd21666632c25035e54

Observation 05607002-a70e-44f7-bb3e-313b32897772 · outbound

This paper cites Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.341911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.341911Z digest=sha256:f7934e2c26830a7526b64adbc6dffe5e338aa0a92740658f79d8399a2a734ea6

Observation 9f802145-6ca0-4ede-bc71-aefee01ef16d · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.354446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.354446Z digest=sha256:8d87a4fa24780be263fc3d322a2ab4f1bb0cc857d1ffc5ee41d616a3b6c75de8

Observation baa8d1cd-9425-4610-b82a-e6bbd84bca61 · outbound

This paper cites Proceedings of the 2018 conference on empirical methods in natural language processing , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.360089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.360089Z digest=sha256:21fe9429e8f3f29d15cf97d04559207871dadc0bac793f44ccbd397686597f76

Observation e5af50b8-21fb-4128-b71b-c30f7aa8743c · outbound

This paper cites Proceedings of the 28th International Conference on Computational Linguistics , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 28th International Conference on Computational Linguistics , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.366573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.366573Z digest=sha256:7a43f1a9f49e892e0e4d69bb6d54bb4912a9e8f416407d1d284670f7b5bc5755

Observation 38266335-f39a-499d-9c1c-93407a11ae7a · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Transactions of the Association for Computational Linguistics , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.376814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.376814Z digest=sha256:d9cdb7efea73d0f6171d686d8d67a835fd0a84980a485b94bfcb19a89117c806

Observation 47c397c8-2827-4ea1-97a6-b37411405097 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.391568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.391568Z digest=sha256:eff08938a5eb6631b32d2aaf9207acc989590adf6177442969a43e4ff9b78bdc

Observation 59f97e98-7067-418a-9b87-e5456bdee721 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.400799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.400799Z digest=sha256:9e19060e9a6e961f60e829daa2b514ff7b66ae3713e294b55c5bfdb128b614b7

Observation 0799257c-2fd8-465b-8d62-1b94be6143c5 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.407335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.407335Z digest=sha256:0765cfc9dd9114facbc159f17823ce051ed22c6233aeeb6ec09f81b99d41263f

Observation e19984f0-4bd1-49d2-a2f6-1f328b391069 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.413689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.413689Z digest=sha256:dee0af430bc1d54696b3cc627901fb1d5344e473e97e797aac954d61c0bd09c1

Observation 19a84209-762e-4208-8078-1552acbeca38 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.418872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.418872Z digest=sha256:7a0d747657d6648de2d9814c4a4b842fc766e0df2630b3d2b34ba01236840a74

Observation ed2c257f-72f6-4040-8532-4305bb8a4b74 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.424814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.424814Z digest=sha256:d23e8dd8f6cae46824b202d3e6507f72e627d2637c27438c3bb02d917fdbe271

Observation 435bca4a-1b8b-4bbf-babb-8c9198a3daf4 · outbound

This paper cites an unresolved cited work.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:53:28.960790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.429457Z digest=sha256:d7781d9ccdb9e51709917f9235c5deedcdcc4a57c6c8f1d108e3dcc49cb19ffb

Observation 9b77dde9-97cb-4019-9b5b-082451aeb98a · outbound

This paper cites Qwen3 Technical Report.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.434211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.434211Z digest=sha256:5aeb218f3fc24fa4f565bf00ce35c37777af4387ca640f1a9c7bd1cbb39f280d

Observation d481ff67-37a9-46e6-9107-65384b2c8bc8 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Kimi K2: Open Agentic Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.438567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.438567Z digest=sha256:c58ac79d021db71ad646fd9dae13f656fd5d2613634f0e44efda35619902ce57

Observation 3763a7ad-38f7-4568-ae6b-fe66c08fd348 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.902064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.444028Z digest=sha256:74463984d57c05323e04f54466103d0aadc2af55598cc78ce4c9fbe962296681

Observation c197bb77-bb0a-48d7-9159-3b7a12768644 · outbound

This paper cites Proceedings of the ACM on Web Conference 2025 , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the ACM on Web Conference 2025 , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.850775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.449543Z digest=sha256:2b656cb7ea7b91176efe596e12b9d1e612106de12045cbf0446de4b612f77026

Observation bbbdf02e-35a0-456f-8d4c-fce5ee880c45 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.454538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.454538Z digest=sha256:6f5e95380fc46eab01c68390fcf3b55c183141c4b3b89611882f7a5a12843bce

Observation 4862341c-860a-4872-984e-94d3f753b5e5 · outbound

This paper cites arXiv preprint arXiv:2601.16725 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2601.16725 , year=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.459799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.459799Z digest=sha256:d4e8ff207142d923206b7938d90999e1f23a2ab6088241e013004d80cc4a71c7

Observation 3057d07f-715f-4706-a0eb-4d67daca02a5 · outbound

This paper cites GPT-4o System Card.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.464858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.464858Z digest=sha256:86be697b9e0c386d38f46600f568a819d89e92a59f31410f991337d36d5fa30e

Observation 08ba990a-7737-4c8c-828d-c4f80992a5dd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.470259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.470259Z digest=sha256:743053690e87c7c932ecf80da48fd0591aef748a9edb21493fcf1494cbf24f52

Observation 1e755faf-f6ac-40fe-96e7-97639f0b215f · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.475258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.475258Z digest=sha256:78393c21caafe54de8a7eab64eb3c15eb4854f09323ab3cefbc2f600977537e7

Observation dc2c9a41-6e4e-4169-98b9-6e3b1d2631e2 · outbound

This paper cites Agentic Reinforced Policy Optimization.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Agentic Reinforced Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.480329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.480329Z digest=sha256:1714a641c7d745d0593482de5773071a3f3aac05d0516f48f2c971006fda8cee

Observation e08336bb-95f8-45eb-ac97-0cc0c28fcf84 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.484747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.484747Z digest=sha256:9bc33bf913c7dca2626e7cec8daa482b4df6c1086e3163b2c402f6733bd8fc00

Observation 572c8c8f-2052-4491-9a37-6cd4287f1e2d · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.511208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.511208Z digest=sha256:ad6794dc9d96ac449be758a38ac4baceaa07a9e9cfb1a7b6e942e90b2a7f990e

Observation eafcfba5-3be3-4aa8-b6e1-d2ba35632447 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.549779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.549779Z digest=sha256:a6c1506585f72176a02cd195b2de81c9358811831990704e5204ad4b0278a42c

Observation 0782cfde-7851-4e2e-ac2c-4037575d4ff3 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.564156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.564156Z digest=sha256:e2319438d867f540690ee3e0fb92aff8da6249059a36c68c11b73c8ee7635106

Observation 8afca331-ee0c-4835-bc1b-a594829a9ecc · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.595205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.595205Z digest=sha256:65a0cbb94c639f7d1f30cf50d725c0d46bdfb5e414012cbaaad216b27d4492dd

Observation 6655eeed-dbc9-4b51-b32e-426f2c580a98 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.619108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.623667Z digest=sha256:225f15a6b9c1864b740bf66e73c5b861772a63d77c3d119c5ed5ab84afd306fa

Observation d506fd67-ed4c-452e-9700-251ccf57922a · outbound

This paper cites 2023 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2023 , eprint=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.647136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.647136Z digest=sha256:41358e04efc44d5fa3b30f502f52f04c9306050788a2e7b414bec569e6e4cf99

Observation 0694b948-1b45-429c-b819-511adca85e8b · outbound

This paper cites 2011 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2011 , eprint=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.671107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.671107Z digest=sha256:e0d6feea2a3291c07e83bbd95cb43fda3103cddd9f5f4fa8da702d52f2d70452

Observation 9d9835cf-a47d-46c9-a800-00dddecdce5e · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.719544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.719544Z digest=sha256:8a6208abae778e5b73a8c244166ef4743de292213f6beb75f107b267f30dded8

Observation 51d3c44d-c768-4c71-8091-f0dc1852d21d · outbound

This paper cites 2019 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2019 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.481273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.744472Z digest=sha256:a9a9a59ebae7c735d897c4a6fb276509145b96f797e8f2bb175dde5afa833b3f

Observation 26c3af54-464c-4908-9818-21de4473cd16 · outbound

This paper cites Mobile-Agent-v3: Fundamental Agents for GUI Automation.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Mobile-Agent-v3: Fundamental Agents for GUI Automation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.765585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.765585Z digest=sha256:c6f6f6f4621de9021e18224fd2aaee17ab29a42d85bfdf1d2b3d3ecc765b9755

Observation 69b7dcf3-89c5-4f58-878c-33c07661a583 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.786093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.786093Z digest=sha256:1e88d05b78a445c5b32440b01dee0a881a9cdbbd8b73cf32f7f64c6c8f929e25

Observation d0db4c0b-5dc5-4c92-861a-cdd6372cfdf0 · outbound

This paper cites 2024 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2024 , eprint=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.791433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.791433Z digest=sha256:6b41f85154198295741b719a400a4ee5ca60221326c214d53dbdf92feeee81af

Observation eee1988e-068f-45b6-b20e-d4002c8314f8 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.796098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.796098Z digest=sha256:a97a4bf863a03649c1caeb7dd20b267efca3c0c112fd07a17fcbdfd29294c9ac

Observation 97c813f3-7da2-420d-a63b-8389b4d8abb0 · outbound

This paper cites 2023 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2023 , eprint=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.800734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.800734Z digest=sha256:cb06f7541c5fbe64031f54d37a6388799b806e9bfe8efea766786d482b5b951e

Observation 53474c2f-3fdb-4541-b82d-b9b86cbf9833 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.805370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.805370Z digest=sha256:89aeef5bacb807a7e4b3a5ccc113d9813ab937fe7c8f0476484d9809be9d1069

Observation 3236c731-949b-4c7e-9279-ebe740505a87 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.317587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.814977Z digest=sha256:c65d4c97386ca80071a203831608433b5e6272a09091bbd12449c66d01ff0616

Observation d6c29359-0c29-4e01-a06c-fa5bb5030e1f · outbound

This paper cites arXiv preprint arXiv:2602.03048 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2602.03048 , year=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.824341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.824341Z digest=sha256:7c78d7ee9339c7219661ce15eb8628fd975a3e65e57f45083fe20fda75b16057

Observation c9e66127-70dc-4a95-9552-615c82320dc8 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.237592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.837906Z digest=sha256:12c9914a144c5950d20fa40447ace7cf83acd8d563f041e60cb3f406dc0f1898

Observation 21728072-1869-488f-9096-e1cef7744cfa · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.206314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.854963Z digest=sha256:4af5903e4e20758e118a1179667f9c38b999f025dd0e1766a0994ca2044cf8f9

Observation 34ea40fb-8d47-4cd8-9d2d-bbc6dac40c63 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.190864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.868092Z digest=sha256:2dc2ef8a836bc78df1252f3d17b2df832bb0e69bf62a334e17f27f4e7b64f17c

Observation 9808c1ba-c370-49b1-9be9-5c1f14521c1d · outbound

This paper cites 2017 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2017 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.872410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.872410Z digest=sha256:ce97132577fd6570c8ef99bdce8cf4a2f33f2232fc9541f83db12b21035148da

Observation af5e4c2e-d7f3-4cad-b2d3-cda22e9a7338 · outbound

This paper cites 2016 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2016 , eprint=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.165787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.877195Z digest=sha256:07c2786fe965ede33578e8ab84853077b0f827cc2443f56c7d03fd76cbc452c7

Observation 571504e8-9896-4662-9772-a7c3a091c430 · outbound

This paper cites 2019 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2019 , eprint=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.149599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.882562Z digest=sha256:dfa8590138509bc80c16fd5140e335b669c5b66f4a326ef90d559c90d79f3859

Observation a7b45576-74ae-44ce-9fa0-e0d633787833 · outbound

This paper cites 2024 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2024 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.134605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.887615Z digest=sha256:62f37542ea7c46ab2f0e296db1d97a73bafa2f548bb14321433e1ee6856b6ed8

Observation d362eef3-4808-431e-8529-c8335ed30d54 · outbound

This paper cites 2025 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2025 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.892424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.892424Z digest=sha256:a1fb96373b81cf97800c1d5aefc3badc1627026cb9344947c7945726ed6842db

Observation 3bc4c876-61cf-4ab1-a78b-c0ae6c7cb9d3 · outbound

This paper cites Journal of the American Statistical Association , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Journal of the American Statistical Association , volume=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.107894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.896700Z digest=sha256:177f40cdd1b7e188131b3bd943f5013802fe78a2c5e3d616fd6607a7bffc4b23

Observation cbb49228-4cbd-44e1-b1c6-5957ad50e682 · outbound

This paper cites The Annals of Mathematical Statistics , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning The Annals of Mathematical Statistics , volume=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.091321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.901640Z digest=sha256:8931158579937882c4309c298c0696ee249ade4bcd603432fba906d9b82ffec8

Observation 0c69cbd7-d25b-445e-a4ff-0cc00bd9f76b · outbound

This paper cites SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.906690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.906690Z digest=sha256:264d3dcbf3020a1667174a7645978cc8b8d50c4971fd5396f462908967f572b7

Observation f139449b-87b8-44c3-8130-b16a48d37b81 · outbound

This paper cites OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:27.087135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.911804Z digest=sha256:0747ad0e899e67185df1a140c598fb71d1adbe477d067a1721c4f3ea617a0477

Observation 711881f7-a651-4888-8f97-894bf50f5a8a · outbound

This paper cites SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:27.063688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.916475Z digest=sha256:ab448c93e0d4b1093e5ce8228db66f1598e990bca50b467b8c3edb574b55e95e

Observation 2822f740-672d-41cc-8588-6f1dfa8c2170 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Self-Distilled Agentic Reinforcement Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.921007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.921007Z digest=sha256:7d9e58372fc599ecc37980e7d6e81edc52ab98d77a0cdc5b3450ce4cd74e3257

Observation fe8762e6-3bb4-47ce-b6a6-1854bb9d4d11 · outbound

This paper cites Artificial Intelligence , volume =.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Artificial Intelligence , volume =

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.926059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.926059Z digest=sha256:9a768f05739b98b7d1590764d5796120fab2ae7fdac23c6f7716f30f0fc31391

Observation 320f4540-0391-40a4-9ae5-48be7875b92c · outbound

This paper cites Journal of Mathematical Analysis and Applications , volume =.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Journal of Mathematical Analysis and Applications , volume =

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.062099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.930651Z digest=sha256:bab212662f061afb91731c9747e03958f3e1b0c2ed058b94081145ebe51adecd

Observation 9a0b8028-fd40-4eb3-8fe3-a26b2292cfc6 · outbound

This paper cites arXiv preprint arXiv:2602.07594 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2602.07594 , year=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.935615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.935615Z digest=sha256:45b965d505f52d3f5d583dfcc15dab93c6cec2264d5b16e4b443ca204d1aa6aa

Observation fe176c02-ad4a-4efa-b45b-4c78bc7b7dff · outbound

This paper cites Look Before You Leap: Autonomous Exploration for LLM Agents.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Look Before You Leap: Autonomous Exploration for LLM Agents

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:26.818924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.940186Z digest=sha256:ad861b81c786b510ea97255bbe1dbe312f60f71fa7fa9ff988c943beb5f6c688

Observation c41e9e68-7cc8-458d-8f78-953acbafccda · outbound

This paper cites arXiv preprint arXiv:2601.14050 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2601.14050 , year=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.959586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.959586Z digest=sha256:98a4d831a0b54d15cbff3adb5a3897f88595e81439618ec7b227591d4ec3cc03

Observation a1c9e453-598e-4be7-979c-de5c7e49473d · outbound

This paper cites Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:26.507641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.999733Z digest=sha256:dc15d6e2062acaca4ed52ac3933301ae72f38fa94f6e79d85ba0a1e49e95ba72

Observation 81cb3325-a230-4c30-89de-3c6bf51209f1 · outbound

This paper cites Memento: Fine-tuning LLM Agents without Fine-tuning LLMs.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.038845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.038845Z digest=sha256:b6dc4ecc8188b93797b92b726670b1259e695729d32ebdf3e56d2aed2c1a2b01

Observation 2972291e-fd12-47ea-83e2-d0fd5007ea01 · outbound

This paper cites Reducing Tool Hallucination via Reliability Alignment.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Reducing Tool Hallucination via Reliability Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.084902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.084902Z digest=sha256:92b50642f5e8459bd98e69013686e7e2945a93c8610245599ccb7ca02f966c6f

Observation d28bd039-8e3e-45fc-a621-9b7d51d947a0 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.141889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.141889Z digest=sha256:d52acf810c78a2367f8723ab3d86e49b7441487cc63529ec167afb0992920f25

Observation 6f511aa7-b0b0-485b-993e-0790a83398b0 · outbound

This paper cites arXiv preprint arXiv:2509.11543 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2509.11543 , year=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.173725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.173725Z digest=sha256:0fd26a65d61bfecef864fdae5cf5c129e9eaa1426dd9dc4a4b3a36dde626711c

Pith citing papers

No inbound Pith citation observations are available.