Pith. sign in

Paper Citation Record · LEDGER

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2608.05987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05987 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:53:26.173725Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82d516c7-7361-4464-8e74-9c0c1715ae86 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:24.992869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:24.992869Z digest=sha256:88d23dd90e18664d5081f843b7caac8524f1e745b42fe4b8e9fc3db4fae34a87

Observation 9b876981-78b5-410b-af7a-3a85d061131b · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:24.998800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:24.998800Z digest=sha256:20308be9554fcf549740de4e5cc9c4554f22930a3d8ed19cf11e2c0eea59a8ae

Observation 942e9632-ec05-48bf-9c02-77f51e92b41e · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.004243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.004243Z digest=sha256:d2d3948c1c0c5b7cd7f43bb557d331eee763a5749817859e332eccd07917ae12

Observation 39cb9bad-9b74-4640-8722-3fa2f343bdf6 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.009238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.009238Z digest=sha256:0d255ffb959368118a5b454fbbdb6b0ff39456415d4daad34bdc72d6e7e40a72

Observation 667a98ec-614e-47d4-8999-26fb44de5b99 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.034126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.034126Z digest=sha256:586161405701fb7e4aebc0f4d3d95933a8590012cef7445a1b7950ca2872cff6

Observation 7d3d36e5-cd90-433b-bacb-6f821ec383a5 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:29.145173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.085810Z digest=sha256:a74536ddefa7cf9a5457270b00b520bbfc9506b8a4cbd5c0958e29e5707bb202

Observation 38f9310c-887e-4e49-82dc-a9f20ac46ec1 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.115986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.115986Z digest=sha256:9589b8d2ee46df10f3a0a3c039a7fd458f85966015936eab1e85b57d7b56e65b

Observation e7c5d4b0-11bd-4d90-8efc-386bb35b325f · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.149185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.149185Z digest=sha256:73d513435e5dfb9412960368ca41b306b22c1a13544bd4b7740acc02653f4e82

Observation b38d946d-e347-49ed-9f92-f7c6a045cff9 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.176942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.176942Z digest=sha256:0f053bcf8508a1c0634cd6a8b9f59589f2666a8306eef26b36a8ac2c163fb941

Observation f6aa2346-0797-4f1c-b7eb-0d3b0690e0f3 · outbound

This paper cites The eleventh international conference on learning representations , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning The eleventh international conference on learning representations , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.189302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.189302Z digest=sha256:43d1b68610c64b82400207c90a494365e8615b1f7e750b3cb81f7e740c3f416f

Observation 46fbc3ad-f81c-4ce0-a59e-7e629cc2bcbf · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.218055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.218055Z digest=sha256:9a3925020c2efc0d45026b9784018478a220c50f5726cf3a50d0dcc87a522945

Observation 93c6b923-0541-4cbf-8dfc-cf2d80c5f27a · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.273982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.273982Z digest=sha256:34ae9a10f0d21ec1c3f6f585613103a8982a28854fa392d15fd5e6c731ff728f

Observation 80ccb4e7-5879-4ca6-8af8-86c1b2414914 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Transactions of the Association for Computational Linguistics , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.325014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.325014Z digest=sha256:af03e3dcb16b6d23484d96a5f758d46dd8a6742f30df190c3dff9d0f0af9371b

Observation 05607002-a70e-44f7-bb3e-313b32897772 · outbound

This paper cites Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.341911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.341911Z digest=sha256:59346584ca064f24ab7ec023ab9ed52340c6cd7096d9014f497d9720fc68bf3d

Observation 9f802145-6ca0-4ede-bc71-aefee01ef16d · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.354446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.354446Z digest=sha256:6d63fb555ddc3b3788330995fcd6640cdb14955ff048dce032dd5a396c5ef02f

Observation baa8d1cd-9425-4610-b82a-e6bbd84bca61 · outbound

This paper cites Proceedings of the 2018 conference on empirical methods in natural language processing , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.360089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.360089Z digest=sha256:34e18afd73544408bddd40fccb3999b70c928d80763bb9faa1b41d4b86faf994

Observation e5af50b8-21fb-4128-b71b-c30f7aa8743c · outbound

This paper cites Proceedings of the 28th International Conference on Computational Linguistics , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 28th International Conference on Computational Linguistics , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.366573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.366573Z digest=sha256:615129ae18cfe1152fdc2d6b1e08223b32d5d224d44ec5fa788e97cd4e537a14

Observation 38266335-f39a-499d-9c1c-93407a11ae7a · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Transactions of the Association for Computational Linguistics , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.376814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.376814Z digest=sha256:35a17418bf795ca6097030754842892ea701ac929aca6451b6064f807fb6a88c

Observation 47c397c8-2827-4ea1-97a6-b37411405097 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.391568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.391568Z digest=sha256:679d58e84c59c7e5ea60569df8a6766594d607512a307cc45e43a21116d9f014

Observation 59f97e98-7067-418a-9b87-e5456bdee721 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.400799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.400799Z digest=sha256:77d4a0374930186bc80221d6013bda96f8855bdbafc79485a7720795119fcf85

Observation 0799257c-2fd8-465b-8d62-1b94be6143c5 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.407335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.407335Z digest=sha256:d0b0efb3512f4b9bd7c39308338ee654abf33940f77cd1a2e16147a44c4d08ad

Observation e19984f0-4bd1-49d2-a2f6-1f328b391069 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.413689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.413689Z digest=sha256:033bfaeb6adeb1462a0fd88c1f39db2bf17703d64b6484d5e02ad905696f88ac

Observation 19a84209-762e-4208-8078-1552acbeca38 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.418872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.418872Z digest=sha256:d9540043a3d754f936a14a5ed997f2bd4115470d1afd3e1ea47b3aa3437fd899

Observation ed2c257f-72f6-4040-8532-4305bb8a4b74 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.424814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.424814Z digest=sha256:17c1ee25ad48a4c4472e4786983c10f2edfdd2cae65c4f39d66ab9279ab43525

Observation 435bca4a-1b8b-4bbf-babb-8c9198a3daf4 · outbound

This paper cites an unresolved cited work.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:53:28.960790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.429457Z digest=sha256:67786af382b015cf3906d87493aa8dd74a73895e9b5b39242f2a90be478772b3

Observation 9b77dde9-97cb-4019-9b5b-082451aeb98a · outbound

This paper cites Qwen3 Technical Report.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.434211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.434211Z digest=sha256:4a4e2b7385b3ea589d29cd294f5ef5ee51791427c83539b9a0c44fda92cc9039

Observation d481ff67-37a9-46e6-9107-65384b2c8bc8 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Kimi K2: Open Agentic Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.438567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.438567Z digest=sha256:2861d3c2e8475251234ad338d94d6a1608387261548f6b4c321f2e87e3cfe514

Observation 3763a7ad-38f7-4568-ae6b-fe66c08fd348 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.902064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.444028Z digest=sha256:179011bcecd2ffe2fc934700d0200369a9fc2582b2c20a9f921e938d72512a3e

Observation c197bb77-bb0a-48d7-9159-3b7a12768644 · outbound

This paper cites Proceedings of the ACM on Web Conference 2025 , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the ACM on Web Conference 2025 , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.850775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.449543Z digest=sha256:f531733322fae69ac3e2e11ba512bd8db852ad05f97e786b999a7b1b1403656c

Observation bbbdf02e-35a0-456f-8d4c-fce5ee880c45 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.454538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.454538Z digest=sha256:306be6d7455978bbb9e58830545b09de5007b2a83588ca7acba4a00acf05e652

Observation 4862341c-860a-4872-984e-94d3f753b5e5 · outbound

This paper cites arXiv preprint arXiv:2601.16725 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2601.16725 , year=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.459799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.459799Z digest=sha256:de08b777cde6ead06ef92b5d029b0b6ec1ecf168e2d1a10f6a047b7c8ba47b7f

Observation 3057d07f-715f-4706-a0eb-4d67daca02a5 · outbound

This paper cites GPT-4o System Card.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.464858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.464858Z digest=sha256:a032a464e6360be0a140815b84c8feb9ac33485adb110c71777ada2703ea672e

Observation 08ba990a-7737-4c8c-828d-c4f80992a5dd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.470259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.470259Z digest=sha256:8817126d7178f72df76897f12ec4c4655fdf5c7b9c21478244f909b0ac1edd20

Observation 1e755faf-f6ac-40fe-96e7-97639f0b215f · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.475258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.475258Z digest=sha256:c47bbc86560fcade5284e6570a0465d9c246b93e9687624dadb8d585f98260d8

Observation dc2c9a41-6e4e-4169-98b9-6e3b1d2631e2 · outbound

This paper cites Agentic Reinforced Policy Optimization.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Agentic Reinforced Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.480329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.480329Z digest=sha256:a564a05560de3c85947c7931cd618e54bc2261db614113b10224e92c2d2fbb04

Observation e08336bb-95f8-45eb-ac97-0cc0c28fcf84 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.484747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.484747Z digest=sha256:b4931ec1c97b01604195955a611bce5f307330a4bb0f328e1389413d6a879dd5

Observation 572c8c8f-2052-4491-9a37-6cd4287f1e2d · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.511208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.511208Z digest=sha256:b474d93d04d9bf6d6921b8ddd020f97d513039f135e480172ba4324e6ab7b9dd

Observation eafcfba5-3be3-4aa8-b6e1-d2ba35632447 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.549779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.549779Z digest=sha256:4e052a34589a98d9d712ec3c2207a452c3dd0761f6db582d3ccac37380867de0

Observation 0782cfde-7851-4e2e-ac2c-4037575d4ff3 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.564156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.564156Z digest=sha256:97881be33ece65f5a175e5a00fe45f21564af2be0203da8d916ddb8760c29fa8

Observation 8afca331-ee0c-4835-bc1b-a594829a9ecc · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.595205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.595205Z digest=sha256:418824ffcfcd5bb32adb186555d4b74e1dcd99c8c911762fca610c19063f75ca

Observation 6655eeed-dbc9-4b51-b32e-426f2c580a98 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.619108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.623667Z digest=sha256:7efee8ea326634871758c007dd140667e9a262cce0837028750d3b98fb9caab8

Observation d506fd67-ed4c-452e-9700-251ccf57922a · outbound

This paper cites 2023 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2023 , eprint=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.647136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.647136Z digest=sha256:ec43b8d5fd481b18ebef0672788d0ff9e3f5af03984701fc778fa86700b0f8e9

Observation 0694b948-1b45-429c-b819-511adca85e8b · outbound

This paper cites 2011 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2011 , eprint=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.671107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.671107Z digest=sha256:f65b485ada196a006a84acd02cc349e51c1b5dcb550014ed78df993de3d25ce1

Observation 9d9835cf-a47d-46c9-a800-00dddecdce5e · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.719544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.719544Z digest=sha256:4cc6e23cf33bfbebb42d5eb76cf41a0c59330efc4cb9d93201796b79965f71ac

Observation 51d3c44d-c768-4c71-8091-f0dc1852d21d · outbound

This paper cites 2019 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2019 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.481273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.744472Z digest=sha256:35a6488a63caeb07b94c05f79667c7e5e0cee9f6200ebc8a509516c9a33d59be

Observation 26c3af54-464c-4908-9818-21de4473cd16 · outbound

This paper cites Mobile-Agent-v3: Fundamental Agents for GUI Automation.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Mobile-Agent-v3: Fundamental Agents for GUI Automation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.765585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.765585Z digest=sha256:b2f5f7e2a7a5cf433de0bae578897a4c38abe0f1183494557876eadcc3ea2ba3

Observation 69b7dcf3-89c5-4f58-878c-33c07661a583 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.786093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.786093Z digest=sha256:a35ed1c3cda1d36de87efe7699b1204a94beee4cfd1e23eb4957343763909600

Observation d0db4c0b-5dc5-4c92-861a-cdd6372cfdf0 · outbound

This paper cites 2024 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2024 , eprint=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.791433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.791433Z digest=sha256:e6269806b2f4b99d3111ff8ca73d777c8c9f0b0654153f46fbc506aeaa532766

Observation eee1988e-068f-45b6-b20e-d4002c8314f8 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.796098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.796098Z digest=sha256:9842c8439fd655a8efbe43412865829873a04f4e4addee7b42ebcbde721ee800

Observation 97c813f3-7da2-420d-a63b-8389b4d8abb0 · outbound

This paper cites 2023 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2023 , eprint=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.800734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.800734Z digest=sha256:7ca7ef9a1da385ae4faa450b30148940f3e15a61db601bc2c10ff5a83985e6bd

Observation 53474c2f-3fdb-4541-b82d-b9b86cbf9833 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.805370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.805370Z digest=sha256:aae93da081b38f2cefac18a56764c2c679cf2c70cd696709b50178fd54ade75c

Observation 3236c731-949b-4c7e-9279-ebe740505a87 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.317587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.814977Z digest=sha256:df369aec05b9c57a21eae8100f79f17f679a2d340b85dd2c0bfb2c3c21a0302b

Observation d6c29359-0c29-4e01-a06c-fa5bb5030e1f · outbound

This paper cites arXiv preprint arXiv:2602.03048 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2602.03048 , year=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.824341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.824341Z digest=sha256:9a3712f592b88c5bf4abe8bc5923c5f86767c92609f9a5968fe0d4c0df7611f8

Observation c9e66127-70dc-4a95-9552-615c82320dc8 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.237592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.837906Z digest=sha256:cbe2faca8608a701f02715519f0c893c69fb9f687670f3f011c4ab3c900b55c5

Observation 21728072-1869-488f-9096-e1cef7744cfa · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.206314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.854963Z digest=sha256:ebe9a132231649dd19fbe083a4967c7b6f143763669426b23edabfaa2b1c7bab

Observation 34ea40fb-8d47-4cd8-9d2d-bbc6dac40c63 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.190864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.868092Z digest=sha256:08983db9b1431ddf57829220b5c57603088dc874174f0c10b02a2fcaada4a58f

Observation 9808c1ba-c370-49b1-9be9-5c1f14521c1d · outbound

This paper cites 2017 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2017 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.872410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.872410Z digest=sha256:e4ee160c08266e2d4729ccd97bb61cb6d0b0eec2535ed3cc0418013e9ca6cf90

Observation af5e4c2e-d7f3-4cad-b2d3-cda22e9a7338 · outbound

This paper cites 2016 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2016 , eprint=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.165787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.877195Z digest=sha256:bd7f547d8bce4d3af7046c281e50a603a6c9f58e585d8313f31b7eec8569301f

Observation 571504e8-9896-4662-9772-a7c3a091c430 · outbound

This paper cites 2019 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2019 , eprint=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.149599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.882562Z digest=sha256:b24bf112c5d896f285e58aec9c68df062e77af62f9f190cfc06ef66c96d2222d

Observation a7b45576-74ae-44ce-9fa0-e0d633787833 · outbound

This paper cites 2024 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2024 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.134605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.887615Z digest=sha256:62074f59499b42d2700346dbff70e7a41ecaffbc68d089621576f4fd83e096ff

Observation d362eef3-4808-431e-8529-c8335ed30d54 · outbound

This paper cites 2025 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2025 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.892424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.892424Z digest=sha256:5608952eee1fa5b24218fab990b73e2405853b0f9ab2a3d1363b4ea1e878c46a

Observation 3bc4c876-61cf-4ab1-a78b-c0ae6c7cb9d3 · outbound

This paper cites Journal of the American Statistical Association , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Journal of the American Statistical Association , volume=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.107894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.896700Z digest=sha256:6afdfbc4e0d086cd6ecacf0128eb52da3e17ba01395ba71bd41ad6973427951b

Observation cbb49228-4cbd-44e1-b1c6-5957ad50e682 · outbound

This paper cites The Annals of Mathematical Statistics , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning The Annals of Mathematical Statistics , volume=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.091321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.901640Z digest=sha256:c83dc8698d50c6bd6f762a60289980318d26afe219e08de0879b2a5cddc4f3cc

Observation 0c69cbd7-d25b-445e-a4ff-0cc00bd9f76b · outbound

This paper cites SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.906690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.906690Z digest=sha256:dad799ffda228f24991847c2b3d4827c5b742d961d0f73747d8c59dd8bc12960

Observation f139449b-87b8-44c3-8130-b16a48d37b81 · outbound

This paper cites OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:27.087135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.911804Z digest=sha256:b8d345ce6c68ecc32f58289f165236892eb790d6b46e949aab4b5f4c45e81ff6

Observation 711881f7-a651-4888-8f97-894bf50f5a8a · outbound

This paper cites SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:27.063688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.916475Z digest=sha256:4d97b36b075e846d2e457ba95b713b9b66cc4f1b8426f2f2cd9c06c8e32bc0b0

Observation 2822f740-672d-41cc-8588-6f1dfa8c2170 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Self-Distilled Agentic Reinforcement Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.921007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.921007Z digest=sha256:504d82385f396b811b32642f23a4c59d130029d05c6a1f9df787c7562c0bff6b

Observation fe8762e6-3bb4-47ce-b6a6-1854bb9d4d11 · outbound

This paper cites Artificial Intelligence , volume =.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Artificial Intelligence , volume =

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.926059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.926059Z digest=sha256:f9e042bf15742593a9ba0f86694d7b0934ee4d0b722d6a0ee5a65edb7f66ef62

Observation 320f4540-0391-40a4-9ae5-48be7875b92c · outbound

This paper cites Journal of Mathematical Analysis and Applications , volume =.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Journal of Mathematical Analysis and Applications , volume =

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.062099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.930651Z digest=sha256:a158319e14ec0cc5649eea926700864d5f4e6f4d2bcaa7db8d2c5c0085502bba

Observation 9a0b8028-fd40-4eb3-8fe3-a26b2292cfc6 · outbound

This paper cites arXiv preprint arXiv:2602.07594 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2602.07594 , year=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.935615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.935615Z digest=sha256:4767ac1898a154cb8718b790a0d1913e0e95b98e7c1cbc903fd96414caa2c2b5

Observation fe176c02-ad4a-4efa-b45b-4c78bc7b7dff · outbound

This paper cites Look Before You Leap: Autonomous Exploration for LLM Agents.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Look Before You Leap: Autonomous Exploration for LLM Agents

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:26.818924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.940186Z digest=sha256:d4e88445df0855a5326f827fe3313ecdd35e835018bac8c0403d3668487299e8

Observation c41e9e68-7cc8-458d-8f78-953acbafccda · outbound

This paper cites arXiv preprint arXiv:2601.14050 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2601.14050 , year=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.959586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.959586Z digest=sha256:c4d8d3d57627c8975aa339afab2379d135255c4988402838719b86a65b2e5448

Observation a1c9e453-598e-4be7-979c-de5c7e49473d · outbound

This paper cites Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:26.507641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.999733Z digest=sha256:1bdffee4cdc68e4e66f029d8e92b9f90524b9dc32b47dd35e7fff77ac29c2a1e

Observation 81cb3325-a230-4c30-89de-3c6bf51209f1 · outbound

This paper cites Memento: Fine-tuning LLM Agents without Fine-tuning LLMs.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.038845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.038845Z digest=sha256:7933dd11a9b36d16593cd8565ff0981f630c839d628ca51660d328dd1a0aeec2

Observation 2972291e-fd12-47ea-83e2-d0fd5007ea01 · outbound

This paper cites Reducing Tool Hallucination via Reliability Alignment.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Reducing Tool Hallucination via Reliability Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.084902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.084902Z digest=sha256:90e94488009b36a523e9f7e70c749b8192d486f21debc36780e0be0923d6bc3a

Observation d28bd039-8e3e-45fc-a621-9b7d51d947a0 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.141889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.141889Z digest=sha256:d4e7352395988c7173b8e978d4cd89d1a4b53ac24fdc8bcd413e657901f8596c

Observation 6f511aa7-b0b0-485b-993e-0790a83398b0 · outbound

This paper cites arXiv preprint arXiv:2509.11543 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2509.11543 , year=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.173725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.173725Z digest=sha256:e33db9539cea43357fd3d200b24dff6766bc3d18cb826836896a071b065bcd98

Pith citing papers

No inbound Pith citation observations are available.