Pith. sign in

Paper Citation Record · LEDGER

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training

As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2606.02355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.02355 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T14:33:00.408984Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T22:51:49.158185Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T07:24:21.967669Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch23

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7c3edcb-85a9-4a69-986e-7df18891d504 · outbound

This paper cites Aho and Jeffrey D.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Aho and Jeffrey D

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:207eaf982458054fb4c85a25e62c75ca28b2c83c77c05b688aaa87e1b15e1d86

Observation f85de3f7-2069-4e57-a1ac-c69ce3ca98f4 · outbound

This paper cites an unresolved cited work.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:ef11bc11195c3f1b64c68919cd6ee1e483a2f04cb68b5b2d9a1b6961e6d86939

Observation 5787217c-216d-4f65-806d-db646afb91c5 · outbound

This paper cites Chandra and Dexter C.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Chandra and Dexter C

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:42:18.091135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:9c955de5bdf68d4f6bda4747e41b2492710a52d7b3afef0cfd37a9f8d7cdb5c2

Observation de959fcc-80c6-4cc3-b73f-2844d00fc798 · outbound

This paper cites Scalable training of.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Scalable training of

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:ed43a01e361555a63524ab31125d746a173dfcb82a144e0594cd7b96a3025372

Observation 23b150b9-93d5-499b-a65b-2b0d6c573242 · outbound

This paper cites an unresolved cited work.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:0242ab138481b5ee863098def16129b1346a893860a1a996c5cab7ade1e29c35

Observation 51c9992b-4ebe-493d-a516-7cb1b015b2b1 · outbound

This paper cites Tetreault , title =.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Tetreault , title =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:2c1862976d8270980a2ca8974cbec7f2141456c291df5fb44c2fbf70d7ad80d8

Observation 4f67183c-b721-4e5f-96f2-f28423392952 · outbound

This paper cites A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:e901cbb7ead19d0636564ae2bd48e38ecddc8a6065044bd1f3439ad6bcf7ce07

Observation 3e420bb1-fd22-4df5-b64c-bd3c5788ae74 · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.940397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:0557d1956488219f4eb82d7b32b82c415830116ea1a92e58fe8e1247dddc1686

Observation 24cbd096-6109-4bb2-9f40-25890e9e900d · outbound

This paper cites Dynamic Dual-Granularity Skill Bank for Agentic RL.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Dynamic Dual-Granularity Skill Bank for Agentic RL

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:16:23.883568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:2520896eb6df4b1d4e2b753c314506603633e667c90c8569529b0ef16dfbab08

Observation ea0ff959-6359-41b4-8645-9ec4e6f7bc0f · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.889061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:307740975a7fef1458858aa56dd891cc220a06d7362924acf356a65845dcb29b

Observation 00affdbf-b378-48f0-a1a9-3c7d616ce018 · outbound

This paper cites Arise: Agent reasoning with intrinsic skill evolution in hierarchical reinforcement learning.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Arise: Agent reasoning with intrinsic skill evolution in hierarchical reinforcement learning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:23.898373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:27916758d39fd81dee54194b37d3ffab23e0ef4acfc5f1142dabc39d0621dbcc

Observation 042e0ba2-0df4-40ae-87c8-2083ffa7034b · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.901767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:8e42aca54df18ca6cc317ac20583f885c608b9b57ee87740fd2a2d45e2ce77fe

Observation 1fa1e855-6224-4b7b-a3dc-ac1b4fa03c0d · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Advances in Neural Information Processing Systems , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:3566a29e099c013f06356b9f3936f717c171b4fe1d1c49f4b1ded500987238b3

Observation afe4a27f-2d1d-4260-9714-9a8e970d022e · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , year=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Proceedings of the AAAI Conference on Artificial Intelligence , year=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:01df3c3ce63b89e531c72c6006d7af7068f4ef42cfef96553bb0701a8caaa5f0

Observation b9bfeafb-591e-4ae8-8461-5c63fa65ed0f · outbound

This paper cites International Conference on Learning Representations , year=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training International Conference on Learning Representations , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:16f6e9f25ab729419f31d88bd3e9c9361ef051341f7bc2e6d76dad1781f1324d

Observation 6ea5c6d2-e63d-4371-85e8-41be3235031e · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Advances in Neural Information Processing Systems , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:c2dd80dba011f6ee0b855b4b38a5084ab334fb580adb679ef9a96eac58a5e889

Observation 583d742e-a60d-475f-aa98-435e268757e2 · outbound

This paper cites AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:23.895486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:57c5749af2a763f5b9f7e9a5e58c4b6347f812db42cd59a5f0c07316cd6334d0

Observation 373e9909-1ed6-4f11-81be-f901b1e7e36d · outbound

This paper cites International Conference on Learning Representations , year=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training International Conference on Learning Representations , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:ba0d0cbe199ebb2705a8eb7de7cc53d670662d5e3969468ca5fccadc3c1915fa

Observation 9995315c-f420-4a84-b9bd-b3cf5ecc40e3 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.892126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:b2e8798d5848ad900161810b080284121d63fe9ca8ff73a0e9488af74eb60234

Observation 65cb10c8-1ace-4f2f-a748-6600542717ee · outbound

This paper cites Tree search for llm agent reinforcement learning.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Tree search for llm agent reinforcement learning

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:23.923897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:a1bf5ba62bd7f4dd6a4086429289e5c7f54a028f40b215d46192947de2cfa318

Observation 2fdd18d3-aac2-43fc-a2a2-7cb9aad31500 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Group-in-Group Policy Optimization for LLM Agent Training

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.874286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:4424a161b24de92e523b73eabfa8ce0917303192790a72ff80646846c378a99f

Observation 24f30cd9-5e82-445f-9802-4b69b2865469 · outbound

This paper cites arXiv preprint arXiv:2603.08754 , year=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training arXiv preprint arXiv:2603.08754 , year=

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.926830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:33872fb6e38bebb8b3b45ca461389fc8b0d05a5887a73a49fade5471457572bf

Observation 2af2748a-b2ed-4162-9698-d1db558cb154 · outbound

This paper cites arXiv preprint arXiv:2603.03078 , year=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training arXiv preprint arXiv:2603.03078 , year=

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.937032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:2abde9139fba526ccc42eb7d04fe4327e5dc56ddcfc874ab2ac0e05dd9435094

Observation 95c90cbf-0e8c-4441-b627-c857aa5d77a3 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.933497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:85fc807f759dace034c74898f5db68733df56132dd8f57551a7a7b613650b666

Observation 4b69b58f-aa39-40de-b262-02a844b3d7b3 · outbound

This paper cites MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.947110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:479b3610b766ea28416ea2fceacc54722802c736d7374b29e51857a36753f06e

Observation d501a210-3cb5-405f-bff5-42b65864e08f · outbound

This paper cites EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.957095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:92dce0f804cf60ad49dbd31f059077a5151bcfeb8d00c28b0cc9842b0403e90b

Observation 83b74d42-2047-4d3b-8f56-9c42c8448584 · outbound

This paper cites SimpleMem: Efficient Lifelong Memory for LLM Agents.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training SimpleMem: Efficient Lifelong Memory for LLM Agents

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.886381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:f329ba655591cf05697ba1994ba705862b765521219ed8de481f129dc68a71a1

Observation 55396aae-7a40-42cb-bfd7-0718b1fea377 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.877335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:410f42a7c7b40328738bc53b69f3bd6620b8965b8d11171284f5a9f55d6ae95e

Observation b0552a8d-3632-4538-96ea-8e02ce667168 · outbound

This paper cites Advances in neural information processing systems , volume=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Advances in neural information processing systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:ff0f06b722fa8520326f9d30b7230e357eb769d8ae9c0f50f07f797e71064d59

Observation c3f5b415-4c40-4d00-a53f-729b798fb1ed · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.904757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:66038dd22f19fa9a74ee16a6c8278ab4e01196cb2518a09648b735f3277a4ccb

Observation 567f7c6c-6417-4799-8242-3ce7723d3ab8 · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:a5379de9c1f7e4aebc7f8600ebedf2bd6cc2296d74a7ba7faf949e1233735b2f

Observation 3b8ec7d5-0b75-4e95-87f4-518b8a547c21 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training WebGPT: Browser-assisted question-answering with human feedback

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.880437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:4c88ddff6e6ce8b201ed68ac9e1689b282044362dcf3c01a1274e25cf3f54233

Observation 5d8d7ca4-61ad-4787-9cb2-05c8e7a187d1 · outbound

This paper cites Advances in neural information processing systems , volume=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Advances in neural information processing systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:11541ab4f5dfa07ca45dc7e6431bff5b791a6d327dddaaefde808a5b3a5de6bb

Observation 4e8dbb42-a2ff-4e2a-b88c-001e7d34f15d · outbound

This paper cites nature , volume=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training nature , volume=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:2e2d7e820a76339fa87cfaeb13674d603f68af8702ad0216414b7b144b157b8b

Observation 38b6f1c7-a6f1-429c-b5cf-0c0208c01539 · outbound

This paper cites Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , pages=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:41f5f3d064f33dfea321fc368e70229f7c07fa2a2540dd38c1a3cdd504f27906

Observation 0101eaa4-0d34-442b-bd2d-57357d55f2e1 · outbound

This paper cites Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:0be4b483abe5fa5165bd72109e6186784301d1bf78f9f5cd255663b09323cc5f

Observation 6638e893-ea36-4086-aeff-aafae6be1d09 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:42caaa5b1709fb532746c8427c2941166d4fe4b67f60303d0ff67aedd37f87da

Observation e73b26fe-012c-41a7-8cf0-a8486627ae88 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:23.950721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:44dce153726edf2f4726134ea476c8e4959da378a550a9673918c75386156794

Observation 6759baf3-9ac4-4d18-a34a-0ade98708a39 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.954039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:4b7acca715394357106821595a3bf95180bf1a65119b229776b5e91a20c37ca1

Observation e0c6c9f2-3dda-46d5-8b59-7eb88bd81ba9 · outbound

This paper cites Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.920593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:7e51bf876ac8e287e6ad5091d48673826c94481ace7973a4334121ba70d66120

Observation 2e23d26e-160b-4cee-8734-d8aaf6893079 · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:23.914071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:14957030ffc96067ef462ad6ccd4b7c05c432c8d48a76643929371d49fd16a30

Observation 0bff4fa0-2046-4936-924d-e42815f36a12 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:3f0181a9e4f58e20e5300de6c2b915a4ff7d4a395fe2617d47a2a8bd1a384e80

Observation 0e8916c8-2fc7-468c-b68b-306bb611548d · outbound

This paper cites From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.917364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:08ac1e59441a6613cc5bacedfb608744aa7bafe3808e061f5147158162f62a4e

Observation d757079b-5053-4254-a5e2-deba08a82383 · outbound

This paper cites A Survey of Temporal Credit Assignment in Deep Reinforcement Learning.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training A Survey of Temporal Credit Assignment in Deep Reinforcement Learning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:23.930228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:8523f485e1f31e2280bf63a2341786bffb656306f46e9da4c87c321572351a83

Observation b1d74dbb-ec2c-4ebe-8bd7-71095a4027d2 · outbound

This paper cites 1998 , publisher=.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training 1998 , publisher=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T14:33:00.408984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:6edd5164d7f7a853657d7acfe27f3832aa48a9471ae0a516103e9180c5886e59

Observation f3a9b7c5-7764-4e36-aac8-39a7f268ae5b · outbound

This paper cites Deep Reinforcement Learning in Large Discrete Action Spaces.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Deep Reinforcement Learning in Large Discrete Action Spaces

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.943730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:67b746c0ef232da6c921966e9f3736809feae969f90a1ebb6f9f790423770ac3

Observation 6b95f177-45a9-479a-bd6b-1707712d481d · outbound

This paper cites Proximal Policy Optimization Algorithms.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Proximal Policy Optimization Algorithms

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.910869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:c1eb0cf13c636fdd2a6d745d7d7a37161250273e86de3a73c1e5b8cb50735b12

Observation 453b1ca1-d603-4e5f-bc28-0ff5c8f192a8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:23.907962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:fd4e8ce0973f36d5b1500b2b1e461bbbb867edd8858824199d5e3fb509a8c691

Pith citing papers

Observation cc2f660e-94ac-46f5-9b76-8877455630b6 · inbound

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation cites this paper.

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:21.969035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T07:15:44.387347Z digest=sha256:e9857648180f9f052a7e677b8ae121b8778f2ca8015d7d9ab92d7362f5dd948d

Observation 23a05fea-0bd5-4aee-93f2-bac8687bf5f0 · inbound

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents cites this paper.

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T22:51:49.158185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:51:49.158185Z digest=sha256:a210408d807c8f20599ba61d456648dd6f0b449d9e88d374983c6bb434f02eb0