Pith. sign in

Paper Citation Record · LEDGER

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

As of 12 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.09555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09555 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:14:31.352194Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 507b7d8a-38a2-42b3-b2b6-0acf7a468d1a · outbound

This paper cites International Conference on Learning Representations , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations , year =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.976180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.976180Z digest=sha256:bd2c169d17c9b26f8d4d57f4c8320ec5e222e6a1e631725e1ac505f6067854a0

Observation ec0e4c63-4d3c-49c9-9246-db7a4424b54a · outbound

This paper cites 2022 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2022 , url =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.980248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.980248Z digest=sha256:85ab5226ff3d4c0c6a91e72c3a21562813c13873b7bba8ed3328f59d7d110c5a

Observation f414b360-34f2-4511-a4ab-4ad9c7342747 · outbound

This paper cites 2017 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2017 , eprint =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.983543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.983543Z digest=sha256:028e321562d3a0154efc8dbf1afb2d98a70ecadb48bd6951c0edb019161d27c0

Observation 9f3f7cbb-3c51-4871-9536-6cd3e0f3067a · outbound

This paper cites International Conference on Learning Representations , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations , year =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.987646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.987646Z digest=sha256:c56e48c02b8ed609ee4df4b549c1303b533c45b0916dd7ca42e0a2d288ba082a

Observation 3efc714c-d80d-41c1-91f0-89e9c2ece75e · outbound

This paper cites 2025 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2025 , url =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.157380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:30.996013Z digest=sha256:5522a4c8e3f6d2ac81d93cfda78241abccb6b54123675d7b90bf75b3dc9db0d4

Observation b4e1fdf8-05e5-4121-bce1-fb5acde55492 · outbound

This paper cites 2025 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2025 , eprint =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.999591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.999591Z digest=sha256:c5fd593f7062c4dfb73b5837dc332b6a020e2170e8525da27a2ddc94746d8d30

Observation 92734b7a-db07-4d6d-9659-552b77650e4b · outbound

This paper cites 2023 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2023 , url =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.003152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.003152Z digest=sha256:0b5e87825e83f6b5b40169130bf8c46a239fc8d1343e4f2b840ffc2c90515681

Observation 3bece669-2d40-414f-bb49-228eee50de32 · outbound

This paper cites International Conference on Learning Representations , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations , year =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.006428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.006428Z digest=sha256:2497a19a45e024102fc8b3c74ef32ff7a510fa44ace10ad0f4925ee906847407

Observation 99017870-c196-4656-8f59-c359729bf9e7 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations (ICLR) , year =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.129503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.010594Z digest=sha256:dcdf0afdfca5d08c80d22b39ed123a1531831c80e59cc6632a7a19542dd12f4c

Observation e40c3af5-00e0-48ec-bc7a-d4b0a9fe2f33 · outbound

This paper cites Second Conference on Language Modeling , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Second Conference on Language Modeling , year =

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.118597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.015953Z digest=sha256:966e055b0d8dcbbab6be24aedfaac123ea51aeb7237d257dd834a0889af3f135

Observation 6dd801f0-3f9e-4817-8d7e-3e1cfa13d9cb · outbound

This paper cites 2026 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2026 , eprint =

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.109156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.020099Z digest=sha256:cb47cfc4d0a1190144055094f1e7dd0c58daae3f83688d386b39adb9f3e9c42f

Observation 735f970c-752e-4b13-9537-801c6497fb0e · outbound

This paper cites 2026 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2026 , eprint =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.054296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.054296Z digest=sha256:8c5988f48a818e2d099ebc7ab818c4e731e820021349c975dc539942f52d1002

Observation d178ae83-6349-47c7-9064-aaad55c6bdf3 · outbound

This paper cites GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.058179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.058179Z digest=sha256:5c5a165c1f315a728f491df5080b5b4ec11f29d0b25d64782c2ff372e7fa8150

Observation 77e13218-f9ee-42fb-94a6-63aad1b95416 · outbound

This paper cites 2023 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2023 , url =

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.092993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.073931Z digest=sha256:b4b369d35059fc149d6e6af89166d050dc786f84f6a89d15c52f1a770d264da7

Observation ea3cc6d0-a8f2-415a-8700-49c02817f0e3 · outbound

This paper cites 2026 , month = may, eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2026 , month = may, eprint =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.082071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.078010Z digest=sha256:82a0ad12482b8d01f98b595ced057cdf346079d55783ef11f4b7dc01b7044d51

Observation afb1506b-0360-48a2-b815-4fcbb778c4b2 · outbound

This paper cites Proceedings of the 43rd International Conference on Machine Learning , volume =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Proceedings of the 43rd International Conference on Machine Learning , volume =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.071353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.093811Z digest=sha256:2809b064dfb57ea88f3ad7c2b4792179476c9ccd57b6d8cc2cf0beede71a71b9

Observation 9625b209-5e08-4a6b-9e98-aa3ff18ffc87 · outbound

This paper cites Proceedings of the 43rd International Conference on Machine Learning , volume =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Proceedings of the 43rd International Conference on Machine Learning , volume =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.060992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.099335Z digest=sha256:296b6e95ebd12f0dd0c75e7b20b88d44d3a37b6bee9752f7cf711880c29b1b22

Observation ec4464c1-269d-46b5-87f9-ef545c6507cb · outbound

This paper cites Advances in neural information processing systems , volume=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Advances in neural information processing systems , volume=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.125179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.125179Z digest=sha256:67910a23597ddb71f9c08237e53e18cb795a177a2534a39d1d7390043f6d506c

Observation c3291273-6664-455b-8734-e00f2be9b645 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.139246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.139246Z digest=sha256:53b1f5969bb47c0a7494f052a6e25fb191ddbee9452f19eb01b6a5e34d4eddc7

Observation 2fadadc6-3f06-4d56-bdb6-f4c7f8fafa4a · outbound

This paper cites Forty-third International Conference on Machine Learning , year=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Forty-third International Conference on Machine Learning , year=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.026672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.142247Z digest=sha256:90f8ff32ba283b4de6904d6a181b4b7aca7dfba73e27aa4808888e32abd2fe5a

Observation 94c8467a-5574-4b64-b110-113fa04877b1 · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.014141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.145122Z digest=sha256:fd8c0ffb63e7384b2caaf44b4476628693fc4874493f6ccc4929e41330305865

Observation e3d8046d-db4a-4506-925d-baf8001abceb · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.003465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.162221Z digest=sha256:d347f8d20f8973aefd0a7286d8080fd225b5d9a4cec469ac069172eb6ec72f6f

Observation c2e3fe46-1ea6-4bd7-a256-7a6315a69c4d · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.992407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.173862Z digest=sha256:1872f5ac455dc1367ba6f8986fef6436bec6a0792a162a15dacd7222f603e772

Observation 0fb8bede-4e9f-4bf5-80ef-f88c3eefe78c · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.177934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.177934Z digest=sha256:4f62eec6ed0b39f4bedd2651604723d2c4956e054d122cd078e87a6d7f8a25aa

Observation b1f247c1-3cb1-4910-83e1-40bf05a79537 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.973251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.183855Z digest=sha256:bcf5a8c470f31bfd740f62c016fb336443a6c20db0eba761d89cb33c7e84a5ba

Observation 7903429c-8806-47f5-980f-186ef96cb48c · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.186850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.186850Z digest=sha256:278574235907c7f71cbbe0e2971fd4c96ded024e115a072cb7c43fcd908f69dd

Observation 4593b1e1-e865-42c5-80b1-38575037517d · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.189966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.189966Z digest=sha256:1544a1f8c33183518ed9372951114638e232f72be1ccbf80f8f84cbe3ac8bdfe

Observation 9f45d545-e925-4af2-bfa7-27ec2a10c16d · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.193019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.193019Z digest=sha256:21d58f44992459ddcdd93407e990d6546f20415845f4f5cd45c75efe316cbfb7

Observation 71d40a3b-ce9c-42d6-80d4-a5be85594e34 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.953664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.196233Z digest=sha256:02a463c35c1b92102cd7d7162eba6a4bc0cb4cffabc3c6beb59c1b863df36086

Observation 6fa0560f-0a04-4218-8b47-06a2655faad2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.199401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.199401Z digest=sha256:28045fed33b16deff24a876d9abe40bdc81f4e0bdd2d0d53c68b4d5153ab875c

Observation ba50a996-3829-4f31-aa03-7ad9e382394e · outbound

This paper cites SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.203282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.203282Z digest=sha256:c2e73eecc7dbf0168ead9d44c54812bc5b02711dc7037ee6891b31fe7e9847cb

Observation 1a5adf60-4924-4522-bbdc-8762f73a4f17 · outbound

This paper cites u botter, J.; L \.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents u botter, J.; L \

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:31.941921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.207650Z digest=sha256:dcf7844893ad5a029b5f838dfa55e9bf2a45899097411d6b808eae820ba1deaf

Observation 332ac8d8-a995-4f21-9355-a4f861712fe7 · outbound

This paper cites SoK: Agentic Skills -- Beyond Tool Use in LLM Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SoK: Agentic Skills -- Beyond Tool Use in LLM Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.212116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.212116Z digest=sha256:f0426442975a44f78eec4e5042bb7f6974b261e00886261a12fe76bcf5e36145

Observation 2aab8971-1b3e-4fcf-8d37-b6971394b381 · outbound

This paper cites EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:14:31.556314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.217336Z digest=sha256:49c9e42102a146934ef34ed37ed7f4ed7a0603e52cf93ab32d742767c8cb0793

Observation ebec9a3e-062a-4e55-a20f-6ed588e00157 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.929641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.221025Z digest=sha256:91f621b307013fe71624d5b1d2c9a9a16e076947f55183c245f69e92d1964982

Observation 7609e894-bf27-43ad-9b10-a3b69843474f · outbound

This paper cites P.; Li, L.; and Li, Y.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents P.; Li, L.; and Li, Y

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:31.918665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.225101Z digest=sha256:ee38bd37ed375b351806e533db36a2ec98394454d6f5cb21a7f2ddf8de3ae2f3

Observation 503145b6-beff-4139-ba22-c6a09da9dca3 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.229838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.229838Z digest=sha256:be917c408bb8e4d5fd42ae40c9a64ad2cd2b28033ec1cd62dc898740b3cb7b7a

Observation 6a4ff8dd-f86f-4355-890f-0676403ec2fb · outbound

This paper cites SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.233000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.233000Z digest=sha256:15a7daa3c43f4ac1195e5a1dcf91a2debba7fb4f8a908967df5eff3e5df69ab2

Observation b2730c46-5290-4f27-89a8-d2e32580b505 · outbound

This paper cites HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.236567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.236567Z digest=sha256:7385ea4ea1c689adc26c1da80d4f74b330cd3e19cb72888f6557d7aae8cd2b96

Observation f0c112e7-467d-4bfd-ba97-f50e376d4d6d · outbound

This paper cites How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.241692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.241692Z digest=sha256:359e9d5ac3f678d07fe661c73085d57a3691b63d9e12224cfcbc902200fa842e

Observation 6bc7b0b5-192d-4e4c-bbdf-54a7ffb1cbc1 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Self-Distilled Agentic Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.246337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.246337Z digest=sha256:1007404d12c63d897d3fd2bba19ed38bd2cc8d2478d6f9be96650166e24a46d9

Observation 5921ee81-82c8-4cdc-b881-177b858037a4 · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.251390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.251390Z digest=sha256:e91ea8e77e41c94be0ee919e5237306c25a0bc2dd485893269346c3b9c50b056

Observation acaa1adb-a251-4f1e-9c7a-a594570efc28 · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.255877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.255877Z digest=sha256:86df30eb84049fa07a01ee38d6c77bcd44fafdaa444ebbd0922189a1afb43bf5

Observation 2fd1d420-36c9-458a-b1cc-3269af44cfce · outbound

This paper cites Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.259533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.259533Z digest=sha256:909a299ca19d4c861fc7fa05879de7bb80b31f8b24a82ca7c443ed4051017d33

Observation cc3f5913-16a2-4ea7-bcb1-91dc3ab8e900 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.264307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.264307Z digest=sha256:db68536b66b27659d689853ba1352e606a186a87d4feae0224d77d263e52e0a1

Observation 9135da0c-edf4-4732-9c47-80ad2c89eb9a · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.268830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.268830Z digest=sha256:cca8306968298b7eecd253dddbff450dc0dc6553a8935c2a40657e87b4d06a0c

Observation 1943922d-9eaf-45a5-892e-17d2320a0630 · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.273003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.273003Z digest=sha256:480c6bc832fab8e3281f77a4e5ab84a2fa06cd593e659334a923e819ee9cf092

Observation af4d37ac-b3d8-42e5-b92e-e11cc2b8acb3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.276375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.276375Z digest=sha256:8ac4e56db041ac1fda8f73529132d0c23da59a7f81718eb0a52eafbe7639f95b

Observation f55933d3-ef4f-4477-af85-e6428f2678e7 · outbound

This paper cites Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.279791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.279791Z digest=sha256:4ba0d9a58ca66b5d252dda1c7e3c4041e18f9ffe69e4679b7a482f115ba08431

Observation 79d3ff01-bf7c-4ee0-8f10-d36bb43efd3a · outbound

This paper cites Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.282967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.282967Z digest=sha256:81159278cd0497d95c6adf2ecb05009dfde083692ea91e6bc99593b425222250

Observation 5925ede0-620e-408d-85f7-5ffcc4a6cd17 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.900861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.285811Z digest=sha256:062d9024f9eb589546b0d96e542db215d89d267b57c157659866ba020be3f0a0

Observation 03d315e8-672a-4aaf-9c97-5582db8c0e50 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.289006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.289006Z digest=sha256:e5283fad6834b79ed01accb19eef0864bd0de5fa7c5b489956ff1c55d41391e6

Observation 65f63379-d7dc-45ae-a4dd-85fe6f5ba98a · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.293116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.293116Z digest=sha256:a87579c07803c8e2b0365c733d44f919e959167d8c89ef92eab64015419d7d8c

Observation 4a6e1d17-bafa-4b86-896e-574f0074307b · outbound

This paper cites TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.296243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.296243Z digest=sha256:a79af38cfcd28f7e572a25102202a264349dd0f701d930f201ca6dec646cf8a5

Observation 696b4123-bccf-47e3-9afe-e2687846f7ec · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.883537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.299306Z digest=sha256:a3a1b69933362f5b6beb7c1c3485e0c7c91da1a664cff76b7cadf034487becc2

Observation a6f46b7f-6602-40b4-bc82-886745447927 · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.302274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.302274Z digest=sha256:78c2536ae09a1704e97848895a2175870568153fb059ebd1fd45e9bb83f81567

Observation f746a172-24c2-4525-ab6c-2d609408231b · outbound

This paper cites Qwen3 Technical Report.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Qwen3 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.305540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.305540Z digest=sha256:79fc550b9077378607d73ffb52766a625d1fd73c7c602d9179b815f410024731

Observation c4a4b54e-0ca0-4b12-9818-4f47142fabfa · outbound

This paper cites Qwen2.5 Technical Report.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Qwen2.5 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.309205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.309205Z digest=sha256:a0fa821515469a278e1da97796a48e819fc0c6fc65fbcb9117e9f6bcc13f585e

Observation 53025679-a086-473e-8a08-d6103d765c5c · outbound

This paper cites Self-Distilled RLVR.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Self-Distilled RLVR

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.312419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.312419Z digest=sha256:ed7050ac26036ad7e73504b167a76017630c06fa002d2ebe8c7ff2c8e94d59a9

Observation f5ef13e7-1098-48c1-96ef-4159d503cbcd · outbound

This paper cites OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.316254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.316254Z digest=sha256:f5a51568f78ccba14f17f7027e7fa8b480b5c23ae7ecca66331a194464f93d7b

Observation 406e0b03-90a7-46b6-93ee-5ce3060a1db3 · outbound

This paper cites OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.319381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.319381Z digest=sha256:cd715b36e53fa54af038f552c9ca4a3d766ea46ce691488d59dc3c50ecdefae4

Observation f08f7ea4-21a0-4269-b632-0ab9f03f23e1 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.872955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.323263Z digest=sha256:398571dc01fa3bd671c4fc177ea490505b7091447baea6c6e8080a16fabd36c3

Observation 88209d9f-c0b7-47df-a118-3fe92ac2f159 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.860620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.328269Z digest=sha256:451dc3c4be6f64038c87791c963f8488dc3cbc63ce8882b1d03f1d2a04ee857c

Observation a9b215d6-3d46-4fa7-bae0-17f50f5065f8 · outbound

This paper cites SkillEvolver: Skill Learning as a Meta-Skill.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillEvolver: Skill Learning as a Meta-Skill

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.332594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.332594Z digest=sha256:4f8ac1ee0028f3c460ca5373eeb4841919654f0c249d4ddae1c53b9bb9aed171

Observation 0be07b05-d407-4886-b3e9-ad59563e7207 · outbound

This paper cites OPSDL: On-Policy Self-Distillation for Long-Context Language Models.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents OPSDL: On-Policy Self-Distillation for Long-Context Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.336027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.336027Z digest=sha256:c004dec34786e23c9922a1043a9060631b81bec8f99c44623d2c1b744d7c324b

Observation 09ad5cc2-8a49-43e4-9f4d-fdf361eb19f7 · outbound

This paper cites StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.339899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.339899Z digest=sha256:3c9283407d3089a9d273dc675e35905dc3818db01a40feb7d7b9dc1839078787

Observation 2a11cf51-f63a-403c-aa6c-549e0cf300b3 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.848387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.343797Z digest=sha256:4f7d1744e5b2dbab60d494f0bf1269cda83af66fe628799a53f3d6a7fb483037

Observation 2782d5f9-46e7-437d-90e8-8929a3f50b8b · outbound

This paper cites SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.348326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.348326Z digest=sha256:60c66f5f741b342216c0426cb074b1bbe14b823837683ad2a6a95acc7f920cbb

Observation 599e9b4b-7b41-4dfb-b197-fbe4ed3e0544 · outbound

This paper cites SOD: Step-wise On-policy Distillation for Small Language Model Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.352194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.352194Z digest=sha256:6ccf3c695d163ffe49e5699a92d544190b051ab72cc0eecfc33602382d4d4362

Pith citing papers

No inbound Pith citation observations are available.