Pith. sign in

Paper Citation Record · LEDGER

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

As of 12 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.09555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09555 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:14:31.352194Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 507b7d8a-38a2-42b3-b2b6-0acf7a468d1a · outbound

This paper cites International Conference on Learning Representations , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations , year =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.976180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.976180Z digest=sha256:a66786e6ff7c72b872fd17edbf861ab4a3f77ddbe9603d5b41372d3eb0c2f343

Observation ec0e4c63-4d3c-49c9-9246-db7a4424b54a · outbound

This paper cites 2022 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2022 , url =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.980248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.980248Z digest=sha256:c1937e0cf9397daf9725bf56cd9531dcd0a7351dd10bb8bb82887a695d5b7eb2

Observation f414b360-34f2-4511-a4ab-4ad9c7342747 · outbound

This paper cites 2017 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2017 , eprint =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.983543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.983543Z digest=sha256:cb06b1c02889b506e1586f83f7731e8550ce1b24cb18d7a2ee57c75ce874793e

Observation 9f3f7cbb-3c51-4871-9536-6cd3e0f3067a · outbound

This paper cites International Conference on Learning Representations , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations , year =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.987646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.987646Z digest=sha256:d4b5b5a8a8add6bd4e002df53a1f10ef2fada74b3635197eba400d4fdde39817

Observation 3efc714c-d80d-41c1-91f0-89e9c2ece75e · outbound

This paper cites 2025 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2025 , url =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.157380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:30.996013Z digest=sha256:b233af01992260b5e98a61882c3447b9457c725baea39f52c0d1c286214d504e

Observation b4e1fdf8-05e5-4121-bce1-fb5acde55492 · outbound

This paper cites 2025 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2025 , eprint =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.999591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.999591Z digest=sha256:5230bb419779eac769680d342ddf0285da1d3b73da22b2b095e8c2b7710d4e36

Observation 92734b7a-db07-4d6d-9659-552b77650e4b · outbound

This paper cites 2023 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2023 , url =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.003152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.003152Z digest=sha256:5b7f351d12df90c5a481de4a088b05b4847f681b8c3c0932b5f2f3b683f1bf00

Observation 3bece669-2d40-414f-bb49-228eee50de32 · outbound

This paper cites International Conference on Learning Representations , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations , year =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.006428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.006428Z digest=sha256:bd45f2de62b6a9aa7027604dda8fc813de5c7f0f6443a5d5816a9bd3a2ba756d

Observation 99017870-c196-4656-8f59-c359729bf9e7 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations (ICLR) , year =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.129503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.010594Z digest=sha256:bf31c6d9f9b7f0b11fa0c207dae4fb4a54d8e8db0be127ed7d0cea40fca31976

Observation e40c3af5-00e0-48ec-bc7a-d4b0a9fe2f33 · outbound

This paper cites Second Conference on Language Modeling , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Second Conference on Language Modeling , year =

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.118597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.015953Z digest=sha256:7d0dd945502fcb37818f5efb4dc063db1dc6987417eeece4031703c9c1afd1fe

Observation 6dd801f0-3f9e-4817-8d7e-3e1cfa13d9cb · outbound

This paper cites 2026 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2026 , eprint =

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.109156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.020099Z digest=sha256:39389a93903ccdc60e26e40597c73f7380a2ff769af0a10b49a15c411cba61d5

Observation 735f970c-752e-4b13-9537-801c6497fb0e · outbound

This paper cites 2026 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2026 , eprint =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.054296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.054296Z digest=sha256:bd9842c29959b88cebe02d68edadaea3b2b743bb8e53a46c62247798e49059f7

Observation d178ae83-6349-47c7-9064-aaad55c6bdf3 · outbound

This paper cites GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.058179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.058179Z digest=sha256:34182587eb747cfc06353e3490b2f436d29357abebaa45706ad99fe1b0269c34

Observation 77e13218-f9ee-42fb-94a6-63aad1b95416 · outbound

This paper cites 2023 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2023 , url =

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.092993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.073931Z digest=sha256:7032119e22cb423230c00e0569e9fca8a2d05b1363d76c9dc006aba375ff24d0

Observation ea3cc6d0-a8f2-415a-8700-49c02817f0e3 · outbound

This paper cites 2026 , month = may, eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2026 , month = may, eprint =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.082071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.078010Z digest=sha256:8d0566ddb08759105d1c873e1472b8fa3cbab7c9d1c6a23c1c3668fa22e9ea3a

Observation afb1506b-0360-48a2-b815-4fcbb778c4b2 · outbound

This paper cites Proceedings of the 43rd International Conference on Machine Learning , volume =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Proceedings of the 43rd International Conference on Machine Learning , volume =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.071353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.093811Z digest=sha256:c8690c8cc8c100f0238b0116482322942f3bf50634b594c50693933143bf59ea

Observation 9625b209-5e08-4a6b-9e98-aa3ff18ffc87 · outbound

This paper cites Proceedings of the 43rd International Conference on Machine Learning , volume =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Proceedings of the 43rd International Conference on Machine Learning , volume =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.060992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.099335Z digest=sha256:2b1a3aedf4fa575680eab41ba4575da2f306d6db2ea3b75868a78e3c4af29fbb

Observation ec4464c1-269d-46b5-87f9-ef545c6507cb · outbound

This paper cites Advances in neural information processing systems , volume=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Advances in neural information processing systems , volume=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.125179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.125179Z digest=sha256:bc2367da7929afde548460cd624046105e49a8099ec96323fb780a1bc922a0f2

Observation c3291273-6664-455b-8734-e00f2be9b645 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.139246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.139246Z digest=sha256:58c1bcbcacacf2cc0cf86f65b3be859178495d6207d8ef078f2b2aedfe93b8e8

Observation 2fadadc6-3f06-4d56-bdb6-f4c7f8fafa4a · outbound

This paper cites Forty-third International Conference on Machine Learning , year=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Forty-third International Conference on Machine Learning , year=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.026672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.142247Z digest=sha256:c2027c68214307ea8c2e49e250305c676f4a4fd353389999f940b3066f0ac104

Observation 94c8467a-5574-4b64-b110-113fa04877b1 · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.014141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.145122Z digest=sha256:b895e6632152c8957efd167dd81988f7ce9d84583d7b194301aeaa3f7ade4b73

Observation e3d8046d-db4a-4506-925d-baf8001abceb · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.003465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.162221Z digest=sha256:8e07dd4c795dfccaf3349fc4f9f4b104d219a3404d837ed4a5646b95c913516a

Observation c2e3fe46-1ea6-4bd7-a256-7a6315a69c4d · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.992407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.173862Z digest=sha256:b5d7a5436d96e46440603c82bacee1126a6c4b620d5e36aefbf3bee59c4b085a

Observation 0fb8bede-4e9f-4bf5-80ef-f88c3eefe78c · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.177934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.177934Z digest=sha256:31e092ecf184bc70d4a8ab4fb7999312c40ce7f81d02056c3c242e7cb6624791

Observation b1f247c1-3cb1-4910-83e1-40bf05a79537 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.973251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.183855Z digest=sha256:db62ae8b92eb05436f0f45227cf4ce72ded857a4022642d3faaddf29f1505e2b

Observation 7903429c-8806-47f5-980f-186ef96cb48c · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.186850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.186850Z digest=sha256:e19f9d40f252eb1554dbf8106fe0686aa8f44ef41c4339a71d009621e82a5087

Observation 4593b1e1-e865-42c5-80b1-38575037517d · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.189966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.189966Z digest=sha256:9607bf10df1b7134e69747eb89d1c8d794295eb41794891a17f5be10e30f3b92

Observation 9f45d545-e925-4af2-bfa7-27ec2a10c16d · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.193019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.193019Z digest=sha256:461631b1a22eee47030caefede903b1d5adcca5b28bae3aa12a309f23c2cd7b9

Observation 71d40a3b-ce9c-42d6-80d4-a5be85594e34 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.953664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.196233Z digest=sha256:7c364eb61e31be37b41392b122787b52910e6d327f91fdd588f96cf38a387a9a

Observation 6fa0560f-0a04-4218-8b47-06a2655faad2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.199401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.199401Z digest=sha256:c2d911d17b5a999994152eadd90466ee1d554816233e2df106ef2bd3cd0b6220

Observation ba50a996-3829-4f31-aa03-7ad9e382394e · outbound

This paper cites SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.203282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.203282Z digest=sha256:3d0256e8d370b7d881f752838e28406594ccea0e58d51f0be955e788b232ccd9

Observation 1a5adf60-4924-4522-bbdc-8762f73a4f17 · outbound

This paper cites u botter, J.; L \.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents u botter, J.; L \

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:31.941921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.207650Z digest=sha256:ecccab42cbf44653dcac05a9871020523d37329e1ef3fc590ba69898cf87b803

Observation 332ac8d8-a995-4f21-9355-a4f861712fe7 · outbound

This paper cites SoK: Agentic Skills -- Beyond Tool Use in LLM Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SoK: Agentic Skills -- Beyond Tool Use in LLM Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.212116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.212116Z digest=sha256:aa855c88f47a0b16db20c03f5aa738ae1a9306cd2740ce09145a2daf7f2c9efe

Observation 2aab8971-1b3e-4fcf-8d37-b6971394b381 · outbound

This paper cites EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:14:31.556314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.217336Z digest=sha256:267c09de682b52adb7e97f6a4bad31077ceb803ab4e703a699aed98eb283961a

Observation ebec9a3e-062a-4e55-a20f-6ed588e00157 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.929641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.221025Z digest=sha256:1fdc627d943f813b0812507a7dae7103f99fbbb96bc1c9f54f9cb928e6a32b4f

Observation 7609e894-bf27-43ad-9b10-a3b69843474f · outbound

This paper cites P.; Li, L.; and Li, Y.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents P.; Li, L.; and Li, Y

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:31.918665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.225101Z digest=sha256:82e0f4884d1c25de48f6ac1785b8056a1d7c519ce19d681c5ac5e26ddf61dfa7

Observation 503145b6-beff-4139-ba22-c6a09da9dca3 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.229838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.229838Z digest=sha256:78e5f32906f6eaafd3e8b9cd4ce77e7f971b592b7278985ca2e4008aee95082d

Observation 6a4ff8dd-f86f-4355-890f-0676403ec2fb · outbound

This paper cites SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.233000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.233000Z digest=sha256:c0a64a1c6ccf0428edd9f6275af36c2cb100b2afd621a6376536ceb70ed10f97

Observation b2730c46-5290-4f27-89a8-d2e32580b505 · outbound

This paper cites HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.236567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.236567Z digest=sha256:7aab798dbe5d7094633fb024d828889c470403a8b80fb83d55b4f6c41f02a639

Observation f0c112e7-467d-4bfd-ba97-f50e376d4d6d · outbound

This paper cites How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.241692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.241692Z digest=sha256:ce6cd9d76fec1f7da8f5a7e7370ce54c0fe7b3f0a3bae3b38b274c64a2e4ff6f

Observation 6bc7b0b5-192d-4e4c-bbdf-54a7ffb1cbc1 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Self-Distilled Agentic Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.246337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.246337Z digest=sha256:21812ac6dde63a3a9851b25a1893a243aa73a08f951fef8063dd1c0b5ac54836

Observation 5921ee81-82c8-4cdc-b881-177b858037a4 · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.251390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.251390Z digest=sha256:c7907fdf5e9b449a44f1818d4c9fa9cdecf3708e77b64dbbf1c83ac70e43a5b3

Observation acaa1adb-a251-4f1e-9c7a-a594570efc28 · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.255877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.255877Z digest=sha256:40cd2cd814d745b014a1178919e7d33387a3664a101e901a8d209252522efd79

Observation 2fd1d420-36c9-458a-b1cc-3269af44cfce · outbound

This paper cites Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.259533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.259533Z digest=sha256:6512d583ef38c263993e34ff1299899a5e0eb803cdb8a1fac321d950afdedf07

Observation cc3f5913-16a2-4ea7-bcb1-91dc3ab8e900 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.264307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.264307Z digest=sha256:2ec2442661c8e71c6b89df83b94e7bd63059921f9a1deca33ea697223b7d4787

Observation 9135da0c-edf4-4732-9c47-80ad2c89eb9a · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.268830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.268830Z digest=sha256:25dcc1be3f2f121f327f517c9aa0c760e621747f77f988bf05367a75ab5f72c2

Observation 1943922d-9eaf-45a5-892e-17d2320a0630 · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.273003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.273003Z digest=sha256:94b540141f907f63d50725c0d14d31625ba56c8527a1339a0314409ad829d47c

Observation af4d37ac-b3d8-42e5-b92e-e11cc2b8acb3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.276375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.276375Z digest=sha256:5d63c9cf5228d508d924ff583d1915b1dfc0de6f29ee32cd18ad8ae149a6a92f

Observation f55933d3-ef4f-4477-af85-e6428f2678e7 · outbound

This paper cites Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.279791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.279791Z digest=sha256:03abf59e5795c74fc6ed1b415f7bc3ad94403bd232428edc1c50837690c60469

Observation 79d3ff01-bf7c-4ee0-8f10-d36bb43efd3a · outbound

This paper cites Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.282967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.282967Z digest=sha256:e25c7685678188db839debf633e269937f16d02593bd9ad759aaba4dcffcb0b1

Observation 5925ede0-620e-408d-85f7-5ffcc4a6cd17 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.900861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.285811Z digest=sha256:3e60a7817c93843952c3d35f72d2a33f5be7db5a32a83a04c0be288429d00129

Observation 03d315e8-672a-4aaf-9c97-5582db8c0e50 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.289006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.289006Z digest=sha256:f667c3ec69d4373d07636a6f12927284b0bd24f79827317d60ada88e1d90f176

Observation 65f63379-d7dc-45ae-a4dd-85fe6f5ba98a · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.293116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.293116Z digest=sha256:4a52d1d52de321e0347886da66e818c99298ed80bc9e4298db2dff357fa9a5d3

Observation 4a6e1d17-bafa-4b86-896e-574f0074307b · outbound

This paper cites TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.296243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.296243Z digest=sha256:f16c9d647a61427e194ab7b8fd278cd3c306ac8c66a438255691d59206694c50

Observation 696b4123-bccf-47e3-9afe-e2687846f7ec · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.883537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.299306Z digest=sha256:743eb221a15d02ba2b397304ca3135906b77814c0d103a52efada4408ab7675c

Observation a6f46b7f-6602-40b4-bc82-886745447927 · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.302274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.302274Z digest=sha256:85aacb132e3af83271930be48589ffc9d8a233190bc30a9653bf5458395d66d1

Observation f746a172-24c2-4525-ab6c-2d609408231b · outbound

This paper cites Qwen3 Technical Report.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Qwen3 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.305540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.305540Z digest=sha256:fc89faa9d9702a65f602bc294837f2e674ed479edf4113c507297e618e4b0d37

Observation c4a4b54e-0ca0-4b12-9818-4f47142fabfa · outbound

This paper cites Qwen2.5 Technical Report.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Qwen2.5 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.309205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.309205Z digest=sha256:df157a03c17deb5e5ba4cbe67a18fa19588cd7fd1143347a878e40d552fd0273

Observation 53025679-a086-473e-8a08-d6103d765c5c · outbound

This paper cites Self-Distilled RLVR.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Self-Distilled RLVR

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.312419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.312419Z digest=sha256:40d240c2e30e0494ec2e15208a83a3fde4b72ba22bd3cc7849a9dcce1b9f977a

Observation f5ef13e7-1098-48c1-96ef-4159d503cbcd · outbound

This paper cites OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.316254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.316254Z digest=sha256:34f1d63c6d0d0c34a77efa70419a4835204c0cf2cf6a4fea7726f210e8b2dd5f

Observation 406e0b03-90a7-46b6-93ee-5ce3060a1db3 · outbound

This paper cites OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.319381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.319381Z digest=sha256:bff0057c8132cdbb5648b94e5894a14847fb42d80ce60380729b7b196f13127c

Observation f08f7ea4-21a0-4269-b632-0ab9f03f23e1 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.872955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.323263Z digest=sha256:44366669a272deaa7aa947a231481cac346c5c1b356f8834d3fc3f186aec7f27

Observation 88209d9f-c0b7-47df-a118-3fe92ac2f159 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.860620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.328269Z digest=sha256:349340d2c450c418c3739b3c7a9478135f5822b60ae8e3446f8e438b566bbe85

Observation a9b215d6-3d46-4fa7-bae0-17f50f5065f8 · outbound

This paper cites SkillEvolver: Skill Learning as a Meta-Skill.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillEvolver: Skill Learning as a Meta-Skill

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.332594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.332594Z digest=sha256:1cc9d32affc081becdd85c9b84e24b74d3952fd2637bb180ce6dd0a52423b269

Observation 0be07b05-d407-4886-b3e9-ad59563e7207 · outbound

This paper cites OPSDL: On-Policy Self-Distillation for Long-Context Language Models.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents OPSDL: On-Policy Self-Distillation for Long-Context Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.336027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.336027Z digest=sha256:c472e4f3554662ae098d81fb22789b01d11217a918118d96b581a92a702fd765

Observation 09ad5cc2-8a49-43e4-9f4d-fdf361eb19f7 · outbound

This paper cites StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.339899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.339899Z digest=sha256:393b26eb5d63da8b20ac5b3cb3016884c39477fa569b1055155148b684b79aab

Observation 2a11cf51-f63a-403c-aa6c-549e0cf300b3 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.848387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.343797Z digest=sha256:9ac598bd2d9859a068283570dee1ca6f835c4f22a2388938632a19f84ceda7be

Observation 2782d5f9-46e7-437d-90e8-8929a3f50b8b · outbound

This paper cites SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.348326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.348326Z digest=sha256:22f4984a0419810b414961ed4afe885608b46f5c0a92411c45fc3734688eabdc

Observation 599e9b4b-7b41-4dfb-b197-fbe4ed3e0544 · outbound

This paper cites SOD: Step-wise On-policy Distillation for Small Language Model Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.352194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.352194Z digest=sha256:581a4c3cd1f6701d9067cef3deeee2d6c87205c32f8cb97e1fe745f2d16b4410

Pith citing papers

No inbound Pith citation observations are available.