Pith. sign in

Paper Citation Record · LEDGER

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2607.28076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28076 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T03:23:33.135384Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e8172bb-79c1-4e07-b25f-fba0ce655d3f · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:28.501637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:28.501637Z digest=sha256:1255e0a517a6f1beb78ad83f9c066d9afdaf25f5790d2e2eac897155f70602e2

Observation fc31096d-43f2-4f8e-9935-2b9f4e82ed6a · outbound

This paper cites International Conference on Learning Representations , volume=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning International Conference on Learning Representations , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:28.555115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:28.555115Z digest=sha256:919ce54a715917b8cc9f6e083f251bb10c5a75ad82f2c00f1c5d0dfb381224a6

Observation ff66f73c-6472-40f4-a603-001d30a933bd · outbound

This paper cites International Conference on Learning Representations , volume=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning International Conference on Learning Representations , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:28.612759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:28.612759Z digest=sha256:bba0fe779ab0affc274498f9a565fa0ce9d33ee1ab80785cbd609d3aebf748a5

Observation 702411f5-9bf0-409c-8c52-aa436248288b · outbound

This paper cites SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:28.686482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:28.686482Z digest=sha256:c340fb695a9ef4ab77fb0a789165a52085b7392f95e17f7933f8b94659de1dba

Observation ecd3a71d-e623-4109-8a4f-04109f6b960b · outbound

This paper cites arXiv preprint arXiv:2511.10643 , year=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2511.10643 , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:28.766509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:28.766509Z digest=sha256:5c74909f0bc6f607e175de906f820fa18185fee25634fc71e77a1a689a936b85

Observation b6733b28-0330-44a9-a61d-f415abe92211 · outbound

This paper cites SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:28.881458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:28.881458Z digest=sha256:3666e31f0b05f3183670f62a7f3855ee8334ca349a162fe3949c02d3f36f321d

Observation b46b6cc4-bf9a-4c43-a2cc-4b3c902ec7fc · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:28.979456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:28.979456Z digest=sha256:cde940bbd3128a7166751d58c4df3d9d51089fd0852e943f7bffd965ad5a5927

Observation 253bba33-4b18-48a6-a081-dd3df9188149 · outbound

This paper cites TIP: Token Importance in On-Policy Distillation.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning TIP: Token Importance in On-Policy Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.066401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.066401Z digest=sha256:b50bd65fc72b6566930711e60664dc61f9b4c454e46edc720502bad40297793a

Observation 9893034f-faf8-49cc-8406-e8ab0adc43ed · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.135910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.135910Z digest=sha256:42cee3d0bbc2a3c7186ff9a0c8e1fd0b533874efecced3281964aa91ca445bb2

Observation b9a2bbab-cb7c-4910-8365-1ea8ff42b21a · outbound

This paper cites Reinforcement-aware Knowledge Distillation for LLM Reasoning.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Reinforcement-aware Knowledge Distillation for LLM Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.261239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.261239Z digest=sha256:7d83e33dbc5c9095f61a81c4f36c59c3968dcd9df544b08b55cb10dd9657f9fe

Observation d403953d-0e33-4607-bf60-1684d2c52883 · outbound

This paper cites Self-Distilled RLVR.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Self-Distilled RLVR

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.390980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.390980Z digest=sha256:fd44f34330445ecd572b89e550865df51a3af57503704c8b7bce019e4453ce1e

Observation 40dd717c-e0a2-4327-93da-c29ad0f50f90 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Self-Distilled Agentic Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.520380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.520380Z digest=sha256:7357cbf21e8c5d56254bee83bf944578b161c9b2442a4c60214977a4fe066d43

Observation 7d5cda3f-e3be-4a2a-bfa1-8534a31fdd6f · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.598520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.598520Z digest=sha256:f993df9f97f408bb83217cca5cb07df5213e7463208bbe49ad4c737c062f6c35

Observation 3c96f8ff-7d79-426f-ad57-19f8d77f4368 · outbound

This paper cites OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.655983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.655983Z digest=sha256:36349332c80cdaa44fd59157892db1cce5160470d9be30939e49c227a2048696

Observation acc7dedf-2d15-4036-b452-5ec1a670beff · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.738704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.738704Z digest=sha256:02ba43ba4579287ed2c3295d33b7e063d60a9591290bf8e0343af51cd430e4ee

Observation 18f0648c-eb94-4cab-b66d-39a6ba55cd0f · outbound

This paper cites Advances in neural information processing systems , volume=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Advances in neural information processing systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.818262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.818262Z digest=sha256:8ed453be55ab7468699a904ec2a7c37a98f6c578bcd3a4b693c8a307af299d75

Observation 60f2afc6-7a34-479a-b357-7e79203781e3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:29.957997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:29.957997Z digest=sha256:b6b7666e6a243c15896f672ba5e1fc9a8fd9e1b961b918ef2453c6d527aadead

Observation 32638830-3b3c-4b9b-97ec-8cc6bf756421 · outbound

This paper cites ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:30.100402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:30.100402Z digest=sha256:ec825ece15ccd5b4728826f01a5ec066b7dc444bc46b1567dffea2e4baf888eb

Observation a3227e63-310a-4bee-9199-0d712b99a2b7 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:30.201855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:30.201855Z digest=sha256:94f5774e46b72b0f4f3980fd0740ebc5872e2e118a870656bf84685f849b6a19

Observation aaa7e3b4-659c-44f7-b5e3-e9842a7e5516 · outbound

This paper cites an unresolved cited work.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:30.324693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:30.324693Z digest=sha256:db410e4a9fb9291e2efea2f3cdc53a77553630a2cfb1006c83cecadb2c692a26

Observation eb8ed1dd-6043-4d29-9610-0bd0ce4ecd10 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:30.474060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:30.474060Z digest=sha256:1ed13327957cee3c87da26058a8bc7120cecc2c329ed13af47510d85d56c7f76

Observation 6e035d40-6089-45b7-8d2a-7ed4e6c53e8b · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:30.573998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:30.573998Z digest=sha256:fe1e20fb5715c6bb59b7c8d729fdb111a4a5dcc07c10493714376b193c97b972

Observation 374434a2-dc85-452e-9588-154ec221e931 · outbound

This paper cites arXiv preprint arXiv:2510.14545 , year=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2510.14545 , year=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:30.669123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:30.669123Z digest=sha256:4bbc8f68d125bfbfb143da35a1cbd1db5a7c4a4414a737087fd2a73a521a853a

Observation 7d5895cc-3b3c-4899-abb6-c04e930f38c5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:30.727278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:30.727278Z digest=sha256:263d947b17b3f91ebd648f1bfe46c47cacfcaf90d6a63733e97e58b75eab17ae

Observation ce1318fd-122b-439d-96c4-43d39b2117ed · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:30.888885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:30.888885Z digest=sha256:971549592d211a4798d835e27fdf331edb5555ff78d27b9d3cdf335b3c1e637b

Observation 317bd219-c488-45fe-9fcb-8719d4ea1565 · outbound

This paper cites arXiv preprint arXiv:2603.08754 , year=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2603.08754 , year=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:30.970519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:30.970519Z digest=sha256:7b270c5098e0750eea4c24b56e40d141eeba16403a0b2ae17b683dad86d4e0b0

Observation 4010f159-17dc-43fa-a917-e3d22f064006 · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:31.058816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:31.058816Z digest=sha256:c25f70af0053f698143eeb6596331ad6258ae1bca4b9bfdfc44606fc4f4d4a4a

Observation 404fb529-99d6-408e-b12d-1e5c0f10486c · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Reinforcement Learning via Self-Distillation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:31.153753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:31.153753Z digest=sha256:8f7b9f4cedd6e954b6ca6c314ed9da487eff9c355e9ba2d5028b54b92f501224

Observation 58e5bc50-b9d5-4850-be04-0ca605e48e6d · outbound

This paper cites UniSD: Towards a Unified Self-Distillation Framework for Large Language Models.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:31.294548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:31.294548Z digest=sha256:ca8603823a2e84d33a8ea41e6a64b4967e1865cdc557cf3169ff212bb209e642

Observation 73831e79-f25f-4256-819f-4233d5cc5bee · outbound

This paper cites OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:31.448481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:31.448481Z digest=sha256:cb0bc67940f4ec91d6314da68b9d4cf60abe56ecb04821d4e7be369b54673025

Observation 33257439-8c7f-44d6-bb77-8d48bc9befd5 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:31.626165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:31.626165Z digest=sha256:7765c81ab2437af537773c4d5812fdd1ebd5e08fa83bb74220e01ddf890c8fb4

Observation cfa5e7ad-9c77-4ef0-954d-56db5f9483dd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:31.813966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:31.813966Z digest=sha256:cd36974c44a487de991f862f06ff897f7120fc7d5a61488e2317c8f011c8eeb7

Observation 9b532ba7-3e9d-4f1e-a0c9-148d7785d855 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Transactions of the Association for Computational Linguistics , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:32.014435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:32.014435Z digest=sha256:dea9e2ea69d1fb818808cd0359659418391720e84182a1f764fa156bec470f55

Observation 16265c6e-03a3-414b-ad8e-cb4ab1e53c32 · outbound

This paper cites Proceedings of the 2018 conference on empirical methods in natural language processing , pages=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:32.099679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:32.099679Z digest=sha256:e480cab143b72590bbf09ebc5cb90749ef9adc50c1db8602aeb5bdced6196def

Observation e0a77114-5bf4-4e4d-9c69-e68211d21594 · outbound

This paper cites Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:32.174800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:32.174800Z digest=sha256:6736e2d1828cca23713dd8acb7ce095e6f10ef95b590051711695d0cecc88256

Observation 58c24a99-b950-4e26-b76d-a5ead428475f · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:32.281138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:32.281138Z digest=sha256:8a0d509ae6bae72b8ff5157b65255050c9402f98b32d36e0f4c01e0b3506c32a

Observation 72c850e6-5948-437d-a8da-a287491891e4 · outbound

This paper cites Proceedings of the 28th International Conference on Computational Linguistics , pages=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Proceedings of the 28th International Conference on Computational Linguistics , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:32.452668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:32.452668Z digest=sha256:53fdfdc517b607f66fd1b481c9974fdd5ae2d1194ea246d1e59579d5ddd93521

Observation ad863716-d0a9-4a4c-b76a-ca7ce78a893a · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Transactions of the Association for Computational Linguistics , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:32.594400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:32.594400Z digest=sha256:884ec2f2fe15ad66daefdb0eb2b1d62864cf1faa565dbf5ebed2bc43f873dd00

Observation be90f18d-87b6-4ebc-b118-4b334f34c5c1 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:32.754593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:32.754593Z digest=sha256:d126bf095b6ddcd7fda3ce7f38bc050323e6a74759cfa3dafbc6cbe8f996346f

Observation 3a001b18-d8d0-4d52-a2f4-b41507844e65 · outbound

This paper cites SOD: Step-wise On-policy Distillation for Small Language Model Agents.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:32.882155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:32.882155Z digest=sha256:48da17fade47b983aa3a51ea9e8af34c1c10f77e042ec01854aebaa2c36bbbb0

Observation 691d5dc1-1590-4148-a46f-56242c5063f8 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:33.011150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:33.011150Z digest=sha256:38074d387f06f0410cbb7ced09c1fa8d9a415c1a954b57f0fe322b24154eb232

Observation 74de7f36-81d0-4e79-9e6c-70ff22a79792 · outbound

This paper cites 2026 , eprint=.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:33.135384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:23:33.135384Z digest=sha256:8149ae708313a2b5d871e779ff282bcb17f83fa7e44be8f7e0e57ce22ead12a8

Pith citing papers

No inbound Pith citation observations are available.