Pith. sign in

Paper Citation Record · LEDGER

ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2402.19446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19446 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:39:37.983121Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:51.549262Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2372b857-3786-4fdf-a12a-237a24cce147 · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.455408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:a6f1fd15b8d9d72f9ed9f755f8b855a1f96a89cd2e3975953bb20c6639dd4a4e

Observation 8b5a8aa2-19e3-4392-a33e-eb36ad3f05c5 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.461524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:f515762639db250a6f7a0c85fc0fb24ca61042755f5675328552c045113aa7c5

Observation 11484c73-dde9-4c93-aa03-456ef3ded975 · inbound

Process Reward Models for LLM Agents: Practical Framework and Directions cites this paper.

Process Reward Models for LLM Agents: Practical Framework and Directions ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.983121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.983121Z digest=sha256:dc84bb88923e2f402f9db0ec21b9b26642fa85a4e4a8b6b112b261e1691262ff

Observation e5c53f99-5b8a-4238-9ef7-6295b93fd574 · inbound

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution cites this paper.

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:59.670045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:59.670045Z digest=sha256:550c4ce6fc9ff91fdcf82500595c3ef7709f6c6360615f40667a24055d90dc85

Observation ac711189-9f6d-43d7-9184-727b8818fe3f · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.408981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.408981Z digest=sha256:d87d35bc041fd0a0833d069319c357a70fc2c792c69012a6383887027f9dc645

Observation b389c14e-e26e-4d02-9c71-a9b0eadee153 · inbound

ARIA: Training Language Agents with Intention-Driven Reward Aggregation cites this paper.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.985006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.985006Z digest=sha256:c8d85068d854679dfde8dd43c04e7f03d4782f5d20881eaeffde9f5cb1a1e804

Observation fe81ee85-bce5-4ea0-9666-744b80f639c5 · inbound

Self-Challenging Language Model Agents cites this paper.

Self-Challenging Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:58.654310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:58.654310Z digest=sha256:561b52c1dd0f537840b9ee60ec1bd03a6f056e6518e7712e7a339b44d879547f

Observation f04a9702-4b69-44c0-ac3b-86915a75c762 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.549942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.549942Z digest=sha256:4440dbd6988870c22f3b0f0d09a70c0930773c502b4bc0c9f4f1bfe035c3ab68

Observation 630fbc79-01bf-4030-8d21-19eaea262789 · inbound

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback cites this paper.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:30.382349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:30.382349Z digest=sha256:bf8169385cb8f3ba4d28cfd2587bdb734049131e7928697f7886cdf75f7e74e6

Observation e863c26d-93fc-486a-968c-e70c501dd59d · inbound

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction cites this paper.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.356815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.356815Z digest=sha256:80415a9320200087e8d53023f9e39b0201494f2c143d23895260c5411f3fb1ea

Observation 42837901-c731-4b89-8ada-6a3b37bb2e11 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:39.085675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:39.085675Z digest=sha256:07976a443b1fa041cf25ddaadefa27e16909dc820ba53d249baaa7aaaf46fadb

Observation 0cfb8b71-f14a-44d6-b8f3-c7a381c3cc12 · inbound

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison cites this paper.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.960285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.960285Z digest=sha256:dff4c262a54f1af621d71d4ff8504966f41c31ce276c215c0e6f67d21ed710a3

Observation 6de5f99d-b149-450f-a171-fb3dcc223bb0 · inbound

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning cites this paper.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.049831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.049831Z digest=sha256:4d70c9bf308ac112e2004a4092679eaa5e854e2308c04ee1c84eba5b55b5fef8

Observation b1682c18-d074-4e3e-ac43-0408273afa93 · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:51.017069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:51.017069Z digest=sha256:dad99a5097c63138b5a31fbea2812596ae2b3417dfaf7547f95334e53927b34d

Observation 535dcb56-38ea-4662-8ec2-dd8827cf14c3 · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.947042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.947042Z digest=sha256:e7894cad60341d31fe16f7be19ef6945b9e2f67d207bddf0e578f20ac75d31f5

Observation bc9e16b3-d475-4011-8cee-c3b82b527ec1 · inbound

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment cites this paper.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T12:27:28.745515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:27:28.745515Z digest=sha256:ef3c6e52a0c9dd33f8e22c1dc478a5a92babc4c26df688960a4716cd140a44da

Observation 95e6a819-f23e-479c-a18a-da58ad239fed · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.014508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:42a90e6946553392b9fd673fa922a528f2bc2aacf6eda8c03e4e7218387e74cc

Observation 46186c15-5bed-4455-bc3b-dccfdcf18c18 · inbound

From History to State: Constant-Context Skill Learning for LLM Agents cites this paper.

From History to State: Constant-Context Skill Learning for LLM Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:01:05.529590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T16:50:43.547830Z digest=sha256:fc6781cac6c650ca4e3dbc69415ac81b10c99c7532dcaa1e771993188ee2f39a

Observation e1984de5-e929-4a2b-94e2-83467b69203e · inbound

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction cites this paper.

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:11:12.171337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T09:59:48.604813Z digest=sha256:95ddf16dba815950ff0a3a9e8b38fff807841f6ff5eee8d9d6777b9b1ae4b8cb

Observation 830e555e-e74d-40ef-84fb-4290fa4652ad · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.475555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:082d6c642ab8c19134e77f40ab7d820f8dde20c6aba2284f7d75ff618828f8f9

Observation 761b8ce8-90d4-4dfd-b99f-1e1288696ce0 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.153287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.153287Z digest=sha256:91f2ba0220a38730ac037b3004779323a042c9edce043fbcbcdeb66a1d7912fb

Observation 4ac2a3c3-7399-4a83-aacd-b7eedd783c24 · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.483436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:751b4c65a7372d88fab5aa7c284c8817a4615298704ca83bd754e1da31cac4fb

Observation 300a6290-65b5-4edd-8337-474164b0d212 · inbound

Unlocking Proactivity in Task-Oriented Dialogue cites this paper.

Unlocking Proactivity in Task-Oriented Dialogue ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.495072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:31:37.171422Z digest=sha256:7b16f721a79ef3c9e44a141b99eae3bbb72feb49aa84ed2dd6a6221b65dfff37

Observation 1d4d9da5-573f-441a-a56d-71f2d1ae5eec · inbound

Unlocking Proactivity in Task-Oriented Dialogue cites this paper.

Unlocking Proactivity in Task-Oriented Dialogue ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.708669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:31:11.074231Z digest=sha256:457d11794bc07630006d761257676e8320bb99776083d0cff0665ddb86b154ed

Observation a3d0c357-7b0e-47d6-ba3f-8aa4ff6a0155 · inbound

ECHO: Terminal Agents Learn World Models for Free cites this paper.

ECHO: Terminal Agents Learn World Models for Free ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T15:04:46.491807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T14:57:03.095107Z digest=sha256:1750b98dd81740c0a559e4ffcca9a7c96fe13e098e2d9e9f9d8e6bd5fdb755b7

Observation ecc25546-622a-4b51-95e1-5097794fe2d9 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:46:14.296998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:d83443715ec96d41da1ff825284fe1f84ce768a66a6d98487369027aede17bc1

Observation e73b26fe-012c-41a7-8cf0-a8486627ae88 · inbound

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training cites this paper.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:23.950721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:867501a1743eb8bf993cf3cae4d1dbabc4a5543d865cf7aa4ec9a1182af74986

Observation d7de7e4d-e539-468c-b87b-9ffc62e5ab04 · inbound

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training cites this paper.

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.932501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T02:17:32.324432Z digest=sha256:b7255f2bdbd9b752627fd9868c3b230bdf391a114be9c2071fc90fbe1feeb8ca

Observation e4197b60-6ae7-4740-8e5a-28a05289d8a5 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.601656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:4a6085b1ff9b0124f8a22bde41be994aac5bbb72450d03b6e7c1572780a8860c

Observation 0e1bac64-edce-4f51-a4c1-d4b607991b0a · inbound

Diagnosing Task Insensitivity in Language Agents cites this paper.

Diagnosing Task Insensitivity in Language Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:51.551580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T04:58:11.929896Z digest=sha256:6a0b2830a37881e7e78ea3bb7e33064af4beb695dc3c67e0a75d852699166c53

Observation 6d680464-8def-4b55-8d3b-4975436da8a8 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 123

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:50d1992fc9b707286a8f988ae91f5216deb47dd3bd2d57fd3611e0fa96bad5bd

Observation ef2ce5bc-eb45-405c-aec1-288c1d48ea6b · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:35.709395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:35.709395Z digest=sha256:6dfd950aaa218606ed43049fc70585fda17a5914208f2247ba324d16ff834385

Observation 7d9fba77-4255-4205-b1ba-41b57be0a765 · inbound

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works cites this paper.

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:01:23.548585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:01:23.548585Z digest=sha256:da04e3304ee51ae0c7be2e9a87c4f77f29e283c503c6714006401ff9b068ce28

Observation 1d592169-b77b-4224-8570-132ba46210a4 · inbound

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning cites this paper.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.221262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.221262Z digest=sha256:8044973a062b342e4ad3b3c4f7a151b4019c838804f748987566889407b65e51

Observation 028e7780-fa07-48ae-a71e-1085e1c681df · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T15:25:49.223723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:25:49.223723Z digest=sha256:598c3b6f9b26c47405ee016711ae4fae992a735b7498139e90ce0408597241aa

Observation 55ea2611-4fe6-4363-b476-5a5e49756d56 · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:15.761164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:15.761164Z digest=sha256:1d35ff7f3bdb9887b7c9aa2a38c78542fb8c10bf6c3950d312d865a91eecaa21