Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2402.19446.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T18:39:37.983121Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:39:51.549262Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 2372b857-3786-4fdf-a12a-237a24cce147 · inbound
Training Language Models to Self-Correct via Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b5a8aa2-19e3-4392-a33e-eb36ad3f05c5 · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 200
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 11484c73-dde9-4c93-aa03-456ef3ded975 · inbound
Process Reward Models for LLM Agents: Practical Framework and Directions ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c53f99-5b8a-4238-9ef7-6295b93fd574 · inbound
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac711189-9f6d-43d7-9184-727b8818fe3f · inbound
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b389c14e-e26e-4d02-9c71-a9b0eadee153 · inbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe81ee85-bce5-4ea0-9666-744b80f639c5 · inbound
Self-Challenging Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f04a9702-4b69-44c0-ac3b-86915a75c762 · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 630fbc79-01bf-4030-8d21-19eaea262789 · inbound
Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e863c26d-93fc-486a-968c-e70c501dd59d · inbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42837901-c731-4b89-8ada-6a3b37bb2e11 · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cfb8b71-f14a-44d6-b8f3-c7a381c3cc12 · inbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de5f99d-b149-450f-a171-fb3dcc223bb0 · inbound
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1682c18-d074-4e3e-ac43-0408273afa93 · inbound
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 535dcb56-38ea-4662-8ec2-dd8827cf14c3 · inbound
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc9e16b3-d475-4011-8cee-c3b82b527ec1 · inbound
Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e6a819-f23e-479c-a18a-da58ad239fed · inbound
From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 46186c15-5bed-4455-bc3b-dccfdcf18c18 · inbound
From History to State: Constant-Context Skill Learning for LLM Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e1984de5-e929-4a2b-94e2-83467b69203e · inbound
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 830e555e-e74d-40ef-84fb-4290fa4652ad · inbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 761b8ce8-90d4-4dfd-b99f-1e1288696ce0 · inbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ac2a3c3-7399-4a83-aacd-b7eedd783c24 · inbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 300a6290-65b5-4edd-8337-474164b0d212 · inbound
Unlocking Proactivity in Task-Oriented Dialogue ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d4d9da5-573f-441a-a56d-71f2d1ae5eec · inbound
Unlocking Proactivity in Task-Oriented Dialogue ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a3d0c357-7b0e-47d6-ba3f-8aa4ff6a0155 · inbound
ECHO: Terminal Agents Learn World Models for Free ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ecc25546-622a-4b51-95e1-5097794fe2d9 · inbound
Trust Region On-Policy Distillation ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e73b26fe-012c-41a7-8cf0-a8486627ae88 · inbound
SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7de7e4d-e539-468c-b87b-9ffc62e5ab04 · inbound
When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e4197b60-6ae7-4740-8e5a-28a05289d8a5 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 279
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e1bac64-edce-4f51-a4c1-d4b607991b0a · inbound
Diagnosing Task Insensitivity in Language Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6d680464-8def-4b55-8d3b-4975436da8a8 · inbound
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 123
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef2ce5bc-eb45-405c-aec1-288c1d48ea6b · inbound
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 117
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d9fba77-4255-4205-b1ba-41b57be0a765 · inbound
The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d592169-b77b-4224-8570-132ba46210a4 · inbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028e7780-fa07-48ae-a71e-1085e1c681df · inbound
Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ea2611-4fe6-4363-b476-5a5e49756d56 · inbound
Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.