Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:17:55.180539Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 29 inbound Pith citation observations for arXiv:2508.03680.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:17:55.180539Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:16:44.178281Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-05T12:30:59.971543Z
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 19b88973-fc7d-4669-8f19-5fec698cc059 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35631d70-40c8-4909-b2ad-40ae625c4f0c · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7bc0da-7fcf-4656-8f12-123cd6233a1d · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Calc-x and calcformers: Empow- ering arithmetical chain-of-thought through interaction with symbolic systems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3600006b-13e1-436f-8fc0-7afee1e242d4 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Calc-X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic Systems
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c207ee94-d154-46b3-91a2-e81ff202fb9f · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Large Language Model-Based Agents for Software Engineering: A Survey
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97ccb916-e052-42f8-adcc-95288bfa8d99 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6447b7ca-4b75-4386-ac06-caad8b5ded18 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f254891-bc51-4c60-ab8e-12a631f3dc87 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e55033c0-d6f1-4607-a3ef-ab798e4e3160 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29adce2f-2fdc-416d-b994-ef7f08d0cc62 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 897f391b-d021-49bd-bf3a-e1519725b756 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc800de8-0a7e-4cd1-8c98-84dda9e21b25 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Retrieve anything to augment large language models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15852412-3112-4d4b-ae06-a47236665171 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Trinity-rft: A general-purpose and unified framework for reinforcement fine-tuning of large language models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4054d7c1-8bf6-4ee3-a37e-43c1898056da · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd184d4-2551-4b11-a715-df97c12b1297 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c27dc36d-5ebe-46fa-ab14-4e34213b8488 · outbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa540f4-64bf-464e-bd6b-1700fed9f2f3 · inbound
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fa5821be-347f-433a-8119-8b008824ceb0 · inbound
OpenTinker: Separating Concerns in Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9829e51-f9e8-4ccf-a841-be6d5c3d5a0e · inbound
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 40aa32a2-bd46-4817-b029-5e0d63017c47 · inbound
MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1aa28de-4ca5-4970-8924-f6dabe9fa547 · inbound
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 99b6cb2c-e8b0-427c-9e6f-a91abd6a64e5 · inbound
Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2ef144ab-c294-4f08-a830-35a4421634c7 · inbound
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2a368e2a-dbf3-432e-b6df-8c406dd9bcd5 · inbound
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd334dc2-348c-456d-ad07-b0cc26576527 · inbound
CastFlow: Learning Role-Specialized Agentic Workflows for Time Series Forecasting Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e22e45f8-ad69-479a-af6a-706478cd2344 · inbound
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 56c63d53-8a28-4005-8cb0-a887fe0fa735 · inbound
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8cfc8208-9155-479e-a9b6-219b10890a65 · inbound
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 91af1536-3d0e-42d9-97a2-ecbd6d1ba5e5 · inbound
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f30601ee-e4e2-44dd-b7b2-4f06a82946c9 · inbound
DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f812e7a-e40b-40f6-b5d4-41bd368033b1 · inbound
DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db17482e-7863-442d-8f95-c5dea6eb5308 · inbound
AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6404d1d6-c93a-4ae6-928f-3f3f9580f487 · inbound
AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda8438b-7141-4fa0-b3cd-50c150ce4b38 · inbound
Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 47db7415-e90d-4548-91be-a657816baa20 · inbound
Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bffb571-cdb1-4ee8-a8dd-004942714109 · inbound
Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 70793426-beb0-4a6a-bde9-e15c730f2ab3 · inbound
TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term Memory Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86a49177-1e79-492c-a7f5-410e2516b81c · inbound
Learning with a Single Rollout via Monte Carlo Pass@k Critic Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 617abad4-a511-40ae-8edb-cad265016600 · inbound
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f4918a-182a-47d9-bff4-716a6e3ab1ac · inbound
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09a285f4-5ea3-4534-88cd-dc32d861ebc8 · inbound
From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25524da6-685d-4160-bd87-83ce41fad8c3 · inbound
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5deec36a-4a7c-410a-955a-0cd152744073 · inbound
Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525c4e31-9b8a-4ab2-8e01-d00e5e1783af · inbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 626ed6e3-f6cf-46fa-98f5-9f5becf81f10 · inbound
Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.