Pith. sign in

Paper Citation Record · LEDGER

Agent Lightning: Train ANY AI Agents with Reinforcement Learning

As of 6 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 29 inbound Pith citation observations for arXiv:2508.03680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.03680 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:17:55.180539Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:16:44.178281Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T12:30:59.971543Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19b88973-fc7d-4669-8f19-5fec698cc059 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.224738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.224738Z digest=sha256:215767b7128fe4bb1b9fba6add4b007524d75b70316aa0afa6d7ffed76c39866

Observation 35631d70-40c8-4909-b2ad-40ae625c4f0c · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.271607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.271607Z digest=sha256:08e15c9e83fc7412e85fb6421605e58ef2218b3941cf8fe14a2ae6ba55f72572

Observation eb7bc0da-7fcf-4656-8f12-123cd6233a1d · outbound

This paper cites Calc-x and calcformers: Empow- ering arithmetical chain-of-thought through interaction with symbolic systems.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Calc-x and calcformers: Empow- ering arithmetical chain-of-thought through interaction with symbolic systems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:56.478147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:17:54.415024Z digest=sha256:601540b12bba92d13e02073ba4b00ddb0b30688f220989dfe746fd83762efde0

Observation 3600006b-13e1-436f-8fc0-7afee1e242d4 · outbound

This paper cites Calc-X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic Systems.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Calc-X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic Systems

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T04:17:55.692183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:17:54.513604Z digest=sha256:99290d14bd6a7cf02d5b996e9aa39bb083482562a4ea2ca16b544c916f12bd0e

Observation c207ee94-d154-46b3-91a2-e81ff202fb9f · outbound

This paper cites Large Language Model-Based Agents for Software Engineering: A Survey.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Large Language Model-Based Agents for Software Engineering: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.598374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.598374Z digest=sha256:290c8a7ee8c9b8ef9379bd176df504d1e725d840171f0d3fb88455bab0de5240

Observation 97ccb916-e052-42f8-adcc-95288bfa8d99 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:56.250159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:17:54.665686Z digest=sha256:e82bb28fffa1cadbe7497ffb27079bca765a20be6782a354631d972e51639631

Observation 6447b7ca-4b75-4386-ac06-caad8b5ded18 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.814258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.814258Z digest=sha256:46a6950a8edc062601db7f968d505f20376bb4cad28dc59a80865a9e086acb0a

Observation 1f254891-bc51-4c60-ab8e-12a631f3dc87 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.881176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.881176Z digest=sha256:6c90a11ca2ae3d7e1a7f7787b27b8627e4927a3c02f6f14abb045fc0dcf78024

Observation e55033c0-d6f1-4607-a3ef-ab798e4e3160 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.950655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.950655Z digest=sha256:5ebdf4e30bb2543c44e7da69097264fc99a8252fd4f2c8c904c79c8fd0792a8e

Observation 29adce2f-2fdc-416d-b994-ef7f08d0cc62 · outbound

This paper cites Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:17:55.985289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:17:55.091229Z digest=sha256:179630e4cb26d733d6b90d263da17c91c8f8fdc9c9b51e04897247e722a86228

Observation 897f391b-d021-49bd-bf3a-e1519725b756 · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:55.008086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:55.008086Z digest=sha256:1c4eea8fb9ee7f43f87044cea90174a38c40bee3a95bda342bce927a3feaadc4

Observation bc800de8-0a7e-4cd1-8c98-84dda9e21b25 · outbound

This paper cites Retrieve anything to augment large language models.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Retrieve anything to augment large language models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:55.180539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:55.180539Z digest=sha256:3b663edc24bd0c122da8654a0ca6fccde6b59f17dd0d4423c399e96f8eaa61c1

Observation 15852412-3112-4d4b-ae06-a47236665171 · outbound

This paper cites Trinity-rft: A general-purpose and unified framework for reinforcement fine-tuning of large language models.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Trinity-rft: A general-purpose and unified framework for reinforcement fine-tuning of large language models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.723625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.723625Z digest=sha256:671f47594d3c832ccd68bd5b0f76af75a20afa6509002734ee04390a97b569c2

Observation 4054d7c1-8bf6-4ee3-a37e-43c1898056da · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.038821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.038821Z digest=sha256:d91894abc0a42050d88c5435946761306ae6203cb03ae85618936fc13b61efb0

Observation cfd184d4-2551-4b11-a715-df97c12b1297 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.349118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.349118Z digest=sha256:fc7a50320a6856a567f959b6f1561b3c46ae7fdecb7dc6088430ad4edb92423d

Observation c27dc36d-5ebe-46fa-ab14-4e34213b8488 · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.142833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.142833Z digest=sha256:b4ce3071fa34c8f45881cd0f742e36b4eccba0a64ceaedeaa84cd88890b0426d

Pith citing papers

Observation 4aa540f4-64bf-464e-bd6b-1700fed9f2f3 · inbound

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse cites this paper.

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:05:39.048068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T02:04:46.559881Z digest=sha256:15b752fefd4de300334e53072d8271724fd0c450167b6cad2daa73dabc18dfe7

Observation fa5821be-347f-433a-8119-8b008824ceb0 · inbound

OpenTinker: Separating Concerns in Agentic Reinforcement Learning cites this paper.

OpenTinker: Separating Concerns in Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:09.429860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:09.429860Z digest=sha256:9bbbab78818ef4707d02c46997375f80762abf3c994ad490bc8df7c509f5a003

Observation b9829e51-f9e8-4ccf-a841-be6d5c3d5a0e · inbound

SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models cites this paper.

SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:08:25.731016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T01:06:42.421062Z digest=sha256:f7eca758620b96c1f08de91e918ac86a1764a701de57b65e29a8c1267bc26d30

Observation 40aa32a2-bd46-4817-b029-5e0d63017c47 · inbound

MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents cites this paper.

MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.598498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:21:15.014067Z digest=sha256:64275cbb1df1b82897d68807ec328d6f4a19701f9e242565c0416bcdab3c297a

Observation e1aa28de-4ca5-4970-8924-f6dabe9fa547 · inbound

From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models cites this paper.

From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:58.434148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:09:36.341574Z digest=sha256:e8dbe2e9bf80b9b4b795f9cb6b264ba39c7273971f1f2dad848980d342d04724

Observation 99b6cb2c-e8b0-427c-9e6f-a91abd6a64e5 · inbound

Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures cites this paper.

Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.991797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T04:31:28.242097Z digest=sha256:a4343eff78bd3a14bf769a2dfe76b3b79a4460b36a1ee3ee5015771d35f392e4

Observation 2ef144ab-c294-4f08-a830-35a4421634c7 · inbound

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning cites this paper.

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:29.833782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T04:26:59.781416Z digest=sha256:c9f70543de490cf1a4edf27bb3d446026b0de1793a9a205cf2f3233f657db3ca

Observation 2a368e2a-dbf3-432e-b6df-8c406dd9bcd5 · inbound

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning cites this paper.

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.972779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-05T12:23:00.487399Z digest=sha256:4c411ab3f191ed8b5c2cfc3e838a7822322559219b1a52546907cfa2d7271ebe

Observation bd334dc2-348c-456d-ad07-b0cc26576527 · inbound

CastFlow: Learning Role-Specialized Agentic Workflows for Time Series Forecasting cites this paper.

CastFlow: Learning Role-Specialized Agentic Workflows for Time Series Forecasting Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:31:30.367492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T05:45:13.969847Z digest=sha256:2925ab3d58386e5283f94c3ccaeef1404f121750e3c0275da418a0e5be863b42

Observation e22e45f8-ad69-479a-af6a-706478cd2344 · inbound

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces cites this paper.

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:15:37.848107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T18:44:27.685266Z digest=sha256:1329fb7d1086e8f1dc77bb2c109a7b506f164f24baed20d0e2268fc61207baff

Observation 56c63d53-8a28-4005-8cb0-a887fe0fa735 · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:18.481920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T10:12:58.421050Z digest=sha256:0b32cdb64368280bd72c56ce7921f8008809c337b0f2abe90231dc2307313a6e

Observation 8cfc8208-9155-479e-a9b6-219b10890a65 · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:54.754916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:56:48.838028Z digest=sha256:5b1fbf14c6da62e805677905213ebabba3ddb39893c48eb900cc9b24bb4060c2

Observation 91af1536-3d0e-42d9-97a2-ecbd6d1ba5e5 · inbound

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction cites this paper.

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:11:12.228122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T09:59:48.604813Z digest=sha256:35a5bfff22b47d0fc8d90874b7d35f1a8e75d9aa88b1709bae8dd5bc8ae5c79c

Observation f30601ee-e4e2-44dd-b7b2-4f06a82946c9 · inbound

DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback cites this paper.

DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:45:57.284002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T02:44:50.866913Z digest=sha256:5719434d033e8a425f95d54cdc410270be9909ce27796a5c023b1c888f20d6ad

Observation 8f812e7a-e40b-40f6-b5d4-41bd368033b1 · inbound

DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback cites this paper.

DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:24:57.148274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T16:15:24.370887Z digest=sha256:94da0ac957ec3125f507f2c973c8a2f16df88979d2cd5632f4b1c57e9b460345

Observation db17482e-7863-442d-8f95-c5dea6eb5308 · inbound

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning cites this paper.

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.903485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T06:26:35.100859Z digest=sha256:3016a00e077acce81fa50844abde7c4428d029d9617fec48d8e46ec49a8c2192

Observation 6404d1d6-c93a-4ae6-928f-3f3f9580f487 · inbound

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning cites this paper.

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T12:27:08.761678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:27:08.761678Z digest=sha256:9c4f1ec56149fc5bfa77686d6ca89c68f4fa6c1f9443941ef4c80fe30979c313

Observation cda8438b-7141-4fa0-b3cd-50c150ce4b38 · inbound

Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models cites this paper.

Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.419557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T01:29:55.885301Z digest=sha256:7ae1263b67c461dcdd84c0fb7313c11696bd3c59efcbbfa6a0d2e20c5b7e5f3b

Observation 47db7415-e90d-4548-91be-a657816baa20 · inbound

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning cites this paper.

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:28.221077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T17:28:58.574865Z digest=sha256:0df809413730ec80558c61dc501c8b30c746b833e2d348f6d0a0115e0adc9c97

Observation 6bffb571-cdb1-4ee8-a8dd-004942714109 · inbound

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning cites this paper.

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:29:45.775369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:44:23.085858Z digest=sha256:c1521abbb1c1690e030edcd7ad12182c93720b9a3a5390aaee302fc06c50d750

Observation 70793426-beb0-4a6a-bde9-e15c730f2ab3 · inbound

TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term Memory cites this paper.

TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term Memory Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:30:01.311829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T22:51:47.190143Z digest=sha256:e833fe041cb45534765e277b16317b8ab1463b3f600115e329d06d068728da16

Observation 86a49177-1e79-492c-a7f5-410e2516b81c · inbound

Learning with a Single Rollout via Monte Carlo Pass@k Critic cites this paper.

Learning with a Single Rollout via Monte Carlo Pass@k Critic Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.370844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-25T20:53:10.016447Z digest=sha256:2e9f591481f94c73fa539a7d96516fc79e30083fa0bd3d992bc761efa1a69965

Observation 617abad4-a511-40ae-8edb-cad265016600 · inbound

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning cites this paper.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.137852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.137852Z digest=sha256:97ac691047b4cfc33bb387a07aaaa5727fc668643fc2998eb076081e6659c64a

Observation 10f4918a-182a-47d9-bff4-716a6e3ab1ac · inbound

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning cites this paper.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:29.589477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:29.589477Z digest=sha256:ec35524f08ef4a330e7ec216e58c48a44801a7139acdc4f4a4e447336b95eeca

Observation 09a285f4-5ea3-4534-88cd-dc32d861ebc8 · inbound

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents cites this paper.

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T22:51:47.527608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:51:47.527608Z digest=sha256:e7efab21052ffe98ca9a015c6577062ddf10b8073e91587d130b7a775199e831

Observation 25524da6-685d-4160-bd87-83ce41fad8c3 · inbound

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers cites this paper.

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:14.237509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:14.237509Z digest=sha256:8384ef554991d2d1698703a7262014c63ff9d162f0f381f2c3b5b9e61292e036

Observation 5deec36a-4a7c-410a-955a-0cd152744073 · inbound

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit cites this paper.

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:32.258349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:32.258349Z digest=sha256:76eb4a1c616700d7c9c45cc0aa10c3a6a3b451ebcaaf82afe20f003b312f9d72

Observation 525c4e31-9b8a-4ab2-8e01-d00e5e1783af · inbound

Agentic Reinforcement Learning with Self-Distilled Reward Shaping cites this paper.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.795278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.795278Z digest=sha256:98fde5d234c5141cefae2ac37f7dc682f6b45c9d74790e8f96bf8a9d4827f1f8

Observation 626ed6e3-f6cf-46fa-98f5-9f5becf81f10 · inbound

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning cites this paper.

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:44.178281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:16:44.178281Z digest=sha256:d785ce1eca2c6bc2502eefb35aad5cc89393b67ab324e4313d336d1618748ec5