Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:23:14.694452Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.01667.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:23:14.694452Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8b721a9c-415e-4ccb-a10e-af5f6890dec9 · outbound
TCPO: Turn-Level Credit Policy Optimization ToRL: Scaling Tool-Integrated RL
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92693cd6-c1ba-428b-99a1-50faf8fae3ce · outbound
TCPO: Turn-Level Credit Policy Optimization HybridFlow: A Flexible and Efficient RLHF Framework
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b0240b-e3f3-48f3-9235-56130f21c21e · outbound
TCPO: Turn-Level Credit Policy Optimization Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 344e4dc4-f6fc-4f6f-a7eb-7302855a467e · outbound
TCPO: Turn-Level Credit Policy Optimization TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b824a36a-6db2-479b-9853-46c1de6bd896 · outbound
TCPO: Turn-Level Credit Policy Optimization Exploiting tree structure for credit assignment in reinforcement learning with large language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fdc7c8d9-5d4f-496c-a375-929f8d2adcdb · outbound
TCPO: Turn-Level Credit Policy Optimization Solving math word problems with process- and outcome-based feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a172f626-5a46-44ad-b41e-6b49a319ce06 · outbound
TCPO: Turn-Level Credit Policy Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efce9824-9194-48da-8992-39b850606df8 · outbound
TCPO: Turn-Level Credit Policy Optimization Rein- forcing multi-turn reasoning in llm agents via turn-level credit assignment
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ded1493-716c-4061-a89a-18136fec6c84 · outbound
TCPO: Turn-Level Credit Policy Optimization Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29aec9b2-098f-4024-ada1-5ec02895ecfb · outbound
TCPO: Turn-Level Credit Policy Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ba18ea4-ed78-43cd-bc54-deb51ad88f44 · outbound
TCPO: Turn-Level Credit Policy Optimization At 2po: Agentic turn-based policy optimization via tree search.arXiv preprint arXiv:2601.04767,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74360757-7012-4b6a-9243-770addcc35bc · outbound
TCPO: Turn-Level Credit Policy Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f450e9f-b910-4ae1-be33-eae85bac52ea · outbound
TCPO: Turn-Level Credit Policy Optimization An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee6508c6-f791-4ad9-923b-8dccee513ef9 · outbound
TCPO: Turn-Level Credit Policy Optimization Let’s verify step by step
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9eac6ccf-3610-4ef0-9123-e2250fd1828d · outbound
TCPO: Turn-Level Credit Policy Optimization Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.