Pith. sign in

Paper Citation Record · LEDGER

TCPO: Turn-Level Credit Policy Optimization

As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.01667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01667 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:23:14.694452Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b721a9c-415e-4ccb-a10e-af5f6890dec9 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

TCPO: Turn-Level Credit Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.645199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.645199Z digest=sha256:60fb1e7064837d55574518dc3bbebd087655e69a9c589ed487ceb6177f9ece5e

Observation 92693cd6-c1ba-428b-99a1-50faf8fae3ce · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

TCPO: Turn-Level Credit Policy Optimization HybridFlow: A Flexible and Efficient RLHF Framework

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.653931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.653931Z digest=sha256:dcd500a8526e99d951b3b43ae9586122ed9e823b8af2cd9a4a6c3be2b80e47bb

Observation a1b0240b-e3f3-48f3-9235-56130f21c21e · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

TCPO: Turn-Level Credit Policy Optimization Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.658329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.658329Z digest=sha256:94a50be1dc1133b50495475a6b0c1d84ac0c36d41938677b30be65357c109e7f

Observation 344e4dc4-f6fc-4f6f-a7eb-7302855a467e · outbound

This paper cites TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents.

TCPO: Turn-Level Credit Policy Optimization TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:23:14.809554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:23:14.662916Z digest=sha256:149457a932e961255352b089420e8849c4a3d7965864a5a1c182dd41f88ebee2

Observation b824a36a-6db2-479b-9853-46c1de6bd896 · outbound

This paper cites Exploiting tree structure for credit assignment in reinforcement learning with large language models.

TCPO: Turn-Level Credit Policy Optimization Exploiting tree structure for credit assignment in reinforcement learning with large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:23:14.927639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:23:14.667862Z digest=sha256:23a5b2153f634dbcfbea28d8f2cab36c0eee168d6bf25edc3bb233b92930d51d

Observation fdc7c8d9-5d4f-496c-a375-929f8d2adcdb · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

TCPO: Turn-Level Credit Policy Optimization Solving math word problems with process- and outcome-based feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.671864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.671864Z digest=sha256:fb781889ce7c37408cc09ac38e2a2ef663c86f57db402ca7e4f1f1512ced81bd

Observation a172f626-5a46-44ad-b41e-6b49a319ce06 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

TCPO: Turn-Level Credit Policy Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.676492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.676492Z digest=sha256:07980f34ba23b341cad4c3eaf0f3d1d36687022101f1a9ff2b8fc329d2edf27e

Observation efce9824-9194-48da-8992-39b850606df8 · outbound

This paper cites Rein- forcing multi-turn reasoning in llm agents via turn-level credit assignment.

TCPO: Turn-Level Credit Policy Optimization Rein- forcing multi-turn reasoning in llm agents via turn-level credit assignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:23:14.912954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:23:14.681238Z digest=sha256:8488172c73aef0fbd605557b89878b8aff734650b2521c1797ce496c5060266c

Observation 8ded1493-716c-4061-a89a-18136fec6c84 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

TCPO: Turn-Level Credit Policy Optimization Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.686054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.686054Z digest=sha256:dbca251eb48d156fc95ae198bb504a2ffe356d8adfdd7c3bf35f95f01bbbe78b

Observation 29aec9b2-098f-4024-ada1-5ec02895ecfb · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

TCPO: Turn-Level Credit Policy Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.690152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.690152Z digest=sha256:7b614ff2f29aa33bb90ce45234be9add8d269f9359a2bb1be0794a9f2232bc88

Observation 2ba18ea4-ed78-43cd-bc54-deb51ad88f44 · outbound

This paper cites At 2po: Agentic turn-based policy optimization via tree search.arXiv preprint arXiv:2601.04767,.

TCPO: Turn-Level Credit Policy Optimization At 2po: Agentic turn-based policy optimization via tree search.arXiv preprint arXiv:2601.04767,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.694452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.694452Z digest=sha256:c88f7d867b07b7fb72592fa7c18362288162d509b280bfe6a240f0bce53e85dd

Observation 74360757-7012-4b6a-9243-770addcc35bc · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

TCPO: Turn-Level Credit Policy Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.629933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.629933Z digest=sha256:6a37c21ab798423c601539bbf2ccc456871f4507d26ebd0886516990505b0679

Observation 3f450e9f-b910-4ae1-be33-eae85bac52ea · outbound

This paper cites An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents.

TCPO: Turn-Level Credit Policy Optimization An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.639942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.639942Z digest=sha256:852453390a0eebb4e6c3bdb7ca2544e1b3a9d45dd4c402e14a177dc3ad8fada5

Observation ee6508c6-f791-4ad9-923b-8dccee513ef9 · outbound

This paper cites Let’s verify step by step.

TCPO: Turn-Level Credit Policy Optimization Let’s verify step by step

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:23:14.941960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:23:14.649732Z digest=sha256:9f8f9828717147548d6570a8d00a5b809da60038195a03c6164bf6c6380812a6

Observation 9eac6ccf-3610-4ef0-9123-e2250fd1828d · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

TCPO: Turn-Level Credit Policy Optimization Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.635075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.635075Z digest=sha256:1e357459af27cb7cc9821e65e3aa7f92fb3f56788609ca2d612aa908583e746b

Pith citing papers

No inbound Pith citation observations are available.