Pith. sign in

Paper Citation Record · LEDGER

TCPO: Turn-Level Credit Policy Optimization

As of 19 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.01667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01667 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:23:14.694452Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b721a9c-415e-4ccb-a10e-af5f6890dec9 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

TCPO: Turn-Level Credit Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.645199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.645199Z digest=sha256:bc7bbcfd776d0d4fab2777ef1488facc47443a9dfebf7d7b21c3f6e7370abd73

Observation 92693cd6-c1ba-428b-99a1-50faf8fae3ce · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

TCPO: Turn-Level Credit Policy Optimization HybridFlow: A Flexible and Efficient RLHF Framework

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.653931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.653931Z digest=sha256:88eccd0f6be448a8851d6a2c6fd7bd2b5545a1127278c57b534e7f7126006950

Observation a1b0240b-e3f3-48f3-9235-56130f21c21e · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

TCPO: Turn-Level Credit Policy Optimization Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.658329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.658329Z digest=sha256:14b908c2732ebd498a0a7c8e445958e0c214b616aa6ced3ece655aa906d86031

Observation 344e4dc4-f6fc-4f6f-a7eb-7302855a467e · outbound

This paper cites TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents.

TCPO: Turn-Level Credit Policy Optimization TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:23:14.809554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:23:14.662916Z digest=sha256:c3a4209fd0c64898d452cb74cd0d9906a27c09d31d90d37fa97d904e7cc49ca7

Observation b824a36a-6db2-479b-9853-46c1de6bd896 · outbound

This paper cites Exploiting tree structure for credit assignment in reinforcement learning with large language models.

TCPO: Turn-Level Credit Policy Optimization Exploiting tree structure for credit assignment in reinforcement learning with large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:23:14.927639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:23:14.667862Z digest=sha256:4f1a97c3320728505dabe66dbc37c4b6c24ca44f71c7a8d457a9414863bf734e

Observation fdc7c8d9-5d4f-496c-a375-929f8d2adcdb · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

TCPO: Turn-Level Credit Policy Optimization Solving math word problems with process- and outcome-based feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.671864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.671864Z digest=sha256:8545fe321e854a140821ead9ba37824d2882ea2be6271e94d033d577fb5072f3

Observation a172f626-5a46-44ad-b41e-6b49a319ce06 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

TCPO: Turn-Level Credit Policy Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.676492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.676492Z digest=sha256:f5696cffa3340e955e38062bcb747f7c06a333bb6fac562abcfcc9885ad8ebd7

Observation efce9824-9194-48da-8992-39b850606df8 · outbound

This paper cites Rein- forcing multi-turn reasoning in llm agents via turn-level credit assignment.

TCPO: Turn-Level Credit Policy Optimization Rein- forcing multi-turn reasoning in llm agents via turn-level credit assignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:23:14.912954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:23:14.681238Z digest=sha256:17072947ea8d36420fe59144e4f42704a645da674eff2212b202f8331e33f95b

Observation 8ded1493-716c-4061-a89a-18136fec6c84 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

TCPO: Turn-Level Credit Policy Optimization Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.686054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.686054Z digest=sha256:af637354c5a710a300ff88b005e432e8c5e1058f9edd27b21765d14a4a2c673b

Observation 29aec9b2-098f-4024-ada1-5ec02895ecfb · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

TCPO: Turn-Level Credit Policy Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.690152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.690152Z digest=sha256:f10ff79df9a039808c60fd8aa7312c66f5d44ea57163163616bfbb01ea8504ae

Observation 2ba18ea4-ed78-43cd-bc54-deb51ad88f44 · outbound

This paper cites At 2po: Agentic turn-based policy optimization via tree search.arXiv preprint arXiv:2601.04767,.

TCPO: Turn-Level Credit Policy Optimization At 2po: Agentic turn-based policy optimization via tree search.arXiv preprint arXiv:2601.04767,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.694452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.694452Z digest=sha256:c25c541ea6794779af7c1ddb9be17c214e750ae91554d52e14e56931ab5c0ac0

Observation 74360757-7012-4b6a-9243-770addcc35bc · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

TCPO: Turn-Level Credit Policy Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.629933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.629933Z digest=sha256:1ae07da0d333bbf831fe9a660102bbec21e7082baf7632799da5c970651f4afc

Observation 3f450e9f-b910-4ae1-be33-eae85bac52ea · outbound

This paper cites An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents.

TCPO: Turn-Level Credit Policy Optimization An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.639942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.639942Z digest=sha256:a901c29ae61983205604ebd9ef84e0ae1ff6f9f34157f9dbe85eb198ff052e7e

Observation ee6508c6-f791-4ad9-923b-8dccee513ef9 · outbound

This paper cites Let’s verify step by step.

TCPO: Turn-Level Credit Policy Optimization Let’s verify step by step

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:23:14.941960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:23:14.649732Z digest=sha256:7d8ca70a72f5140e2d333a140c26f36c0d34e9b02d80904179ccce7260ef48c7

Observation 9eac6ccf-3610-4ef0-9123-e2250fd1828d · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

TCPO: Turn-Level Credit Policy Optimization Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.635075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.635075Z digest=sha256:3d9f670ec0f63a775fd305613b892ae3274780fb13946d821640abb685605eaf

Pith citing papers

No inbound Pith citation observations are available.