Pith. sign in

Paper Citation Record · LEDGER

Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization

As of 23 July 2026, this Paper Citation Record lists 5 of 5 outbound references and 6 inbound Pith citation observations for arXiv:2604.07165.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.07165 v2

Coverage vector

measured 5 of 5 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:42:37.065795Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-22T06:31:00.163083+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T16:03:26.866407Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T20:47:23.391702Z

Reference resolution

5 of 5 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2136d82-fca9-4b9c-845f-e4684b6422a3 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization Group-in-Group Policy Optimization for LLM Agent Training

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:15:09.492454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-10T18:42:37.065795Z digest=sha256:af0f14c770d1f017f645222e178defe0bc391471af03b52d22d8b20ff702ab3d

Observation dce55c31-fb94-4643-8380-ddd97fda185c · outbound

This paper cites arXiv preprint arXiv:2509.09284 , year=.

Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization arXiv preprint arXiv:2509.09284 , year=

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:49.978273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-10T18:42:37.065795Z digest=sha256:4b40eec5de1473e7a754de95056087073c1286e7e4c039c0f82f10ccd77f461b

Observation cc0dfc80-e603-459e-aafc-9d4590d40378 · outbound

This paper cites Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space.

Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:05:49.983836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-10T18:42:37.065795Z digest=sha256:95437c0a0879b3e44db180ee87e60d12f5b21ace3003c2ebcb9bee8c2e30be10

Observation 15473c44-3a79-4937-baf1-6dd658ce8f99 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T00:05:49.989867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-10T18:42:37.065795Z digest=sha256:5a8ce265a568a133d301bce62fff9ebb7975733df1d21d74b1923c12b50b0aac

Observation 7e08366f-f6c2-460f-bb88-9c68f0ec2935 · outbound

This paper cites assembly required.

Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization assembly required

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:13:05.252413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-10T18:42:37.065795Z digest=sha256:eb34b357224d125ce61a9d921944e1db228d545237aa1e6ae94975b453a663eb

Pith citing papers

Observation 11925afe-9ba5-4e32-83f5-4ea16ec3fd39 · inbound

NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search cites this paper.

NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T15:51:43.080098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=arxiv_source observed=2026-05-09T19:08:50.518219Z digest=sha256:2c0729d95435166ebc88b2a2cb55309da58226f987d41f720098e2ce53284aa3

Observation 1c8b7a9d-e82c-4743-9563-556c4959924d · inbound

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning cites this paper.

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:11:16.990640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-22T08:10:55.720464Z digest=sha256:2bd33b2a0f5c8b4ddc2574122b6e25e65f5aa99c07e7336622da7e912e5130c5

Observation 2d916b6f-2bd0-4a88-b0fa-d45a06792c15 · inbound

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning cites this paper.

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:50:24.355871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-25T05:47:33.925413Z digest=sha256:f48d71ae98de1eebea791b7917ab8ea5112257f9130a51ce06a29ccea5a0b52b

Observation d900fad7-4b4f-40a2-9b43-c2258b6e6922 · inbound

A Note on Stability for Orthogonalized Matrix Momentum with Client Sampling cites this paper.

A Note on Stability for Orthogonalized Matrix Momentum with Client Sampling Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:56:15.636973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T16:03:26.866407Z digest=sha256:588e46e15beb9926fc23742f2b5095a2836fc072fc94a3763499e61200aa0b95

Observation 6241198f-c8ad-4dd2-b082-203b58cf95ec · inbound

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction cites this paper.

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.774163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:06:31.471481Z digest=sha256:909b679b774c891ebf05ed711f137b9af21c4c3acdc7f5b1801f9eef41be5cde

Observation 3d7b86d3-f1ce-4a28-b794-d8dd0834cbe6 · inbound

Customer-Agent: Overcoming Context Limitations in Ultra-Long Shopping Trajectories via Tool-Augmented Agents and RLVR cites this paper.

Customer-Agent: Overcoming Context Limitations in Ultra-Long Shopping Trajectories via Tool-Augmented Agents and RLVR Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T20:47:23.393991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=arxiv_source observed=2026-06-27T20:05:34.326966Z digest=sha256:94a434f3c6af76dfbc5c04b21e8debc919e7ce18db2a2c7435d78e291da402f3