Pith. sign in

Paper Citation Record · LEDGER

RRO: LLM Agent Optimization Through Rising Reward Trajectories

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2505.20737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20737 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:43.605901Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0faff669-6ec0-4b39-a99e-fe973c712998 · outbound

This paper cites write newline.

RRO: LLM Agent Optimization Through Rising Reward Trajectories write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.015662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.015662Z digest=sha256:7ded0aef3037c4d09eea1bfee38f3c8113239f37ab7824262a33cb7dd68d8c85

Observation 4c2f4002-50e6-489d-adf2-d869cb36b850 · outbound

This paper cites Language models are few-shot learners.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.148103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.148103Z digest=sha256:b5ea721068a83139573ae25d07e092db719e12145db7646cdfa150e92da0c7ba

Observation 34e23431-e069-445d-90ab-5951361985e9 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

RRO: LLM Agent Optimization Through Rising Reward Trajectories FireAct: Toward Language Agent Fine-tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.263415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.263415Z digest=sha256:4e737343155943f17a590376a352b22a439e23fc2dd17d44307e94498c973efe

Observation bd212bbe-d3f0-4958-8578-56ae96444fd3 · outbound

This paper cites an unresolved cited work.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:51:44.799382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:51:41.347777Z digest=sha256:df2d4d8c6c8c1c7a51d980268951449e0f376899e7b1996500a691c242c9bd1c

Observation 0ae589bc-8a04-4b7e-92cb-0bc154a358af · outbound

This paper cites Deep reinforcement learning from human preferences.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.442326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.442326Z digest=sha256:208e63daad035b3120c4af74dd6cbf035e084fb469cbe8922eef6e34482d072d

Observation 9651b20e-3ac5-41c8-a74d-50a085c7d770 · outbound

This paper cites Complexity-based prompting for multi-step reasoning.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Complexity-based prompting for multi-step reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.616698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:51:41.525532Z digest=sha256:5e889e79430afd9aaa26ce98c3c8647cb27d707eaeaed7c37a15837c1bc76cf4

Observation 55751f3d-ce1b-40cb-a013-31ebdee19dab · outbound

This paper cites Large language models are zero-shot reasoners.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Large language models are zero-shot reasoners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.647817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.647817Z digest=sha256:f02d1fcbb308c22fbbaeede74687fb65e5a9a6a6648da217bfdcd0f3fa47e1f8

Observation 0fa6e4be-6a92-465d-9c0b-61f713909791 · outbound

This paper cites Let's verify step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.770139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.770139Z digest=sha256:4975a0e2c36d9d51c08a0ebb6bc3877ae90fca11e386074ddafcdb417f35600e

Observation bfbee95e-0654-4dbc-9ef7-d123955fa6bf · outbound

This paper cites Let's verify step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.851235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.851235Z digest=sha256:7043e4f1eee7884a62f44b21a1804d7b3007c24c3890ae8f955da5468716617b

Observation e3ca2ecf-ebce-4e01-8414-30e0e656cd66 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.974215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.974215Z digest=sha256:f852c24a158803cee5cb0a644f26ffc09c9d7b758e196ed2bcac3e2e6aa81c2f

Observation ba0de3c6-a8df-4580-8830-c8cb4ec0ff69 · outbound

This paper cites Training language models to follow instructions with human feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Training language models to follow instructions with human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.072863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.072863Z digest=sha256:ffb9d16e39c2bf34bbff77722215c4d623f7a14f130624a434ff917dde30fabb

Observation c7627424-038d-4c6c-b72c-cd0a68fa5f26 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Direct preference optimization: Your language model is secretly a reward model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.132509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.132509Z digest=sha256:b2771277ccb482262ca64d736599cbd42cee6667ce5973c032848b6cf8e52fcd

Observation 0a5508d8-bd6f-4483-a7d7-4feb62be20b3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.247181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.247181Z digest=sha256:615284eb13e2ac53c6c58273cc072ffb0894b4b06ec48877f59bbd24d19076af

Observation ca9dc9fc-7c7f-414a-9001-8d0c75cf5792 · outbound

This paper cites Trial and error: Exploration-based trajectory optimization of LLM agents.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Trial and error: Exploration-based trajectory optimization of LLM agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.374197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.374197Z digest=sha256:ef1abceb9204f74d50b0da01178a279a1ff29b43d7f73b95447a35a5d6b81698

Observation 0b254d60-96f2-4a93-8945-e698f000a5d8 · outbound

This paper cites Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.452605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.452605Z digest=sha256:07e0d504578fc69505155893c0efc234d4ba60272e389892d29c93b84a90376a

Observation ad41f962-9107-4136-8949-6e47c652b4d7 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.417785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:51:42.540673Z digest=sha256:691eff5961970a72492e56799c6772bae4ed69fae37bae4059cbd22be4c022e1

Observation 742ac0a8-38d1-4a25-b960-0002ac15c9f1 · outbound

This paper cites Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.634639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.634639Z digest=sha256:0a429289a16b5b2513623bb7cd628066183ba3f4361a721e4818ec0a373963bd

Observation 21b930cb-cdb7-4214-ab74-04d336c0e2a8 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Chain-of-thought prompting elicits reasoning in large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.713789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.713789Z digest=sha256:a59b99e22a62c252dd35afe147a2aa3ed1d9c804dfdc8f36d34fa1d9486abc34

Observation 9eb5afc0-64c0-48da-b24b-8ccaca91f770 · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.816272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.816272Z digest=sha256:a34bd87d20fbd8bc91f003ea38c0062a0e9c6aa90db982aa5d1925c17d468b1c

Observation 168df566-aeae-4b62-bb23-f970c537a8b9 · outbound

This paper cites Intercode: Standardizing and benchmarking interactive coding with execution feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Intercode: Standardizing and benchmarking interactive coding with execution feedback

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.176909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:51:42.879826Z digest=sha256:bf6dc62aed9190060bc644805fabc7c0e144ca2ec2319c5a9b2748bcbea2392f

Observation 26ca5ac3-7e97-4b7c-8ede-2f9a8aa11425 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.009609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.009609Z digest=sha256:4888f0353128e56fcab75c1b2c1a7c63052bab37531dce2ec79b07f401c5dd87

Observation 8a53a5e1-3352-4d9a-9141-eb749838ead0 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories React: Synergizing reasoning and acting in language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.080279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.080279Z digest=sha256:c27092dce7c80a74852ccb514801868cdc3cd7c109cda90cf82c7cb573b649c9

Observation eb73a048-7729-42ef-8dd7-71ebfd3c7c38 · outbound

This paper cites Debug like a human: A large language model debugger via verifying runtime execution step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Debug like a human: A large language model debugger via verifying runtime execution step by step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.139184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.139184Z digest=sha256:75796875a7ae5a04dc1bb6ca37ddeff038e34686a7285b41c7038fec7cec5b70

Observation c9aaadfe-1183-415f-8367-027decb7b1da · outbound

This paper cites Least-to-most prompting enables complex reasoning in large language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Least-to-most prompting enables complex reasoning in large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:43.901711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:51:43.257177Z digest=sha256:ac84181ad3c587fc1f6f9784931f2186d0cc5137ec6250145f9abd7d680cbef8

Observation 89da84bb-3e8e-4bba-bb80-1899d29d02a5 · outbound

This paper cites @esa (Ref.

RRO: LLM Agent Optimization Through Rising Reward Trajectories @esa (Ref

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.371191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.371191Z digest=sha256:00549d9fb3b3f0d31218a0e6cca7fa5de5ccaa9f00fb46f9d9607c3a7d421b15

Observation 3a0a0449-7562-4032-8f2a-51fe5889d20f · outbound

This paper cites an unresolved cited work.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.499214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.499214Z digest=sha256:18d0a15f43e90d28563703526b97cd20dc27c7b13a05ee93b968ed728ffbb6d9

Observation 6c2b883c-d0da-4453-b261-85cf64a77b70 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.605901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.605901Z digest=sha256:dd94c1c044b287c31df8841b085e9c441ec3b7a2897f3e4b8b5ca5a4a0cbc93b

Pith citing papers

No inbound Pith citation observations are available.