Pith. sign in

Paper Citation Record · LEDGER

RRO: LLM Agent Optimization Through Rising Reward Trajectories

As of 21 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2505.20737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20737 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:43.605901Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0faff669-6ec0-4b39-a99e-fe973c712998 · outbound

This paper cites write newline.

RRO: LLM Agent Optimization Through Rising Reward Trajectories write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.015662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.015662Z digest=sha256:122237ba3b59f08a90d32c26db03b46654cda647d86e60c44f03701ff3dacd71

Observation 4c2f4002-50e6-489d-adf2-d869cb36b850 · outbound

This paper cites Language models are few-shot learners.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.148103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.148103Z digest=sha256:0764403d00fb008be708ba148ab1ff01c5f7d1c9b49b3080f4415bd42e398fb4

Observation 34e23431-e069-445d-90ab-5951361985e9 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

RRO: LLM Agent Optimization Through Rising Reward Trajectories FireAct: Toward Language Agent Fine-tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.263415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.263415Z digest=sha256:472625aa75c74cf8d731845e7671b9aa4c07c3a54b8f7edb312f7901fe5b09c4

Observation bd212bbe-d3f0-4958-8578-56ae96444fd3 · outbound

This paper cites an unresolved cited work.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:51:44.799382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T13:51:41.347777Z digest=sha256:bba103fd687ab7eff49401068632dd3e027e7c455b42ab14231d27a029aa3c7b

Observation 0ae589bc-8a04-4b7e-92cb-0bc154a358af · outbound

This paper cites Deep reinforcement learning from human preferences.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.442326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.442326Z digest=sha256:d4d22d279e6761048881bd4d86540240a5d004de14e4f493808d40de076f3b92

Observation 9651b20e-3ac5-41c8-a74d-50a085c7d770 · outbound

This paper cites Complexity-based prompting for multi-step reasoning.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Complexity-based prompting for multi-step reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.616698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T13:51:41.525532Z digest=sha256:120f76af6147820b043f47b04840b3da8b10f6c2776a5b634a164140a2f7569c

Observation 55751f3d-ce1b-40cb-a013-31ebdee19dab · outbound

This paper cites Large language models are zero-shot reasoners.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Large language models are zero-shot reasoners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.647817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.647817Z digest=sha256:3094a261b619c0df668b72e3cb412e8df9d9d7be807dc49649878c5893b6f6cb

Observation 0fa6e4be-6a92-465d-9c0b-61f713909791 · outbound

This paper cites Let's verify step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.770139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.770139Z digest=sha256:9e74b5141723d1a6423aaa0470273f59bfe2e6d388c54b70cd018c8ff388b267

Observation bfbee95e-0654-4dbc-9ef7-d123955fa6bf · outbound

This paper cites Let's verify step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.851235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.851235Z digest=sha256:279b24ee89421ff40a06ed8b8a4dfed87c74f105076b9cd4d9c558234b76e61d

Observation e3ca2ecf-ebce-4e01-8414-30e0e656cd66 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:41.974215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:41.974215Z digest=sha256:b82424c4bb84a02456b7bc21701948513e3bda67846935b853bdc3b12a5e3e58

Observation ba0de3c6-a8df-4580-8830-c8cb4ec0ff69 · outbound

This paper cites Training language models to follow instructions with human feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Training language models to follow instructions with human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.072863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.072863Z digest=sha256:119eb457458d9783078b8e68236c7d6832992470f54f551be7d9e3da420c01ee

Observation c7627424-038d-4c6c-b72c-cd0a68fa5f26 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Direct preference optimization: Your language model is secretly a reward model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.132509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.132509Z digest=sha256:7c1d3d89191e1701b2d1f3c47e750ffd869c3f003fccfa94c40f99ff6e21eebc

Observation 0a5508d8-bd6f-4483-a7d7-4feb62be20b3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.247181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.247181Z digest=sha256:66845e7405aacc240fa382aa7c6fa2eb47189feb602da12135c2c3e1615cb583

Observation ca9dc9fc-7c7f-414a-9001-8d0c75cf5792 · outbound

This paper cites Trial and error: Exploration-based trajectory optimization of LLM agents.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Trial and error: Exploration-based trajectory optimization of LLM agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.374197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.374197Z digest=sha256:8897acfb6301b44e63bc5b2e09b7fe86321e66b87c497febb5b5bccf95486bb9

Observation 0b254d60-96f2-4a93-8945-e698f000a5d8 · outbound

This paper cites Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.452605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.452605Z digest=sha256:61029992b5f8a9d5bbed89a61dc05965097232bc4c417236b05a4ebe0c28360e

Observation ad41f962-9107-4136-8949-6e47c652b4d7 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.417785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T13:51:42.540673Z digest=sha256:2275edc54b51399d970606b80b87a2730008d03e6a447218eee3f4826796208c

Observation 742ac0a8-38d1-4a25-b960-0002ac15c9f1 · outbound

This paper cites Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.634639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.634639Z digest=sha256:e8041390c34490d5afcaf88afb79297b49f54938b47cba27ee7ce0e903ff1d03

Observation 21b930cb-cdb7-4214-ab74-04d336c0e2a8 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Chain-of-thought prompting elicits reasoning in large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.713789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.713789Z digest=sha256:e2049a1a91520849ec82c4051bd154fd7187212aeffdcfa48636ec46d4ad15a9

Observation 9eb5afc0-64c0-48da-b24b-8ccaca91f770 · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:42.816272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:42.816272Z digest=sha256:0b4dcc1fc5a133845080234b54137dfaaf75aad904d793ca65811c7b112cbacc

Observation 168df566-aeae-4b62-bb23-f970c537a8b9 · outbound

This paper cites Intercode: Standardizing and benchmarking interactive coding with execution feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Intercode: Standardizing and benchmarking interactive coding with execution feedback

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:44.176909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T13:51:42.879826Z digest=sha256:a243b3f320cacaa292070598bc78fbc78ac2b685035a705b0d4e3640f69fe75d

Observation 26ca5ac3-7e97-4b7c-8ede-2f9a8aa11425 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.009609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.009609Z digest=sha256:5361056aa134637577854a378115767e35637c257ed70a7fd9d9825f43d227b9

Observation 8a53a5e1-3352-4d9a-9141-eb749838ead0 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories React: Synergizing reasoning and acting in language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.080279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.080279Z digest=sha256:2bf6895a289e293b7cdc4cec64b212de5456d1fef81f0f71613bd08cc99d9e23

Observation eb73a048-7729-42ef-8dd7-71ebfd3c7c38 · outbound

This paper cites Debug like a human: A large language model debugger via verifying runtime execution step by step.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Debug like a human: A large language model debugger via verifying runtime execution step by step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.139184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.139184Z digest=sha256:981dd4c0e3d3bfa32c8679474b087919cc7ccf0d20cb47b8eca5ac2ec60e7eb6

Observation c9aaadfe-1183-415f-8367-027decb7b1da · outbound

This paper cites Least-to-most prompting enables complex reasoning in large language models.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Least-to-most prompting enables complex reasoning in large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:43.901711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T13:51:43.257177Z digest=sha256:39a55944588f02e3c814da166733be2a79e103f4fbb02773588c15aba89cbbf7

Observation 89da84bb-3e8e-4bba-bb80-1899d29d02a5 · outbound

This paper cites @esa (Ref.

RRO: LLM Agent Optimization Through Rising Reward Trajectories @esa (Ref

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.371191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.371191Z digest=sha256:234d47ca016976f0fb1cc66d8427de642fd11329ed39f28a06731edf9e493795

Observation 3a0a0449-7562-4032-8f2a-51fe5889d20f · outbound

This paper cites an unresolved cited work.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.499214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.499214Z digest=sha256:7d56b1aa479fe1c1b590ac3d223367b83737f314d4dbf67a9ab15335c8d13c9f

Observation 6c2b883c-d0da-4453-b261-85cf64a77b70 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

RRO: LLM Agent Optimization Through Rising Reward Trajectories Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.605901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:43.605901Z digest=sha256:5c328510614f841f9ed198c4d270ba3b5c73dc820498938b63352b96cd4a24d6

Pith citing papers

No inbound Pith citation observations are available.