Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:43.605901Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2505.20737.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:43.605901Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0faff669-6ec0-4b39-a99e-fe973c712998 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c2f4002-50e6-489d-adf2-d869cb36b850 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Language models are few-shot learners
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34e23431-e069-445d-90ab-5951361985e9 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories FireAct: Toward Language Agent Fine-tuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd212bbe-d3f0-4958-8578-56ae96444fd3 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ae589bc-8a04-4b7e-92cb-0bc154a358af · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Deep reinforcement learning from human preferences
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9651b20e-3ac5-41c8-a74d-50a085c7d770 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Complexity-based prompting for multi-step reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55751f3d-ce1b-40cb-a013-31ebdee19dab · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Large language models are zero-shot reasoners
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa6e4be-6a92-465d-9c0b-61f713909791 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfbee95e-0654-4dbc-9ef7-d123955fa6bf · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Let's verify step by step
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3ca2ecf-ebce-4e01-8414-30e0e656cd66 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0de3c6-a8df-4580-8830-c8cb4ec0ff69 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Training language models to follow instructions with human feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7627424-038d-4c6c-b72c-cd0a68fa5f26 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Direct preference optimization: Your language model is secretly a reward model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5508d8-bd6f-4483-a7d7-4feb62be20b3 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Proximal Policy Optimization Algorithms
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca9dc9fc-7c7f-414a-9001-8d0c75cf5792 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Trial and error: Exploration-based trajectory optimization of LLM agents
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b254d60-96f2-4a93-8945-e698f000a5d8 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad41f962-9107-4136-8949-6e47c652b4d7 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 742ac0a8-38d1-4a25-b960-0002ac15c9f1 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21b930cb-cdb7-4214-ab74-04d336c0e2a8 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Chain-of-thought prompting elicits reasoning in large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eb5afc0-64c0-48da-b24b-8ccaca91f770 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 168df566-aeae-4b62-bb23-f970c537a8b9 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Intercode: Standardizing and benchmarking interactive coding with execution feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26ca5ac3-7e97-4b7c-8ede-2f9a8aa11425 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Webshop: Towards scalable real-world web interaction with grounded language agents
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a53a5e1-3352-4d9a-9141-eb749838ead0 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories React: Synergizing reasoning and acting in language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb73a048-7729-42ef-8dd7-71ebfd3c7c38 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Debug like a human: A large language model debugger via verifying runtime execution step by step
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9aaadfe-1183-415f-8367-027decb7b1da · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Least-to-most prompting enables complex reasoning in large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89da84bb-3e8e-4bba-bb80-1899d29d02a5 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories @esa (Ref
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a0a0449-7562-4032-8f2a-51fe5889d20f · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c2b883c-d0da-4453-b261-85cf64a77b70 · outbound
RRO: LLM Agent Optimization Through Rising Reward Trajectories Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.