Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:20:14.079797Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2411.11681.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:20:14.079797Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4064787b-f7f6-443e-8f16-7f71e4de827f · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment G.; Guo, Z
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aa7c7899-4785-4ec0-a6aa-cfd9ab8ec6e0 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f80b4755-1396-47f1-9daa-b465b198ac61 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e1e333-0396-420c-ae1e-1cf16f1ec4c3 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0454d7f-5243-4c11-bf9b-2248ba093ae7 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8043d3c7-1bee-4bd9-afa5-387d339eff64 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaccdb49-ae77-4a11-a633-333202d96141 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9497a753-89b6-4f82-9ce7-188687de12d1 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f71042d-a3e2-432a-9b13-8d46bdd26366 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1ac934fb-3d04-40f6-91a5-53573afec9ee · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment The Impact of Reasoning Step Length on Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17919b9b-4c53-43bc-83b3-d6b809a27bc3 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Large Language Models are Zero-Shot Reasoners
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf05dcc1-26d7-43c1-8876-d01f3cb28f33 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4802d0c-712b-4a43-b151-360f2d971be9 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 80bd7a93-1d16-4484-a643-d5174506a6a4 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 288e386e-b59c-470a-8802-5dd726898d8c · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Let's Verify Step by Step
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3637dcc3-b26e-4069-b076-782c3ec85f84 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 935a371f-ddf4-4ee1-b8bc-22d547689031 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Let's reward step by step: Step-Level reward model as the Navigators for Reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69447051-366a-465c-8477-f75bb9e5cf60 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment L.; Bari, M
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0c50e244-fe21-46ca-8f14-cea0f08fe4e8 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a045a4e-846e-477d-a9fe-30b6ebb88054 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment D.; Ermon, S.; and Finn, C
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23db6937-3951-40e4-9360-c451fe2ac361 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Proximal Policy Optimization Algorithms
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9da5c1f5-4f36-48fc-b445-30169a388486 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69932421-385d-4f3a-9afc-315608672f4b · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Solving math word problems with process- and outcome-based feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac11809-c912-46ac-a4e3-d6e3c29ad4a2 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8d09c2-9393-499f-af82-283e414df92d · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment V.; Chi, E
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d195a577-dc7f-4dc9-bf04-853b4077e306 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Aligning Large Language Models with Human: A Survey
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7a366ea-27b8-4c33-909a-2a9edaa7432d · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 573e6d72-f60c-4d78-8631-f0b8bf4e563e · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment H.; Le, Q
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 79d8f286-c094-44d7-87c7-74217c82891a · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Qwen2 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df8da85d-3057-4c60-847e-8f10c225b51b · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ec768db9-f861-464d-ae29-35ae60d00303 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e065f485-159f-40ea-b431-68dd1d3c3f96 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 948fde31-2f2d-4ef5-8106-19293324a369 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e0b9875-2783-407a-a813-3cbad5e146a6 · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment , " * write output.state after.block = add.period write newline
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6d25ca-9a10-465f-b98d-d47e637beaef · outbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment write newline
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.