Pith. sign in

Paper Citation Record · LEDGER

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment

As of 14 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2411.11681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11681 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:20:14.079797Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4064787b-f7f6-443e-8f16-7f71e4de827f · outbound

This paper cites G.; Guo, Z.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment G.; Guo, Z

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.762028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.855145Z digest=sha256:192b8d4d37d14a4c1fa6d728f5285f61a9e0d128d3d4e510a33e2421abc61e3f

Observation aa7c7899-4785-4ec0-a6aa-cfd9ab8ec6e0 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.743311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.862914Z digest=sha256:a05513f71eccf9c79d627a206bd8aba744c5494cf5b35d267489c9fca207112d

Observation f80b4755-1396-47f1-9daa-b465b198ac61 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.869596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.869596Z digest=sha256:a61a36db4999affa1cdf23a80bb28dd6b670aa115cf1a776c6d28131ae703bcc

Observation 06e1e333-0396-420c-ae1e-1cf16f1ec4c3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.876021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.876021Z digest=sha256:b8d40fa1079ea23a4481181763ff1fb13dde571e899f97df36b8176dabbc8bee

Observation a0454d7f-5243-4c11-bf9b-2248ba093ae7 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.883203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.883203Z digest=sha256:20bc61a52557bec0f9f88eebd963c65273b242f8d5ce3e216f980b9aa4030b43

Observation 8043d3c7-1bee-4bd9-afa5-387d339eff64 · outbound

This paper cites The Llama 3 Herd of Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.890056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.890056Z digest=sha256:d8bb08d7353cc3fdcb650205f227cd2977c07acabe64c110dff6d441c95296e5

Observation eaccdb49-ae77-4a11-a633-333202d96141 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.699714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.896845Z digest=sha256:b8fe073ed3afbf2e9849cbd2654261d56016ace8a0f338427022a4d684c9e211

Observation 9497a753-89b6-4f82-9ce7-188687de12d1 · outbound

This paper cites J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.902327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.902327Z digest=sha256:ce4aaa3d2635a1ef51c8753f91e99b18165c8e1a2d63aa45f23529f7c1e8991b

Observation 0f71042d-a3e2-432a-9b13-8d46bdd26366 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.666411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.907171Z digest=sha256:ad718c48d66ec6c960d0ddbfcc1b43da1847d53c373fd6bd18fd0510940ca1d6

Observation 1ac934fb-3d04-40f6-91a5-53573afec9ee · outbound

This paper cites The Impact of Reasoning Step Length on Large Language Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment The Impact of Reasoning Step Length on Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.912332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.912332Z digest=sha256:5a38db0e605229d11a0613863462fcf2afb6f17a18db0271ed9f7b833cb8a1f4

Observation 17919b9b-4c53-43bc-83b3-d6b809a27bc3 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Large Language Models are Zero-Shot Reasoners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.918672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.918672Z digest=sha256:8fa055fa1dda08e8e0f615a61332bd7daab214c6b1e1c15e8f464222c10ce6d2

Observation bf05dcc1-26d7-43c1-8876-d01f3cb28f33 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.925085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.925085Z digest=sha256:98aa7a687c3b4ebb2a841c8d54720f0b54fe721913d25805899af343364647ea

Observation c4802d0c-712b-4a43-b151-360f2d971be9 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.648304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.931585Z digest=sha256:83c0d712495c0423d681fcc5511a0408bf45e4ccfc1913b2e2b80dbb75485c2a

Observation 80bd7a93-1d16-4484-a643-d5174506a6a4 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.626816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.937717Z digest=sha256:18b6c97044b79944f81eba36f4fd48a3d101c63564e051193ab6ffd8048874b4

Observation 288e386e-b59c-470a-8802-5dd726898d8c · outbound

This paper cites Let's Verify Step by Step.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.943239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.943239Z digest=sha256:b6ef3fe3936cd97fbbda5b6ab23d863e591061f9c9b4b7a4042dc1c2d00ba885

Observation 3637dcc3-b26e-4069-b076-782c3ec85f84 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.949599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.949599Z digest=sha256:5e7c55d7de6f6c661cbd39dc561a84aeeff4ac0b4d222bc04c6ec25bfb44d0b8

Observation 935a371f-ddf4-4ee1-b8bc-22d547689031 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.956530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.956530Z digest=sha256:d49a80f43c917ec2e62154f1851381b03e5274a63b07c0e0c2d82a24537e15c9

Observation 69447051-366a-465c-8477-f75bb9e5cf60 · outbound

This paper cites L.; Bari, M.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment L.; Bari, M

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.605398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.964313Z digest=sha256:0e0f75ef507315319e803da147e6e2b9afde53d9c31cf7bececac4aa855e29f2

Observation 0c50e244-fe21-46ca-8f14-cea0f08fe4e8 · outbound

This paper cites L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.971334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.971334Z digest=sha256:80ec97529f426174b5f39fd22720aeafe6fadbb3a62ebb01899180ca571d7f59

Observation 9a045a4e-846e-477d-a9fe-30b6ebb88054 · outbound

This paper cites D.; Ermon, S.; and Finn, C.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment D.; Ermon, S.; and Finn, C

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.978109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.978109Z digest=sha256:0a653c558e67850fe708513ea53f5d7d5495d5fc34f4ff26193f5ac191ecce9d

Observation 23db6937-3951-40e4-9360-c451fe2ac361 · outbound

This paper cites Proximal Policy Optimization Algorithms.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.984199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.984199Z digest=sha256:d1189e38f424716a2671f96bc79e06d829eb469e15fc053bc04dd64d48dd327f

Observation 9da5c1f5-4f36-48fc-b445-30169a388486 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.990185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.990185Z digest=sha256:588e6537c455725ac6eee14641e197a59ca64f282e24c23d78230e0af6a2e897

Observation 69932421-385d-4f3a-9afc-315608672f4b · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Solving math word problems with process- and outcome-based feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.001444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.001444Z digest=sha256:9fc384b658c8e1d1d10c73f3f7170ec91280848df4b0d10dfb0214eed546ff77

Observation 7ac11809-c912-46ac-a4e3-d6e3c29ad4a2 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.014172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.014172Z digest=sha256:8c289f985cc61e2dbbf0b65f43b617d29f5de8ab12d4067bd57dbe06298e5906

Observation 7a8d09c2-9393-499f-af82-283e414df92d · outbound

This paper cites V.; Chi, E.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment V.; Chi, E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.560218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:14.020060Z digest=sha256:a6e42b78b76d19f6ff96ad7a2b665d5a116035fc21a762fa1b64116cf08fd88f

Observation d195a577-dc7f-4dc9-bf04-853b4077e306 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Aligning Large Language Models with Human: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.025705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.025705Z digest=sha256:12a47099046fe80b7f9817f678e64b11ee3a90cb0706b6c84c190b90918bff0c

Observation c7a366ea-27b8-4c33-909a-2a9edaa7432d · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.032820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.032820Z digest=sha256:ef3d2140b8e90290857e0921858467a0128dd0df9720ba511eb79ff9cf5db47f

Observation 573e6d72-f60c-4d78-8631-f0b8bf4e563e · outbound

This paper cites H.; Le, Q.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment H.; Le, Q

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.542206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:14.038622Z digest=sha256:d112c159cc1b84e8843b87512ea82e5279c52b6e4699f9a3e333cb190f56979e

Observation 79d8f286-c094-44d7-87c7-74217c82891a · outbound

This paper cites Qwen2 Technical Report.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Qwen2 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.043754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.043754Z digest=sha256:33d965715c798f03aefd9d8ee6e3b8eb4d6ffeae180ace16211777287c347b65

Observation df8da85d-3057-4c60-847e-8f10c225b51b · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.520712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:14.049470Z digest=sha256:0c20b12af48176e6d21cc52422157210407ed95d67391da3b94d286d4859d88a

Observation ec768db9-f861-464d-ae29-35ae60d00303 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.055687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.055687Z digest=sha256:422e9a9c697f56ca966783a01389611704fbd3d957a4c2dbfc24107d0dd4c736

Observation e065f485-159f-40ea-b431-68dd1d3c3f96 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.061179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.061179Z digest=sha256:4b4f2ebb88eea9ccaf115616cb1ba626a99613bba58585f2eb4676962f058970

Observation 948fde31-2f2d-4ef5-8106-19293324a369 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.066939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.066939Z digest=sha256:53ae2c6f7e5f59d194b83078de68b5ce56a9b234a9b03621fe780f12bc84796b

Observation 3e0b9875-2783-407a-a813-3cbad5e146a6 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment , " * write output.state after.block = add.period write newline

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.073328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.073328Z digest=sha256:763085871cec3ffc280bb1f2e75232d325f01bd523cb0679726fe7beaa561ee4

Observation de6d25ca-9a10-465f-b98d-d47e637beaef · outbound

This paper cites write newline.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.079797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.079797Z digest=sha256:f89a65ff86546664929a0d1ee7c134ee1afa8ed8ae275bd28cf340296ea98280

Pith citing papers

No inbound Pith citation observations are available.