Pith. sign in

Paper Citation Record · LEDGER

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment

As of 14 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2411.11681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11681 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:20:14.079797Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4064787b-f7f6-443e-8f16-7f71e4de827f · outbound

This paper cites G.; Guo, Z.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment G.; Guo, Z

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.762028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.855145Z digest=sha256:5d050fa7e01baa973aba7ebb2589ffcaa36d23920cc351f8d03d8422ebef2464

Observation aa7c7899-4785-4ec0-a6aa-cfd9ab8ec6e0 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.743311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.862914Z digest=sha256:85d2bc7364bfcdead3aad5696dde1d6e00b1cdd7c1309e2c545cd64f4aed1cc3

Observation f80b4755-1396-47f1-9daa-b465b198ac61 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.869596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.869596Z digest=sha256:47afda0b386e5f8e91054d7c4c862a3b44b5fb1c1f4cec49d17598bab4cfa309

Observation 06e1e333-0396-420c-ae1e-1cf16f1ec4c3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.876021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.876021Z digest=sha256:045598c6462807435cacee094418e5818ee2317b57593eb9f66ffc48d3638e52

Observation a0454d7f-5243-4c11-bf9b-2248ba093ae7 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.883203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.883203Z digest=sha256:914f7fbe989b2cec370fac858365f3b7155c09da43815f563f070bad16f3deee

Observation 8043d3c7-1bee-4bd9-afa5-387d339eff64 · outbound

This paper cites The Llama 3 Herd of Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.890056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.890056Z digest=sha256:b94efdff92f2d66c780aff5f397745fd53eb6ce1a5c951196430aa1bcee153f9

Observation eaccdb49-ae77-4a11-a633-333202d96141 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.699714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.896845Z digest=sha256:343ff55e79368ad84e153331ca431eaf8f7e8fc4e2cc76566e994f63177e5994

Observation 9497a753-89b6-4f82-9ce7-188687de12d1 · outbound

This paper cites J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.902327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.902327Z digest=sha256:8ab455ddea26b74e04d94fb048122b80435478b3d179dc29137d1253fe240c98

Observation 0f71042d-a3e2-432a-9b13-8d46bdd26366 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.666411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.907171Z digest=sha256:aa421631ba3c8b6573cadb2ebf60bd0283ae2ed535d03ca80f95f2c1d1024bd0

Observation 1ac934fb-3d04-40f6-91a5-53573afec9ee · outbound

This paper cites The Impact of Reasoning Step Length on Large Language Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment The Impact of Reasoning Step Length on Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.912332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.912332Z digest=sha256:97db0e76ea261e43f1d1d4be05428cc5872c40efe1d7b7447fee67ad139ec9f3

Observation 17919b9b-4c53-43bc-83b3-d6b809a27bc3 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Large Language Models are Zero-Shot Reasoners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.918672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.918672Z digest=sha256:ef51074b64437a00d4f640f33a67edff6033213dba1cc052b039a7b984cfed41

Observation bf05dcc1-26d7-43c1-8876-d01f3cb28f33 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.925085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.925085Z digest=sha256:cfbee768624bddf7def93fd21bf1d365f2eb579d9a8e5badaa4269c2fc5f17e3

Observation c4802d0c-712b-4a43-b151-360f2d971be9 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.648304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.931585Z digest=sha256:910ea35c17dd6dae7c336a35b05d7f90251185f47d716dd2de1a0eb660a876d9

Observation 80bd7a93-1d16-4484-a643-d5174506a6a4 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.626816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.937717Z digest=sha256:ab605af74e6eec05d4cd5522b292b4f562c49632e5fae83c547f48b9a0e794fe

Observation 288e386e-b59c-470a-8802-5dd726898d8c · outbound

This paper cites Let's Verify Step by Step.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.943239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.943239Z digest=sha256:b228f2272f5577bffa9d105dcb9a8b886a88f2f4d568937921a220ce344ab45c

Observation 3637dcc3-b26e-4069-b076-782c3ec85f84 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.949599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.949599Z digest=sha256:dea0e3ae3b437bbe891b8ca7277eb78a867b26b12ffb01d4abc52627f170d021

Observation 935a371f-ddf4-4ee1-b8bc-22d547689031 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.956530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.956530Z digest=sha256:6ce5b6461fc475e06c7a270d8a819e6bfefd4063088bf16c29778a656c8cd1b4

Observation 69447051-366a-465c-8477-f75bb9e5cf60 · outbound

This paper cites L.; Bari, M.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment L.; Bari, M

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.605398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.964313Z digest=sha256:883921add6581a1780fbece6ff12d9b38e7780863d6f0a25a96f3187c02cb39a

Observation 0c50e244-fe21-46ca-8f14-cea0f08fe4e8 · outbound

This paper cites L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.971334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.971334Z digest=sha256:30def00bd99a09755c812d7fd6070915afcbbf58a53d7fc013b4605715039d1c

Observation 9a045a4e-846e-477d-a9fe-30b6ebb88054 · outbound

This paper cites D.; Ermon, S.; and Finn, C.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment D.; Ermon, S.; and Finn, C

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.978109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.978109Z digest=sha256:39e642886f51f2217a24d6777e11a151214cd28a73ff5d0cac84ccf07d262b80

Observation 23db6937-3951-40e4-9360-c451fe2ac361 · outbound

This paper cites Proximal Policy Optimization Algorithms.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.984199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.984199Z digest=sha256:c91cd7f2a2901743da92f137f551fdcc9693bf4af987b077c3d54a7820d2a8f3

Observation 9da5c1f5-4f36-48fc-b445-30169a388486 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.990185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.990185Z digest=sha256:bfc666ce4c97c0e9237a46c1d96546c3943ca0c8836c0dacf639f0f77a4c0d76

Observation 69932421-385d-4f3a-9afc-315608672f4b · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Solving math word problems with process- and outcome-based feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.001444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.001444Z digest=sha256:d24e23994de6f0e6b2c1af9457f5c830c3b1594626d39737b6a00afcb3b127cf

Observation 7ac11809-c912-46ac-a4e3-d6e3c29ad4a2 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.014172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.014172Z digest=sha256:23cd945d46b1c785138f81a4ea8e81641ec6a960ee592bc496253a8a83707cb7

Observation 7a8d09c2-9393-499f-af82-283e414df92d · outbound

This paper cites V.; Chi, E.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment V.; Chi, E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.560218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:14.020060Z digest=sha256:198a234053adac1bca15dd58f915cd9ea5e1979596660e09baf921a1b34344da

Observation d195a577-dc7f-4dc9-bf04-853b4077e306 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Aligning Large Language Models with Human: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.025705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.025705Z digest=sha256:ea3b4cade80ca4df4260e67a3a0cc3abb6bcfea2358f9ef35fba1bf820ab48ff

Observation c7a366ea-27b8-4c33-909a-2a9edaa7432d · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.032820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.032820Z digest=sha256:0bf89ef7309385caec764a702c554dc25874a2aed4da75bc73b0bfbfc8fd3a20

Observation 573e6d72-f60c-4d78-8631-f0b8bf4e563e · outbound

This paper cites H.; Le, Q.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment H.; Le, Q

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.542206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:14.038622Z digest=sha256:b398e346e9302d22a726026d0fa90888926023d210f7e0679367c25caf8063f9

Observation 79d8f286-c094-44d7-87c7-74217c82891a · outbound

This paper cites Qwen2 Technical Report.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Qwen2 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.043754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.043754Z digest=sha256:f12f6749b0dcde70d1cd5ebc1d9efae2d04a767fbb77c601a3a6f138ef2b3546

Observation df8da85d-3057-4c60-847e-8f10c225b51b · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.520712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T18:20:14.049470Z digest=sha256:c2529174686d3142b68432097d1431743b6e5280dbe2c94d73de10b8d4d21ff4

Observation ec768db9-f861-464d-ae29-35ae60d00303 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.055687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.055687Z digest=sha256:c3909b324810aeba6c32648f6a075724db28890e75eead4f2a9cd93faa8422fc

Observation e065f485-159f-40ea-b431-68dd1d3c3f96 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.061179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.061179Z digest=sha256:57b80c2658f5c797151f4dff148219fa72d392367192be3d9983a8b9af2031c7

Observation 948fde31-2f2d-4ef5-8106-19293324a369 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.066939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.066939Z digest=sha256:b7b605cd62aa988730cfa191b1961f3e209ea731ff533b6c79d1de33d672f567

Observation 3e0b9875-2783-407a-a813-3cbad5e146a6 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment , " * write output.state after.block = add.period write newline

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.073328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.073328Z digest=sha256:4fbf87ac9d67009ff76009bb948a9fd258f4273c2fab4363d13d35ef6ac0cdaf

Observation de6d25ca-9a10-465f-b98d-d47e637beaef · outbound

This paper cites write newline.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.079797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.079797Z digest=sha256:664c360234796af240baf516f6ae166c7415afbaca8533eb8ffb8fc8eaf0879d

Pith citing papers

No inbound Pith citation observations are available.