Pith. sign in

Paper Citation Record · LEDGER

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

As of 7 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2605.01327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.01327 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T14:44:31.160543Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact14
  • verified fuzzy2
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57500d56-dcb2-4521-9ae0-83162d278553 · outbound

This paper cites Qwen2.5-VL Technical Report.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.841335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:acf8d18c3fa5de7f9df1fefc8932b5a5cde0506b1547b155f74dee3d088f2db0

Observation c3192f1e-9139-4f3a-a8a9-fdae2d14ec60 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.612219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:621aa0e0058ffe331b4087c01093754524629a51469b0386786afcd7d2fb9b59

Observation fad68340-1666-48f6-87af-96c5bdc877ea · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.852419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:40acebb545a3fae46d271e1a6d9b1ec02d2bb4d1246d474ed30339be3b803d66

Observation df3831d5-8104-445d-8732-2eb5f1bc24f0 · outbound

This paper cites Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:07.834936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:81ebadcdc3f2f8146c89fa6d9f85dcc9e18b21d13b41d3af7d0ea72a57eaa403

Observation e6e99dd3-3679-4fcf-826d-619c11f07f40 · outbound

This paper cites Rectifying LLM Thought from Lens of Optimization.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Rectifying LLM Thought from Lens of Optimization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.826343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:429a8ec14e81b1a8f90e2e536a84004e7592944ad1b87af68f0f699d470b53d9

Observation bc31d786-8a70-4b1c-8076-a10e29f79c4e · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:08.012769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:d05db9841eaee3e1be4a60b0504b3c958e30f2bffe02b15f10ce96ef59710ca1

Observation 8c7bb46e-d9c8-4dd6-9a0d-e94ff0e7a0e0 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.951371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:773514fb85026c296897164f557e946d6c86d133bfeb0f833d4c6759312fa3ac

Observation dd2360ba-092e-4baa-9a4b-5c808e68b6a5 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:00:08.752459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:ab53a90307782462a63d358cb0a1f0046513c7d34bdc26b5df0e6c48894e8bcb

Observation fff18f11-d21d-43f3-bc31-21a0e6f38277 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:04:44.937279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:8f72bcefa029eccc01c1f10af8205d6171103f1cbe082c51b476dd6fb20d0b16

Observation b91e4203-3f45-4d49-a053-a7a36d8defcf · outbound

This paper cites Proximal Policy Optimization Algorithms.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Proximal Policy Optimization Algorithms

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.878376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:85d0fbf1decb418415aacd92ad0333d82af49c146a05b42b35f693fee3ab29f1

Observation 488272d9-e371-4749-a8f5-851bd352c3b3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.904346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:dff3dcb1ca384c37660ad24f6c5d7186d138365d772dfb82486f674f24c1b11d

Observation 7f8b7a38-3b6b-4ba0-80a8-754214783e0c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.868222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:f7e23930f88f7883eb93a294e34d4378139e51a08ba8b6c2095326dbe1f0de90

Observation 83deb363-64c7-4970-958d-fc6bd40d1a98 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:07.861678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:e379be85edcf01021cb2a2c1841aeb2e5d024fe63c8ba280d383e232e498e7f8

Observation d78e37a4-71cf-47eb-93eb-2689954b05e7 · outbound

This paper cites Sutton, Doina Precup, and Satinder Singh.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Sutton, Doina Precup, and Satinder Singh

Reference 14

Resolution
metadata mismatch
doi, observed 2026-05-09T22:18:59.682350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:b369bd395c231ff80aa3e71d0cb096ae6f3c3ddeb4794b07df44a7d889b5cec9

Observation ee970ee1-02c0-41a9-88f0-c77b1c286bc6 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:47:26.275168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:363633c052426021ad1ca1940791dbceaf2f3f76b4f2a33c96259c3e20edc2d0

Observation 910260f2-c59d-4abd-ad8c-d6f7827ee00f · outbound

This paper cites Single-stream policy optimization.arXiv preprint arXiv:2509.13232.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Single-stream policy optimization.arXiv preprint arXiv:2509.13232

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:07.964146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:f9c7cd813594d5ab87e03a58bb027e6ceb5067ab8276bc5aa0e018719f32835f

Observation 2fe4635e-d5d1-4b5f-99aa-f632a0fb50c4 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:07.895095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:274535a388a19b86383d12b15175a9cd532101e4fdaef49b1eb0ac2bda420cea

Observation 6b88b11e-4781-4203-9539-f00e70edb1cd · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:4984f4d373b0effb96aaae632f54a249cdaa6cd3efd537b19bbcdce79ea3b116

Observation ee07488f-8eb9-4377-b331-e03faf20c0ca · outbound

This paper cites Stolfo, A., Balachandran, V ., Yousefi, S., Horvitz, E., and Nushi, B.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Stolfo, A., Balachandran, V ., Yousefi, S., Horvitz, E., and Nushi, B

Reference 19

Resolution
metadata mismatch
doi, observed 2026-05-09T22:18:59.686726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:b9c4bb023ed27b1ebbfb85ac9aae974f01d4eb3ab2098911addf71a226cdab72

Observation 2d067ac6-72f4-4e55-9123-4af0e02a0bf3 · outbound

This paper cites emnlp-main.668/.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning emnlp-main.668/

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:47:16.637643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:ae3e4cab578aea56747d51f38f3dec8d3449aec96349d7f67955ae164c8ba35e

Observation dd39c7b7-2c8d-45cc-a046-8e1008d9c602 · outbound

This paper cites Group Sequence Policy Optimization.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Group Sequence Policy Optimization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.932042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:0e634f8bd6ff785ddf13b2427bed47e0c5e81bba9e859b76c04cbe055ab4d5d0

Observation 574323eb-0efb-4a20-8e7b-e5dc35346274 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:07.917445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:f57fd75268b17e96c926080486b6bae0dac189caae57c66c6392a301c59c81bc

Observation 69ba39f0-073a-4adb-ad41-becb0ca9c2d7 · outbound

This paper cites Benchmark Settings.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Benchmark Settings

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:47:16.636649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:ed328687e789b914334d9a1d8bf65d0dbeccc1197153e7de4e81bf44efbc6534

Observation ccb2e44b-165b-4b55-9977-63aad079ea22 · outbound

This paper cites an unresolved cited work.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-26T01:47:16.646518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:ca4169a4fe1a072ea7add6d89966c35c69309562ac19ab23c17e760360171b44

Observation 46acefa1-ba61-4f9c-828d-f82f9f4b4520 · outbound

This paper cites an unresolved cited work.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-26T01:47:16.642592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:86e93ade09ef7f80d87c23acb799eb6d6cf766d357bf65b0c883c68349aace97

Pith citing papers

No inbound Pith citation observations are available.