Pith. sign in

Paper Citation Record · LEDGER

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping

As of 13 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2606.08708.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.08708 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:25:05.437142Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:20:41.136332Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0c59fd0c-4c60-4a4a-905a-157bd73d63ca · outbound

This paper cites The original policy distribution is P=π θ(·|I, c), and the perturbed distribution is P ′ =π θ(·| ˜I, c), where ˜I=P(I).

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping The original policy distribution is P=π θ(·|I, c), and the perturbed distribution is P ′ =π θ(·| ˜I, c), where ˜I=P(I)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:40cc2894ccdbd0fee4b55c05241a0c6fdbff89a8db9172e2787389c60dae2363

Observation 3abfb000-4ab2-435d-b3fb-e975f3ce46de · outbound

This paper cites While St captures semantic dependency, it cannot distinguish whether a high KL value arises from robust visual grounding or brittle numerical over-sensitivity.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping While St captures semantic dependency, it cannot distinguish whether a high KL value arises from robust visual grounding or brittle numerical over-sensitivity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:a634e27b5e2837b589f6092247649061e3727af5871a43d243f5440aba1d3cf7

Observation 12474a78-fd79-4949-b646-1ee88d3ddc20 · outbound

This paper cites -∠ADC= 26 ◦.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping -∠ADC= 26 ◦

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:1013d4209a5557464aeb3dd2fcee5a4824bb06655d5c1850c4d1be01440dc1ad

Observation 79a7781e-b404-49ff-b5d6-707eec5c1352 · outbound

This paper cites Therefore,∠ACD= 2×angle at the center= 2×∠AOD.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Therefore,∠ACD= 2×angle at the center= 2×∠AOD

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:f7361284670d1df8107127826d3f5dd317da292f59181e96dab55c2dd5bf6695

Observation 84f1b899-3603-454d-b645-9352b02b34fc · outbound

This paper cites - Solving for∠CAB, we get∠CAB= 90 ◦ −52 ◦ = 38 ◦.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping - Solving for∠CAB, we get∠CAB= 90 ◦ −52 ◦ = 38 ◦

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:0158c90a327ca3f2c8f5fa5888139071ed178228704ebe32be97b6f832c42bff

Observation a80a4ed9-3939-4000-bc88-e53dd59e7b80 · outbound

This paper cites an unresolved cited work.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:10f18d10f7e8b0d4ecf2face3bacb4cb5d2fc13fc06407836a38e3f2afec0ced

Observation 0a955f9e-e5e8-4e82-b11b-0da16b4b8d60 · outbound

This paper cites Therefore, ∠ABC=∠ADC= 26 ◦.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Therefore, ∠ABC=∠ADC= 26 ◦

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:0475c23edf847871a3b8d3852011b19cdc3f17e31a59a4e9d8490e48dc8922a3

Observation a31642fc-8351-49b8-8ec2-e9d16d10441f · outbound

This paper cites an unresolved cited work.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:8b59a3aa19556629dd63584404f52005be3b5eb1927ab6b386fad2d336840ff5

Observation 9b3f8e6a-3e6c-4ce4-9b1c-5a9b46fbcbfd · outbound

This paper cites an unresolved cited work.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:f15b753207db84deae539a84b61c1ac74a39f56d7be2763ecbe4c1a03dbb5570

Observation f08289dc-3c55-46fb-bef3-dc262ad11724 · outbound

This paper cites an unresolved cited work.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:6b6be5f814b035de4f94ec01a7a7fdf5cf97eb5f6f9ceebceceff22d4c8d75a4

Observation 00d98220-b5bd-4ccb-bdaa-558d7ce28259 · outbound

This paper cites an unresolved cited work.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:d269acf152e33422183b626cb869188136392315e7bf6fb5f42e9aaa6848e170

Observation 36620014-1c30-4a78-b4b5-48aefdc8c169 · outbound

This paper cites Let’s examine the rotations step by step: - The first shape rotates to form the second shape.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Let’s examine the rotations step by step: - The first shape rotates to form the second shape

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:6efa2dc8f9164a4c03fcd689bac74c939702c60c8a02563e8329619cf2bd09e6

Observation 511f6055-4e56-4659-be39-09852d307960 · outbound

This paper cites an unresolved cited work.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:21a88812a3db995175e328d207b74fdb71076eadd726ec42d97bf3d8b6f1bcb8

Observation a13b746c-a546-4f68-8580-76ad8964b637 · outbound

This paper cites an unresolved cited work.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:95ac2744913c20779274eb76844d02163eb89e1461bf65bc8bf79c51f457bb1c

Observation c3a7cf6a-a452-4694-ae70-25335b2d2635 · outbound

This paper cites an unresolved cited work.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:2280d0b53a9a4eb263c42d93b2b8b78ffd8a4bb1ed020063e9fa8a03c8f41c34

Observation 216e310f-1ee7-4f66-962d-e8e0ee09e1ea · outbound

This paper cites Let’s rotate the third shape (the shape at the bottom of the given sequence) 90 de- grees clockwise: - The third shape is a L-shaped configuration of cubes.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Let’s rotate the third shape (the shape at the bottom of the given sequence) 90 de- grees clockwise: - The third shape is a L-shaped configuration of cubes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:0b3e0d13794f926c229b8e5cc8a6ecdd030efc5418d07ad67327f48d6f488d2c

Observation de63150c-40ff-450f-afbd-5d3f2e4efa9c · outbound

This paper cites A color with high saturation is a pure hue, while a color with low saturation is a light, grayed-out version of that hue (like a pastel color).

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping A color with high saturation is a pure hue, while a color with low saturation is a light, grayed-out version of that hue (like a pastel color)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:dc905bbb663d9ff7c9abffaed3072c9b0cc5c7968241164ad0c0860e227d24c6

Observation 88bf90f1-c38a-4c48-aed0-6ad953236d0d · outbound

This paper cites The saturation decreases as you move inward from the outer edge towards the center of the circle.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping The saturation decreases as you move inward from the outer edge towards the center of the circle

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:8eb15cfc0ed91e981a3ddba5c362d76a7c9fee123f322f14f1bd6de731ab6175

Observation ff209a59-7a26-464c-96d4-77e65db45a3b · outbound

This paper cites - Color B is located in the middle of the circle, closer to the center.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping - Color B is located in the middle of the circle, closer to the center

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:7eed4180ca3d7e7fef4b3512d7a2c6034185a0de2c055c433545e5abd1b498b4

Observation 0b779a91-4f80-43c4-983d-70d6edfc73a0 · outbound

This paper cites Limitations.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Limitations

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:9856828d3dd8d476b096ebb867aaddfe5cb1b2ca73fc6559a49b6bac1748af6b

Observation 8174ccb4-08ae-4280-9b07-cf35412d16b9 · outbound

This paper cites Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects.

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T18:25:05.437142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:25:05.437142Z digest=sha256:43143494777262a7b5094ff4bd6c705899247ff04ca14914bcdd318f27304772

Pith citing papers

Observation 4e9af9d2-bf31-4dcb-87b6-d3754817e9f1 · inbound

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning cites this paper.

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T03:20:41.136332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:20:41.136332Z digest=sha256:c22806a07b782ed42370c45e70dfbb535b95a83d17334068b0b0f79074284a15