Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T18:25:05.437142Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2606.08708.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T18:25:05.437142Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T03:20:41.136332Z
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0c59fd0c-4c60-4a4a-905a-157bd73d63ca · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping The original policy distribution is P=π θ(·|I, c), and the perturbed distribution is P ′ =π θ(·| ˜I, c), where ˜I=P(I)
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3abfb000-4ab2-435d-b3fb-e975f3ce46de · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping While St captures semantic dependency, it cannot distinguish whether a high KL value arises from robust visual grounding or brittle numerical over-sensitivity
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12474a78-fd79-4949-b646-1ee88d3ddc20 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping -∠ADC= 26 ◦
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79a7781e-b404-49ff-b5d6-707eec5c1352 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Therefore,∠ACD= 2×angle at the center= 2×∠AOD
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f1b899-3603-454d-b645-9352b02b34fc · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping - Solving for∠CAB, we get∠CAB= 90 ◦ −52 ◦ = 38 ◦
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80a4ed9-3939-4000-bc88-e53dd59e7b80 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a955f9e-e5e8-4e82-b11b-0da16b4b8d60 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Therefore, ∠ABC=∠ADC= 26 ◦
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31642fc-8351-49b8-8ec2-e9d16d10441f · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3f8e6a-3e6c-4ce4-9b1c-5a9b46fbcbfd · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f08289dc-3c55-46fb-bef3-dc262ad11724 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d98220-b5bd-4ccb-bdaa-558d7ce28259 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36620014-1c30-4a78-b4b5-48aefdc8c169 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Let’s examine the rotations step by step: - The first shape rotates to form the second shape
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 511f6055-4e56-4659-be39-09852d307960 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a13b746c-a546-4f68-8580-76ad8964b637 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3a7cf6a-a452-4694-ae70-25335b2d2635 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 216e310f-1ee7-4f66-962d-e8e0ee09e1ea · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Let’s rotate the third shape (the shape at the bottom of the given sequence) 90 de- grees clockwise: - The third shape is a L-shaped configuration of cubes
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de63150c-40ff-450f-afbd-5d3f2e4efa9c · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping A color with high saturation is a pure hue, while a color with low saturation is a light, grayed-out version of that hue (like a pastel color)
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88bf90f1-c38a-4c48-aed0-6ad953236d0d · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping The saturation decreases as you move inward from the outer edge towards the center of the circle
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff209a59-7a26-464c-96d4-77e65db45a3b · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping - Color B is located in the middle of the circle, closer to the center
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b779a91-4f80-43c4-983d-70d6edfc73a0 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Limitations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8174ccb4-08ae-4280-9b07-cf35412d16b9 · outbound
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9af9d2-bf31-4dcb-87b6-d3754817e9f1 · inbound
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.