Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.19595.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:31:02.541976Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:49:46.153282Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 7adeef12-ffdc-4e0e-8f99-e1f659d5405b · inbound
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5a03000-ddb6-4db8-a58a-639361ea3c9c · inbound
Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc56e706-31a6-4160-9c14-1dbfeae88be1 · inbound
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e997e85-8055-440b-a0ec-1515799f464c · inbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d78bad3-7a1d-4923-8247-095ae001384d · inbound
Outcome-based Exploration for LLM Reasoning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d88b1852-3a90-4dcb-bae8-f1db68cd280a · inbound
Polychromic Objectives for Reinforcement Learning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a0054f3-e5f2-4003-b4dd-a3e9efec7692 · inbound
Compute Aligned Training: Optimizing for Test Time Inference Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05744ee3-eb1e-4fbe-a30f-504c9100dca0 · inbound
Compute Aligned Training: Optimizing for Test Time Inference Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be2fa0bb-81f2-437a-a626-b5ae1935083e · inbound
What should post-training optimize? A test-time scaling law perspective Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bdf7ce1-4a7e-44ec-a1fb-7202912bab31 · inbound
Finite-Time Regret Analysis of Retry-Aware Bandits Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21aa595f-745b-4e8d-88c1-dd996626182f · inbound
REVES: REvision and VErification--Augmented Training for Test-Time Scaling Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 448581c1-4c99-47da-9b30-d6f176b30e11 · inbound
SPIRAL: Learning to Search and Aggregate Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b568f87-3394-4b10-8ef4-3182ee70016c · inbound
Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77e2b571-845e-4398-8209-3bd9b587979a · inbound
DecompRL: Solving Harder Problems by Learning Modular Code Generation Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.