Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T09:08:13.233840Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2606.22938.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T09:08:13.233840Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9fa85c59-7dd1-4ea6-b22c-6d09d0269ebf · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 92c70472-23b5-4c9f-a708-67a086483e8f · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 459b3dad-a283-4b71-b0a6-6dfa33a41460 · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e3b78d-e0c0-4aa0-9595-bf9717e659e9 · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently That is, ˜g(s) :=E[τf] where τf := min{t≥0 :s 0 =s,head(s t) =f}
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41ed3a94-9fc1-4355-af21-eb9ebaef0026 · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently both states are absorbing)
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71c59b06-81c8-4722-8296-47db34befad8 · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently In particular, this means that hx(s) =g(s) +q(s)H f
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7fd09d5-7cad-41f6-bcb6-26b1627d7afd · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently In other words, µs is the expected number of visits to statesduring a target branch attempt
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33be9614-87d8-48ce-8ed1-1a1d425efdac · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently In other words, ˜µs is the expected number of visits to statesduring a non-target branch attempt until it goes back tof
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f2de60b-c686-4cf8-a5d3-8a50a7a1ac21 · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06bfd5a6-4855-4f81-ba41-f64501295113 · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 976dee28-bc8d-4e9c-bfcd-0bd1b94fec05 · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5222742-6725-47e5-8d85-0c5c726d302f · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19884a58-adca-48b8-b4a4-18a93de5ffd2 · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99a898e6-aab4-47cf-91ca-84a7d0c4a2a0 · outbound
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently reached the state with headt i), letg i denote the expected time of first entry into R− K+1−i, and fi denote the expected time of first entry into L− K+1−i
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.