Pith. sign in

Paper Citation Record · LEDGER

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently

As of 5 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2606.22938.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.22938 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T09:08:13.233840Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fa85c59-7dd1-4ea6-b22c-6d09d0269ebf · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:09:44.765381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:a5237e4b80a4548e364ef79678234eb2a1daf739e5f6ec4c0e79b86ff6f414b7

Observation 92c70472-23b5-4c9f-a708-67a086483e8f · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:09:44.762957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:e31be08b7b4767c9d8fb44d211430ab00b3534db3dc66839f7aeec2df19badc2

Observation 459b3dad-a283-4b71-b0a6-6dfa33a41460 · outbound

This paper cites an unresolved cited work.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:13e77e7c71e4a304c20f436ba05684a6360cd6c1b4a5a5270f27aee606c46bd4

Observation 28e3b78d-e0c0-4aa0-9595-bf9717e659e9 · outbound

This paper cites That is, ˜g(s) :=E[τf] where τf := min{t≥0 :s 0 =s,head(s t) =f}.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently That is, ˜g(s) :=E[τf] where τf := min{t≥0 :s 0 =s,head(s t) =f}

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:7e68a63db7e4a877fd584f8e7a61b9c942da048abefce0b612db986fe7f723f2

Observation 41ed3a94-9fc1-4355-af21-eb9ebaef0026 · outbound

This paper cites both states are absorbing).

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently both states are absorbing)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:28a8a03f0db73d6cd55aca853167617cdbfa0ce5060e2b1f79deedfc28f1f98b

Observation 71c59b06-81c8-4722-8296-47db34befad8 · outbound

This paper cites In particular, this means that hx(s) =g(s) +q(s)H f.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently In particular, this means that hx(s) =g(s) +q(s)H f

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:0a3a58a5a8e639a914e1f5660cbded5556dc00cfac52e5a7a5fa0efc6cd3b65c

Observation a7fd09d5-7cad-41f6-bcb6-26b1627d7afd · outbound

This paper cites In other words, µs is the expected number of visits to statesduring a target branch attempt.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently In other words, µs is the expected number of visits to statesduring a target branch attempt

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:14fa085ce031eee6b915f55df761eb3672d172288cadf4887f5011ebc25f0bc6

Observation 33be9614-87d8-48ce-8ed1-1a1d425efdac · outbound

This paper cites In other words, ˜µs is the expected number of visits to statesduring a non-target branch attempt until it goes back tof.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently In other words, ˜µs is the expected number of visits to statesduring a non-target branch attempt until it goes back tof

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:b0898722ac2e9e20b0292616b6003057d55d767498fd6fa7ebc525792c640ad3

Observation 9f2de60b-c686-4cf8-a5d3-8a50a7a1ac21 · outbound

This paper cites an unresolved cited work.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:e1fe7ed12db7bcf646e0a1b975bcc0bfc4fb3318dbb059486cb0fb95f284e3de

Observation 06bfd5a6-4855-4f81-ba41-f64501295113 · outbound

This paper cites an unresolved cited work.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:7a17fa9582239843749a755e07ee539b2b4ab02307c3234ad02df97c2cab5291

Observation 976dee28-bc8d-4e9c-bfcd-0bd1b94fec05 · outbound

This paper cites an unresolved cited work.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:7786cb5723530f12b61fea7f3a281d0c03e89f1d1ba0ce6165fc35511e99f3ff

Observation d5222742-6725-47e5-8d85-0c5c726d302f · outbound

This paper cites an unresolved cited work.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:8389a2144f79303c3f995f7de44094a8154f6644bcf82f9af4f292880fcfa7d6

Observation 19884a58-adca-48b8-b4a4-18a93de5ffd2 · outbound

This paper cites an unresolved cited work.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:bbe5f17f9ec2d5480a6854567eb3b513a54581118ffa7adfad363ee7af25f47d

Observation 99a898e6-aab4-47cf-91ca-84a7d0c4a2a0 · outbound

This paper cites reached the state with headt i), letg i denote the expected time of first entry into R− K+1−i, and fi denote the expected time of first entry into L− K+1−i.

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently reached the state with headt i), letg i denote the expected time of first entry into R− K+1−i, and fi denote the expected time of first entry into L− K+1−i

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T09:08:13.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:08:13.233840Z digest=sha256:0634810f52da61cfa77566d71bbe1b61ed4455dd9759d8b6d957bdd04ae8a861

Pith citing papers

No inbound Pith citation observations are available.