Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T06:48:41.372507Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.28457.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T06:48:41.372507Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a825b8a0-ae52-4f03-be07-a13e398d6c2e · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1866da13-62ae-419f-8316-fa105773e787 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f97217-0f6e-4d72-b4da-786b3a6a79a2 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60bf378f-42f9-4f5b-8588-235043699475 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0bab8b-2f5b-4b63-a1f9-71c082b6835c · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc1fbba-3563-430b-8356-e8d315eb3bd6 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7772c9f4-63dd-41c3-879e-f713757fdd92 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7307b0b4-e442-4238-bc34-fa070b144c78 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c29a4b29-bf6f-4119-bda5-724646d05362 · outbound
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6597c73-e57e-46a6-8bd4-6d90be0b67a2 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 547217c4-d886-4f21-bce1-a19450646a4b · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a89943fd-80d0-4c0b-a257-ea385b6cc01d · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d1fe7f-29fd-43ee-af3c-a4747163b302 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5dfbd5f-c099-4e32-b559-2fb2dc834d7a · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dd1ccbf-491f-4809-9220-d1e08955db23 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91b87691-aeb7-4b45-aa10-3717a8b52558 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99e235a5-2182-4e8c-9e4e-ab0c208aa240 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8690b733-71d7-42d6-8f5f-f3e1c0a19c52 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 931698e5-3d3b-4bf7-a4f7-17ee223f0808 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36908b5e-647a-435f-935f-0a0d4693b699 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15dc2459-51eb-4145-8da7-549ff3b58b24 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0642e10a-d52f-46ca-8c09-f5f437884a98 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1361378-cf90-4eeb-a66b-945b5aae723e · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a29516-b769-4f04-b7b8-94af82f3eb99 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5599348-6e7e-4a5c-9cc8-0105cdfacdb6 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669cfb3c-f3aa-4e79-8ee6-c4d087d527ad · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a14ef52f-9d3d-4df4-bf9c-759e86481854 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdca9a63-aad3-4285-995c-87118c8e3c97 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Process Supervision of Confidence Margin for Calibrated LLM Reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365bfda1-39b9-4cda-9048-a94ec618026c · outbound
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5beffa7d-7656-4943-b9ae-fb73df30609f · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8cb41a4-2e09-4d63-aa68-afd154732010 · outbound
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f15e6020-67ec-470c-be4e-24586c11e5b9 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e09cf6b-017e-4025-93c8-2c079fc9a959 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5f91e8-3d50-4bb6-9a3a-b23ec8d8f839 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaed60f1-e0d4-4c17-9a79-efd6bcacb137 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75c8d8c-e57a-4643-b252-50faf96d5b89 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b431275-1549-4c2e-bd65-bde46ca7c758 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c16beda9-ad19-464a-9e24-01e347fc532c · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Group Sequence Policy Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6bb76ba-9a96-425b-baf0-a5a718f980af · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f988685-4a34-4678-969f-324a920868e9 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute The maximum completion lengths used at evaluation are 800tokens for Countdown,1200for GSM8K,2048for MATH500 and AMC23,3072for AIME26 and MinervaMath, and4096for Olympiad- Bench
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ca0fa7-283b-436c-8bdf-7354f5bc85a7 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Adaptive inference uses greedy decoding with𝛾= 0.85and 𝑇max = 10, so the observed variation primarily reflects training stochasticity rather than decoding randomness
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0fb69d2-8dad-4010-a0bf-7464807a44b1 · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Language Models (Mostly) Know What They Know
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd618804-d1ea-4a5a-a40e-50c5451a5c1d · outbound
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Efficient PRM Training Data Synthesis via Formal Verification
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.