Pith. sign in

Paper Citation Record · LEDGER

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.15512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15512 v3

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:37:24.974376Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved52
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3fb0fe32-d866-4d42-a53a-1af59abac6f5 · outbound

This paper cites online" 'onlinestring :=.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.848566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.848566Z digest=sha256:acb11b06c2714fbfdb1998a0237fe02eb827b27e566df9c5b2ed66346284c566

Observation 1e472068-85f0-41fa-85f2-924d6c0ea33c · outbound

This paper cites write newline.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.853986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.853986Z digest=sha256:986e8fe62e0558701f529bc1a091c49e451a8eb299724d5c664ed1e60e81f9c2

Observation 06ed4b1f-d4ed-4cc3-b07b-41f36c918a98 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.856899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.856899Z digest=sha256:c015b7ed91b1a31449178c683dffc2685b82486a4b39c59ca4b2ea753fc4722c

Observation 00d2d16e-e232-4f97-b42b-9ff633586a45 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.860769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.860769Z digest=sha256:4dfbf5d85029adfbf6259a6c419d6e25238eebb75d1ef6f8cd759a16ba2e95bd

Observation 0888cbf8-91aa-4517-a4cf-70e55c111605 · outbound

This paper cites Efficient Prompting Methods for Large Language Models: A Survey.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Efficient Prompting Methods for Large Language Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.863525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.863525Z digest=sha256:df638afac0abb3e79752eb3c50bfe30c18d3de8eaef818d7dc056205841bbcc1

Observation fec87c8d-ba97-4723-9c28-e1dca317d7b0 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.865956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.865956Z digest=sha256:a1fa19eb12f8dd0bb5b4e6b41ec5087ff07fb6bb51d13af8e2514051925c542a

Observation 387a8b27-0780-4018-8b0b-d7d4ae9d039c · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.869021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.869021Z digest=sha256:fff07ff856c7f0f6afa17db8405793de77e448b4849cbd15c0ea32425bb3dd7b

Observation cf468863-344c-42fa-8c3d-fbdcf602386d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.872325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.872325Z digest=sha256:6f731faaf0e022284089c50f0773410be28362b41213e274bf6f7b5ba32185b6

Observation bf547cc7-471a-431c-a0fa-dd6c01112d2b · outbound

This paper cites Dynamic Parallel Tree Search for Efficient LLM Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Dynamic Parallel Tree Search for Efficient LLM Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.874870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.874870Z digest=sha256:f4897cbb9b9594afb120dfa834073da47ee6feda80656a72f8b572fef250adff

Observation b919a394-ae99-48ee-a1d0-fda14ca763de · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.287141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.877196Z digest=sha256:4ce5d7c6548987ec27395638fe942f1ddea2f253125ff78a11654f4871cde16b

Observation 79fffff3-dc27-4602-9fb6-b8e1030a7516 · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.879489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.879489Z digest=sha256:267de5d73564dea557522ae7cc3b46b363ca851dbe62c5e42997228c7a9e6c72

Observation 8794fd29-cfe3-42de-8b6e-43d1c798a70c · outbound

This paper cites The Llama 3 Herd of Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.882178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.882178Z digest=sha256:122275466af56e3c9da52518da1349de883920758d30b42f6197fc194db88850

Observation fd5cbf11-bbb8-4adf-9c00-05a41dab1f82 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.884204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.884204Z digest=sha256:edfb9ff26bd726120b40ea654be7934fd8db4d4c40993c0799ed06431d423e83

Observation 270f58dd-1586-44d9-b353-99dc74c54396 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.280708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.886512Z digest=sha256:0b70d2abaf1c8d1ac8a3a3882cce3d66f4a6b1037097b003e8652f60ee3b957a

Observation 32281b08-23e6-4e8c-8a35-9f3f0d922fae · outbound

This paper cites A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.888836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.888836Z digest=sha256:bb652eef079dbf4f7bd454329ca7be6d6045d2c9de7d0ade9234e5c1c690fac5

Observation 1f219167-10ac-4d6a-82cf-37dfdece25e0 · outbound

This paper cites MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.891071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.891071Z digest=sha256:c62c0750686f427e1b2035a04d8e4a9a94694e7cefd03ed3c39f731dce3e2fc1

Observation 122bc742-36ce-4876-a508-ca2ceaf498f3 · outbound

This paper cites Scaling Laws for Neural Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.893285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.893285Z digest=sha256:ca71afbaa5058ea90a5d286cebf37f69d5a80d0ebd75ee8ce7f942e19bf650a9

Observation cd30e8d6-b71f-4d93-9b37-cce2ecf9e573 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.895591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.895591Z digest=sha256:da0e14d8e312c2ff8d282bc487e26cc5d62b05c1e9c65e9a2b194bdc5368b137

Observation 25fa7974-b8ca-4780-b4e5-9fc7d7d2fc39 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.897708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.897708Z digest=sha256:4ee22b711618a7a57eb479ab61ce84501d0fe5592db764f52ec7578682b32951

Observation 93694884-1d37-48aa-9d6c-fe9a1fb6f47f · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.900251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.900251Z digest=sha256:386cf7243f8a83dd74da4c191dad20da733f44ac2f8c761eb066eb7fa0ab0661

Observation 94a41992-67de-4659-88cf-986d501d1dac · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.266832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.902306Z digest=sha256:905493c0b63fe506d1cee91462dac7eab52d502c873e17eb5367097b6338bf7b

Observation 7d89afc0-0d57-4e87-8861-e1e16611743d · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.259854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.904368Z digest=sha256:b7a821a4d85fb4cb73b40e1c1dc3fd84262b86d4a90ff815966e4adfe4fbc75f

Observation 0dc4419a-3029-4923-8ceb-d323da4ce8fe · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.253476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.906307Z digest=sha256:85f8a3700210ca2fdbf53079e451fbb159774aa27247cbc190d48314e59c176b

Observation 8f289958-4208-4114-9432-419e4611a362 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.246885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.908421Z digest=sha256:7c87b14a9ae34fb076130f45a6479a5b3c07e5eff9500a80613c7d284fc2167a

Observation af9f2b31-e30b-4a7a-8262-590c58eb3025 · outbound

This paper cites s1: Simple test-time scaling.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.910293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.910293Z digest=sha256:0ffe3cffe88b9e81037b5e2a7a3e4e19cc956c2bda1ea7cf5837594e58959615

Observation de22a193-3e29-4776-8967-80c4c52146bb · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.240689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.912448Z digest=sha256:2ca906de88f8fb755b614692011608bd1a370ce97fa456dd634c309efa325ef7

Observation 54efe568-fdcf-4759-b2bb-ba8298b473d7 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.233762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.914826Z digest=sha256:a3e6473eac549dc6a1950a77c5fb92d1f761a0141959292dc7b9024ffd1631e3

Observation 3d9c8bdf-a42b-4db3-b90f-a95684e72fd2 · outbound

This paper cites MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.916693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.916693Z digest=sha256:6203b0cbe680f59c3f4bfdec58df0ab1ec80e790e02858308d4dc33de491ac12

Observation ded48cee-0d04-43ff-beca-400b55807270 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.919131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.919131Z digest=sha256:72668a3e161430ac357207fa032c718cef446cdaba8018036359da864b66461a

Observation 8fc70004-692f-4544-9468-1df91cef1faf · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.223708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.921365Z digest=sha256:6177149330223ab32ecd879f39f4cfc5be50a5fe69cc67884a5e7a6c3d634683

Observation 56eda410-0fa7-473f-936f-6e655336c53b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.923893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.923893Z digest=sha256:3cba94fb6622e4487095424c6658aa6e5a53e8e6808c1e427a3b097e7c0c1b92

Observation 255745b6-cfa4-4de7-8ab4-231be8d982ba · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.925962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.925962Z digest=sha256:e090e8589486fb72a71ba4c5822fa97151a7309b3cf08d7cdbc7acf36dec5511

Observation cf571c63-c71a-4caa-946a-eea22154689b · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.928392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.928392Z digest=sha256:d05db23c06d939910424313ca04a7b08cc27ec6c52e241fd97b12533eef3f6bb

Observation 31c72ada-d685-449e-8af8-0f8c0db39c97 · outbound

This paper cites Gemma 3 Technical Report.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Gemma 3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.930558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.930558Z digest=sha256:604ddcf8f1bf026ce19d1437b3e2fea7e84608fce65e56d2dbd60b6c457d3153

Observation 942872f9-e760-4cfb-8a5f-5c709f276273 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Solving math word problems with process- and outcome-based feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.932824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.932824Z digest=sha256:31dd4d6907f5149e107f33143adc4cada43384ed0eec61c5ec6405d36a4baae6

Observation fe7517a8-eff5-442b-81c5-822196e81a48 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.217484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.934824Z digest=sha256:28afc8f8e0bfdc9bc93700fd167a52682543c8f3a46d1497dae3dcd72267bc0f

Observation 3ea2de8a-6282-446c-b3c8-dcb133f67084 · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:37:24.936924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.936924Z digest=sha256:46dfc57976e6122b08b057b85e85f929eb666f5dc310a4f717bcbb3d0926ad06

Observation c1d1a203-5b25-4b88-83ee-0776460c306d · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.939339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.939339Z digest=sha256:159796f1acfbb4ace9c8a8b9edcedd28d550d465bf0407c51e921bb403a6bf28

Observation 15a8f3ad-d36f-491d-b35c-69e2c695b8dd · outbound

This paper cites Le, Ed H.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Le, Ed H

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:37:25.211239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.941418Z digest=sha256:3efb95fd41fcc4050f542f35e5bc535ba9a629d5b81d020a3dbf55aa05954867

Observation 01f4e93f-2cbf-4427-8840-5990a60af7c4 · outbound

This paper cites Chi, Quoc V.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Chi, Quoc V

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.943275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.943275Z digest=sha256:a20b01026a777a3142f16b5a61f6ef457bad3b30e43acb51b8265cbc08d1bc1a

Observation 079a00e7-b7f0-40a1-8453-95ada4905eb1 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.945481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.945481Z digest=sha256:2847b9e18997114da1ea46ef34b719174c3a51d8da7e6dae5c99e85f6956c06c

Observation 888b3b45-63a5-4bf8-a25e-c57afb5d0ea5 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.200649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.947418Z digest=sha256:5a9ee30e726af03e481869fce6cd4390a0f7fdaaf8c16458c7281f265674a743

Observation 47daccca-fc97-4a3f-a989-a7a8b31a624a · outbound

This paper cites Foundations of Large Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Foundations of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.949552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.949552Z digest=sha256:e8f25e608fca7254fa9e89d837110871fec831a98cba119aab98355bb00f0a44

Observation 663df73f-68df-4176-9922-391a9287f81a · outbound

This paper cites Qwen2.5 Technical Report.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Qwen2.5 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.951647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.951647Z digest=sha256:51ae5692da6ba7c2916864f12f09ba3ae0d218615ba36b57310a87058ad7fdb7

Observation a27e4034-1814-4d89-afa5-11b468fa5a14 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.193983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.953990Z digest=sha256:db3b9c60e3fa8962ba615dd583b6ea70320f66f16817e2d267c54ec45ff821c6

Observation 40a8dd28-cd25-4b90-86d0-9a9856adc807 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.955844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.955844Z digest=sha256:ef97a7fca7ee274c953bb4b109f9881882ab2f746caacb4df3d0d94a2a85f258

Observation a1ecf1ba-7fac-4019-a561-0bda22fb1dce · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.957990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.957990Z digest=sha256:157d795504588451611389db8b45995374543ad95fc62e51f6d8c640a90b259c

Observation 98a9c10f-eba1-46ef-b98e-3ff54639620e · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.187756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.960189Z digest=sha256:8c64c7b95b35171329d51cd4c87880dd7761c752ca5b43cc0284cd736e92457f

Observation 1bfd7f65-1aeb-45a8-a77f-4127ea42aed0 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.181254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.962130Z digest=sha256:357b412dcbfd5e1edecded142b112b63aebde240f71ae429213c431c737d790a

Observation 109efe03-ab17-4f2a-b832-038edf186e46 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.964766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.964766Z digest=sha256:f6c1c4050092b76a19d86adb90f6cd7042da52b301c26ee1d6f75ec041aeaf60

Observation 8f21145a-2da8-4827-b183-28c5bb86f8d7 · outbound

This paper cites Small Language Models Need Strong Verifiers to Self-Correct Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Small Language Models Need Strong Verifiers to Self-Correct Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.967055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.967055Z digest=sha256:c19dead9d85d805eb9a389a195e2677b830390f012db1300e1fdc6cfc008191c

Observation d07800da-c014-48f5-a681-d6d9f3143dd2 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.969856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.969856Z digest=sha256:3e79767cbce7e39b0cf6bfee7d9058ecf699c35b6f46658fafe46862e127c272

Observation afad08a3-61b1-4782-8c08-1af9a7c37026 · outbound

This paper cites Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.972019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.972019Z digest=sha256:6551fe8b898e84b99d5409ede1400be779e5c729a82ece56f00590ee6e80fd1e

Observation cb2dd002-b98d-4d15-a457-1b9382094597 · outbound

This paper cites Learning to Reason via Mixture-of-Thought for Logical Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Learning to Reason via Mixture-of-Thought for Logical Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.974376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.974376Z digest=sha256:e13e5a356a67855b14ccf0df25f75834c5ba21b6c2026f7c7259e7fb06c6233d

Pith citing papers

No inbound Pith citation observations are available.