Pith. sign in

Paper Citation Record · LEDGER

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models

As of 20 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.15512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15512 v3

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:37:24.974376Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved52
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3fb0fe32-d866-4d42-a53a-1af59abac6f5 · outbound

This paper cites online" 'onlinestring :=.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.848566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.848566Z digest=sha256:01bd11f55d0f5654deda05298ba068c45ceb5908a80900e1b67b99150386bf91

Observation 1e472068-85f0-41fa-85f2-924d6c0ea33c · outbound

This paper cites write newline.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.853986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.853986Z digest=sha256:da2291726b6b4a47d1f71669214d9eeb90d92eed264cbc358db352133b9f3ebf

Observation 06ed4b1f-d4ed-4cc3-b07b-41f36c918a98 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.856899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.856899Z digest=sha256:e3f471a044a5f69915964092783fc623b6fa7565f9e0cdf4270366ffc5254935

Observation 00d2d16e-e232-4f97-b42b-9ff633586a45 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.860769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.860769Z digest=sha256:cc7904d24a9ecb6c61a0f3be5bdd63f1b5642d1df706fa3855d029a801626ec3

Observation 0888cbf8-91aa-4517-a4cf-70e55c111605 · outbound

This paper cites Efficient Prompting Methods for Large Language Models: A Survey.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Efficient Prompting Methods for Large Language Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.863525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.863525Z digest=sha256:38f55864cb92b16a5f66f80afac40794b2e12466da44f85ac63962681d154a9b

Observation fec87c8d-ba97-4723-9c28-e1dca317d7b0 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.865956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.865956Z digest=sha256:3768ecd7b33a4878ded1878a974ad7cbceeac466a16d335ffd39f506d96498d5

Observation 387a8b27-0780-4018-8b0b-d7d4ae9d039c · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.869021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.869021Z digest=sha256:8954a2b3dbfc406e2078cb02ad4a12f0ec93ede98e8bbcf3c949384eb8b7f60b

Observation cf468863-344c-42fa-8c3d-fbdcf602386d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.872325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.872325Z digest=sha256:ccdaa3a0c0240d1a9755448dc50f0f1b56fcf231dca9a355b929193dadaf8810

Observation bf547cc7-471a-431c-a0fa-dd6c01112d2b · outbound

This paper cites Dynamic Parallel Tree Search for Efficient LLM Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Dynamic Parallel Tree Search for Efficient LLM Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.874870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.874870Z digest=sha256:e10360405d079f1c0f7cc6d121a728df01529363a8f188016bdc64b546183f4d

Observation b919a394-ae99-48ee-a1d0-fda14ca763de · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.287141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.877196Z digest=sha256:cf5220cd6baa6b349c76ab3d53d48fc5ffdd715f86704052355daef8fcc6254c

Observation 79fffff3-dc27-4602-9fb6-b8e1030a7516 · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.879489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.879489Z digest=sha256:22fc45e58e94a40ce8b32698d89632973b19609bb024ac40b1854c0a42c68871

Observation 8794fd29-cfe3-42de-8b6e-43d1c798a70c · outbound

This paper cites The Llama 3 Herd of Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.882178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.882178Z digest=sha256:adc8cec9117b5cb5812008897773f611c9485e59dd19c5f1c0726bf2ef670240

Observation fd5cbf11-bbb8-4adf-9c00-05a41dab1f82 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.884204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.884204Z digest=sha256:45c5bea4674f561cc50dc7f3db2bc861296d7074f76984705a2ec2efcb62de11

Observation 270f58dd-1586-44d9-b353-99dc74c54396 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.280708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.886512Z digest=sha256:793e9035fd9e99b9a473722a0cee1f404830d783f9368ebb07d1eaa9053b805f

Observation 32281b08-23e6-4e8c-8a35-9f3f0d922fae · outbound

This paper cites A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.888836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.888836Z digest=sha256:327d8f490cb068b54c8029bc20191b2beca0e5f5656fbc3154e411568d465bca

Observation 1f219167-10ac-4d6a-82cf-37dfdece25e0 · outbound

This paper cites MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.891071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.891071Z digest=sha256:504b661624231a54e4f02bb4fac2649455ac688617d12293bb9fe2902da3a4d4

Observation 122bc742-36ce-4876-a508-ca2ceaf498f3 · outbound

This paper cites Scaling Laws for Neural Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.893285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.893285Z digest=sha256:604d4b8e9692b01f4cef7adc6e2bc2c6dfda5756ef7593a9cdf7bf6f61a8d845

Observation cd30e8d6-b71f-4d93-9b37-cce2ecf9e573 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.895591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.895591Z digest=sha256:3a951a4d0aa54ce161470ba1350e876215f171cfb7be99eddb871b546c832129

Observation 25fa7974-b8ca-4780-b4e5-9fc7d7d2fc39 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.897708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.897708Z digest=sha256:604cdbb36a2d54e6e20d4700a5bcb62a89afc6d97c158f295f1de05d79d02d91

Observation 93694884-1d37-48aa-9d6c-fe9a1fb6f47f · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.900251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.900251Z digest=sha256:cf56743f5180815add579e306c75a1beb2d6e6384bedbfaed4921c1083210be4

Observation 94a41992-67de-4659-88cf-986d501d1dac · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.266832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.902306Z digest=sha256:ea25f66cc65e1042f29105ac345943468805312810feb9dce3b074ca834f9e60

Observation 7d89afc0-0d57-4e87-8861-e1e16611743d · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.259854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.904368Z digest=sha256:078fb457099cf4952534d09f999a5f7212f8e7fe471485a071be3d1f3bcb867c

Observation 0dc4419a-3029-4923-8ceb-d323da4ce8fe · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.253476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.906307Z digest=sha256:9b0ccbf82cca0e3e0927f306f9884e9cef80423ddc625c3cf05003893399e1f0

Observation 8f289958-4208-4114-9432-419e4611a362 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.246885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.908421Z digest=sha256:5a25d8743569f8226b4d2e0ea24d45cd7dbb42385f3c4558715dcfafe05bc6f2

Observation af9f2b31-e30b-4a7a-8262-590c58eb3025 · outbound

This paper cites s1: Simple test-time scaling.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.910293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.910293Z digest=sha256:ae77c10734eb378a6136e617b767096e7d9fb8dad0746b22530517215e7b87be

Observation de22a193-3e29-4776-8967-80c4c52146bb · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.240689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.912448Z digest=sha256:b9e7bb8112c4edbca3af7b9b413eec7051e83f010912570502d58cb75e996a84

Observation 54efe568-fdcf-4759-b2bb-ba8298b473d7 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.233762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.914826Z digest=sha256:69cda977a9df56b5998cb7414ee29af2e19ff401de4dc69b389a915a46957e26

Observation 3d9c8bdf-a42b-4db3-b90f-a95684e72fd2 · outbound

This paper cites MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.916693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.916693Z digest=sha256:1f30522b60de5aaeb6b386740851cd2a337f0e4d08d46076a6839e6e86d3a84a

Observation ded48cee-0d04-43ff-beca-400b55807270 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.919131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.919131Z digest=sha256:9b5cbcaff6c98699acdbb68788b0c728f96a2b4a0f850c49a8e0cc45c6d75d1b

Observation 8fc70004-692f-4544-9468-1df91cef1faf · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.223708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.921365Z digest=sha256:60941f36ba17cf567251db4596136f4988835073fa7abe56571cf7e565cf142a

Observation 56eda410-0fa7-473f-936f-6e655336c53b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.923893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.923893Z digest=sha256:b9c98085a3b3135acab5fe3a95bda0164f3ead32988e7d7eca27fb5b8ec50d69

Observation 255745b6-cfa4-4de7-8ab4-231be8d982ba · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.925962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.925962Z digest=sha256:459694679712e3b36e8bb074fb7d3414bc7384a743e03d3af82d42d44a93f660

Observation cf571c63-c71a-4caa-946a-eea22154689b · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.928392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.928392Z digest=sha256:d7135595b26aac5444a4325e2dad1bbd76c2e63be4fcfbef8d8a9abe932c6e85

Observation 31c72ada-d685-449e-8af8-0f8c0db39c97 · outbound

This paper cites Gemma 3 Technical Report.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Gemma 3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.930558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.930558Z digest=sha256:d4719846861e09a4403908c018a93a1270a08576957f8bccd32731658a0c4f75

Observation 942872f9-e760-4cfb-8a5f-5c709f276273 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Solving math word problems with process- and outcome-based feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.932824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.932824Z digest=sha256:69d4d7ead29fac92132b60905b870314ce0f5bff43b3f010ca336ccab847f31d

Observation fe7517a8-eff5-442b-81c5-822196e81a48 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.217484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.934824Z digest=sha256:e2641fc4e4fee0103488a93d86b7ed984ea7edbd75b7a31045e9a1c1254a14a3

Observation 3ea2de8a-6282-446c-b3c8-dcb133f67084 · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:37:24.936924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.936924Z digest=sha256:75b83a0ba20544b6465a4cfd42500d39542a0a021914ea736cbc815bae31724d

Observation c1d1a203-5b25-4b88-83ee-0776460c306d · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.939339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.939339Z digest=sha256:dece5defc46ca54881e3b5f326f94274ab10582832161619755a7f092ad48155

Observation 15a8f3ad-d36f-491d-b35c-69e2c695b8dd · outbound

This paper cites Le, Ed H.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Le, Ed H

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:37:25.211239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.941418Z digest=sha256:d7e08df6fd8a9ea829930fecaf6e8b222b32f891c3d68092a178c2db23fa3239

Observation 01f4e93f-2cbf-4427-8840-5990a60af7c4 · outbound

This paper cites Chi, Quoc V.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Chi, Quoc V

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.943275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.943275Z digest=sha256:0b6866eeadaf5ecf97183b813f75cb0703891eaebb649358573ee3d0a07af82b

Observation 079a00e7-b7f0-40a1-8453-95ada4905eb1 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.945481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.945481Z digest=sha256:b158bac2cdc5676beea32d269e3baaed981e3d61e4533132b68cac445af7162f

Observation 888b3b45-63a5-4bf8-a25e-c57afb5d0ea5 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.200649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.947418Z digest=sha256:b482dc33622b731053c43a53d62cf9eafce5cd0f1bd1eabcae10757099683fe2

Observation 47daccca-fc97-4a3f-a989-a7a8b31a624a · outbound

This paper cites Foundations of Large Language Models.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Foundations of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.949552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.949552Z digest=sha256:b2670d088a3baf32bc07962711763a55ea89685b5a581843c3cde1b90df66d48

Observation 663df73f-68df-4176-9922-391a9287f81a · outbound

This paper cites Qwen2.5 Technical Report.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Qwen2.5 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.951647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.951647Z digest=sha256:1964f5668e92d6154ff421ac0c5c6a6554ee7f8aa7970542e71ae2cf0b520016

Observation a27e4034-1814-4d89-afa5-11b468fa5a14 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.193983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.953990Z digest=sha256:4a99f8ac19309bb517babcaff2beefe448d675981f633f37df657c1e5666bfdf

Observation 40a8dd28-cd25-4b90-86d0-9a9856adc807 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.955844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.955844Z digest=sha256:3c8f7616b6730c362873d6112f91310ba797fcce38ebca69c695e03cbd8db55f

Observation a1ecf1ba-7fac-4019-a561-0bda22fb1dce · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.957990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.957990Z digest=sha256:33e9ead9f35cddd008cf1b6e2f1c85b01524a3a81ba20eaa59d2a0cb6723d2f4

Observation 98a9c10f-eba1-46ef-b98e-3ff54639620e · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.187756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.960189Z digest=sha256:1c9bc160c7d867eb370e28db327ca162b2bab4e653a3541fb41348062d53dd7d

Observation 1bfd7f65-1aeb-45a8-a77f-4127ea42aed0 · outbound

This paper cites an unresolved cited work.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:37:25.181254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T15:37:24.962130Z digest=sha256:2b6680414ef9c4f063b19c9df536a6f85d65a96e8c4dfff30d3342b3687171c2

Observation 109efe03-ab17-4f2a-b832-038edf186e46 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.964766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.964766Z digest=sha256:d3dacca81b1e0e1f18e6e04fb3f570e6571bbdb1415bc92c9104e1e6b1a2637c

Observation 8f21145a-2da8-4827-b183-28c5bb86f8d7 · outbound

This paper cites Small Language Models Need Strong Verifiers to Self-Correct Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Small Language Models Need Strong Verifiers to Self-Correct Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.967055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.967055Z digest=sha256:8d112444401297522595ba1ca827cb189c88a2f38d799ab5ee508d070d8f3562

Observation d07800da-c014-48f5-a681-d6d9f3143dd2 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.969856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.969856Z digest=sha256:39e726d48be841f6f68489ee7c696efc2237ec244156712b9fc05e5479d9bd84

Observation afad08a3-61b1-4782-8c08-1af9a7c37026 · outbound

This paper cites Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.972019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.972019Z digest=sha256:c1fe765a138c3f7dd133bbe7dae1461edcdc3d65c165cecbbc4772f20d77d4a9

Observation cb2dd002-b98d-4d15-a457-1b9382094597 · outbound

This paper cites Learning to Reason via Mixture-of-Thought for Logical Reasoning.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Learning to Reason via Mixture-of-Thought for Logical Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.974376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.974376Z digest=sha256:2725994439637585fdab194074c7abc03141f06eeba36d5473abb9cc95e74e67

Pith citing papers

No inbound Pith citation observations are available.