Pith. sign in

Paper Citation Record · LEDGER

Reward Model Generalization for Compute-Aware Test-Time Reasoning

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2505.18065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18065 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:27.675920Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d07cc6f-051c-4d59-b313-62ffd2c6aba3 · outbound

This paper cites Le, and Denny Zhou.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, and Denny Zhou

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:33.454821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:24.081236Z digest=sha256:7724bd090aade0a83f57edb4075fb99e87b570a055e47f692c020b86478db59b

Observation 05fdb4ef-c227-4761-88b5-7b958e0d5059 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:24.152490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:24.152490Z digest=sha256:bc627ae6c09bfbd8c3a6c172c87d72212a0b88d29caf7a082e383ea03ad15800

Observation 727b13c4-1c25-4be9-9086-42328e456b6d · outbound

This paper cites Griffiths, Yuan Cao, and Karthik R.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Griffiths, Yuan Cao, and Karthik R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:33.163810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:24.250753Z digest=sha256:6291291ec4cd706c60ffd746f5b0584d5caab5ba392271c1bbac1147ff3901c8

Observation 18a474a5-bc81-4dc5-9c91-06cbda94bfb9 · outbound

This paper cites Learning to reason with llms.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Learning to reason with llms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.923839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:24.362104Z digest=sha256:5bacd07878404623ec90428df944ddb33b01ffedb8336686256122506406676b

Observation c3a4d3ef-8136-4f88-b5b5-ae8006654fbf · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, January 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, January 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:24.426613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:24.426613Z digest=sha256:c04b1c8548d36e350db7947764803a0164ad3d53a812b3c3d2dcc299a104eaeb

Observation 13428347-59ee-4da7-aba4-b08a2574e3ff · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:32.681900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:24.534100Z digest=sha256:ec83b40b1db2d909fa52ddd01fd210d61ee7644f664e4482e9112746c084bb03

Observation a0dfc639-8053-4ea3-b46b-f059cd7739c1 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.593348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:24.600772Z digest=sha256:8299355cfa8a49a351957d036bde27801b5064ca05b0e6d2ec15287ca721b893

Observation 6ecda57c-0ee5-49bb-b990-d2760e9246bb · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling, February 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling, February 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.491605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:24.687363Z digest=sha256:84eb192580a1eaac50eac58fb9856e42c7adcab93d13856af3b265fc1fc9e8f8

Observation fc930b81-756e-43f1-ab68-5a3c5057776d · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.372697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:24.806743Z digest=sha256:f7f07e8f317d938593fbdb53b5b78197e01214e18ad0743e829aaaf631254987

Observation 8d6b8157-deac-40a5-9ae5-9e66c590011a · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Graph of thoughts: Solving elaborate problems with large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:31.491532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:24.911167Z digest=sha256:7d738e1dda66f4d08aff4c9716b4aa7fb78a7177e8db4471f7e880337701023b

Observation 22167700-eacc-421b-b6dd-ad8d092ee810 · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:30.980423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:24.986744Z digest=sha256:d9159eb5610003b85b6d1a90ad98953c72431b41780049958606faccf580ca0f

Observation 4d3043e1-1d95-4555-9bfe-6683a1065dcb · outbound

This paper cites Let’s Verify Step by Step, May 2023.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Let’s Verify Step by Step, May 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.665499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:25.116168Z digest=sha256:14a544916e365261904b9f0e33e826847ae638cc74b6f82ed7cd7e2944ed96a4

Observation c38bed93-daca-4fd4-b267-9bb599ca8ab4 · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:30.551034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:25.192718Z digest=sha256:64f73eb6863002b5c2fdac3c45834b502fe1e0d596b160d9d2de32f0d02d2660

Observation 5a24e95c-eec4-45e1-bb29-4a6f6118c9f5 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Measuring mathematical problem solving with the math dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.297848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.297848Z digest=sha256:d7d9c27b55572b65d422ebbc0062af135ae101057766611ac22cb0aff9ad4c01

Observation 0f2507de-19be-4ef7-bfb3-b53c6b5ec71e · outbound

This paper cites AIME 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning AIME 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.371400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:25.365975Z digest=sha256:17fed9ac582f92497431a2b6604f27c58b77587aba5a247144a1d5b024e6c0e8

Observation ed7f689d-b065-4e0f-b891-59a0eedcc988 · outbound

This paper cites Qwen2.5 Technical Report.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.458387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.458387Z digest=sha256:95856dfd604707137f2109ca702420ad36c4408a8198d0d332b3d894b2039b38

Observation 12682f0f-c37d-4d95-9637-82eaa88d51f1 · outbound

This paper cites The Llama 3 Herd of Models.

Reward Model Generalization for Compute-Aware Test-Time Reasoning The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.571342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.571342Z digest=sha256:f073ea45f58ab0bdb1aca89e20aefbc3b581ef41d656d18711ed838f13edf431

Observation 97dcf570-6c88-4f66-abed-e1c38a9228d1 · outbound

This paper cites Llama 3 to connect 2024: Vision, edge, and mobile devices.https://ai.meta.com/ blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Llama 3 to connect 2024: Vision, edge, and mobile devices.https://ai.meta.com/ blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.257185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:25.638982Z digest=sha256:b0496fe109ae5a6cc41422a48e8044b375c07cf4c829e1fb9acd806459fadea2

Observation 111de10c-420a-4d6e-bf88-52694d5cc6c8 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning, January 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning, January 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.155577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:25.709602Z digest=sha256:aed77dabd438b40094737d8db3349faac13bedff80466da009f5ebc4df4ceae4

Observation 5bfd9ae4-d6fb-48ab-9074-776f40b33358 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Reward Model Generalization for Compute-Aware Test-Time Reasoning RLHF Workflow: From Reward Modeling to Online RLHF

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.802972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.802972Z digest=sha256:0e185341d839b977477d154c1781db564bdb116146b9f985b66f349083d8a393

Observation 9b4d4216-0c01-49c7-bc2e-deb80b3ec527 · outbound

This paper cites Skywork-o1 open series.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Skywork-o1 open series

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.991155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:25.912177Z digest=sha256:d82938f3882c2c41f33f2e08bb73fe678642248ca705b258b5cec10aacd679f5

Observation 63e3dcf3-98ff-48b8-b892-fd046de0b922 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models, April 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models, April 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.885755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.045623Z digest=sha256:7739cadc095e21df9912ac783b8941249d9cea23a6665c55f4560d316b168e2d

Observation b298c93a-8765-4701-b57a-9268218fcdb4 · outbound

This paper cites Self- Refine: Iterative Refinement with Self-Feedback.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Self- Refine: Iterative Refinement with Self-Feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.779220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.134417Z digest=sha256:1d0547cdb03539b59bd9f0dde5e80c5367cebaf99ee538662bfe7c771f6a9852

Observation 79eae5de-7d27-441b-9245-a15d3ef2d0c8 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Self-critiquing models for assisting human evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.240819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.240819Z digest=sha256:41c30398ebf7834901001d0520f7c0009e3c9dfd81b26475d36cc685019fb458

Observation 4ab4ba77-f6a9-464a-a491-97db88620f98 · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:29.651910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.313976Z digest=sha256:19e78008d67694879bda671ba003c299e56c60035332f00f776e9eb314824b63

Observation d1580805-140a-400e-af05-3e50ae1c20c2 · outbound

This paper cites Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.538636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.414898Z digest=sha256:41ac83e39448dc6b7b865f720e4b13b5387c949b99f8b4e781556d9e885403c8

Observation 23457b43-4bf6-4ab9-986a-95249c5d837d · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models, October 2022.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Automatic Chain of Thought Prompting in Large Language Models, October 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.403676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.484880Z digest=sha256:d73cd845582086d2e10c1f305db11ac7e593669879f50b184225f286681c96c3

Observation 3e53c68c-5b28-483c-8688-904cf476686a · outbound

This paper cites Le, Christopher Ré, and Azalia Mirhoseini.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, Christopher Ré, and Azalia Mirhoseini

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.235347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.586366Z digest=sha256:c151817ac1b688d07dde02b5448d810d4ebc742fc1253f22366f0fdc58b11d22

Observation 80aeda13-4e57-474a-830d-9b7989e22c95 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Solving math word problems with process- and outcome-based feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.655098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.655098Z digest=sha256:6451d6543a685d26d6a2a9b08ce26516b11d3cd37fdb0c00083c1d8a2ec8910c

Observation 3f49e629-983f-402e-8de3-c6bd1a6974a4 · outbound

This paper cites Self-Evaluation Guided Beam Search for Reasoning, October 2023.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Self-Evaluation Guided Beam Search for Reasoning, October 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.066930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.752346Z digest=sha256:b016b51994a326ebce828851d4823897ac22261ec46d10c486c80b2ac2a89b5a

Observation 5989dbf9-6351-427a-96f7-a2b83417e589 · outbound

This paper cites Le, Ed H.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, Ed H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.888280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.821415Z digest=sha256:ef0eb2b567051fc96e67b66b871316f7a63d6b39c293810ad47e7b3c5ccd9dd5

Observation 0f63f736-7ba2-4871-bd45-e37bdb4fc507 · outbound

This paper cites Sparsity-aware generalization theory for deep neural networks.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Sparsity-aware generalization theory for deep neural networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.710906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.878196Z digest=sha256:c6c5c2c934b947f717114b26e6a76c7c8770cdc0812bbd84fd9439ee90b6e078

Observation bcd5bd05-065f-458f-8156-47c1acb6cf61 · outbound

This paper cites Pac-bayes compression bounds so tight that they can explain generalization.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayes compression bounds so tight that they can explain generalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.529245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:26.941110Z digest=sha256:1b7767cd2ef9856e46ac65422985e38002cf8dedf76678bd423e0901e281871d

Observation 4310f865-1e8a-40e8-8302-6476b744c6a7 · outbound

This paper cites Efficient content-based sparse attention with routing transformers.Transactions of the Association for Computational Linguistics, 9:53–68, 2021.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Efficient content-based sparse attention with routing transformers.Transactions of the Association for Computational Linguistics, 9:53–68, 2021

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.981219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.981219Z digest=sha256:747568a42053084f693547774cf824387d31d9f7238be21088ca768b76c7a5c8

Observation aeec097a-a6a7-4685-b0a1-af1c8d334661 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Reward Model Generalization for Compute-Aware Test-Time Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.030634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.030634Z digest=sha256:fb333d90bb42f3d9ec95bd952bea163a96e6a7a8c3ca015acef67518cec95e7a

Observation ae4a9032-b346-4db8-893c-4db624af8ef8 · outbound

This paper cites Policy gradient meth- ods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Policy gradient meth- ods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.086932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.086932Z digest=sha256:1cc03d81700510634f06ca701393a1c81948506cb813aed5995b66e3925269ff

Observation 8ab56fe0-8d44-4325-bf39-5dde9dc87c86 · outbound

This paper cites Pac-bayesian model averaging.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayesian model averaging

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.135988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.135988Z digest=sha256:931215d63289919fdda5eeb60bcf845204808429a5a1918e925d46f73395cc11

Observation 908e7d2b-f6af-4cf3-ae26-283ddea24f6a · outbound

This paper cites Pac-bayesian generalisation error bounds for gaussian process classification.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayesian generalisation error bounds for gaussian process classification

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.359430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:27.172308Z digest=sha256:793c9e62d7b0bd292b978074e49aebbe1852e65a95a1379b21409b4bcdbf7cc9

Observation bd957e38-9971-47c0-9ac3-95acd361c0e5 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Training Verifiers to Solve Math Word Problems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.222407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.222407Z digest=sha256:3f4607e55b7f3170f17e1c8ec34ae13dd782c8285460886eadcb2cf88d52d8af

Observation 88c52a6c-370f-4d0c-8894-c2d94f5e85d4 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.263006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.263006Z digest=sha256:aa7b68e349f5f906bd611c0b1a9ff92c3a67d5d5f0cc421962008fca2ba0c8d4

Observation f9b003d7-32fe-4b32-ab05-8497d578fab5 · outbound

This paper cites Enhancing LLM Reasoning with Reward-guided Tree Search.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Enhancing LLM Reasoning with Reward-guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.309575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.309575Z digest=sha256:32dfdcb42b59c3ac15caee0eb97c903cb6790e023aff0f45e119346d0d594ca3

Observation ff489b9b-8686-4b99-929b-ebd174b217a5 · outbound

This paper cites LiteSearch: Efficacious Tree Search for LLM.

Reward Model Generalization for Compute-Aware Test-Time Reasoning LiteSearch: Efficacious Tree Search for LLM

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.388786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.388786Z digest=sha256:f89654cb788fafa91ddd8e6168a7b9ac173b3d2d09801119a13ebcbcb8052400

Observation 5deda4ba-41c6-44cd-8c92-812854e08f0d · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Reward Model Generalization for Compute-Aware Test-Time Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.432302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.432302Z digest=sha256:298e4a6c9688724bf9fd60d4bd4c1ce307a5a23e91dffee5fdf6741bfbe8fe02

Observation b0e8a27f-7ab7-4583-9ee6-bbe4a596520d · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.461851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.461851Z digest=sha256:3ef08b22def829544b2a12b1c2a213b818ac39b0108fa722ba35120f3cc2e885

Observation 4ded9f65-0ab6-491b-815f-f82eeb5b2c6c · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.488154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.488154Z digest=sha256:7f91473099b2e66d9808b5208e7df3ea191ed797318d34f2445eaa03f025c89c

Observation b23991e7-dce8-4463-9a82-445e0a635091 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.539519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.539519Z digest=sha256:47ca1c5456827bf778cde8d0aa2f0231ac940bc3df248285b4d6a78b3f225677

Observation 534d5fd5-3165-403d-b888-4f26d9ddf483 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.571108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.571108Z digest=sha256:fc3fed6b5663f0c6a76422ea0f79050e54a12cff5043f561b4f5279c11699187

Observation 55695bc1-c234-4fde-aef1-1f7a950991d9 · outbound

This paper cites BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.601318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.601318Z digest=sha256:0f698e6cb1b450052ca8261f4e2db4103af7d82f0b845e97efbf6f81d172140f

Observation da4596c2-6956-4a5a-8812-5bfe8b0fe509 · outbound

This paper cites Subtracting from 1 yields the bound Equation 6.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Subtracting from 1 yields the bound Equation 6

Reference 49

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:42:27.897972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:42:27.675920Z digest=sha256:bc6359890cf85e3666b250d19b3ca5ee08b6e84acea52a13351d19f22b20ae07

Pith citing papers

No inbound Pith citation observations are available.