Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:27.675920Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2505.18065.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:27.675920Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3d07cc6f-051c-4d59-b313-62ffd2c6aba3 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, and Denny Zhou
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05fdb4ef-c227-4761-88b5-7b958e0d5059 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727b13c4-1c25-4be9-9086-42328e456b6d · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Griffiths, Yuan Cao, and Karthik R
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 18a474a5-bc81-4dc5-9c91-06cbda94bfb9 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Learning to reason with llms
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3a4d3ef-8136-4f88-b5b5-ae8006654fbf · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, January 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13428347-59ee-4da7-aba4-b08a2574e3ff · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0dfc639-8053-4ea3-b46b-f059cd7739c1 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ecda57c-0ee5-49bb-b990-d2760e9246bb · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling, February 2025
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc930b81-756e-43f1-ab68-5a3c5057776d · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d6b8157-deac-40a5-9ae5-9e66c590011a · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Graph of thoughts: Solving elaborate problems with large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22167700-eacc-421b-b6dd-ad8d092ee810 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d3043e1-1d95-4555-9bfe-6683a1065dcb · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Let’s Verify Step by Step, May 2023
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c38bed93-daca-4fd4-b267-9bb599ca8ab4 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a24e95c-eec4-45e1-bb29-4a6f6118c9f5 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Measuring mathematical problem solving with the math dataset
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f2507de-19be-4ef7-bfb3-b53c6b5ec71e · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning AIME 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed7f689d-b065-4e0f-b891-59a0eedcc988 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Qwen2.5 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12682f0f-c37d-4d95-9637-82eaa88d51f1 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning The Llama 3 Herd of Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97dcf570-6c88-4f66-abed-e1c38a9228d1 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Llama 3 to connect 2024: Vision, edge, and mobile devices.https://ai.meta.com/ blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 111de10c-420a-4d6e-bf88-52694d5cc6c8 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning, January 2025
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5bfd9ae4-d6fb-48ab-9074-776f40b33358 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning RLHF Workflow: From Reward Modeling to Online RLHF
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b4d4216-0c01-49c7-bc2e-deb80b3ec527 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Skywork-o1 open series
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63e3dcf3-98ff-48b8-b892-fd046de0b922 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models, April 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b298c93a-8765-4701-b57a-9268218fcdb4 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Self- Refine: Iterative Refinement with Self-Feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79eae5de-7d27-441b-9245-a15d3ef2d0c8 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Self-critiquing models for assisting human evaluators
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ab4ba77-f6a9-464a-a491-97db88620f98 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1580805-140a-400e-af05-3e50ae1c20c2 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23457b43-4bf6-4ab9-986a-95249c5d837d · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Automatic Chain of Thought Prompting in Large Language Models, October 2022
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e53c68c-5b28-483c-8688-904cf476686a · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, Christopher Ré, and Azalia Mirhoseini
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80aeda13-4e57-474a-830d-9b7989e22c95 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Solving math word problems with process- and outcome-based feedback
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f49e629-983f-402e-8de3-c6bd1a6974a4 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Self-Evaluation Guided Beam Search for Reasoning, October 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5989dbf9-6351-427a-96f7-a2b83417e589 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, Ed H
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f63f736-7ba2-4871-bd45-e37bdb4fc507 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Sparsity-aware generalization theory for deep neural networks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bcd5bd05-065f-458f-8156-47c1acb6cf61 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayes compression bounds so tight that they can explain generalization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4310f865-1e8a-40e8-8302-6476b744c6a7 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Efficient content-based sparse attention with routing transformers.Transactions of the Association for Computational Linguistics, 9:53–68, 2021
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeec097a-a6a7-4685-b0a1-af1c8d334661 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae4a9032-b346-4db8-893c-4db624af8ef8 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Policy gradient meth- ods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab56fe0-8d44-4325-bf39-5dde9dc87c86 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayesian model averaging
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908e7d2b-f6af-4cf3-ae26-283ddea24f6a · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayesian generalisation error bounds for gaussian process classification
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd957e38-9971-47c0-9ac3-95acd361c0e5 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Training Verifiers to Solve Math Word Problems
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c52a6c-370f-4d0c-8894-c2d94f5e85d4 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b003d7-32fe-4b32-ab05-8497d578fab5 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Enhancing LLM Reasoning with Reward-guided Tree Search
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff489b9b-8686-4b99-929b-ebd174b217a5 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning LiteSearch: Efficacious Tree Search for LLM
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5deda4ba-41c6-44cd-8c92-812854e08f0d · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning AlphaMath Almost Zero: Process Supervision without Process
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e8a27f-7ab7-4583-9ee6-bbe4a596520d · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ded9f65-0ab6-491b-815f-f82eeb5b2c6c · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23991e7-dce8-4463-9a82-445e0a635091 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 534d5fd5-3165-403d-b888-4f26d9ddf483 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55695bc1-c234-4fde-aef1-1f7a950991d9 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4596c2-6956-4a5a-8812-5bfe8b0fe509 · outbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning Subtracting from 1 yields the bound Equation 6
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.