Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:40:36.160535Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 58 inbound Pith citation observations for arXiv:2502.06703.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:40:36.160535Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.391391Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T21:00:08.398075Z
77 of 77 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9b234873-5a17-4cbc-9950-59c4831a6423 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Aime 2024, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 06ce97ab-dcdd-4aa0-9a79-f9045c9869d8 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Introducing Claude , 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ddd5ca04-ed60-4b7d-af4b-aef9152c1f26 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Jiang, Jia Deng, Stella Biderman, and Sean Welleck
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 712a3c0b-25bf-45d1-b9cb-a3beea0be966 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Scaling test-time compute with open models, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91d79cbd-262d-4ba7-9547-f5012acd2b88 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e25032-ec12-4d41-9dde-cba02e654965 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Alphamath almost zero: Process supervision without process
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 30793fe3-4fea-493f-b778-cf1d4dc31f23 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b3be7772-8dda-4476-b35e-0a398b107fd9 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Process Reinforcement through Implicit Rewards
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a1a20f-3ff6-4919-a5d1-d5ea3cf29412 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3b29857-2dab-48a0-bc8f-de5c5cfb570e · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566da9e1-339e-4e73-b2ab-52de7e17902c · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling PAL : Program-aided language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d82bdf7d-f831-4cb8-b1ac-5e61b2fba3d5 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling To RA : A tool-integrated reasoning agent for mathematical problem solving
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7e1abdd1-50fb-48bf-8a14-e2f02454c213 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5840cf9b-0d50-4fa1-9a3b-262f2902e37d · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Reinforced Self-Training (ReST) for Language Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab3ca706-af8b-4e5c-9fd9-cf4d4b816374 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Measuring mathematical problem solving with the MATH dataset
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0bc30d98-3359-4a0b-a26f-831a39a02904 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling V-STaR: Training Verifiers for Self-Taught Reasoners
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 624dac24-83cc-4a27-be2e-2c9f3a305274 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce7b896e-a227-4b1f-922d-b56f2204c105 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling GPT-4o System Card
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57320b16-f3c1-4e58-b6cc-4950d0f23936 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Mistral 7B
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1acd9cc6-35c3-4e0a-af46-d6b722437adf · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a654a2-13ab-401c-8f11-08bb223f5ebf · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling ARGS : Alignment as reward-guided search
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0741e2ef-f8c8-4360-9f98-9db2a4ba181f · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling k0-math, November 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 064fd4a5-66a7-4b43-b591-a837bb24d4f9 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af72ee27-cc45-4415-a83f-ff4808dc3417 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Training Language Models to Self-Correct via Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c47f1cf-0b92-47ba-8bef-8f498b9b7289 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling CoMAT : Chain of mathematically annotated thought improves mathematical reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c3df4bf-21af-4c82-87ea-c8af838614f8 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Process Reward Model with Q-Value Rankings
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce9f88a-7f18-41eb-84d3-b7974d12d9fe · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Let's verify step by step
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a81314a-c661-4bb7-b89d-1de1b1eb2365 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Autopsv: Automated process-supervised verifier
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b70475cd-8263-46b3-bae9-cc6442650a4e · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad645c84-0867-4c71-a2c1-fb03097333fe · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a638610-5dad-438b-982e-cef6dbf8a1db · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Self-refine: Iterative refinement with self-feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 28dbae4a-9763-4ab2-b62d-6bb555d23956 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c24f4a-71ea-486c-8656-7176fd0f95f9 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling GPT-4 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c56bb5-23de-4d70-8814-04cc66d0faa1 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Learning to reason with llms, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25dd2f91-ba65-4050-8ee3-fe1df5e68c3c · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling O1 Replication Journey: A Strategic Progress Report -- Part 1
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c303f902-d6a2-4476-b52b-c8661b8c323e · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Recursive introspection: Teaching language model agents how to self-improve
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 742df8c8-f473-4602-b00a-13436d3260fb · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a82e362-277a-4120-910f-1cc2d62ec641 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a86e0f-6e88-4f8c-86c4-3fd1c3c5815a · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1916c2ad-1322-4db7-876a-fbef78dd2aee · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dce8e4f-0f9d-4d11-a579-5c0aabe028d2 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23018d0f-dec7-41ab-a100-76ea8108a30e · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Skywork-o1, November 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9dde4a2d-d7d3-45af-8c16-eb539b2c6a83 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Skywork-o1 open series
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4c3df50b-c127-466c-96e4-8db1aa666aa1 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37177012-0210-4d98-88e6-f2e2220604b7 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b87581-3c4a-4909-b64a-c04607f3f0c2 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Reinforcement learning: An introduction
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38f11997-da93-40d2-90d1-71f742bd1619 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling M ath S cale: Scaling instruction tuning for mathematical reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e3d29d45-a60e-4d05-87e1-a8126d48b419 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling DART -math: Difficulty-aware rejection tuning for mathematical problem-solving
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0789a4ea-4b89-427d-9fbe-7739eac0e165 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26a9589-09f1-4d73-a331-db9acceeb84c · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Reft: Reasoning with reinforced fine-tuning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation daa40684-8410-4522-937b-42520a90b9e0 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Solving math word problems with process- and outcome-based feedback
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff50eb6-6c28-4bf0-9a17-79734fe149c4 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling A lpha Z ero-like tree-search can guide large language model decoding and training
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2fea5eae-ac56-4f25-a10a-7eb4351bbeac · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774282b2-a9c4-4325-8fa5-87dd9a34556d · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d5603a6b-9393-46dd-8c6e-5c5d0a5c7a09 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9624cf7a-6058-47e1-9e9c-3df7d3a45a81 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 055280d8-4604-453b-b2a2-8e3e5f71105f · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Chain-of-thought prompting elicits reasoning in large language models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 118fdae9-c98e-4e55-9036-0c2d304eb8ae · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Large language models are better reasoners with self-verification
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8e022e89-257e-4389-9a15-22aac81bde78 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe2ff11-667c-4b94-8c94-fa569d2a000e · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Self-evaluation guided beam search for reasoning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e42eec14-34b8-4c59-9232-6f56194215e5 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling An implementation of generative prm
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7804413-2afd-4129-858b-7d5094ef3eeb · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Qwen2 Technical Report
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eca76b55-733a-4676-ad93-7da8f5b93698 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Qwen2.5 Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee1a504-814b-4a98-a9c1-003f89a534a6 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9af62b3a-bb44-4bfb-bd5e-d53ed81bd928 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Tree of thoughts: Deliberate problem solving with large language models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8cf6ffe1-ee6e-4e76-a1ae-24dbc07fd35f · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling MetaMath : Bootstrap your own mathematical questions for large language models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b22cc61a-0819-4ca9-b55b-4430dfc7bc60 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Free Process Rewards without Process Labels
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09937fb9-fcca-4823-b78e-9e8096432330 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling STaR : Bootstrapping reasoning with reasoning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0976b0af-aa03-4cb9-92b9-b9d325a6ba1f · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Quiet- ST ar: Language models can teach themselves to think before speaking
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 36d41b5b-d316-4c83-b945-3fead7233752 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dfb2880-ecef-4fdd-a85e-775e78307e49 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b525bb1b-7f58-491e-b1e7-c250e11dfd5b · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Re ST - MCTS *: LLM self-training via process reward guided tree search
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation df41818f-15c1-457f-8453-228ca05e8ec2 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Entropy-regularized process reward model
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8985249a-65d0-45ed-9783-3e917694a695 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling The Lessons of Developing Process Reward Models in Mathematical Reasoning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b308f711-f1d7-44e4-b73a-f08d37d36706 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8732e1ea-9b23-4cf6-a522-b7934a693e48 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling ProcessBench: Identifying Process Errors in Mathematical Reasoning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bd7b484-084e-43ef-b060-f8173872cc84 · outbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 12e23caf-5740-48e4-a599-552c1dc2ba0d · inbound
Process-Supervised Reward Models for Verifying Clinical Note Generation: A Scalable Approach Guided by Domain Expertise Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe093e0f-d652-441b-b5fa-38c9472106c5 · inbound
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation de9a547a-9609-43e4-a615-cc6cdb1b7bca · inbound
Generative AI Act II: Test Time Scaling Drives Cognition Engineering Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 198
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75491ede-3eb2-43fb-a5ac-1803552c43bc · inbound
Rethinking Inference-Time Scaling: Efficiency Limits and Linguistic Signals Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d34214-fdf5-4c87-a187-334209e2a6f2 · inbound
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6566f18c-c441-4f8c-95c5-1b5bb6c878d8 · inbound
Retrieval Augmented Learning: A Retrial-based Large Language Model Self-Supervised Learning and Autonomous Knowledge Generation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4e6488e-6282-4356-b486-808a8578c478 · inbound
A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ca01e10-82d6-4c5b-bc2c-171388286475 · inbound
SLOT: Sample-specific Language Model Optimization at Test-time Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4935d78-a4ba-474c-b0b3-7ad6ed3f10c7 · inbound
Multilingual Test-Time Scaling via Initial Thought Transfer Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cec42c1f-e556-47dc-b89a-fa22ea3aa623 · inbound
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 061d8941-df23-4968-95e9-82c3a7d639e5 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 770b0951-6d1e-4e04-80c8-1b81e777a340 · inbound
Faster and Better LLMs via Latency-Aware Test-Time Scaling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9834023-d4e3-4163-b2fb-b9b35d000c63 · inbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c218e16-140b-455b-92f5-e2e6dc5776f2 · inbound
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08a3bb2-51a6-464e-bd4b-23b5a2357fc3 · inbound
Can Past Experience Accelerate LLM Reasoning? Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b51122-ee9a-4f4c-a762-8a2f07fc51e9 · inbound
PixelThink: Towards Efficient Chain-of-Pixel Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df98069d-c01a-4be8-a355-bf5788b300b8 · inbound
EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70bd4eff-189c-4e66-ba8b-fcf6c0b35422 · inbound
CyberV: Cybernetics for Test-time Scaling in Video Understanding Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b527d30a-272e-4fb3-9d5c-64ecec4f111a · inbound
Scaling Test-time Compute for LLM Agents Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1dffc48-5b3e-485f-844b-22ccc0ae2bac · inbound
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bb8fc79-a48f-435b-ba00-1ec1647935fc · inbound
Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24c08bd8-c10a-4603-a5cf-ceaf7891c0a7 · inbound
Reasoning in machine vision by learning fast and slow thinking Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47967841-d013-48fe-a5c4-a9853eeb27be · inbound
Test-Time Scaling with Reflective Generative Model Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d5599b-be23-49fd-9f3b-148432ac69ae · inbound
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53431e79-949f-4e23-b2ac-162677dd400c · inbound
Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acd9282a-53af-4db8-a978-38b29d59fe79 · inbound
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f4e7dcf-2c7d-4e7e-aa2a-2a3989bf4133 · inbound
Thinking Isn't an Illusion: Overcoming the Limitations of Reasoning Models via Tool Augmentations Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c1a11d3-de30-4c11-9006-604f787b15e8 · inbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0972242c-07d0-4677-81ae-206435b91c3f · inbound
SSRL: Self-Search Reinforcement Learning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bff4313-4b00-48ed-aa9e-c23e1d52487a · inbound
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6267bea-60f9-4002-9d17-e79cf6229fd4 · inbound
Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e3eaa87b-24a6-41d4-be51-6f9b9d280004 · inbound
Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a58f66e-2eb4-4c11-a9b8-328ef81b1a4e · inbound
Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a18a124-832f-43ac-b30b-6e04511b983b · inbound
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 92240cde-4213-4933-91e7-3328797739ac · inbound
Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bbda081e-6dd4-4094-a0ad-2de16fb6ff00 · inbound
Decoding the Critique Mechanism in Large Reasoning Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 54d004f6-27d7-4e08-9bce-fb0331d17a3d · inbound
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 126d36a2-d261-418c-8477-6cc26e49bbea · inbound
Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e27a3869-9ed8-4ef0-9efb-add1b299e1c8 · inbound
TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7ad84d25-93e2-4f8e-b738-589f602625db · inbound
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3b38060d-072b-466c-b341-4eeb9f518978 · inbound
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a195219e-a949-4335-852b-361d6126e051 · inbound
Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 11ce872f-0adc-4370-8d30-fe73dc63e331 · inbound
Stream-T1: Test-Time Scaling for Streaming Video Generation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bf6069d9-8c6b-41c2-b8be-3d5599c9b936 · inbound
CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation caf4bf73-971a-4969-b53a-d38ec486c73f · inbound
Unsupervised Process Reward Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bee8b60e-40cd-4275-a5e4-987132bfa689 · inbound
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 03950951-7438-428e-bfa6-432eab0f39d4 · inbound
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3be9bfea-ecf3-44d4-bf02-409477c54888 · inbound
Trust Region On-Policy Distillation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6388b362-79f2-43ae-9054-63e432943e5d · inbound
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7170760e-940f-4c37-902f-f3141149314a · inbound
Data Selection Through Iterative Self-Filtering for Vision-Language Settings Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ec2d4e04-69fa-49d4-ab3d-60fa6008b35b · inbound
On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ff9bd608-6a8d-4d92-9f84-fce98c499edf · inbound
When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fb27912b-109e-4283-9f1d-4c4727f952d5 · inbound
Test-Time Scaling for Small VLMs on Multilingual Visual MCQ Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 479f5bd3-9708-4dd4-bb78-add836d6214f · inbound
LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c322afa-c81b-4eb6-ae47-174fb3224b10 · inbound
LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71089cd9-0a78-4511-b41f-ac42a2359ce4 · inbound
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e0c3839-c60d-4d1e-808f-c23fbca65622 · inbound
Interpretable Adaptive Sampling for LLM Test-Time Scaling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d7ffdc-ce11-42f4-96f2-5e42db65a721 · inbound
Hidden Language Consistency Phenomena in Reasoning LLMs Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 278
Source-reported events for the cited work
Unavailable: canonical work link unavailable.