Pith. sign in

Paper Citation Record · LEDGER

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

As of 23 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 58 inbound Pith citation observations for arXiv:2502.06703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06703 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:40:36.160535Z

measured 135 of 135 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.391391Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:00:08.398075Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b234873-5a17-4cbc-9950-59c4831a6423 · outbound

This paper cites Aime 2024, 2024.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Aime 2024, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.140448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.899577Z digest=sha256:eec5969d41458dd08ab65c902b1c44e8fc8a642cdfba1ff0a9cdbe2e5b2e5d3e

Observation 06ce97ab-dcdd-4aa0-9a79-f9045c9869d8 · outbound

This paper cites Introducing Claude , 2023.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Introducing Claude , 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.130065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.903281Z digest=sha256:36c01936f3c44f2ad70e7b10b2f38ac2ee3949f5b1dc3703e2dee3a81b01c64a

Observation ddd5ca04-ed60-4b7d-af4b-aef9152c1f26 · outbound

This paper cites Jiang, Jia Deng, Stella Biderman, and Sean Welleck.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Jiang, Jia Deng, Stella Biderman, and Sean Welleck

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.119494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.906437Z digest=sha256:69a88b8aa228808e9f1e0643a7b715fb2e1be9662f7ccf491377a98ed46c7f60

Observation 712a3c0b-25bf-45d1-b9cb-a3beea0be966 · outbound

This paper cites Scaling test-time compute with open models, 2024.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Scaling test-time compute with open models, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.909681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.909681Z digest=sha256:9cb580b1b8416cfe25acba9ecc886f443803bc29ac042a2471d34afdb35e83d6

Observation 91d79cbd-262d-4ba7-9547-f5012acd2b88 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.912801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.912801Z digest=sha256:f1ffab27ef7e86974238f16218d6f034f9ae13669872bc46b46738e1d2ab5ea1

Observation 37e25032-ec12-4d41-9dde-cba02e654965 · outbound

This paper cites Alphamath almost zero: Process supervision without process.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Alphamath almost zero: Process supervision without process

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.101228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.916332Z digest=sha256:0a9c507559a3b3e92f958eb1f33604a298ed226ec0effe86753163c3fa2f104f

Observation 30793fe3-4fea-493f-b778-cf1d4dc31f23 · outbound

This paper cites an unresolved cited work.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:40:37.089675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.919626Z digest=sha256:bb0401ce9279f7a32e74bc3cc7fea41c258acca8bb276f72c5e020c4217ca6d1

Observation b3be7772-8dda-4476-b35e-0a398b107fd9 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Process Reinforcement through Implicit Rewards

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.922961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.922961Z digest=sha256:d14adc1de34bc81e7b6ec35f0f3c59b46b546a55e41762aa7c1567db205aea4e

Observation 45a1a20f-3ff6-4919-a5d1-d5ea3cf29412 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.926575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.926575Z digest=sha256:0cdc23e020c5e30b35bc3010442becd72ddbe9bea0a359e866f3563bc4212bcd

Observation a3b29857-2dab-48a0-bc8f-de5c5cfb570e · outbound

This paper cites The Llama 3 Herd of Models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.930952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.930952Z digest=sha256:89b199c836d41e08f25fa226bdd2dd5db093387acb33319ce8f5d5ca6d73a3b5

Observation 566da9e1-339e-4e73-b2ab-52de7e17902c · outbound

This paper cites PAL : Program-aided language models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling PAL : Program-aided language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.078105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.934485Z digest=sha256:e9b6a3b18cfbf39c7ad2e327b9f363542dd912a74181d7345c045c2b0a25ea9b

Observation d82bdf7d-f831-4cb8-b1ac-5e61b2fba3d5 · outbound

This paper cites To RA : A tool-integrated reasoning agent for mathematical problem solving.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling To RA : A tool-integrated reasoning agent for mathematical problem solving

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.066886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.937991Z digest=sha256:7f27fe7b167e9cb264384398fa1a30765aa12b15ecd49ee75d4d89e010242be3

Observation 7e1abdd1-50fb-48bf-8a14-e2f02454c213 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.941372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.941372Z digest=sha256:3dfe6d43c21e0ff1ab9a34c60563186da0029268d27a7201bb60d26f428cfce1

Observation 5840cf9b-0d50-4fa1-9a3b-262f2902e37d · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Reinforced Self-Training (ReST) for Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.944896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.944896Z digest=sha256:68342c3ec346fb6c408412bb82a8af73f572fda471fb3e9c66d9a2d5a1fef068

Observation ab3ca706-af8b-4e5c-9fd9-cf4d4b816374 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Measuring mathematical problem solving with the MATH dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.055719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.948498Z digest=sha256:6b63bf6a93b3138653441b9be370b0bc59a8bf6488883a1992f91d1ed223d1a5

Observation 0bc30d98-3359-4a0b-a26f-831a39a02904 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.951932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.951932Z digest=sha256:29af82978a19b4f184240c52fe92f6cc3de6a61188f8945605fb7d2998f1a8a5

Observation 624dac24-83cc-4a27-be2e-2c9f3a305274 · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.955396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.955396Z digest=sha256:b04b1864baf33f7c1699247c1d1a4fceecafc2e10fa42175d2663439c5535adb

Observation ce7b896e-a227-4b1f-922d-b56f2204c105 · outbound

This paper cites GPT-4o System Card.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.959414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.959414Z digest=sha256:7d5a660d9ae58622705a9e7aad4124d0ee281a38c0e80feadfca14708ecd4325

Observation 57320b16-f3c1-4e58-b6cc-4950d0f23936 · outbound

This paper cites Mistral 7B.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Mistral 7B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.963873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.963873Z digest=sha256:86cf13d214fb3ce2b14ac503569dd02c1f90c285b0ec18f92f0008b399cfa9d1

Observation 1acd9cc6-35c3-4e0a-af46-d6b722437adf · outbound

This paper cites MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.967949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.967949Z digest=sha256:76c6e5436d34a522305e78d9540fdfc53154071768257041db792def7015e145

Observation d4a654a2-13ab-401c-8f11-08bb223f5ebf · outbound

This paper cites ARGS : Alignment as reward-guided search.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling ARGS : Alignment as reward-guided search

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.044614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.971685Z digest=sha256:6a75507249884f775819644e5153423c193fb1d489a4df275b3c65ed7235f3ad

Observation 0741e2ef-f8c8-4360-9f98-9db2a4ba181f · outbound

This paper cites k0-math, November 2024.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling k0-math, November 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.033274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.975103Z digest=sha256:ba9630e270fcdf95ef84975e97ef72695dad7d01434a5385ce91addfd9ceec57

Observation 064fd4a5-66a7-4b43-b591-a837bb24d4f9 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.978584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.978584Z digest=sha256:57fe36c2abfa4c41aa91c238855120ddcaa7e3357bb032fe1a67648791442784

Observation af72ee27-cc45-4415-a83f-ff4808dc3417 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Training Language Models to Self-Correct via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.982068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.982068Z digest=sha256:5c06a3aa69aeb6090b59da2018ab868c7e5f54ccb30a8cc668814ee3e06ec707

Observation 7c47f1cf-0b92-47ba-8bef-8f498b9b7289 · outbound

This paper cites CoMAT : Chain of mathematically annotated thought improves mathematical reasoning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling CoMAT : Chain of mathematically annotated thought improves mathematical reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.985577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.985577Z digest=sha256:57ed4d72e5177a698fd90de1c63da8216d9eab75346f716ff48ea8dd02d60765

Observation 8c3df4bf-21af-4c82-87ea-c8af838614f8 · outbound

This paper cites Process Reward Model with Q-Value Rankings.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Process Reward Model with Q-Value Rankings

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.988761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.988761Z digest=sha256:1320019215ddc49699137ccae682ba141c3d5a7c2c71b9831f17702335f54349

Observation 0ce9f88a-7f18-41eb-84d3-b7974d12d9fe · outbound

This paper cites Let's verify step by step.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Let's verify step by step

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.991827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.991827Z digest=sha256:3577b59cce8ac6be99f48decf27fba1f8b3b2a982d53ea3fe0195eca0fb4e3ca

Observation 8a81314a-c661-4bb7-b89d-1de1b1eb2365 · outbound

This paper cites Autopsv: Automated process-supervised verifier.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Autopsv: Automated process-supervised verifier

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.014738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:35.994620Z digest=sha256:cbcb0b968d784913e504fb87fd2ea95117aae636cec6881800ac8ba77f8bf970

Observation b70475cd-8263-46b3-bae9-cc6442650a4e · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.997449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.997449Z digest=sha256:6cc89de97a46f0bfcc999e6d8d30fe6e03a9b01e2790ffd7f4641d8c7128d026

Observation ad645c84-0867-4c71-a2c1-fb03097333fe · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.000646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.000646Z digest=sha256:7a43de3b1e3db706aa837813b659cc214eef37fd112e0b34c016474f25451b9b

Observation 5a638610-5dad-438b-982e-cef6dbf8a1db · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Self-refine: Iterative refinement with self-feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:37.002315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.003933Z digest=sha256:2f1a320ea4525eff88d159dd7a879ee4aaac874ab86423bb0a21903938d897ad

Observation 28dbae4a-9763-4ab2-b62d-6bb555d23956 · outbound

This paper cites Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.007327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.007327Z digest=sha256:327be9d5fb6408be88c68bd690cdb937271c849316c3e06dc0bf1c031c379e5a

Observation a4c24f4a-71ea-486c-8656-7176fd0f95f9 · outbound

This paper cites GPT-4 Technical Report.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling GPT-4 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.010978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.010978Z digest=sha256:837b382d2db720f190774e085fb87dbedd9ccc073c77ab3d32f728839b552d69

Observation f2c56bb5-23de-4d70-8814-04cc66d0faa1 · outbound

This paper cites Learning to reason with llms, 2024.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Learning to reason with llms, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.014673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.014673Z digest=sha256:f11abd86172c344abbeca2f3b5f4ae6fa016bd221445e954dc801f708c491f8c

Observation 25dd2f91-ba65-4050-8ee3-fe1df5e68c3c · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.017882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.017882Z digest=sha256:93b7cc3d0ea65235ce425cb832caedfeec10d7e02907d0f2a96962847564c4a5

Observation c303f902-d6a2-4476-b52b-c8661b8c323e · outbound

This paper cites Recursive introspection: Teaching language model agents how to self-improve.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Recursive introspection: Teaching language model agents how to self-improve

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.984757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.021431Z digest=sha256:313548dc0b0d270b62200f2171dfca162c2ef56c1f4288bdcc7c979d8a896f6e

Observation 742df8c8-f473-4602-b00a-13436d3260fb · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.972908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.024620Z digest=sha256:4a58be7acf43d31d373cf54a7e1a9467ff786977cfc00095becabf9c07bdc735

Observation 2a82e362-277a-4120-910f-1cc2d62ec641 · outbound

This paper cites RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.027707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.027707Z digest=sha256:3ff5715f257239dfe22f9f43c8594d1d3ab08a08215924bd5f45945ebb61af47

Observation c4a86e0f-6e88-4f8c-86c4-3fd1c3c5815a · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.031174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.031174Z digest=sha256:30d0614bd3ba460d0d90d5c54f3b0fca4292ba431ca9c39f2af4a0e28cb2c14b

Observation 1916c2ad-1322-4db7-876a-fbef78dd2aee · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.034863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.034863Z digest=sha256:fa95ae0ee6b722eb172c7aefe9b3b30d2fc4eaca1e8ab7ed18a2173b73a7e25d

Observation 3dce8e4f-0f9d-4d11-a579-5c0aabe028d2 · outbound

This paper cites Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.038438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.038438Z digest=sha256:3bb0f36085d86771812ea4e58b0a605c79198b472fdc599832a734acd2610d67

Observation 23018d0f-dec7-41ab-a100-76ea8108a30e · outbound

This paper cites Skywork-o1, November 2024.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Skywork-o1, November 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.960840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.042118Z digest=sha256:f13b178d556ff98542116ca6d29a9e2d232a92871bbd56918e3ec6acb01b26cd

Observation 9dde4a2d-d7d3-45af-8c16-eb539b2c6a83 · outbound

This paper cites Skywork-o1 open series.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Skywork-o1 open series

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.949886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.045536Z digest=sha256:114d66d8230193e9821cfeccffdee705b7cc79a5e177281bd9e08f09d591a0c5

Observation 4c3df50b-c127-466c-96e4-8db1aa666aa1 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.048931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.048931Z digest=sha256:fe5ae418a57d85491ce3af7cd459c8926f3394420621aa17432ae9ee1db9f223

Observation 37177012-0210-4d98-88e6-f2e2220604b7 · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.052680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.052680Z digest=sha256:9819c41e6d25614754b7d8be01c9de91522c68fab603a9016b0726836abad67a

Observation 53b87581-3c4a-4909-b64a-c04607f3f0c2 · outbound

This paper cites Reinforcement learning: An introduction.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Reinforcement learning: An introduction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.056441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.056441Z digest=sha256:ea082f85c9fc39afa5c28254b6c353d2efc61be7c961d581ed75ec66d40a5426

Observation 38f11997-da93-40d2-90d1-71f742bd1619 · outbound

This paper cites M ath S cale: Scaling instruction tuning for mathematical reasoning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling M ath S cale: Scaling instruction tuning for mathematical reasoning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.932605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.059886Z digest=sha256:d3815bda93d1d09283c45798e09e205b84fa8a7ee1598120d8d3523507bf5fcd

Observation e3d29d45-a60e-4d05-87e1-a8126d48b419 · outbound

This paper cites DART -math: Difficulty-aware rejection tuning for mathematical problem-solving.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling DART -math: Difficulty-aware rejection tuning for mathematical problem-solving

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.921249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.063273Z digest=sha256:aac18f55f45ec60776e306557145c1be0412eb2a1bcec88d11f7d9b1941fe6ce

Observation 0789a4ea-4b89-427d-9fbe-7739eac0e165 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.066620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.066620Z digest=sha256:f1c7a60ad4cfef924b5cf5b9be5eebcb43a2aaa6023e6777b49fc4f83131ccd8

Observation b26a9589-09f1-4d73-a331-db9acceeb84c · outbound

This paper cites Reft: Reasoning with reinforced fine-tuning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Reft: Reasoning with reinforced fine-tuning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.909963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.071452Z digest=sha256:353f871d5db1e400199b4e1486ef003e40ffa5232458770d7d07dcb5bd13102f

Observation daa40684-8410-4522-937b-42520a90b9e0 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Solving math word problems with process- and outcome-based feedback

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.074931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.074931Z digest=sha256:2b275883c27f66ac42c722f1bb70aa0ff81e6772fa6bb4bb2cccc02a2ea006dc

Observation 8ff50eb6-6c28-4bf0-9a17-79734fe149c4 · outbound

This paper cites A lpha Z ero-like tree-search can guide large language model decoding and training.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling A lpha Z ero-like tree-search can guide large language model decoding and training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.898379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.078652Z digest=sha256:8dd75b8bfa8f438b956a6c5c52be66d4b4106e8365a07c67b620e891814533b5

Observation 2fea5eae-ac56-4f25-a10a-7eb4351bbeac · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.081619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.081619Z digest=sha256:7e22fbaffd39e1ff18fb0bfdfcc459137f5a9c6cbb03c44f26db6f0d047bb1fd

Observation 774282b2-a9c4-4325-8fa5-87dd9a34556d · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.885802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.085554Z digest=sha256:85f8a73397180c4d2f79045baa24bab2f548f7a26abd3cb4e85394fdbf50a00b

Observation d5603a6b-9393-46dd-8c6e-5c5d0a5c7a09 · outbound

This paper cites Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.089127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.089127Z digest=sha256:7b5e1a8ae972fe5bf4472342a7d0ddd1d17717e3e6374a527e1b5c9b776eddec

Observation 9624cf7a-6058-47e1-9e9c-3df7d3a45a81 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.873687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.092625Z digest=sha256:8fe14196f55b229818b7ccb3614b51381340a43b2cf900b298909868d1f651a9

Observation 055280d8-4604-453b-b2a2-8e3e5f71105f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Chain-of-thought prompting elicits reasoning in large language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.862420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.095724Z digest=sha256:bb0d765010c07a8c17a6bf10e25699b59f7bc927f065db79ad938c629640cea0

Observation 118fdae9-c98e-4e55-9036-0c2d304eb8ae · outbound

This paper cites Large language models are better reasoners with self-verification.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Large language models are better reasoners with self-verification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.851289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.099470Z digest=sha256:27dd3f27f1f4bcbdd0d3483a8f93175bc002e8eeb174a846cf7a48a5dd96cdc6

Observation 8e022e89-257e-4389-9a15-22aac81bde78 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.103088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.103088Z digest=sha256:4291d4abc4ae61b0fa6f9cd04715935abdcdb4aa900f29c6f2f0194256081218

Observation 9fe2ff11-667c-4b94-8c94-fa569d2a000e · outbound

This paper cites Self-evaluation guided beam search for reasoning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Self-evaluation guided beam search for reasoning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.839516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.106511Z digest=sha256:ec3da583bc3b74cc98372d81d688667c2a3671143ff39c3f1c522f77df72a900

Observation e42eec14-34b8-4c59-9232-6f56194215e5 · outbound

This paper cites An implementation of generative prm.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling An implementation of generative prm

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.109732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.109732Z digest=sha256:a01bfc1f65091cc283b7c9758a351e54e2553afe73d19f4355c113d9200f5c3b

Observation f7804413-2afd-4129-858b-7d5094ef3eeb · outbound

This paper cites Qwen2 Technical Report.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Qwen2 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.112622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.112622Z digest=sha256:689ea18c7705e9b28292b8a8dd16fdf7ed634644f613aa17c497f9a85d5f8559

Observation eca76b55-733a-4676-ad93-7da8f5b93698 · outbound

This paper cites Qwen2.5 Technical Report.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Qwen2.5 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.115695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.115695Z digest=sha256:f22be923cc1b84cf2d4b14f0acbd3480c0a576f23b2d5a18132f170c7bee14da

Observation bee1a504-814b-4a98-a9c1-003f89a534a6 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.118755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.118755Z digest=sha256:b4de5e325c8f15e28e399d97484bfcede2c760a4c1aa9f0d54c714a9576996dc

Observation 9af62b3a-bb44-4bfb-bd5e-d53ed81bd928 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Tree of thoughts: Deliberate problem solving with large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.821051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.121948Z digest=sha256:4d06c3e9740266bc8f0800250deb361a390018b063ba2924cc1f880200791e6b

Observation 8cf6ffe1-ee6e-4e76-a1ae-24dbc07fd35f · outbound

This paper cites MetaMath : Bootstrap your own mathematical questions for large language models.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling MetaMath : Bootstrap your own mathematical questions for large language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.809196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.125025Z digest=sha256:62b02b762c6f777184183f2c1032724bc634adae55491db20262900412aca8e2

Observation b22cc61a-0819-4ca9-b55b-4430dfc7bc60 · outbound

This paper cites Free Process Rewards without Process Labels.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Free Process Rewards without Process Labels

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.127934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.127934Z digest=sha256:1b98a9813a60a4809a20222cb13f9faf967bb6695b2bfe48019ba9755877c646

Observation 09937fb9-fcca-4823-b78e-9e8096432330 · outbound

This paper cites STaR : Bootstrapping reasoning with reasoning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling STaR : Bootstrapping reasoning with reasoning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.797361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.131103Z digest=sha256:60023c039be98fdf6d50c1ba0537574a61bf078cbaebe43f9c863bcb5e62db5f

Observation 0976b0af-aa03-4cb9-92b9-b9d325a6ba1f · outbound

This paper cites Quiet- ST ar: Language models can teach themselves to think before speaking.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Quiet- ST ar: Language models can teach themselves to think before speaking

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.785727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.133970Z digest=sha256:cda74a2810c0d2762911b8178ce7c7f47d813b464104e85b927028f2ddf9ed00

Observation 36d41b5b-d316-4c83-b945-3fead7233752 · outbound

This paper cites Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.136935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.136935Z digest=sha256:f4e0d485021a42734332dbe6045d397a8a4d92b96f2281f98d2fb73cc460989e

Observation 3dfb2880-ecef-4fdd-a85e-775e78307e49 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.140097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.140097Z digest=sha256:44c44d14b4a709c27bcb96296a4ffa8b2c0ebfe3e215907a52adb5c910235c3b

Observation b525bb1b-7f58-491e-b1e7-c250e11dfd5b · outbound

This paper cites Re ST - MCTS *: LLM self-training via process reward guided tree search.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Re ST - MCTS *: LLM self-training via process reward guided tree search

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.767667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.143337Z digest=sha256:2fa140be308d45e84d1def7f5c6e282545732e81682376071d318b5a9fa4dfd4

Observation df41818f-15c1-457f-8453-228ca05e8ec2 · outbound

This paper cites Entropy-regularized process reward model.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Entropy-regularized process reward model

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.146362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.146362Z digest=sha256:f52ae7c8d291029a23043036c76c237ee178087adf94221ab9be52b74ee1161a

Observation 8985249a-65d0-45ed-9783-3e917694a695 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.149694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.149694Z digest=sha256:06d750a7a3ddaffb24cf2a35588ec73985d93f8e4f543d1c34f09a024ea041dc

Observation b308f711-f1d7-44e4-b73a-f08d37d36706 · outbound

This paper cites Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.153233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.153233Z digest=sha256:f4c18fb2f4aaed365b5d411ee4b3845bcfebdcabfbc3eead4bc6cbcabe4417ba

Observation 8732e1ea-9b23-4cf6-a522-b7934a693e48 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.156886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.156886Z digest=sha256:07f0e150a0b6471bac9fdb586e367f95f2cd1640bdc4fff03ef479a08e3e28fd

Observation 8bd7b484-084e-43ef-b060-f8173872cc84 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:40:36.755459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T14:40:36.160535Z digest=sha256:e4852fcc571149c700b9200c4cba3035774949ecf40bf5f4a04dc4d9e27a09d4

Pith citing papers

Observation 12e23caf-5740-48e4-a599-552c1dc2ba0d · inbound

Process-Supervised Reward Models for Verifying Clinical Note Generation: A Scalable Approach Guided by Domain Expertise cites this paper.

Process-Supervised Reward Models for Verifying Clinical Note Generation: A Scalable Approach Guided by Domain Expertise Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T14:03:18.881721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:03:18.881721Z digest=sha256:414da7091786da8981d87b714e00ab25b8e94e8528955f710e6517d649111998

Observation fe093e0f-d652-441b-b5fa-38c9472106c5 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:59:03.190673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:877648529e14393603dc2515fb3f5c112c94d55e4d4503f30dc5e6bb0b3bf927

Observation de9a547a-9609-43e4-a615-cc6cdb1b7bca · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 198

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.391391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.391391Z digest=sha256:ae3ed589ccf486228f97acafec238b0b9becee668421feaa1adbffc0ace4b2fd

Observation 75491ede-3eb2-43fb-a5ac-1803552c43bc · inbound

Rethinking Inference-Time Scaling: Efficiency Limits and Linguistic Signals cites this paper.

Rethinking Inference-Time Scaling: Efficiency Limits and Linguistic Signals Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:01:24.978443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:01:24.978443Z digest=sha256:fbb52dbcb92a35a0b4f915f5ebcaabea078178ea538243a06fa27d29a4c492a1

Observation 04d34214-fdf5-4c87-a187-334209e2a6f2 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:11.989033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:11.989033Z digest=sha256:a39909dbd5dcb8c823e2c05247f64a84185780bf4bdfc3cc1ba4e30fcf0ee916

Observation 6566f18c-c441-4f8c-95c5-1b5bb6c878d8 · inbound

Retrieval Augmented Learning: A Retrial-based Large Language Model Self-Supervised Learning and Autonomous Knowledge Generation cites this paper.

Retrieval Augmented Learning: A Retrial-based Large Language Model Self-Supervised Learning and Autonomous Knowledge Generation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:32:20.445343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:32:20.445343Z digest=sha256:7902d86d6e2ff4e6d97647c4bf0e6ad461a9276d8eb779747845328be7436d6b

Observation b4e6488e-6282-4356-b486-808a8578c478 · inbound

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law cites this paper.

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:48:17.937441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:48:17.937441Z digest=sha256:974c6dd697c43e318e51c3b2e6391199940610346a69daa7800b9e835d8d83b3

Observation 7ca01e10-82d6-4c5b-bc2c-171388286475 · inbound

SLOT: Sample-specific Language Model Optimization at Test-time cites this paper.

SLOT: Sample-specific Language Model Optimization at Test-time Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:16.941819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:16.941819Z digest=sha256:8023d5e3474b1c82cae9d822e74cef228b1fbcda4a4b0ccdffee43ded96a5b70

Observation d4935d78-a4ba-474c-b0b3-7ad6ed3f10c7 · inbound

Multilingual Test-Time Scaling via Initial Thought Transfer cites this paper.

Multilingual Test-Time Scaling via Initial Thought Transfer Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:21.815453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:21.815453Z digest=sha256:cbb71c4361d9d1b1212cecb2f906221e06d3e72595a9d413f1efa0cef9fb64c4

Observation cec42c1f-e556-47dc-b89a-fa22ea3aa623 · inbound

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning cites this paper.

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:17.775098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:45:17.775098Z digest=sha256:d3a7ff187ce39a37d2dc908043ed5e288cdd40ce5c3f57cde01a12559bbabf3f

Observation 061d8941-df23-4968-95e9-82c3a7d639e5 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.154734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.154734Z digest=sha256:26c69a1549193e71934574c060b5b80b6ad8eafa0d89e9afea629e62697b6a23

Observation 770b0951-6d1e-4e04-80c8-1b81e777a340 · inbound

Faster and Better LLMs via Latency-Aware Test-Time Scaling cites this paper.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.357992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.357992Z digest=sha256:f00aeaabb00e347fcaf68e7815109f1555cd6f7e6b0fc2af1208391511943041

Observation c9834023-d4e3-4163-b2fb-b9b35d000c63 · inbound

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision cites this paper.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.428465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.428465Z digest=sha256:c00a605e1c5e6fb36a5b758c162e869df80b352838f91f92edcf78daa2ead0b5

Observation 0c218e16-140b-455b-92f5-e2e6dc5776f2 · inbound

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence cites this paper.

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:06.969200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:06.969200Z digest=sha256:0b052e3abcf2194fbbee8b3c8ecf4e2fb5db54c6f6cb40b6b6288b31b7c37ed3

Observation c08a3bb2-51a6-464e-bd4b-23b5a2357fc3 · inbound

Can Past Experience Accelerate LLM Reasoning? cites this paper.

Can Past Experience Accelerate LLM Reasoning? Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.290694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.290694Z digest=sha256:b8ee1d6b1c207fcd8240fb60a89b95afab6761aa03509e40ffd68ee7009efc4b

Observation e8b51122-ee9a-4f4c-a762-8a2f07fc51e9 · inbound

PixelThink: Towards Efficient Chain-of-Pixel Reasoning cites this paper.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.144862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.144862Z digest=sha256:6100b9f4c06971d2bc6842de3d8c700dd885cc840c00b62d1c38a951658b3d25

Observation df98069d-c01a-4be8-a355-bf5788b300b8 · inbound

EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation cites this paper.

EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:49.678833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:49.678833Z digest=sha256:7c0bcb0c67f08c69316dec53b0ed60075b6440d6ccd7b2654ca7c6d951f8690a

Observation 70bd4eff-189c-4e66-ba8b-fcf6c0b35422 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.172225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.172225Z digest=sha256:bfd4066a38dec1dc73cea46c206706c091191af1f8438611aa70ed7354331cfc

Observation b527d30a-272e-4fb3-9d5c-64ecec4f111a · inbound

Scaling Test-time Compute for LLM Agents cites this paper.

Scaling Test-time Compute for LLM Agents Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:48.364617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:48.364617Z digest=sha256:4d8f64d6102e901572f50c1e01d7abf6f3a0c87db9faf668d3fc329fab948d66

Observation e1dffc48-5b3e-485f-844b-22ccc0ae2bac · inbound

DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling cites this paper.

DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:23.585240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:23.585240Z digest=sha256:3b42e0c69b380cebea192e5e4c7321096b033288413bbd82b20bedad23cd0cae

Observation 2bb8fc79-a48f-435b-ba00-1ec1647935fc · inbound

Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations cites this paper.

Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:59:14.428010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:59:14.428010Z digest=sha256:99536f625369138e43bdf5e0a9d7a7d57518fe34e0799f62e5f33197dbd8ce2a

Observation 24c08bd8-c10a-4603-a5cf-ceaf7891c0a7 · inbound

Reasoning in machine vision by learning fast and slow thinking cites this paper.

Reasoning in machine vision by learning fast and slow thinking Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:18.935562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:18.935562Z digest=sha256:eb663ddd4735e45789f1688d76741ef5abb92a188cf7f7e2b0b8ce6146584afc

Observation 47967841-d013-48fe-a5c4-a9853eeb27be · inbound

Test-Time Scaling with Reflective Generative Model cites this paper.

Test-Time Scaling with Reflective Generative Model Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.778485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.778485Z digest=sha256:0cdf8b37eead61fa0ab24cd9e9fea5a52d3a1fe89939815862b17cd457095efa

Observation 19d5599b-be23-49fd-9f3b-148432ac69ae · inbound

Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs cites this paper.

Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:11.344018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:11.344018Z digest=sha256:f4f7816a4aa55a1e2c65821b8c492dcb600278b86647ad6dd661f6910126c4fa

Observation 53431e79-949f-4e23-b2ac-162677dd400c · inbound

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need cites this paper.

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:44.141518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:44.141518Z digest=sha256:1d075666901cf7bc7ba5cfaef8c896c5b56eca8562de2a277fcaaf95d25dd752

Observation acd9282a-53af-4db8-a978-38b29d59fe79 · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:24:26.286832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:7d38a605f8a7ac0dfd3886dd9df602ffbf8c47eb9116cc17712ccb50f6e213e2

Observation 0f4e7dcf-2c7d-4e7e-aa2a-2a3989bf4133 · inbound

Thinking Isn't an Illusion: Overcoming the Limitations of Reasoning Models via Tool Augmentations cites this paper.

Thinking Isn't an Illusion: Overcoming the Limitations of Reasoning Models via Tool Augmentations Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:59.459727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:59.459727Z digest=sha256:fd966bbf7e6656d98bfa77f8d7f61ff8c050ab150f61863cb17d70b860d23491

Observation 9c1a11d3-de30-4c11-9006-604f787b15e8 · inbound

Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction cites this paper.

Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:49:56.372971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:49:56.372971Z digest=sha256:c774f455fe60735ea3446fa4f6f9b565288663d87c27939a3b42ba32021343f3

Observation 0972242c-07d0-4677-81ae-206435b91c3f · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.690804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.690804Z digest=sha256:303341291271b43274d2977ccd1191a5a8d597dcd98a7f8e25a6f2b72359b189

Observation 7bff4313-4b00-48ed-aa9e-c23e1d52487a · inbound

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism cites this paper.

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:02.266502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:02.266502Z digest=sha256:79ed03f087adb891bb782ab84c0958a5c5a88ccaedecbfc475d925fba73cca73

Observation d6267bea-60f9-4002-9d17-e79cf6229fd4 · inbound

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models cites this paper.

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:41:52.929951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T22:41:26.047957Z digest=sha256:e4208535fdfb44a956265886224ea34bec39faaca4d81dd3231af1ef2ff44130

Observation e3eaa87b-24a6-41d4-be51-6f9b9d280004 · inbound

Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness cites this paper.

Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T17:31:41.548743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T17:31:28.644151Z digest=sha256:9c9c5574485f0a780a77bd7f923393bb07ed2e82be02970309ce501ca2a1ae1c

Observation 2a58f66e-2eb4-4c11-a9b8-328ef81b1a4e · inbound

Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction cites this paper.

Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:43:41.263738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:43:41.263738Z digest=sha256:9bc6c095010b5d24102e97249f30a61b473c20210adc576bfe95891253af0a07

Observation 2a18a124-832f-43ac-b30b-6e04511b983b · inbound

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL cites this paper.

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:30:35.518062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T20:29:25.874620Z digest=sha256:1fe18aea24fbf5e739557af908e363fec2664e898df7a965e11eb234f69ee66d

Observation 92240cde-4213-4933-91e7-3328797739ac · inbound

Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling cites this paper.

Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:57:43.015774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T09:53:14.045466Z digest=sha256:c1d9d90e370a31e88eb8fe734536e56e2b622c720b9ceb3f9db5e63f28e0f4ea

Observation bbda081e-6dd4-4094-a0ad-2de16fb6ff00 · inbound

Decoding the Critique Mechanism in Large Reasoning Models cites this paper.

Decoding the Critique Mechanism in Large Reasoning Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:50:28.456759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T06:48:20.712294Z digest=sha256:c3f2abceba639c01141b9140bf376df159aa73e79da36bd72ba049ff73c30e01

Observation 54d004f6-27d7-4e08-9bce-fb0331d17a3d · inbound

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators cites this paper.

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:52.070140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:04:17.725111Z digest=sha256:04eeb65a8c577137bec29c8404d89e5d09fc76b198fbb2a67da7b390e4336bd8

Observation 126d36a2-d261-418c-8477-6cc26e49bbea · inbound

Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization cites this paper.

Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:00:22.130974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:57:44.680423Z digest=sha256:e856dc670804f8305ffcfc37b4725d3ac8bb72c5bcef702b66ee294808a43f2d

Observation e27a3869-9ed8-4ef0-9efb-add1b299e1c8 · inbound

TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens cites this paper.

TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:20:11.264650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:11:45.176020Z digest=sha256:20f4692e1915dd331f4c3339f35cbe0d3685ebcbe8fff6ad1aff06814a016d1b

Observation 7ad84d25-93e2-4f8e-b738-589f602625db · inbound

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis cites this paper.

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:15.853807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:47:34.897401Z digest=sha256:b29bfbfdbadad81b435ce2796cdbc57c1f1902accbcb90c6ff0512a6ca3a328a

Observation 3b38060d-072b-466c-b341-4eeb9f518978 · inbound

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis cites this paper.

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:15:43.492180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T09:13:20.071265Z digest=sha256:3489b0cb10b38de9e311d8262a57ad7e03155df18a0a0b18e02c9ce336f3b0e1

Observation a195219e-a949-4335-852b-361d6126e051 · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:09.074348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:3d90d85ab56e949c32148da3f92d600cfd83c30c9cede929da120f34bcf342ad

Observation 11ce872f-0adc-4370-8d30-fe73dc63e331 · inbound

Stream-T1: Test-Time Scaling for Streaming Video Generation cites this paper.

Stream-T1: Test-Time Scaling for Streaming Video Generation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:39.485808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T18:16:14.985693Z digest=sha256:acdbc259823363732f097f21e80f9d1de56e4712e4e8595202b7f420cbf9f819

Observation bf6069d9-8c6b-41c2-b8be-3d5599c9b936 · inbound

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation cites this paper.

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:25:58.231082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T02:28:20.366674Z digest=sha256:d57ebae18feb56fd700ebff5ca8c85c07f0d73632718ecabefccb62f8e5570ab

Observation caf4bf73-971a-4969-b53a-d38ec486c73f · inbound

Unsupervised Process Reward Models cites this paper.

Unsupervised Process Reward Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:21:18.916734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T03:19:04.275069Z digest=sha256:587250aa264bfc4d41ec6ee61c0eea11e838f93f495d8c6f93e0623b916c1eb0

Observation bee8b60e-40cd-4275-a5e4-987132bfa689 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:51:30.089678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:0778bfac3d92c0fa838b2f6fccd1874512a2d0791b810f76ca5a923ef07fe159

Observation 03950951-7438-428e-bfa6-432eab0f39d4 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:15:03.365580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:02c5628b9279cb1661d5215cc305d663bb735ce3329f5c4d4aa3863453221b9c

Observation 3be9bfea-ecf3-44d4-bf02-409477c54888 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.733864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:f464ea2523294d880957685b76590c88e0b473497b0c95d641c42c55dcd3f0fe

Observation 6388b362-79f2-43ae-9054-63e432943e5d · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:27:26.378161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:7b2dd597a3683b9d1fe740c1172bcfb611ec3b72aa9ba4eefd6ea54258885462

Observation 7170760e-940f-4c37-902f-f3141149314a · inbound

Data Selection Through Iterative Self-Filtering for Vision-Language Settings cites this paper.

Data Selection Through Iterative Self-Filtering for Vision-Language Settings Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:49:44.883889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T09:22:47.537137Z digest=sha256:e969f110bd3547f44666cde48aefd3389b29ab4590729ac7a324345e818ce355

Observation ec2d4e04-69fa-49d4-ab3d-60fa6008b35b · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.400710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:0e8459608e3f1cc76314cb0faa628d2431cd4e7424a44e77c32b4779e0f0476c

Observation ff9bd608-6a8d-4d92-9f84-fce98c499edf · inbound

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling cites this paper.

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:44:37.092805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T09:44:27.786630Z digest=sha256:27ee06d737720c7972e6cefab69888b6479bd84bad7791a6b0f8506f563dda16

Observation fb27912b-109e-4283-9f1d-4c4727f952d5 · inbound

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ cites this paper.

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T03:00:51.318412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:00:51.318412Z digest=sha256:9f8af218f24e0de7d88344bb3f1dcb77b8d4199b13e8cc7750df779b72d82794

Observation 479f5bd3-9708-4dd4-bb78-add836d6214f · inbound

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning cites this paper.

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T14:00:54.863716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:00:54.863716Z digest=sha256:664e4e35731d43759aa59d50be42ade2e1e1e116840760ee6912da1bf30d2ba0

Observation 1c322afa-c81b-4eb6-ae47-174fb3224b10 · inbound

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning cites this paper.

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T07:28:23.264118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:28:23.264118Z digest=sha256:106a1b2ec2b7d33da37c7b2c23257a845242e0544d9e0507eb611f9e3d6e93a7

Observation 71089cd9-0a78-4511-b41f-ac42a2359ce4 · inbound

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals cites this paper.

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T05:09:00.865375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:09:00.865375Z digest=sha256:4cea1306e4dc4f7e1a9a25900ba28df0874b702c41d8254ec10eede6658038fd

Observation 5e0c3839-c60d-4d1e-808f-c23fbca65622 · inbound

Interpretable Adaptive Sampling for LLM Test-Time Scaling cites this paper.

Interpretable Adaptive Sampling for LLM Test-Time Scaling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:10.707287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:02:10.707287Z digest=sha256:2900a8b8dd08a3f169c44112f65d3f3f3b3a7db6a0d8b3982ff5fb02b6568300

Observation 93d7ffdc-ce11-42f4-96f2-5e42db65a721 · inbound

Hidden Language Consistency Phenomena in Reasoning LLMs cites this paper.

Hidden Language Consistency Phenomena in Reasoning LLMs Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 278

Resolution
unresolved
no resolver link, observed 2026-08-14T04:40:10.392620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:40:10.392620Z digest=sha256:dd5dd6ec2c0f25044b72b382f865166d6dc336d8a68c5b1248a4cb0f7f05f5ab