Pith. sign in

Paper Citation Record · LEDGER

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

As of 16 August 2026, this Paper Citation Record lists 100 of 283 outbound references and 0 inbound Pith citation observations for arXiv:2608.04001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04001 v1

Coverage vector

measured 100 of 283 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:49:04.602845Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 283 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved95
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 716c193e-8f4e-4fe9-b004-ee536ac34a12 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Curious Case of Neural Text Degeneration

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.053337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.053337Z digest=sha256:f176f1be8b2525994bc883aa70d96d452eaef0c9478c96f6ece630b80ab292d1

Observation 574b94b4-bd18-42be-989f-9f0fdd92393f · outbound

This paper cites Optimizing Large Language Model Hyperparameters for Code Generation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Optimizing Large Language Model Hyperparameters for Code Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.064748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.064748Z digest=sha256:088aaaa152b7b20e97787d288fd95acfc1f270206a4079a5897cbf43edece18f

Observation 2ec0beb8-9236-414e-bf08-9e8fec269b03 · outbound

This paper cites arXiv preprint arXiv:2407.01082 , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility arXiv preprint arXiv:2407.01082 , year=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.078009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.078009Z digest=sha256:4bbcca4172b860d13d410f60dca3b80631cff0553e5eb568826e9b6d89b28593

Observation 5211dacc-8084-4023-b3be-f4ce55c42205 · outbound

This paper cites Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.090744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.090744Z digest=sha256:2444c086f479fe28b4390366a0a320664c9ff4237b7ce3d5903485a979a274c0

Observation cfe11679-bd56-40d2-8b17-5901f17e2972 · outbound

This paper cites A Thorough Examination of Decoding Methods in the Era of LLMs.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility A Thorough Examination of Decoding Methods in the Era of LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.109318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.109318Z digest=sha256:0ef31a6bf19c4610c101c33f30894ad461f29b11ba602ad0d05a530dd7474655

Observation c69370a4-72d4-494f-9e74-37073bd86801 · outbound

This paper cites Closing the Curious Case of Neural Text Degeneration.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Closing the Curious Case of Neural Text Degeneration

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.120745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.120745Z digest=sha256:461ebe3dd0f86ed3f4c4e122dd3e43c58c475b12d9f7814082b2751bcf59e8bf

Observation deec60cd-dbb6-4a14-8959-22d97b844292 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.142677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.142677Z digest=sha256:353ad5e24f3ee3a226aabef4adedf06da68b193043cb8cce58d76145ead0c490

Observation 1b8a3d81-e2c2-4880-a60d-ad3efc81c50f · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.161173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.161173Z digest=sha256:5fc2454935cecc33f8533b8897f07088457376b2eff4bf98e026ddaa3776b926

Observation b887cf36-7211-4919-810c-389fa2787d5d · outbound

This paper cites B leu: a Method for Automatic Evaluation of Machine Translation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility B leu: a Method for Automatic Evaluation of Machine Translation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.174835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.174835Z digest=sha256:dd775ef156411b669016e31c9c00637e7240f0bff794fa8ca65eafb312ed6371

Observation 3f1b430f-64f8-4a92-a944-e457a4588547 · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.188195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.188195Z digest=sha256:d57c512f9b9fc38b18f7e5c035b181608793f22dedb6f30f8db6a937e6d1bff4

Observation b723fd48-a3b5-4b29-9a75-ec7392e48ea7 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.205102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.205102Z digest=sha256:b2ffb6bbf26eaae08b5a632f97fc0c43242622086dcd24eef1ade1663fd626a2

Observation e0020c01-3cba-4f4a-9b7c-296a91bfd21c · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.240825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.240825Z digest=sha256:b727730a448517315ce3892fc5cc963372f1b2c207be81cacfa011403263e17e

Observation 3863b5a4-b9cc-40b4-abc3-35f5ece7cf82 · outbound

This paper cites Advances in neural information processing systems , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in neural information processing systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.253273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.253273Z digest=sha256:d079bd36891f8895aba5eca7f00a71dc58820a7e9c1ce1b11c5ddede1c2f61b6

Observation b0d85b71-c52f-4c4b-883f-78f86182d36f · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.265846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.265846Z digest=sha256:20d24a58a71cb0f8ac65fde4cd785cfc9be9690408c051805f275384c8cfd6df

Observation b83f2adc-29e9-45d2-87ac-e8ed9c14eda2 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.276331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.276331Z digest=sha256:7a40a707d354369c2df0dfb5262a1070b67d3ede99f708bef36864b6c9d8bc88

Observation 4fce2481-a609-4069-9706-5fbf63403382 · outbound

This paper cites Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.288313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.288313Z digest=sha256:3f560be27eca9d71889696c74a83da9246ee04a97b6a117debc42b0c9473d4b3

Observation d741da6b-030a-48f8-a57b-f465dc501039 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.301220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.301220Z digest=sha256:cbd4cdbda14be25cdc9820e2bf9f67a7d59df6f0cf5e01da6e389e837daf945b

Observation bcd421e1-7f19-4c41-a407-3ff5fa179cb6 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Fourteenth International Conference on Learning Representations , year=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.316125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.316125Z digest=sha256:3027f6062dd08297cc13fb46439bed7de3911c5a16475b958561c779ac4de4d8

Observation 535e6139-6e96-4916-b79b-def2626253b9 · outbound

This paper cites The Annals of Statistics , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Annals of Statistics , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.341311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.341311Z digest=sha256:89810dd384b405f4911f42f659864d67b65030191ec8e8f6e21e7602466bed56

Observation 6a62d4e9-d075-4824-93f2-a765cf32bd0a · outbound

This paper cites Forty-first International Conference on Machine Learning , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Forty-first International Conference on Machine Learning , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.353746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.353746Z digest=sha256:d0664512190e8a0426bb27bbb66467f2a2884b6a0fe0d97af93e7b3ac82dc4fe

Observation d044c412-1ca5-406c-b8dc-a71f4dfdd7ef · outbound

This paper cites A Statistical Framework for Ranking LLM-Based Chatbots.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility A Statistical Framework for Ranking LLM-Based Chatbots

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.364792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.364792Z digest=sha256:621a1e62b3ea51670df072d90655221ca4d8a30bb18d234f8727e02a9aeb7571

Observation c035fa7d-3389-46df-8c86-5a2f9fbb9662 · outbound

This paper cites Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.377159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.377159Z digest=sha256:adf0d8631d352bded20d374bcc061088088d92e42b1a54eef3874b1a14d0ad13

Observation deb6767e-a044-4401-b2cf-9598a2fce834 · outbound

This paper cites 2nd Workshop on Models of Human Feedback for AI Alignment , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2nd Workshop on Models of Human Feedback for AI Alignment , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.398215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.398215Z digest=sha256:a3469dc0b379126ea1697c59652e035a02c1ce4c6a8e366f117e0c31d92524cb

Observation be1aded2-6283-4387-927a-e5c1f6e35417 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.417664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.417664Z digest=sha256:4cd5f9b9efb0969b4c86125795d3c325978b9e8aa3897c9dda51e93c1db6ba8d

Observation aeb66974-4ffe-4bb5-818c-cc2bb35510fe · outbound

This paper cites Improving Reproducibility in Machine Learning Research (A Report from the.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Improving Reproducibility in Machine Learning Research (A Report from the

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.426647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.426647Z digest=sha256:03d198fbf2499bd68c926e90b132cf42591bebb651a98bd35c42a81b7f69fe7f

Observation ad7954dd-7d55-4389-ba08-db58a481ca6e · outbound

This paper cites Proceedings of Machine Learning and Systems , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of Machine Learning and Systems , volume =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.448648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.448648Z digest=sha256:588ebc13e54aaffef52acb7a25dd2d264f125e65a954cb1a9ffb830e69ae022d

Observation 7b8f2e42-ac33-4335-86d9-61c10e30b6d9 · outbound

This paper cites Second Conference on Language Modeling , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Second Conference on Language Modeling , year =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.468322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.468322Z digest=sha256:f57544ec807604191b34ecd72195a47c2e52e6716f43ded42d2c0a6078d1a5c4

Observation fb471102-fd09-4971-ba7f-609d09589d2d · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.476117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.476117Z digest=sha256:5f9351d627245697742d0889e1babc8f6abd4320eb28666bacbad1283c37a40c

Observation 21dfafae-aaf9-41de-9244-23aa91f4628d · outbound

This paper cites Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.483178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.483178Z digest=sha256:e3445c0e5c072fd4c6d18b50621c2712b9f8b1d43faed65239b28620efb6b38e

Observation e4f141b8-ff04-4dae-b1bb-35acbed02698 · outbound

This paper cites , booktitle=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility , booktitle=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.502063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.502063Z digest=sha256:2ab12f6d5fe9b6341f521bdc6aced20711a35ea8b782b9f7411b276e973d9c66

Observation 15c95db8-f6c4-45c2-aabe-cdeae021d0a0 · outbound

This paper cites 2025 , archivePrefix=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2025 , archivePrefix=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.508897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.508897Z digest=sha256:0ea5b399325519c0c23c0f594547a491e39a18a7d3fa0fbca8a1089f5301bad8

Observation 744ab445-8875-4d6c-af4f-b3b486e8406c · outbound

This paper cites Transactions on Machine Learning Research , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Transactions on Machine Learning Research , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.518460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.518460Z digest=sha256:44a71b1c1a5e114ec8020264861b440750fdd488edeee5dbaa39b227cc004c4c

Observation febafdcc-4389-4721-8159-ec2b15fd27bc · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.530899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.530899Z digest=sha256:714e8e0efc4e0956afd81c2db43f7a2923641f92860fbc6f043ba6d77827c080

Observation 2da06781-b7d6-4afb-bed0-80fbdb395e4a · outbound

This paper cites International Conference on Learning Representations , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility International Conference on Learning Representations , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.538941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.538941Z digest=sha256:c8f5a7b0f24c6f0b1cf20deb36fa8d821ce772345d7021af7621160db8b3c27b

Observation dfabeeec-4c98-4079-8907-19d85a39c81e · outbound

This paper cites Nature , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Nature , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.557934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.557934Z digest=sha256:61c6ab7fbb67a8ba34dc700bf447cc6e8eb4ab8f64f65f84cc6ddecda607d627

Observation 859ed12d-2a8c-4e0f-95e8-38b563a9cc09 · outbound

This paper cites International Conference on Machine Learning , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility International Conference on Machine Learning , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.571013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.571013Z digest=sha256:5628ae462a779b23a0f3113f7720f4e3dfdc8b475c0384b4c4fea0322ce208da

Observation 140c94f3-3238-4563-8827-73dfb990a5a2 · outbound

This paper cites International Conference on Learning Representations , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility International Conference on Learning Representations , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.583744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.583744Z digest=sha256:9cc47549096b6bc62085c0f17120eb04de8f7820637132817f095ac80380d998

Observation 5e5f0879-9b08-4e83-8dca-8a90ad9890d3 · outbound

This paper cites First Conference on Language Modeling , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility First Conference on Language Modeling , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.594268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.594268Z digest=sha256:a7e224f0a12eddcac04c255df69b770b6ea43f266247e31c4f0aafa5aacd78f2

Observation aafa9cc9-6c4f-4924-80b9-bbf6fd078530 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Instruction-Following Evaluation for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.604718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.604718Z digest=sha256:4e113680a8918f274ac60490e4ae634a32143ace2f07b07973eea00514589300

Observation 4ae6fece-e2b3-4f0c-b451-ee0c051d3c09 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.620654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.620654Z digest=sha256:d2260c1ac7b319b5d7715f4e8d79523acee762703ee28c9c696706b3857b6feb

Observation b91102af-3bb0-4ed8-8aa6-b13203c9c064 · outbound

This paper cites MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.637202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.637202Z digest=sha256:173b0cf2d42d1236ebe8cdb912d8cdc2fbcd31435355eb1dfd260d529a0eb690

Observation 465eeb61-4db3-405f-9654-1203bd1759fc · outbound

This paper cites International Conference on Learning Representations , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility International Conference on Learning Representations , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.650955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.650955Z digest=sha256:1a03a30a434d38fe71ac16c3a8b6c7b9c2d8e63ea4da0addc81ce0bd3ef259e1

Observation a43964eb-31bf-490f-94d5-0a9d68e2461d · outbound

This paper cites arXiv preprint arXiv:2602.10367 , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility arXiv preprint arXiv:2602.10367 , year=

Reference 51

Resolution
verified exact
doi, observed 2026-08-15T14:49:10.171236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:49:03.665804Z digest=sha256:1d0f2c36940b940e1abd78a62fa2b8fa96408f974cc8fb8b82668b0273d4c6e7

Observation cf1eea6e-0a93-4652-a4a6-1969a1400f5e · outbound

This paper cites Challenging.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Challenging

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.677534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.677534Z digest=sha256:1be993f5db661c1863b6f7c1e95f03ce8cd2962f1c44e058980cde09b0cb5435

Observation 187f9a1b-360c-4700-9c77-76e077d89bc5 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Measuring Mathematical Problem Solving With the MATH Dataset

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.690322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.690322Z digest=sha256:f55f0c9b87dd28d81aa8e59cdde79a633498374c70ba9cd62a4e6e96b78aa719

Observation 7e4a299d-4bec-466d-ab03-e131870d2d61 · outbound

This paper cites 2025 , archivePrefix=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2025 , archivePrefix=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.705214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.705214Z digest=sha256:df11acd8828ad37e053c1184242d2e757ec3abaa9c8d08a1f8e7bb7cdfb86cfb

Observation b7df270b-d100-489a-8eec-9c32b809d689 · outbound

This paper cites 2026 , archivePrefix=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2026 , archivePrefix=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.714669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.714669Z digest=sha256:2b6c3feffedf23ed9bfc51a314da3cde3a084370b43fc7e9a07f56346c5dfa96

Observation ca48349d-395d-4946-98d5-122b2d80fe5b · outbound

This paper cites 2026 , doi =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2026 , doi =

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.722278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.722278Z digest=sha256:1b2d37556d5ff95d6f976a89fd98d0bb5276c14ab3d9b5e764d8fce164ab1142

Observation 7b8a6c81-ce65-4b6e-a8b4-fddee27edba2 · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.748051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.748051Z digest=sha256:b8e4e9ea109883bb6d3f8a9d95f0840321317b45f56a83ead0efe84b0508632e

Observation 31dc3ccd-c9b6-4ff9-ba90-214a4f4677ab · outbound

This paper cites Qwen2.5 Technical Report.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Qwen2.5 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.759711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.759711Z digest=sha256:651f8ccdde488674901d3c649e6c28562c543556f2f6f3ce7619a629381d51ef

Observation 8afa8fae-5397-41e3-b0b5-aac3e2051f8c · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.772795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.772795Z digest=sha256:552a4ad8835926a5fc51fd6e96e5fa6d554074bdbb01fb566bb90cd12a634ff5

Observation bb3c1410-5ba5-4936-98cd-b4636762f556 · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.789033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.789033Z digest=sha256:b36de5a89218979b6d8f99a0c915c51fd28aa8ce2fc3c6338f33975185e5aa83

Observation a6195a13-6e25-4577-b6c2-23e65d0920ae · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.810806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.810806Z digest=sha256:ed297e53d45b6c980eb266288729e11141f0c1999f6fc095fe13d76f6fa3fdc6

Observation 9897b4f5-f0cd-4f79-8b44-358fd6a9cffa · outbound

This paper cites S*: Test Time Scaling for Code Generation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility S*: Test Time Scaling for Code Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.822238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.822238Z digest=sha256:32765a125de9c0ef43aa5ac72aaa4375da79f9b22808214a781d84489d1a9f67

Observation 9467578a-49db-46af-970a-74d3768ec7a7 · outbound

This paper cites LIMA: Less Is More for Alignment.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility LIMA: Less Is More for Alignment

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.850897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.850897Z digest=sha256:2b3d5277621db0ccc9a85afc341f5988aabba31cbfaf83125852c0a791a7a281

Observation a8024bfb-8703-4df7-be90-c56f4140dd78 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility LIMR: Less is More for RL Scaling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.884561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.884561Z digest=sha256:ca2d19729194b09bc66cba495a4196fc5aba8befb5b96afd405379f2fbcc7d1a

Observation bf26976d-e44e-426b-a78b-5bcac61826df · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.892518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.892518Z digest=sha256:d94592905fd0f6c13aabeacd5d4b8da69b876b48d95bcbe125ad4bb525714421

Observation 2aeb59e9-da0e-4608-9a67-7518b7e2697b · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.900136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.900136Z digest=sha256:7bc3345c1bd390547a8d251198f7a934951317d513c142c24c8da45782733c17

Observation 380d3236-b9ad-4d63-933f-42094778f7fd · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.912354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.912354Z digest=sha256:a83bfeb67153c4005d14c21118f03b462bdd0b9c7d4113b5e7f3a876b69179de

Observation e023607b-b724-4b2f-93c9-1862ccaa2d65 · outbound

This paper cites Knowledge Fusion of Large Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Knowledge Fusion of Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.926672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.926672Z digest=sha256:5aed3bca454c32c15246bcb9e232f752f9e5f2ddf0bca8595b34fa69b2594925

Observation 332bdf63-fe15-4976-9ed2-77e40cc2b11d · outbound

This paper cites FuseChat: Knowledge Fusion of Chat Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility FuseChat: Knowledge Fusion of Chat Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.942405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.942405Z digest=sha256:24259b8a124405c1fc6f87243d3ea4a5fba8f81f962fe1afc9b3034fe945ec13

Observation a3bd99cc-ceeb-449e-981c-92e22641e042 · outbound

This paper cites FuseChat-3.0: Preference Optimization Meets Heterogeneous Model Fusion.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility FuseChat-3.0: Preference Optimization Meets Heterogeneous Model Fusion

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.966561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.966561Z digest=sha256:41e0f9c0b1af4c5e500ceb635b342718ea90814683a4068beec7e853c5813e1d

Observation 1eafb797-f179-47da-97b1-418540031a12 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.983309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.983309Z digest=sha256:aa2113c07b0b27ec38411157aa3d756251e66f9bfc8249d893f53a0d52ec11d5

Observation df092afa-de0a-4486-af45-a42e10b4aa6a · outbound

This paper cites TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.000351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.000351Z digest=sha256:30ae963cf7babe6258bf70f59214c200cd6147ebe09f6b330ee6648de73b2351

Observation 809b3a86-bf6f-4c0f-9fce-ae93cedd94bc · outbound

This paper cites The Twelfth International Conference on Learning Representations , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Twelfth International Conference on Learning Representations , year =

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.026475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.026475Z digest=sha256:3e37455f409db62584e3a4cff7fe0f3a5eae42807fa4f4bc1bb358bdbdbbd891

Observation 478271a6-d874-498f-87f3-4be5a1de9795 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Fourteenth International Conference on Learning Representations , year =

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.036564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.036564Z digest=sha256:c0578d42adfdd4fd317c8f0380c82b76fc3a2c5c8c0905b3fb38bc038fb88732

Observation 2ce0b4eb-d728-45e2-9ffb-879216c3baa0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume =

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.052419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.052419Z digest=sha256:7cb08d0adfd529b8477f652baf42cfaa1e9dc0d17719dd3a114ca8ad95db1f5e

Observation fdfb8f6c-12eb-4cea-91eb-3be546c48d71 · outbound

This paper cites Optimal Aggregation of.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Optimal Aggregation of

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.065643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.065643Z digest=sha256:0d03397205807aef3f3f38de5dcfb37e83d5d950b86b5f54ac6e79e4de53bbef

Observation dacdbd11-2b29-464c-8d42-666e13e8adb3 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume =

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.073304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.073304Z digest=sha256:ef07c1f0956db9a646c3e0dcfad8e0641a46904bf71d73722b3e4dbe8c196eb5

Observation 125c9ea7-6e1c-4a21-ae1f-7eb909cd7fd2 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Fourteenth International Conference on Learning Representations , year =

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.084816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.084816Z digest=sha256:08a7b8c563763c6e333c0d0f61b6d450984a03e2e8be27c89cdeec7171bccf71

Observation 77f8467b-9cc6-4d44-b7ca-fd51bec82f1e · outbound

This paper cites Transactions on Machine Learning Research , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Transactions on Machine Learning Research , year =

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.100667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.100667Z digest=sha256:6f249de5841a8917956aefc1e613081dfc573f15bbc77c31cb8f27bf053716da

Observation 71f7c46e-1a9b-40fe-906c-5556af1ea1ae · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning (ICML) , series =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the 40th International Conference on Machine Learning (ICML) , series =

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.110984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.110984Z digest=sha256:43960aa1f836fc80c8f85e2b295a2af42340ebc0f0c41ce31401ec2415deb938

Observation 5387bdf4-bcf9-4feb-a438-ad11c386f66b · outbound

This paper cites Proceedings of the 42nd International Conference on Machine Learning (ICML) , series =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the 42nd International Conference on Machine Learning (ICML) , series =

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.136460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.136460Z digest=sha256:2c7bf59980b0ad28881941144048a9f6c890c2cf40871480bd5b99496990a8ce

Observation c56500bb-cc81-4324-9519-2830e00b443c · outbound

This paper cites Soft Best-of-n Sampling for Model Alignment.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Soft Best-of-n Sampling for Model Alignment

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.163654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.163654Z digest=sha256:090b568d2a8b9007c75d863ce9c4255f8d1c971a4122f17422e33119a59131ce

Observation 92d5ee9e-8602-4389-b1ed-c0cb7aed0ca3 · outbound

This paper cites It's MBR All the Way Down: Modern Generation Techniques Through the Lens of Minimum Bayes Risk.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility It's MBR All the Way Down: Modern Generation Techniques Through the Lens of Minimum Bayes Risk

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:49:09.925826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:49:04.177578Z digest=sha256:bfd64f66c4a3d1b0ac90a22a77f36d888112150940b026499256323b56745468

Observation c673f3cc-00dc-4e2c-9e02-904c41bb1739 · outbound

This paper cites Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL) , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL) , year =

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.188855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.188855Z digest=sha256:1ab7b44893972c26fecd29e26344f7dcd344c21b4c0e460ead061cb519b6690c

Observation d436f7b9-b20f-4fee-9250-0b5715addc7e · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Transactions of the Association for Computational Linguistics , volume =

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.197812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.197812Z digest=sha256:62504017042555518e6b905df70573dd36035e2b19acce00007cd01418830f6f

Observation 01295ef8-cac2-4f6d-9a0f-264133eb254b · outbound

This paper cites Better Instruction-Following Through Minimum Bayes Risk.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Better Instruction-Following Through Minimum Bayes Risk

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.204164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.204164Z digest=sha256:a980bce764163f071f557920ed4471a30f4367ab4166d95c96b82fd952bd104c

Observation 07c3911c-76bf-493e-8842-6f83b1d2d69e · outbound

This paper cites Faster Minimum Bayes Risk Decoding with Confidence-based Pruning.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Faster Minimum Bayes Risk Decoding with Confidence-based Pruning

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:49:09.780154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:49:04.212031Z digest=sha256:a1007b7c4d02b89c7c9e129933be0e9189576615e86c43b74bdec11d8ac3d939

Observation ef98f7cc-dce0-4a33-82d8-643987a11e78 · outbound

This paper cites Hyperparameter-Free Approach for Faster Minimum Bayes Risk Decoding.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Hyperparameter-Free Approach for Faster Minimum Bayes Risk Decoding

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:49:09.689468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:49:04.223920Z digest=sha256:eb8969a7c32074ac8ee02c8178318a69ef0f1ada3c005b1a80b447b6f34b0804

Observation e659d0c6-c363-4598-b954-9181caf8ad6f · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.267671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.267671Z digest=sha256:6f3b2112c123f27bcf58e80614f829bb0b03131639b9e8a358d6f47f7c1e604b

Observation 6f644802-bea2-40f0-b29c-f90a475d5cf4 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.283055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.283055Z digest=sha256:906d5efe6920c38ac1b456d88220f41669d6ca3cbd5fc77ebcef91944c338b94

Observation 04540a5a-9b77-48e7-973a-0bae09f5332b · outbound

This paper cites Universal Self-Consistency for Large Language Model Generation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Universal Self-Consistency for Large Language Model Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.294789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.294789Z digest=sha256:7071a53b53d6aad779a3e7f75d3c672a65581426db4105b261bcb0ab3112d417

Observation 5203ef2c-225a-4bd0-83bd-a693844b4bb2 · outbound

This paper cites Ranked Voting based Self-Consistency of Large Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Ranked Voting based Self-Consistency of Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.410743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.410743Z digest=sha256:52a78613281f43b5902bfc97bb2077cd8b80a845cd322590a888aa83221d3c51

Observation 4bf09981-9f9c-44a9-8920-c77510fc6931 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Training Verifiers to Solve Math Word Problems

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.448798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.448798Z digest=sha256:255c8ed026ede8163e2df35d3a87449c6c4a7f578e8f5bd05819f5c6770a0afe

Observation 982ad22c-feed-4b18-939b-2f880fec4c88 · outbound

This paper cites Let's Verify Step by Step.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Let's Verify Step by Step

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.469228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.469228Z digest=sha256:07b1c88a30aa1b4e99d6c4c489a1bc094e8097d40c92709f6e6d961076359b3d

Observation 61a36d49-167c-4402-8516-20f0e84765a4 · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.482328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.482328Z digest=sha256:89324e02e2fe22cc142d23194b6e4b5a29a4b9750d42cfe0473a822e16cc709c

Observation 18fdb6ed-7ca2-453d-a9a1-399b5c03c232 · outbound

This paper cites Variational Best-of-N Alignment.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Variational Best-of-N Alignment

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.493890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.493890Z digest=sha256:84cd8ff0feda85439522045b1d96bec43043c0ad81f268abc4f547f11f8ed419

Observation 9c4f0e19-75d9-4289-b4e4-913bfb9ff9c3 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.507240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.507240Z digest=sha256:1ec9eb46c6a0228f3798f1bcd4b87d2ba3aa9462e32170471e9d68ea5a786473

Observation de297c5c-2d83-4dc2-a45f-fdb00483a7e6 · outbound

This paper cites 2024 , month = jul, publisher =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2024 , month = jul, publisher =

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.514114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.514114Z digest=sha256:cc33f180dbf7712cb79c3f45b4d3fe08f00d210b4dd1ef3c2be6064c0f989b34

Observation 7b9ed27f-c2d5-49e1-886f-7968f8de9626 · outbound

This paper cites 2024 , month = jun, publisher =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2024 , month = jun, publisher =

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.520452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.520452Z digest=sha256:abbab1c80c9b9d51970f4330009733b43fe140a951ee096435e1bd5c4fa7c934

Observation ea09fda9-5d18-4c4a-b8e4-1fff5b46108a · outbound

This paper cites Policy Guided Tree Search for Enhanced.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Policy Guided Tree Search for Enhanced

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.528030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.528030Z digest=sha256:7fcabfe799a49ccf7044b1b2d3d0eaad6e62d05c6aa379b65c13d98a64e9a8f9

Observation 88595a3f-38b0-42c4-85c3-636c7d31c86f · outbound

This paper cites Transactions on Machine Learning Research , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Transactions on Machine Learning Research , year =

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.536806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.536806Z digest=sha256:ab9215cdfb31e05d6244323baa5cc114d892fabb4e29ab37e32b956748265180

Observation e987dfd2-6f6a-4940-8dec-f459990e9a28 · outbound

This paper cites 2026 , eprint =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2026 , eprint =

Reference 107

Resolution
verified exact
doi, observed 2026-08-15T14:49:09.041995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:49:04.543603Z digest=sha256:713e400a34f43f813d8da12b9dd4c3506a77a1b4b7dbd747dfe1f9bbeb78b4fd

Observation 6b0eb9fe-4132-4c92-9fd7-a1b39b19e39e · outbound

This paper cites 2026 , note =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2026 , note =

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.550606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.550606Z digest=sha256:0d086551864b4276a02c2f62beb23872f9793b9ceb6abb4e462facd45a665fba

Observation 70cacda5-6998-49f4-ba8a-15bb4988c998 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , year =

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.558411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.558411Z digest=sha256:c25a3c73b3190254d07056f50ee024780c57c9e0e509f88bffc8d53128e1ffa0

Observation 28e6f3fe-ca38-4975-81f4-a9f712a87b4a · outbound

This paper cites DEFT: Decoding with Flash Tree-Attention for Efficient Tree-Structured.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility DEFT: Decoding with Flash Tree-Attention for Efficient Tree-Structured

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.567845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.567845Z digest=sha256:dfe62bfb514d951ac11605e47ddcd121542c24ed03615e5e44879d62f1e61882

Observation 73740463-84e4-4c4b-9342-7730f1b5dce3 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , year =

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.576930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.576930Z digest=sha256:d11b1aa38f86f09149cfb1ca3a61105f50dc72aacf67a1a74e415b590d89fe12

Observation 021f8166-9156-4f2d-a5a3-4a5afde8ccc2 · outbound

This paper cites Bandit Based.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Bandit Based

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.587472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.587472Z digest=sha256:e1582daa29fa5093e636bc8d59ccbfbad1da88ac342768217fc745bffd587591

Observation a2bcafe9-21f7-4526-bda4-f0ce31176a44 · outbound

This paper cites and Powley, Edward and Whitehouse, Daniel and Lucas, Simon M.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility and Powley, Edward and Whitehouse, Daniel and Lucas, Simon M

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.602845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.602845Z digest=sha256:d9e48436936d6ff896b1a5e5d25d3573bfbff185efec68538cc58bdc1f06a797

Pith citing papers

No inbound Pith citation observations are available.