Pith. sign in

Paper Citation Record · LEDGER

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.15388.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15388 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:34:10.768149Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 946d907e-8464-449c-b318-dbdc1bfefe59 · outbound

This paper cites How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:05.670060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:05.670060Z digest=sha256:c6f27f25f7ee4c9db3388affb80d747ce322de713b85153b6235db32d361b83a

Observation 41556648-636e-46e7-9950-5324951d6638 · outbound

This paper cites Benchmarks saturate when the model gets smarter than the judge.arXiv preprint arXiv:2601.19532, 2026.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Benchmarks saturate when the model gets smarter than the judge.arXiv preprint arXiv:2601.19532, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:05.720372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:05.720372Z digest=sha256:98697b0be3d4535b03f632cb69a7bf2846612fbf742940404bd44bada2db329e

Observation a403b750-4dea-4d47-873b-ac7062439b01 · outbound

This paper cites xverify: Efficient answer verifier for reasoning model evaluations.arXiv preprint arXiv:2504.10481, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning xverify: Efficient answer verifier for reasoning model evaluations.arXiv preprint arXiv:2504.10481, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.070973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.070973Z digest=sha256:4e7787f163f6a879f67604993c8d2dd9f56da0c163fe93f15f8a64faaa8a4f7c

Observation 846e75f2-a91c-47a2-a240-9bde89501526 · outbound

This paper cites A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46, 1960.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46, 1960

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.218181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.218181Z digest=sha256:938f8e223e40c4cee336e063bd76a8d8c31828d9b334ffa5d0edfbed879dfe2b

Observation f4893f67-e0f0-4aa1-8f5e-89e3ef673b17 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.309736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.309736Z digest=sha256:f7727244d9bf812febbe9ca2ae668f94b9f7f65114dfdc000fa2e7a61b904eed

Observation 5f3ef200-f229-4067-a7f3-498e6f885606 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.382219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.382219Z digest=sha256:968d43f0758ddddc3eba13c87089f99b1a20af92b41f6228109ea06c9a3c2b6a

Observation 21176412-77ba-4a5c-a2ab-0f877c91a028 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Large Language Models Cannot Self-Correct Reasoning Yet

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.435035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.435035Z digest=sha256:c1738943371e8a6d0d1716a2f0c323937eda53b3d1be77024235007970db27f0

Observation 30eceb0e-c75e-4f61-917c-707457da1c27 · outbound

This paper cites AI Agents That Matter.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning AI Agents That Matter

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.567758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.567758Z digest=sha256:939d1dd789d9d9ee4bb9772c215f71255ee63a7c5f6ead16bc83e39e2db5cf41

Observation aabdee1f-1d04-43fe-b11a-61957d61f5fd · outbound

This paper cites Towards a Science of Scaling Agent Systems.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Towards a Science of Scaling Agent Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.985088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.985088Z digest=sha256:677fb2b4b590579399c4a95491b5823904695fe31c8d01b96f96449b2f25959f

Observation 42ff4daa-3c5a-46ba-9b67-c49dd8300c2e · outbound

This paper cites Richard Landis and Gary G.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Richard Landis and Gary G

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.101341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.101341Z digest=sha256:50f1becd097c751f74459a2fec6229bb9b818c8396bb7d821e7dcc11b611aa92

Observation 1a7c4be0-789c-49f5-8b13-e97999bf956f · outbound

This paper cites CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.165520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.165520Z digest=sha256:af8657bb277fd427447d3cb1a861d68bbdee537fdf3a91115917c627cb89d688

Observation 341b1928-7631-4c62-bdb3-adb4cdaf3eca · outbound

This paper cites Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis.arXiv preprint arXiv:2601.00828, 2026.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis.arXiv preprint arXiv:2601.00828, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.241798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.241798Z digest=sha256:4ac1c2708b3bc54629fc6da7a876072f913f58d6564d188951848676fb996c29

Observation 70139a04-5810-41de-a13e-81fe8a44fcc1 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning AgentBench: Evaluating LLMs as Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.318575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.318575Z digest=sha256:11196254fcb6b307230699b18c4b5e6d1150c9af4f5199211890b1e6be19cea3

Observation 9beaba57-822b-4468-95b9-3594164b9aa8 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Self-Refine: Iterative Refinement with Self-Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.388941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.388941Z digest=sha256:e1c00cb33765c4781a64d7f1b018ac891c42f47f47eb0655a5e89eccf07c901d

Observation 6bd452a2-4c47-4f12-b6e6-65ba3ebb23a6 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.475071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.475071Z digest=sha256:b5927cbbcdba2b6c1bb4472b09cb2273ce5b684a87652278248361cceca381df

Observation 86a8efee-ccfb-4b45-9811-c868114bd03e · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.799710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.799710Z digest=sha256:8ec21558307e65fe855ac6792411fda0b60602e4b171db792f012531f26c409e

Observation 067ab163-1a96-4334-9b77-ddb1c0f14e19 · outbound

This paper cites On Evaluating the Integration of Reasoning and Action in LLM Agents with Database Question Answering.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning On Evaluating the Integration of Reasoning and Action in LLM Agents with Database Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.962606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.962606Z digest=sha256:8d15b2e8280d9176ea055206357a2eb31c3cc2264326c3f708a332013be2e84e

Observation 3529d686-f8c3-4ff9-b29c-b384223e5562 · outbound

This paper cites gpt-oss-120b model, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning gpt-oss-120b model, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.166111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.166111Z digest=sha256:e56090a1f61b1acbb1c32807fb1caefbb07982e8ff2180c17a1009664a8dca91

Observation 84e8b31b-0595-4817-a0af-95f4cc49235e · outbound

This paper cites Introducing gpt-oss, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Introducing gpt-oss, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.305502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.305502Z digest=sha256:b43504d0fef73e20fe6819c012d030516203451e360ac0e91e09e79940ee1d35

Observation a3640d7a-eadc-4e2f-a5c7-afa68282f0d8 · outbound

This paper cites Towards scientific intelligence: A survey of llm-based scientific agents.arXiv preprint arXiv:2503.24047, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Towards scientific intelligence: A survey of llm-based scientific agents.arXiv preprint arXiv:2503.24047, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.446253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.446253Z digest=sha256:dc833cdbeff8cbff3b425a4820c3c66958a103309419a03de31a3f6f98a077d4

Observation 1cb9daf5-7438-4253-9abd-ac506e51c305 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.642754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.642754Z digest=sha256:0e13f490a365e944aec272b90fb61939a4b7c6a533e4a73dbfda4730fb094b77

Observation 0c60c6a2-bd0e-49b2-98f6-9a1564af8107 · outbound

This paper cites Can llms correct them- selves? a benchmark of self-correction in llms.arXiv preprint arXiv:2510.16062, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Can llms correct them- selves? a benchmark of self-correction in llms.arXiv preprint arXiv:2510.16062, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.734542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.734542Z digest=sha256:43d3340a7410882d58954528ab92b7f675ad50246e613ecd7d5c4f0c071b0a7e

Observation d8195945-d22a-4014-9600-36b09c49f097 · outbound

This paper cites Multi-Agent Collaboration Mechanisms: A Survey of LLMs.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Multi-Agent Collaboration Mechanisms: A Survey of LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.868842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.868842Z digest=sha256:8538fc3d109c13f7d3583c1c33e6a742e85c9ed212b1fc9f63bd41a4aecdb133

Observation 03aecc9f-9a42-46e1-b96c-81e31afb5bb8 · outbound

This paper cites Self-correction bench: Uncovering and addressing the self-correction blind spot in large language models.OpenReview, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Self-correction bench: Uncovering and addressing the self-correction blind spot in large language models.OpenReview, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.999850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.999850Z digest=sha256:4cb35ca312a1875bb3c34e0b1ba01747182264ba9c1f94c94e551de8db39358e

Observation c5fd9750-3627-44b5-8ad7-cbd3003bf917 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.101806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.101806Z digest=sha256:446847c9b26b9e33dc9b00aac5700c5724c223e5784be3ca36acefa3c696a681

Observation bc3468bd-8320-4030-b3f9-f885ad0a43be · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.206984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.206984Z digest=sha256:5dfb4f51bd545b1f2f7083ee5d8a463241dcb33162deff7df828526bbc890065

Observation 920b6124-5843-427c-84b4-11e20deecfc5 · outbound

This paper cites Can LLM agents really debate? a controlled study of multi-agent debate in logical reasoning.arXiv preprint arXiv:2511.07784, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Can LLM agents really debate? a controlled study of multi-agent debate in logical reasoning.arXiv preprint arXiv:2511.07784, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.317994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.317994Z digest=sha256:4014d4ad816f0eacc91412419dec1008e05b4f12d1fa7dab9b6e0086942a6afd

Observation ee57ea8a-2eb3-48ca-91d1-7d5419098410 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.454099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.454099Z digest=sha256:f4e01815fa3981bfd352dfd6d258842d65a43a0c46eb3915e4cf949a14323711

Observation 0968b8c3-d473-4d4f-a221-e804f64a17f1 · outbound

This paper cites Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.539719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.539719Z digest=sha256:069983843c5f6d6ba75ac65b722ab787acf8ee4adb96ae6f44abd82cccfa3694

Observation 10043a02-738e-4f71-8cc2-46530b4140f0 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.629909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.629909Z digest=sha256:81fb2e22843eec7084813f2bfa5c10af397c1931232fd57864e3b6d3e1e4c614

Observation c10dd3e5-58fc-489d-94eb-b793c9554495 · outbound

This paper cites Survey on Evaluation of LLM-based Agents.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Survey on Evaluation of LLM-based Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.710058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.710058Z digest=sha256:b08f8337c2f04dfdad5bea097cd071c8f94d2d50244e7b13f77a4d5b6f410736

Observation 7dbf8087-3485-4e26-9fd2-531ae83c4e6d · outbound

This paper cites Reinforce LLM Reasoning through Multi-Agent Reflection.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Reinforce LLM Reasoning through Multi-Agent Reflection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.797820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.797820Z digest=sha256:a13453b87f03427ba0644411c77607fd52a3177eba7e99ff5669351749fe9013

Observation c3f52520-79bc-49b9-9ab5-59ebf80e1509 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.881076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.881076Z digest=sha256:8a3226e10f868149118a51c6e9c82d4b74ed064660c3615ae8306d181c7de786

Observation 7ccc4a33-b726-4b50-b95a-1191dffcb24e · outbound

This paper cites Small language models need strong verifiers to self-correct reasoning.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Small language models need strong verifiers to self-correct reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.964586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.964586Z digest=sha256:4a35078debd841e4c98d51e1d439118995a5d5f2212be2c58ecb25a843559dcb

Observation cf7f5c99-4b5e-4db2-9b3a-f92db4a927ac · outbound

This paper cites Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.038013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.038013Z digest=sha256:22a798f570833d4dfb10f14b9a9d5365537b1722a30dacf9046890ca7b910821

Observation ad9430a8-1636-4380-a69e-32bdca129aca · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-01T23:34:10.122008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.122008Z digest=sha256:d0f8bf3509e185dd858575dd4b01532bc9ad355a6dfaf1ea0fc694917881c793

Observation d8c4e30e-7b1d-4397-b872-0e12201c7fc5 · outbound

This paper cites useful” and “misleading.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning useful” and “misleading

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.218624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.218624Z digest=sha256:8ad32842569bf0a419d93eae4af83967b645b1fcdc960c1c0a315f930feadaa4

Observation cad18796-d777-4396-957f-386b92c24658 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.309632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.309632Z digest=sha256:0c3edcfb71386eb181ce78845054663c023b2581de1312da5261900ae73ed931

Observation e4471bb6-c271-4c7c-99c8-d05ff01e1582 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.388391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.388391Z digest=sha256:99b64f34c3d5a8cbc72e239b792cb41f6b348b5b7a858286b273d9ec70f6e618

Observation 9dd02216-17b5-4e81-b764-9de6638cc68a · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.478432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.478432Z digest=sha256:150a4ce5a50c2e5568381505ea47dc0b472086a7a8cb4c3b2ab3681698578e49

Observation 9ad57cec-5bdb-47a0-9059-68e4061bf554 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.578951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.578951Z digest=sha256:10012a19ef7b84377cd95299dd4842db750cf4b687f9ebd1299bea5bbb10937c

Observation d68631ab-0d3e-48e9-9b92-62d01dc6c367 · outbound

This paper cites wrong → same wrong answer,.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning wrong → same wrong answer,

Reference 47

Resolution
malformed identifier
no resolver link, observed 2026-08-01T23:34:10.677408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.677408Z digest=sha256:34b2d60b6ea53e539c018118dbf0ed35c90696b9a831a70cb4f0251de16d000c

Observation 8c1ad37e-21ba-4226-bfeb-c4654a353f8a · outbound

This paper cites Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.768149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.768149Z digest=sha256:73fd4012aa4a5453c831286cb501e8fae4c1886917deed955ca8b57169227e6b

Observation d3cedc27-8928-457c-9ae1-94ed70a3bb02 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.685579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.685579Z digest=sha256:39b911f109d3bb53434df8343104766fb2efc0744e14d436b74f70e4b7e76b92

Observation 30cadc61-c180-4eb7-9a91-0179bf43d24b · outbound

This paper cites On scalable oversight with weak LLMs judging strong LLMs.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning On scalable oversight with weak LLMs judging strong LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.878529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.878529Z digest=sha256:2b8cc9d56cac4e53d4247dd34557ea61f7928f3256464de217ee2c06b3145bfb

Observation e508252e-3f21-4cb6-bdd4-1aad689f6499 · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Why Do Multi-Agent LLM Systems Fail?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:05.935527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:05.935527Z digest=sha256:33b9be595e21050bbf39c3e765e32fccd4703047c39278d0c85fa7e0f4a7a47e

Pith citing papers

No inbound Pith citation observations are available.