Pith. sign in

Paper Citation Record · LEDGER

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

As of 27 July 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2605.01566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.01566 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T14:07:43.507761Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-27T06:30:09.085275+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T03:36:57.168246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.534976Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact5
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch9

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1d3d7c6-86b6-46d1-9156-6c8fed48d02f · outbound

This paper cites Stay Focused: Problem Drift in Multi-Agent Debate.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Stay Focused: Problem Drift in Multi-Agent Debate

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:06:03.954920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:994f2cdebff6b708321c17de2ce123d3d3418df2518df2846f82c29688712fff

Observation 327d15b7-e6a5-469b-b23c-5fb5bac26362 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:06:03.958989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:292ca37d346fc93d0ebf4d85a1604cf4a547f310acc0dd7e093b5fa79ac191b4

Observation 94f9400c-5db8-4f85-842a-4162020c0d57 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:45.301651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:89cdbb41832e232360b03c07fa5da5a61a46db8f2987962942ff93e624bbf189

Observation 3d6190e3-2adc-4199-8762-fdb03b7bf5ef · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Large Language Models Cannot Self-Correct Reasoning Yet

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:48:28.335998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:73514d16999fa1d90c9defdfda06d1e41f3a21937029f5c6f8aa574aec57a36d

Observation 11b868db-ae8d-4e5e-a1c4-d667044d409c · outbound

This paper cites InProceedings of the Fourth Workshop on Scholarly Document Processing (SDP 2024), pages 105–119, Bangkok, Thailand.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling InProceedings of the Fourth Workshop on Scholarly Document Processing (SDP 2024), pages 105–119, Bangkok, Thailand

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T02:26:30.158161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:bac69201670bd25cb3a7446d1b8efe38f0e47b20a0bb83a6ccad084799f4b210

Observation cd757e76-605d-481c-a94f-b81911b71bce · outbound

This paper cites Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:06:03.926913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:f49f1d395de9bd0ba6e233f55451c9e2bc280a53de26211678fe4662e1563420

Observation 325bfc76-8678-4d41-946d-5c2757f585bf · outbound

This paper cites Scaling Laws for Neural Language Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Scaling Laws for Neural Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:06:03.962458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:84adffaacffb66924839e73129566ae91804bec955550c7afdc7dc08f6fd0446

Observation 39c0b790-40d8-401f-9589-4db5ccafcdf4 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Self-Refine: Iterative Refinement with Self-Feedback

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:06:03.965474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:201b573ee1ef594dac491b287f911874da2568d97fb2df20259ccc84c158f98c

Observation eb943f7a-b03a-46bd-9a42-0283e59703eb · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling The Llama 3 Herd of Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:06:03.932948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:84a8ff7b6d94c555767b2af1dcefee2b9be2ba80c1491877079c318566923663

Observation 5039904a-1f1f-4379-809d-79398ce67b24 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:06:03.950400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:a171b24b5c6c7cdf997d98bdae79985d3f04c1f80aedb98ab266cdab3c8d4b4a

Observation 7552b230-5f20-47e3-8b64-b7e88137f3b4 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:01:10.451911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:38eef52d90d08ba258785e04be6ec526dd26ff706fc8337e0ebc6e15959c9bef

Observation b205710e-600d-4d3e-b26c-2923fc9e67b3 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:29:34.516338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:e2755e2fb436dd760256513a127b916ec60a7d232cb253c41ffdd93256296897

Observation 42617628-578f-44d7-9742-9f36916819c3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:01:10.436307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:a4e55c931625059569e6b28654ea769c13960b1f2f001607a3403159c95e24b6

Observation d8f3b0b5-94f1-4391-9da1-2a8cdff683a3 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:01:10.429053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:c006d6157380acbb93ae7c3f06343faa0da65a3050754073e09f27f5aa63d688

Observation 334b9a48-f8cd-4efd-8a9f-efb92cf0e7ad · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:38:37.536547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:55f9afcb2819e2619c603ef33d181ed31812c5ab2edc0d2f2d9fcd716d2613aa

Observation 69bdb243-3903-40f2-b014-874be9a2d89a · outbound

This paper cites FLOPs are calculated with the formulas from Table 4 of Appendix B.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling FLOPs are calculated with the formulas from Table 4 of Appendix B

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T02:26:30.149968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:201ffb8f548564bd9e5ec755f78211092be1d659ac316020b1b9651f6817a6a4

Observation 37cc218c-9373-4643-a7c3-e62d04852c65 · outbound

This paper cites FLOPs are estimated by assuming that each parameter is used for one multiplication and one addition per token.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling FLOPs are estimated by assuming that each parameter is used for one multiplication and one addition per token

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T02:26:30.146256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:1ce59ae6f9530e6d9fbe1e445c46ed4c7d0c24dd10c499f7e40cc07589d44010

Observation 8ca8adbe-9cf7-43b5-8ded-9f3a74a0f64d · outbound

This paper cites an unresolved cited work.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Unresolved cited work

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T02:26:30.154013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:2e6a34d7d4f2282184db7fafec345434ec0b12da3d18b5dd9a5e5cabfe953265

Pith citing papers

Observation 532cbc44-7530-4413-97c4-723818559cab · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.536724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:61c73f22e83a1ab437fd8a1a74b6a08275f87080d7ea2f8ba62939b3f74c7adf