Pith. sign in

Paper Citation Record · LEDGER

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

As of 9 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 6 inbound Pith citation observations for arXiv:2505.22960.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22960 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:45.317745Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T17:15:01.617748Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:09:33.482951Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50165dfc-02ec-42c5-b4e2-4589e29b9d64 · outbound

This paper cites Critique-out-Loud Reward Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Critique-out-Loud Reward Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:41.917221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:41.917221Z digest=sha256:68cd60163997e86a24b2528b9bdd2b20fd97e48572c642b8ad74077d32552ecb

Observation f7bcbcfb-5cb3-40a2-a7fc-a9c9d564f61c · outbound

This paper cites AIME Problems and Solutions, 2025.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness AIME Problems and Solutions, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.974827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:41.963916Z digest=sha256:b6d68e02271834b99cfd6b4fe572970e96dd831bad58b093b59acb41aafeac27

Observation a3736bdf-ec6e-4833-ac82-cf84d1106980 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.027532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.027532Z digest=sha256:18c57cf54f8ba365f9e323a4ec02ee0fda58a9ef945c4240a2ec3dcee469d8a1

Observation eb61c74b-3ac3-42f5-a418-7f9e56dd0bfa · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.095621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.095621Z digest=sha256:ea3ee949ce4140b8f1fe74a07ed962a597ebd49fa44ac2c586ad4784c21151b6

Observation 7b8ba3e2-d821-4e8a-8e6a-24cf24afd817 · outbound

This paper cites Reconcile: Round-table conference improves reasoning via consensus among diverse llms.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Reconcile: Round-table conference improves reasoning via consensus among diverse llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.822189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:42.153098Z digest=sha256:5f175837613461b4d143d09cf1a95ca25692379621343c8831165d267b7e29b5

Observation 24d4a73d-fccf-4b25-9b77-7150adbbbef6 · outbound

This paper cites Combating Adversarial Attacks with Multi-Agent Debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Combating Adversarial Attacks with Multi-Agent Debate

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.213315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.213315Z digest=sha256:4f7a9918fb29f3ea32f0b512d87e9d469e1045fc133bd0226f4a6a0437854009

Observation ef59e14e-97b5-44fc-83c7-1e88397b29d5 · outbound

This paper cites Enhancing LLM Performance Through Debate: An Empirical Study on Multi-Agent Debate for Coding Tasks.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Enhancing LLM Performance Through Debate: An Empirical Study on Multi-Agent Debate for Coding Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.317472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.317472Z digest=sha256:89652b9f4eaf3e9c85a87edb9f34fd51f8aaa57ecf1e4c7d70a2048388b566c8

Observation 0da4f06b-ae12-41e0-b5d5-8edcd3c37a3e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.391929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.391929Z digest=sha256:11bf8781f5a8f604cbd0e5018abb7f912ce74a847814493fe0c6f1406c2f9432

Observation 51a40f0e-8052-4add-906c-6f89b7618adc · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Multilingual Jailbreak Challenges in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.474919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.474919Z digest=sha256:2154e96b09cb9528d83879d4197ae71092ec44170d1a156a95877752fbb69704

Observation 157f11ae-88c1-4fce-959d-54bc007b4d75 · outbound

This paper cites Improving factuality and reasoning in language models through multiagent debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Improving factuality and reasoning in language models through multiagent debate

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.617588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:42.545982Z digest=sha256:2d6783c89d2ea5d1e24afa302a3e698f8f9b61f9e864268d8c26b0e9e3742be0

Observation 71e6a7c1-d19e-4968-9609-84f644783179 · outbound

This paper cites Multi-LLM debate: Framework, principals, and interventions.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Multi-LLM debate: Framework, principals, and interventions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.448146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:42.619790Z digest=sha256:35cf47a5cfecccc86b42f99f89eb09127f8eda863dc1b1e7cbb11e4e8af2be4b

Observation 9fa65451-ec23-49c2-b288-95a87761aeac · outbound

This paper cites The Llama 3 Herd of Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.686078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.686078Z digest=sha256:f50f6d912b57bc7861bd466d1c7904b483ff37f1b8f8e3cf997bb220cb629e2c

Observation 28101e25-432d-44fc-bff6-3da30de146ba · outbound

This paper cites An empirical analysis of compute-optimal large language model training.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness An empirical analysis of compute-optimal large language model training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.781403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.781403Z digest=sha256:faf804c7aae5b5a7da2a7c1c8a5d1f8b73cc53d7199f1466eb2421180555d77a

Observation 7cc32020-f4cf-4c49-9778-ccfe62600de2 · outbound

This paper cites The curious case of neural text degeneration.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness The curious case of neural text degeneration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.205048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:42.875033Z digest=sha256:a3d4249410634b2e339f5e4aa225507678c7c3fa01b1f1ec67c7780c30eb8b5f

Observation 91b8e223-e1c5-4e6c-bfcf-6fe735e1549b · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.962354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.962354Z digest=sha256:7e57c57679a7df054909cc908059fdbf7f01158bb79c587252331d893eb93bc4

Observation 11a8cff6-5d82-4963-a225-0b949976c3b5 · outbound

This paper cites GPT-4o System Card.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.051383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.051383Z digest=sha256:a4683841255a66f80c710b91086018dee01b8073240597301de8fb3327e5bcfe

Observation cb22c870-4a60-4254-a16f-2110294f554c · outbound

This paper cites Scaling Laws for Neural Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.128633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.128633Z digest=sha256:83ef3cf92a0595b36cd70018e30604255b106a7c7e4cd77a19947aabcb8a6ff8

Observation b8eb9045-ce00-42b4-8d31-5e0cb97648ca · outbound

This paper cites Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.211841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.211841Z digest=sha256:d779e35588d1eeb68ea75da74113031e73bce95c0d51114ab5d2a2f04c3377a4

Observation f4ad240d-ab0d-47f3-89c7-cd5e72f2eb31 · outbound

This paper cites A Simple Model of Inference Scaling Laws.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness A Simple Model of Inference Scaling Laws

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.310559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.310559Z digest=sha256:fa1ab4951a9fb8ffe6112d3cd6a29599da5aff9a58718d931601e199efbff8ae

Observation caef8f91-cf52-4165-ae02-bf8114cb98fe · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Encouraging divergent thinking in large language models through multi-agent debate

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.390182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.390182Z digest=sha256:0daa6fec6b15f0b0b5d107e651d491ee223649fb3a4dcee70c61bd6122eb5af2

Observation 9576c0c3-7c93-4a84-b6aa-2f631dc9a626 · outbound

This paper cites Let’s verify step by step.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Let’s verify step by step

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.995037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:43.461797Z digest=sha256:e23b323562c0a1c634512a595b3f4cb288729a133047472738b6c90b2a4c3b33

Observation 991a5605-1452-4318-974a-5ebd74924ed7 · outbound

This paper cites Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.514780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.514780Z digest=sha256:0fcfbe4c297beee067184c1fb2b8dbf93b2ef4f59c06d6a21ed96a774c9c1da6

Observation 6892ea3c-2473-4033-95cb-f5b7f5aa43d2 · outbound

This paper cites Breaking mental set to improve reasoning through diverse multi-agent debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Breaking mental set to improve reasoning through diverse multi-agent debate

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.806699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:43.622627Z digest=sha256:026247b8093f5fd3fea72f996221765540926cfacaeb58f22383b05817934620

Observation 3f9e51d1-fd16-4550-bf19-7f0c113d4437 · outbound

This paper cites Large Language Model Guided Tree-of-Thought.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Large Language Model Guided Tree-of-Thought

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.689015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.689015Z digest=sha256:b7303d3be92674246841c5d157d2c3a102fd2a0e9f4dc27adaddbdeae0f04280

Observation eed94867-4db9-458d-afd1-9f9147eb6b66 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Self-refine: Iterative refinement with self-feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.629374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:43.809755Z digest=sha256:5c6e4b7e1ad1bcbc52ec26b7ac69357935d8c20eeced28d59a5da21001ef6c19

Observation b05f67d7-eaa3-46a8-beea-016adb6c588f · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.876587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.876587Z digest=sha256:718d925641c88d1bce1eb0e8202d08676a124975f739094bb0a1615eed69c467

Observation b63599e5-2e99-4f04-8f44-090da8d8c1a9 · outbound

This paper cites Should we be going mad? a look at multi-agent debate strategies for llms.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Should we be going mad? a look at multi-agent debate strategies for llms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.361841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:43.943342Z digest=sha256:b1b84baf4ac7a0ed45778f83a23b9d4277b33b9c2abd6705736c5d48d3224ecd

Observation 0447339c-e3d6-426b-8e53-d9e5eb0858da · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.017112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.017112Z digest=sha256:e279d1168bb63bd4eeb60133f7576e0596cff4fa2d3544aedd17437c9740d563

Observation 3649c79e-4a22-416f-8074-3946eb383c4b · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Gemma 2: Improving Open Language Models at a Practical Size

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.090457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.090457Z digest=sha256:d43dcdf2fdfdfda35f3f4745bf3a7becaaf362782ae65b225ce289a332a93a1e

Observation 4e7a4394-8d00-4085-a812-8ebbb20ad4f0 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.201364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.201364Z digest=sha256:ea34e1b10fdf5989f19af33b02af60ce75b6f991e318f1cdbc09bb41c65bcbc7

Observation 43b7cf01-5e1c-41af-a77e-d9579df47a4a · outbound

This paper cites an unresolved cited work.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.323471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.323471Z digest=sha256:cc676373236382f302861c53d08970c56645c60f1ce75f5a56c5394db2e81c7b

Observation fe252aa3-127c-4995-a1a4-238077d50c11 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.400972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.400972Z digest=sha256:6b257f6fd97db6c29476480e5c2aaac74df20ce0e87f93a05bf6e6d1874a25e3

Observation 771059ea-f15e-46f0-9a36-fc16aa64452d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Chain-of-thought prompting elicits reasoning in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.474136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.474136Z digest=sha256:7577261a4250fcdff69b12a50b5e45b94938de33c2cce0b9383e4a281620e62a

Observation aec96bbc-035d-436b-b99f-9d49c1763475 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.548349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.548349Z digest=sha256:a051c1c1e3c31554c6f8ce1ad6bdba11af19a6d79461ed434e663836f677a252

Observation 712a0928-6fb9-4e55-9655-c1e19a20b67b · outbound

This paper cites Self-evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618–41650, 2023.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Self-evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618–41650, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.645060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.645060Z digest=sha256:2db6dc3246710f34153683530e2eb7a1af6fe96cf27deaee352b19ed017428a6

Observation 645603a2-e112-4f1b-98b8-34c3f1155150 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.713305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.713305Z digest=sha256:9cd8433a556ba03919f44bfc75ef382a2fc4e911424aa8f093edf7d66edac48a

Observation 86185fd4-e7d6-4378-96c3-c525adff82e7 · outbound

This paper cites Qwen2.5 Technical Report.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Qwen2.5 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.890454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.890454Z digest=sha256:2dce80dad554fd251357ad17b3bf84a40a6dc40801746a9da06e04d75412656b

Observation 3f4324b8-7572-49a2-bbf9-870a6f3f9d87 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Tree of thoughts: Deliberate problem solving with large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.091050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:45.024427Z digest=sha256:8094c1d8630677a74dabb6fdb5f0df65efb5ec83c8b8c48fe5c0fdcb60730cae

Observation 0fc19b89-482a-495c-89d1-d4e62f7c113d · outbound

This paper cites Csrt: Evaluation and analysis of llms using code-switching red-teaming dataset.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Csrt: Evaluation and analysis of llms using code-switching red-teaming dataset

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:46.896637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:45.101351Z digest=sha256:71feb26874bc75f180c4ce02dcb2e23ba7be6ebf3e104d40847e28f322c43290

Observation 35c4c82a-a13b-483a-8835-afb609c51de4 · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:45.177061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:45.177061Z digest=sha256:7827abe8c1294c551744a5ab439655113692ffc93df9c9656b5c47c0001a2cb4

Observation 5210efc6-e165-4d9a-829c-9dfb60662f0b · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:45.265378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:45.265378Z digest=sha256:0845da9c88dad26c0acfc0c58bae9e449e2e3f7d532a87179c8b85f986437140

Observation 3c178dd5-8622-43ff-a3da-77473e2e7b62 · outbound

This paper cites I’m sorry, but I can’t assist with creating content that promotes hate, racism, or any form of discrimination.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness I’m sorry, but I can’t assist with creating content that promotes hate, racism, or any form of discrimination

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:46.675825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:45.317745Z digest=sha256:a4f644d84fecaaa1165209f0729290d12a81fb5af3fbb16fe2dd9e2ae5da9cd8

Pith citing papers

Observation 2a2a4fa0-013f-4913-9e00-fa38553442cf · inbound

Free-MAD: Consensus-Free Multi-Agent Debate cites this paper.

Free-MAD: Consensus-Free Multi-Agent Debate Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T17:15:01.617748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:15:01.617748Z digest=sha256:ce6ad91fdedf8e4b632678779f4c1b5d560885c0f2caced3012154a1693e8e84

Observation eb9a11b9-1f9e-449a-a1d0-539a9aea92c3 · inbound

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning cites this paper.

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:03.999666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T15:52:43.274993Z digest=sha256:552b0180c1a3f53a79b30afdfcaa4c2b3b132a91ca96787ad1e5dc267b2419e1

Observation a4491ec6-89d0-45db-8255-668369926b0a · inbound

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate cites this paper.

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:52:35.927613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T18:46:50.409655Z digest=sha256:7ed68c0897c675dea2f323cefd62fda43e5d6ab200a12d034925be95dda33604

Observation ad444918-9aa2-49bb-8725-23f1dae9f482 · inbound

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size cites this paper.

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:15.586924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T16:04:37.244398Z digest=sha256:bb9dec325ec5ddaa4d4ea9b51e44d89e3d86b4790e4cca9e3fb00b989b27e543

Observation fb4142b4-7843-4053-a9fa-48531c202511 · inbound

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning cites this paper.

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:36:23.855401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:05:49.696342Z digest=sha256:8b3735d2b3d81f2351cf60b25751701c88293ad6950a41c74fc5c080d639ea7b

Observation b2ac482a-8485-4275-a117-6a62dac863b2 · inbound

Heterogeneous LLM Debate Under Adversarial Peers: Honest Gains, Replacement Costs, and Resilience cites this paper.

Heterogeneous LLM Debate Under Adversarial Peers: Honest Gains, Replacement Costs, and Resilience Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:33.490797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T17:17:05.107942Z digest=sha256:168410fbdf81918a0ec1c4051d212a505071361183d67db1bd580e2963fa2396