Pith. sign in

Paper Citation Record · LEDGER

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

As of 21 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 7 inbound Pith citation observations for arXiv:2505.22960.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22960 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:45.317745Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:17:50.828950Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:09:33.482951Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50165dfc-02ec-42c5-b4e2-4589e29b9d64 · outbound

This paper cites Critique-out-Loud Reward Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Critique-out-Loud Reward Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:41.917221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:41.917221Z digest=sha256:9722caeb6a1e73469c452476649bdd453de1bb3676d3ad77c31a342310c489a4

Observation f7bcbcfb-5cb3-40a2-a7fc-a9c9d564f61c · outbound

This paper cites AIME Problems and Solutions, 2025.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness AIME Problems and Solutions, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.974827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:41.963916Z digest=sha256:a9e799be8e674d6a97e4719a27f8c18f0724ba18375b0bded7f94ab5a8f1f661

Observation a3736bdf-ec6e-4833-ac82-cf84d1106980 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.027532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.027532Z digest=sha256:d1df026c9d647f0af1cbdedbabe615f723674d19917c665d35454b5d717d3625

Observation eb61c74b-3ac3-42f5-a418-7f9e56dd0bfa · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.095621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.095621Z digest=sha256:04f34362c6b8c056d8d4056806d897591930a245a5323130f9a802a6b118e2a2

Observation 7b8ba3e2-d821-4e8a-8e6a-24cf24afd817 · outbound

This paper cites Reconcile: Round-table conference improves reasoning via consensus among diverse llms.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Reconcile: Round-table conference improves reasoning via consensus among diverse llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.822189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:42.153098Z digest=sha256:84e577da627228cb8f3758a09ecd475539582ef55136e55bc3033a2ab753421c

Observation 24d4a73d-fccf-4b25-9b77-7150adbbbef6 · outbound

This paper cites Combating Adversarial Attacks with Multi-Agent Debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Combating Adversarial Attacks with Multi-Agent Debate

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.213315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.213315Z digest=sha256:417f5cd3d589e856e26865afd66c6f725080b946acb62d2b9eeaf4eadab30530

Observation ef59e14e-97b5-44fc-83c7-1e88397b29d5 · outbound

This paper cites Enhancing LLM Performance Through Debate: An Empirical Study on Multi-Agent Debate for Coding Tasks.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Enhancing LLM Performance Through Debate: An Empirical Study on Multi-Agent Debate for Coding Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.317472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.317472Z digest=sha256:77e99b2b03758eace3841924579e203ef783e7bc2b0074590a4f153a978d95e9

Observation 0da4f06b-ae12-41e0-b5d5-8edcd3c37a3e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.391929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.391929Z digest=sha256:d6e819889afacd1c36c7b6738f0d5dd635b69e2627e6ab438047b9c09d520ea5

Observation 51a40f0e-8052-4add-906c-6f89b7618adc · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Multilingual Jailbreak Challenges in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.474919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.474919Z digest=sha256:39653848bbbef442f0769f6bf3e7cf0d72dea11694ac0e714f01e2160cc6e290

Observation 157f11ae-88c1-4fce-959d-54bc007b4d75 · outbound

This paper cites Improving factuality and reasoning in language models through multiagent debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Improving factuality and reasoning in language models through multiagent debate

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.617588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:42.545982Z digest=sha256:523707a6407cadb621425b20b53a336e3f9b16c178584479165135059e32ac19

Observation 71e6a7c1-d19e-4968-9609-84f644783179 · outbound

This paper cites Multi-LLM debate: Framework, principals, and interventions.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Multi-LLM debate: Framework, principals, and interventions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.448146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:42.619790Z digest=sha256:0a42eb0be0b827c24a51366c470249e387f3a21c96ffdc94d40b9ac68281d86b

Observation 9fa65451-ec23-49c2-b288-95a87761aeac · outbound

This paper cites The Llama 3 Herd of Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.686078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.686078Z digest=sha256:4b04460776d31c0f8db6d181cde19181ea4e3c10adc6506a8fe365c994e61f52

Observation 28101e25-432d-44fc-bff6-3da30de146ba · outbound

This paper cites An empirical analysis of compute-optimal large language model training.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness An empirical analysis of compute-optimal large language model training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.781403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.781403Z digest=sha256:3465c40131499667a697a225c66ab7a42072a3252d5282d1a80c1cbc6b94f2f4

Observation 7cc32020-f4cf-4c49-9778-ccfe62600de2 · outbound

This paper cites The curious case of neural text degeneration.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness The curious case of neural text degeneration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.205048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:42.875033Z digest=sha256:d51d5b3fea85bd23d398ced47491dd0edeedba626492c60b4eb98c2598be99e3

Observation 91b8e223-e1c5-4e6c-bfcf-6fe735e1549b · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.962354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.962354Z digest=sha256:4eea242dee417763e7dabe18ebcfec077c2f3cd4f4c150c7d4c28b28af814cd2

Observation 11a8cff6-5d82-4963-a225-0b949976c3b5 · outbound

This paper cites GPT-4o System Card.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.051383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.051383Z digest=sha256:99464a1d252337c92c057565c29d6cb55d702103f149d756a19726512a5db4f5

Observation cb22c870-4a60-4254-a16f-2110294f554c · outbound

This paper cites Scaling Laws for Neural Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.128633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.128633Z digest=sha256:099f582e8c466668f28a478c65363bfa742fdf2617873e9d259f273f702a6e39

Observation b8eb9045-ce00-42b4-8d31-5e0cb97648ca · outbound

This paper cites Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.211841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.211841Z digest=sha256:ad91ef876355ea668b78d129afbc126a94f5ed875fdc73cfacc3c024e29086ab

Observation f4ad240d-ab0d-47f3-89c7-cd5e72f2eb31 · outbound

This paper cites A Simple Model of Inference Scaling Laws.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness A Simple Model of Inference Scaling Laws

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.310559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.310559Z digest=sha256:41f73f056bcb25c9db9fe9ac18b6cc991ffb17f923bb9237a1d1b85be8e6893e

Observation caef8f91-cf52-4165-ae02-bf8114cb98fe · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Encouraging divergent thinking in large language models through multi-agent debate

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.390182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.390182Z digest=sha256:b307c0049439e3186b58a265005b37865085ea573fa612f00d5b46c2f40eeefc

Observation 9576c0c3-7c93-4a84-b6aa-2f631dc9a626 · outbound

This paper cites Let’s verify step by step.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Let’s verify step by step

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.995037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:43.461797Z digest=sha256:99c93f0e7cf6b4a77f72e65b9d61340d18efa1e8d6f598f14b0752a281ceb257

Observation 991a5605-1452-4318-974a-5ebd74924ed7 · outbound

This paper cites Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.514780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.514780Z digest=sha256:aeb366e38726fe8928c42940f20c411129fe2bce8631bd4083ff264d973d9f77

Observation 6892ea3c-2473-4033-95cb-f5b7f5aa43d2 · outbound

This paper cites Breaking mental set to improve reasoning through diverse multi-agent debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Breaking mental set to improve reasoning through diverse multi-agent debate

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.806699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:43.622627Z digest=sha256:c94d6b9cc23eaec4bb0b0bfa4f4bf70286bbf832ab4186f640040b2b3a0e8b65

Observation 3f9e51d1-fd16-4550-bf19-7f0c113d4437 · outbound

This paper cites Large Language Model Guided Tree-of-Thought.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Large Language Model Guided Tree-of-Thought

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.689015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.689015Z digest=sha256:6b869a87fcf32c641f2f9211c3841e10d301165012c15d688dc28428b58e7ac6

Observation eed94867-4db9-458d-afd1-9f9147eb6b66 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Self-refine: Iterative refinement with self-feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.629374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:43.809755Z digest=sha256:1ef6eed382653d0c2ca59cedcddaef974196974e164e04ab3ceed91c57fceb93

Observation b05f67d7-eaa3-46a8-beea-016adb6c588f · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.876587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.876587Z digest=sha256:82020508a52c851c477a27fe73c1bf44c809ac577f14f7cb121d5a9ed5fb837a

Observation b63599e5-2e99-4f04-8f44-090da8d8c1a9 · outbound

This paper cites Should we be going mad? a look at multi-agent debate strategies for llms.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Should we be going mad? a look at multi-agent debate strategies for llms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.361841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:43.943342Z digest=sha256:49a7a60fa3aecb0d121bcfbc529e7ab5fd39506992fe080eda571fdf8c2ef013

Observation 0447339c-e3d6-426b-8e53-d9e5eb0858da · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.017112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.017112Z digest=sha256:76767a0fa9c7bf53647b78272792a7051fa2bb4fd3a3cccef080302ff5aac7f8

Observation 3649c79e-4a22-416f-8074-3946eb383c4b · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Gemma 2: Improving Open Language Models at a Practical Size

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.090457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.090457Z digest=sha256:4bd46f3985e8872d4c8c6e1c9997555c557ecb1735a37b8a7f9c1dd73aab820b

Observation 4e7a4394-8d00-4085-a812-8ebbb20ad4f0 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.201364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.201364Z digest=sha256:490fcbbd1c5b692a3a3ae2fb2994446314802233c184ca090a21296d5e77324a

Observation 43b7cf01-5e1c-41af-a77e-d9579df47a4a · outbound

This paper cites an unresolved cited work.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.323471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.323471Z digest=sha256:4ec1168bf3cc0f725fcb1559f6b72f663d807e7ab6f2e3b031ef72559bac5c5d

Observation fe252aa3-127c-4995-a1a4-238077d50c11 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.400972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.400972Z digest=sha256:b9eeadd9ed5143782d6e830bde31a2e6283ef88bbae40a31f73a80225de19938

Observation 771059ea-f15e-46f0-9a36-fc16aa64452d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Chain-of-thought prompting elicits reasoning in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.474136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.474136Z digest=sha256:4c3b4cc3b74509ccc26c82742a0b42f770f1657879f8cccc4d4b52da1051219f

Observation aec96bbc-035d-436b-b99f-9d49c1763475 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.548349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.548349Z digest=sha256:90cdc1187d6ffc6e3fbe1f01ac9fe9e36ba3b8150285650915a9146bb452e7e9

Observation 712a0928-6fb9-4e55-9655-c1e19a20b67b · outbound

This paper cites Self-evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618–41650, 2023.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Self-evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618–41650, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.645060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.645060Z digest=sha256:d6402190e85e47eef2095e915b6361552b9cf8966371432b0fef676c67ba424e

Observation 645603a2-e112-4f1b-98b8-34c3f1155150 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.713305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.713305Z digest=sha256:86567bdfe2f22e3cac2ece37ed2a57cf2b9c5c895708a06b04ceeaf5a09be269

Observation 86185fd4-e7d6-4378-96c3-c525adff82e7 · outbound

This paper cites Qwen2.5 Technical Report.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Qwen2.5 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.890454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.890454Z digest=sha256:03e957b2c2960a72f24bed6f7fef4795847e09f35ddfd6e29c1a08b0c907d4b6

Observation 3f4324b8-7572-49a2-bbf9-870a6f3f9d87 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Tree of thoughts: Deliberate problem solving with large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.091050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:45.024427Z digest=sha256:83fe455ddd68d35fbde961f996b6da4cc80561d78dcd8d4bfba2bcf90701cc93

Observation 0fc19b89-482a-495c-89d1-d4e62f7c113d · outbound

This paper cites Csrt: Evaluation and analysis of llms using code-switching red-teaming dataset.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Csrt: Evaluation and analysis of llms using code-switching red-teaming dataset

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:46.896637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:45.101351Z digest=sha256:1697b47a40fcd3a4a28ac171a4ce7e4c22fa78192c11fa87ef59a5c0200255f4

Observation 35c4c82a-a13b-483a-8835-afb609c51de4 · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:45.177061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:45.177061Z digest=sha256:42e3fa7524f4871dadf1320d06d9f5f1910c530aac8790f88cf833547de0e948

Observation 5210efc6-e165-4d9a-829c-9dfb60662f0b · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:45.265378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:45.265378Z digest=sha256:c2bd4683270a2728a00dafd202854b86331c573eb58001e27919fa15f5600bf6

Observation 3c178dd5-8622-43ff-a3da-77473e2e7b62 · outbound

This paper cites I’m sorry, but I can’t assist with creating content that promotes hate, racism, or any form of discrimination.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness I’m sorry, but I can’t assist with creating content that promotes hate, racism, or any form of discrimination

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:46.675825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:02:45.317745Z digest=sha256:9dc8dc9fcf929006c96cee977859afbc31f3ffa25a952dca89ed794839ec68d4

Pith citing papers

Observation 2a2a4fa0-013f-4913-9e00-fa38553442cf · inbound

Free-MAD: Consensus-Free Multi-Agent Debate cites this paper.

Free-MAD: Consensus-Free Multi-Agent Debate Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T17:15:01.617748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:15:01.617748Z digest=sha256:6d4481cb9eb741fd4cf06117eba92d5dc417b090ecae074149e348baec2c70a3

Observation eb9a11b9-1f9e-449a-a1d0-539a9aea92c3 · inbound

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning cites this paper.

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:03.999666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T15:52:43.274993Z digest=sha256:73be0a7e77d652aacdde28326542451a03a726c60c6e3b56b7d5d26464b57e56

Observation a4491ec6-89d0-45db-8255-668369926b0a · inbound

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate cites this paper.

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:52:35.927613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T18:46:50.409655Z digest=sha256:810b56b26a22f7ff11225794bb22397c042e16b26c27b3350a0421f589240689

Observation ad444918-9aa2-49bb-8725-23f1dae9f482 · inbound

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size cites this paper.

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:15.586924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T16:04:37.244398Z digest=sha256:0c8ad9953850cc49f70afc299deb0e09dcc0f04f2934f959b129404581a171af

Observation fb4142b4-7843-4053-a9fa-48531c202511 · inbound

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning cites this paper.

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:36:23.855401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T14:05:49.696342Z digest=sha256:0fa674c0ef386416eb44ffdf0d6d086813f6ee38a04f6c5baa7e71b10c1be483

Observation b2ac482a-8485-4275-a117-6a62dac863b2 · inbound

Heterogeneous LLM Debate Under Adversarial Peers: Honest Gains, Replacement Costs, and Resilience cites this paper.

Heterogeneous LLM Debate Under Adversarial Peers: Honest Gains, Replacement Costs, and Resilience Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:33.490797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T17:17:05.107942Z digest=sha256:c4a635bd111ab5ebe7a53a608a046abdf2de3a46f8b015565fc3c494d22fb717

Observation 852dd9ca-6b3f-4588-8ce2-253d9e9171e9 · inbound

Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology cites this paper.

Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:17:50.828950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:17:50.828950Z digest=sha256:67c90b8c6b89178894f1e3a02def97957f0459184e1d51c4e6295ad91377ed83