Pith. sign in

Paper Citation Record · LEDGER

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs

As of 19 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 2 inbound Pith citation observations for arXiv:2507.08960.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08960 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:13:00.348063Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T22:47:51.676289Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:56:37.756820Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a4260ea4-7abf-4edc-b541-3b2a19422f04 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Graph of thoughts: Solving elaborate problems with large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.005763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.005763Z digest=sha256:abc02c41c8061b5ec700c38502a6f4d2cbd91576d98560a0068c5710e19af783

Observation bde69cd1-60e0-4620-90ec-6d3d5dcda0ab · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs On the Opportunities and Risks of Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.029371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.029371Z digest=sha256:56442657cc3b3190ba9eac198f6a4b5f7aab978b601031348115670bd9c9b70e

Observation 4a557c60-7ca1-4d22-b4a3-e3bae4c1f77a · outbound

This paper cites Language Models are Few-Shot Learners.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.051900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.051900Z digest=sha256:f297682c7d7da56d3c1fb295b227771b66f03d7125c66ee11a56314fccbc2bc9

Observation a67cc61c-e0fe-4db0-8a0b-f7eff4fda5f7 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.078990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.078990Z digest=sha256:19a3efff57329cb0c705fa0a47dd9451d0c1e35458538442e49136ccd74a608f

Observation 12af781f-5054-43ed-9f3b-ab547c44287a · outbound

This paper cites SocraSynth: Multi-LLM Reasoning with Conditional Statistics.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SocraSynth: Multi-LLM Reasoning with Conditional Statistics

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.107269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.107269Z digest=sha256:c764e644a690f4732a7c3ef7e43f7d0edcfa9faba38024633a65d552d85e6c9f

Observation a3fcc3d6-af83-4d07-9272-a7a0c28e18cb · outbound

This paper cites ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.136451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.136451Z digest=sha256:b82a4655e012ac79c23b1ec4c931793c303ced430c4ccede0f6b5c2058a0a37b

Observation f61326d0-7de9-4c33-887e-afeeaeab4906 · outbound

This paper cites Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.172268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.172268Z digest=sha256:c4b68b97f9588d39da409b0473d5b5e0ebcd4276fc13fc36c07c239e819a1fd3

Observation 60908b6d-180d-4b21-a7e5-e772302339b5 · outbound

This paper cites Universal Self-Consistency for Large Language Model Generation.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Universal Self-Consistency for Large Language Model Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.200045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.200045Z digest=sha256:66ad31904e6556bdcdba3c1fa7adbb6c873edab8daedcd168afa046270402191

Observation 5ef96167-bebf-4553-b765-e9a53fc887f1 · outbound

This paper cites Cost-Effective Online Multi-LLM Selection with Versatile Reward Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Cost-Effective Online Multi-LLM Selection with Versatile Reward Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.230490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.230490Z digest=sha256:41018e0e47daadc0f27b7d05640e49ce09e677c2b981a2ccf450d0bc2a115739

Observation 4048bd57-f74b-4197-8b87-0f335d1a653b · outbound

This paper cites Improving factuality and reasoning in language models through multiagent debate.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Improving factuality and reasoning in language models through multiagent debate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.254960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.254960Z digest=sha256:62cc06c45935af68c1db4c1b7f74568039899ec3f468f5bf60012a64bf9971ac

Observation 1808ef03-445f-432a-b6b7-ec0674130b18 · outbound

This paper cites Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.282695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.282695Z digest=sha256:00ffe03f7a56ed43ab5f056dde501bcedb32f93cabad0d5d14e4d191a11dae5b

Observation 2add6b6c-d084-4bb7-91b5-4ef3434c6652 · outbound

This paper cites Multi-llm debate: Framework, principals, and interventions.Advancesin Neural Information Processing Systems, 37:28938–28964, 2024.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Multi-llm debate: Framework, principals, and interventions.Advancesin Neural Information Processing Systems, 37:28938–28964, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:03.384428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:12:58.316555Z digest=sha256:af9c7247b1df607069350919c3d8097f08561e85ff71a9d2ac52d4443b430516

Observation 5cb5f25c-4895-4e1c-b7b0-8109771efe23 · outbound

This paper cites ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.333530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.333530Z digest=sha256:c0787a2b8efaefafeb3bb13e29ca1f61a3291e3d9245dcf97c372b1fea256cf7

Observation ceb1668a-586b-4453-b857-772546007899 · outbound

This paper cites Acc-collab: An actor-critic approach to multi-agent llm collaboration, 2024.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Acc-collab: An actor-critic approach to multi-agent llm collaboration, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:03.218235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:12:58.359750Z digest=sha256:525f698386096d28a17d8f7c0be7a113dee3d5d7ac3d4d5e68d5bdba56d83c07

Observation 4c63ecf6-b762-4b69-a2a7-99d65e5c23fd · outbound

This paper cites Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.377961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.377961Z digest=sha256:ebd59218227f1a605ce07f4520a1fe689dc1738196663c9b1901b3418c98fdad

Observation 62cfc9c8-6a77-45ce-9136-e555f92dca6e · outbound

This paper cites Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.398205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.398205Z digest=sha256:523ae322b651633df6742b66319f846e15e16590cb87cb90f6aecf89514e93c0

Observation 0ab1ff3e-dcc9-448b-a654-72d8a3ac5e3a · outbound

This paper cites When One LLM Drools, Multi-LLM Collaboration Rules.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs When One LLM Drools, Multi-LLM Collaboration Rules

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.412984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.412984Z digest=sha256:ec2ee3afde6c137e5296c0d19ec8818711db57f83742328792310b46e7d91f01

Observation 98037f10-d385-4b10-96be-593571730124 · outbound

This paper cites Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.440669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.440669Z digest=sha256:2652708636f5e8ffb5af27a61f30cb6a7e6dda81e2ee23e37cf62b64441ddd0f

Observation 0c44c63f-9bf1-481e-9251-6cf1b33cc292 · outbound

This paper cites The Llama 3 Herd of Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.461508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.461508Z digest=sha256:c968d8e3ab92ae6a688c43cf4b343c5f9b69bf6fa28f0f76f1bf094f691d05bd

Observation de01ddc7-bbd2-4760-9bb2-e1a489f7016b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.485112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.485112Z digest=sha256:12bee363ec1f61d49281e593c16e55e8c12d893062ae7b691bc65bbe7b161b18

Observation 8bc2de8f-deff-42d0-a3d8-50001e383221 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.517427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.517427Z digest=sha256:91db47c7869e0af96a52c18d20023b4b5e6b1b04173272360075ca76f5de9c59

Observation d07b6f32-3df4-4d8e-80bf-56647a1cfb0d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.548499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.548499Z digest=sha256:3c36801f8bb83f3ff1841c28d3508acbde32bcebea47f83073775df6e82423ea

Observation 644f457c-aea3-4539-905c-75dd4b6023b2 · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.583192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.583192Z digest=sha256:52ca53ee8d296b7e9cee3a8eb8193b0cfcb13297fdf23f456361a0a3dc85508b

Observation 597c834d-ee7b-4050-a9dc-87ef7e557153 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.605891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.605891Z digest=sha256:bbdbd771c3fb25c313c7cad5019bf76f7b10b31c5e277a809e3b7116c3f84260

Observation 99bc94e4-dcf3-486f-a5bd-b69e6a12db7c · outbound

This paper cites Ensemble learning for heterogeneous large language models with deep parallel collaboration.Advancesin Neural Information Processing Systems, 37:119838–119860, 2024.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Ensemble learning for heterogeneous large language models with deep parallel collaboration.Advancesin Neural Information Processing Systems, 37:119838–119860, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:03.050076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:12:58.630989Z digest=sha256:878e4ff74ff7cd8f921200040012a5c7dfbb42f2d703e7cf1f36ebaaeb8c6299

Observation f2669eb8-0f89-4653-984a-3a5a2f1c3b18 · outbound

This paper cites OpenAI o1 System Card.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs OpenAI o1 System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.649279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.649279Z digest=sha256:89e902234a123b107ff7a9587671b2b0c343da61f971d90603f44f442fd021d1

Observation 908b419e-cb3d-45e8-b949-634ed9bfb66c · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.673023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.673023Z digest=sha256:2d50483fcdef78ae687ab9d86f0051501b6eb8275901638cccc174be126c4e3a

Observation f49e5ddd-dd47-4ff4-bbfa-4f50ecc3c14d · outbound

This paper cites Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.703046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.703046Z digest=sha256:850b49335d064ba9ae321113f45b9f6d3891ce2728f7cff8f73663a10ee587e1

Observation 12f9681d-b8c2-4344-b6c2-8546819eb797 · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.734477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.734477Z digest=sha256:6fa7429d14e6cf860f76cd58251456753068e6bedd0a6e08fd15364323d0fcb0

Observation af1bd348-84c9-48e1-bb11-b38be3e4ea4f · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Training Language Models to Self-Correct via Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.743208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.743208Z digest=sha256:061e651a002518188eed67ea2d3977bcc17d97dca79a61ab7b1ea368abb84ffe

Observation 065edb2c-6337-4c27-b62e-6ffbc0d040f6 · outbound

This paper cites SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.768620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.768620Z digest=sha256:ccb4f78a5ba617a1be26cf55747c68e976d3540f4726f2481193dcd85413d85c

Observation b06ebd05-8852-4536-be17-d922927703b6 · outbound

This paper cites Two heads are better than one: Dual-model verbal reflection at inference-time.arXiv preprint arXiv:2502.19230, 2025.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Two heads are better than one: Dual-model verbal reflection at inference-time.arXiv preprint arXiv:2502.19230, 2025

Reference 32

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:13:01.369146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:12:58.791281Z digest=sha256:94a958753574dc3462b69262ca8840a3e897c8d3a5f818cd49e2283484f5dee2

Observation 7db4046e-e50e-45f8-82f3-bc13de926a5a · outbound

This paper cites PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.824124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.824124Z digest=sha256:768932429c00642213938453425dcdcb4d6b3e06a74ee2bde52544e3eaeddf4b

Observation 700d64f3-8551-4e27-b970-ba6321acecd4 · outbound

This paper cites From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.843340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.843340Z digest=sha256:df809fad6c803d0efc686a69cd46c20c33612321122a2911c54087be7677c09c

Observation 139d4ef8-f847-45a9-8eb3-94ccb2f62650 · outbound

This paper cites Improving Multi-Agent Debate with Sparse Communication Topology.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Improving Multi-Agent Debate with Sparse Communication Topology

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.866755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.866755Z digest=sha256:cd0f48d8f07850442fc966229c784b50a6c30ab6d0389d90f1259cb3b2b9b236

Observation 51167679-277f-4271-a834-e45d56d82b17 · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.901843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.901843Z digest=sha256:980e9bf9c4b9280d0db92e1a71fc2d8f2493ae2227ad3509a5e9a05ee51c9c5d

Observation a90916ee-2ddb-465e-bff1-f8c54a8c7fda · outbound

This paper cites MARFT: Multi-Agent Reinforcement Fine-Tuning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.937766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.937766Z digest=sha256:6ca5814580aa88d6e29a9b560e0512dcf19eebcce3757411b0625ee84c97a023

Observation ed24df84-fe46-4fae-8838-6e7361642413 · outbound

This paper cites Groupdebate: Enhancing the efficiency of multi-agent debate using group discussion.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Groupdebate: Enhancing the efficiency of multi-agent debate using group discussion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.973810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.973810Z digest=sha256:28fe28ae900661dfbb164f06df5f9f8483a34c1ed0df4806c449c8a6f72f5bfa

Observation 341a368d-dece-4147-a542-a14dfe331d4a · outbound

This paper cites Towards Hierarchical Multi-Agent Workflows for Zero-Shot Prompt Optimization.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Towards Hierarchical Multi-Agent Workflows for Zero-Shot Prompt Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.009997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.009997Z digest=sha256:baa29dcd917312f37e14bc76606b0ebaded595f80089eee2ea1a0c3978f15498

Observation 5dc73893-7337-4db7-ae30-b87f74102874 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Understanding R1-Zero-Like Training: A Critical Perspective

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.047193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.047193Z digest=sha256:e9413224e05ba6648f24fd7674bfc996cf50e5723be96f98668ec7585fc2df02

Observation 9b776a90-25db-4ee1-a943-321d814100c8 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.901335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:12:59.082303Z digest=sha256:e49cec012bd9c2998e45d44966bfea5f4c53e94ac9c7658553243d9635d4ef03

Observation 88dd62b2-115b-4d69-99f7-df26b77abd66 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.115363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.115363Z digest=sha256:3ec20843088fdbd39973faefe804bd23e74d2e6d64d096396fe5b9a38d187529

Observation 7c3f980c-8035-4390-b9e6-2d0ab0c2195e · outbound

This paper cites SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.145103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.145103Z digest=sha256:ad15bd740e6aa2f08f95b75c35a04e215dc2fcd2b21532ed4d497e0cf7e563ca

Observation 97d93bbb-9545-4b74-b30d-9c455ceb48e4 · outbound

This paper cites Beyond accuracy: Evaluating the reasoning behavior of large language models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Beyond accuracy: Evaluating the reasoning behavior of large language models

Reference 44

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:13:01.062647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:12:59.181998Z digest=sha256:30b1d0e336341c7d1afb2afb5dbb71134b06d9577a85983f6f39bca9213f1b5a

Observation e7966f11-a285-4ac9-b61d-d8d1ac569a1b · outbound

This paper cites Motwani, Chandler Smith, Rocktim J.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Motwani, Chandler Smith, Rocktim J

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.216898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.216898Z digest=sha256:d365f479309c59934154e95611e8c80bf04dc686f26ec334cce2941dc6de9248

Observation 933c70b5-5273-4ec6-b898-39f5003ff99f · outbound

This paper cites MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.245532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.245532Z digest=sha256:997e25af7c8ce1550aebefb5fcfc97038a8b7d6546b9034f75e8ad4efa40350f

Observation fcab31e2-8508-4fe2-9861-2d58435f8910 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.282290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.282290Z digest=sha256:f84429a69871c8cf2157b415de4c9f38e1db4b7360cd289a1a6343c2c7d45487

Observation 6f94231f-a55b-4841-ab74-3b74cc6f3f99 · outbound

This paper cites Towards Collaborative Intelligence: Propagating Intentions and Reasoning for Multi-Agent Coordination with Large Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Towards Collaborative Intelligence: Propagating Intentions and Reasoning for Multi-Agent Coordination with Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.318146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.318146Z digest=sha256:c65d0a1e8d3577afaf9b32b2a283b0a7c57c5a9966aad1ccf9990407456ad58e

Observation 290f5497-d606-43da-af7d-05ffa0e50a00 · outbound

This paper cites Qwen2.5 Technical Report.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.353123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.353123Z digest=sha256:71b1ba26cabd6edf23a83df0db5708c2002ccad861158cb5a28322deadeeedb3

Observation 35070d5d-222d-4fbd-9321-603e556d56d2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Proximal Policy Optimization Algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.389070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.389070Z digest=sha256:ac01a779634c511ee4d773f67050f7bd0fa0122f8b8096a1137cf99916f3367f

Observation 6f8fdf02-dab1-4ce9-8149-e33179bf2602 · outbound

This paper cites MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.444849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.444849Z digest=sha256:8c76a46e660122c00c45be13cbf7944f3458d8b7e067115c46da1a8ff5aac410

Observation afbc721e-c949-4498-8155-5eb6cdfc1040 · outbound

This paper cites Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.473520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.473520Z digest=sha256:fe54f7c601599bf1649ad2789d19dbf2dcc8472a80f42b404ec5798a6581fa75

Observation b08e34d0-5cb3-47a4-a98d-741b4b28bd2a · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.492882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.492882Z digest=sha256:86b333d19aa27746b85f14cd7e3c9e43beecc443e5f1f9255481f34285dd47bd

Observation 42978e99-da8c-48d2-b500-364fcb55df68 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.819843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:12:59.524423Z digest=sha256:93188d9355caed34f51d766a97de1ef0cbc8a9837895e6d7bc8da7f7849df798

Observation ec833267-3708-47c0-93a1-ed34501fb33a · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.592171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.592171Z digest=sha256:5015f6572503e18bc04ac41c1422791ffc0fea01bc08ebc916570d7c3945730f

Observation bee4dd27-919b-47ab-a771-c8f20354d44b · outbound

This paper cites ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.617501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.617501Z digest=sha256:2272388ce6beb4a9eb5b5d794aec0e194c347b0b23a59d36b5575894c0b58a4c

Observation 69d857ca-a367-4744-b812-1b5a3b746474 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.646978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.646978Z digest=sha256:165b461c808add290a127501c63a2989eb9b30158762974106ca9aa2f971477d

Observation 456c1972-61bd-4e5d-8cb0-60ac7f7abce5 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.678125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.678125Z digest=sha256:51e1fd7390f12e539c2193a5bb284d5c91982deeefcfb6e1cd80cf48a980e6c4

Observation c6f7f492-6fe2-4edf-b028-048728c7dd10 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.703873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.703873Z digest=sha256:18994ee20e4175edced684882cc4738cba8e89b6c2fc0e8a0cd9f23ce88bf829

Observation 7ba63848-f0ec-402a-82db-1d4aac63c743 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.735760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.735760Z digest=sha256:e3286a482ecfac20b21dd22d664ba742590c6d6f0125bfb132f7cf48c7a2cb51

Observation a60e3e36-e01e-43d2-b319-9f9fcb5d416a · outbound

This paper cites Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.763216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.763216Z digest=sha256:f33646072d46f5f5abe7f9f97a28c9129750d6602a4bda441b557b5576b1ee92

Observation 8e1bdd83-af97-4041-b0e1-521d2b1b7820 · outbound

This paper cites Qwen3 Technical Report.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Qwen3 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.793527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.793527Z digest=sha256:1bc1102efadb9002024302f0c3a19483234294b7aef5bcfa5df95ec668b93a5f

Observation 608823d2-2b24-4911-9c2b-a6368629f91a · outbound

This paper cites Multi-LLM Collaborative Search for Complex Problem Solving.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Multi-LLM Collaborative Search for Complex Problem Solving

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.822314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.822314Z digest=sha256:6979606fecfe2459fbc50fae0c0287dd3d6430c5eb65694654cccda13f97ba81

Observation eca49996-f6e0-415e-9532-a107d19292dd · outbound

This paper cites AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.855011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.855011Z digest=sha256:65edeb545da94eaa4351fefde243d78529c4f26ee86f342c2d9780ef779e159f

Observation 2873cd1d-8021-4c09-a3b6-8f98b5d851d0 · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.892203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.892203Z digest=sha256:934553abc15b98b586a54c50667df724cf740dd6916239421861549cd8db51a5

Observation d1dd632c-f6eb-455b-9ab8-cdd6c3d1d633 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.917749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.917749Z digest=sha256:275b29977c4a5aa30ddda6d6ba0f105450d30853b83021862333852e5d3909ce

Observation 3a0deef6-b5a5-46c0-893e-2f70b92aef8b · outbound

This paper cites X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.939425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.939425Z digest=sha256:bf222e32caa7decaaab632bcc7a5d9c3e76514f8e938b8580816ae3e02f5ce9e

Observation 98fead6a-7a30-415d-8930-bc0ea65e6463 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.958609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.958609Z digest=sha256:60b6a7aa8f7f0091b965f05c42112e553e05bdbe25c031e09638d93b4414b682

Observation 3afbd129-6a71-4eba-ac6e-3003fa257a68 · outbound

This paper cites Chain of agents: Large language models collaborating on long-context tasks.Advances in Neural Information Processing Systems, 37: 132208–132237, 2024.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Chain of agents: Large language models collaborating on long-context tasks.Advances in Neural Information Processing Systems, 37: 132208–132237, 2024

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.727265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:12:59.992721Z digest=sha256:ece2272aba587ac967bb77e698451cdaa18c9998586eb981de52e0169c0af09c

Observation 633bb555-fc2c-4003-9188-8f313ed67ce6 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:13:00.021058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:13:00.021058Z digest=sha256:d4246216569e28b4a8603fc0d5dc4293d408420fe87cc72ceeeddd3726672d20

Observation 062c7dff-3ce0-47f1-a21d-290ef93ee002 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T18:13:00.055272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:13:00.055272Z digest=sha256:730db273f5f4b2bfab9f8b6db14dfcaf889ed15f3737d89f200e13d9c7c66a35

Observation 0a111599-c2f5-4e98-98f3-4d3ed6bff955 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.606895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:13:00.086857Z digest=sha256:6ea0d1b2deae5dfd299302cce877811b1325ebd31165c40ff5170690cc671bef

Observation 1cbd8334-b5ec-46dd-9f99-393fcc307ac6 · outbound

This paper cites Address each question raised where relevant.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Address each question raised where relevant

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.462769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:13:00.122225Z digest=sha256:02469eda592cd409b715ff381b46414e410befbfdfb9ce29d94d02cf79297768

Observation 671f858a-d84e-4cca-bd6e-490bd5d1cb56 · outbound

This paper cites Regardless of the approach, always conclude with: 25 Therefore, the final answer is: $\boxed{[answer]}$.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Regardless of the approach, always conclude with: 25 Therefore, the final answer is: $\boxed{[answer]}$

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.402514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:13:00.149962Z digest=sha256:692890a4e5337d7b3f4d725f4032cc6896b244b6d9cd3f59162f60db828b2c85

Observation add2d73d-c7af-4113-8882-af81644c6a15 · outbound

This paper cites - End the answer with: Therefore, the final answer is: $\boxed{[answer]}$.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs - End the answer with: Therefore, the final answer is: $\boxed{[answer]}$

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.222163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:13:00.212428Z digest=sha256:a1dbb424889c1646d6fd0429aadac569873c94648e8c258e8ddb3f9b8baf49e4

Observation 5d7ebf72-172c-4525-afe0-5910b2999433 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.140011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:13:00.238645Z digest=sha256:0cdddf8ba7c6f905b129ee2487cc6dcdf8f88b62bab95222d131b755a43515d7

Observation 53a71cad-ead2-4e6d-871b-a79d06efe8c2 · outbound

This paper cites These should be carefully reviewed for mistakes.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs These should be carefully reviewed for mistakes

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.058682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:13:00.261331Z digest=sha256:531b715c03af5105b5412dc67da17422ba80e32840e92f29c71c8a8186bb8935

Observation 26db6557-6743-4503-9159-44a94183d0c3 · outbound

This paper cites Wait, that doesn’t seem right.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Wait, that doesn’t seem right

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:01.989643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:13:00.284581Z digest=sha256:a1d03259f325758c299d23bd6d7e348ec25678337b7a92929244c84fb335d97f

Observation f1997a2b-2b83-42f9-9fda-666bbccea7ed · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.326585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:13:00.320125Z digest=sha256:7688d4dd8f1b9d4a22d006fde60d491f19457c81e01d83ed16205fdba64201a5

Observation 3cb29c6f-d52a-47f0-903e-2ddab44a79c9 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:01.896301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:13:00.348063Z digest=sha256:bc1119981a1e1c82250c35da754859c41fabb4618e8f51a9f9d566560c674992

Observation 20202767-143f-4d82-9975-74f515d63946 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Gemma 2: Improving Open Language Models at a Practical Size

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.556457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.556457Z digest=sha256:aa39608aad23dce0a8f6d01f45bb647992d9f7521071666ee65830c77b86e62f

Pith citing papers

Observation 6fddc6f9-9dac-44d6-9fe6-1bb5a442007d · inbound

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection cites this paper.

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.554324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T06:31:44.815249Z digest=sha256:ec10fa7c5cdf533db77efd6b6c11f39b0531040e1391c4e6b8cbaa10d6fd094b

Observation 4b4c0f79-3c6a-4168-bb90-8a1d61ec4fc3 · inbound

Mathematical methods of reinforcement learning cites this paper.

Mathematical methods of reinforcement learning How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs

Reference 116

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.758040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:f5b831628bbef22f4f94c88adeccd7547c270ba5017a827213f6868f68870507