Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

As of 18 August 2026, this Paper Citation Record lists 100 of 210 outbound references and 17 inbound Pith citation observations for arXiv:2508.13167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13167 v1

Coverage vector

measured 100 of 210 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:57:38.586633Z

measured 117 of 117 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:55:37.907213Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

100 of 210 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03a0143b-e32f-42b7-8b9a-3f79642a9777 · outbound

This paper cites Towards Effective Code-Integrated Reasoning.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Towards Effective Code-Integrated Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:31.528964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:31.528964Z digest=sha256:4d09929c4def9082c59af9bc23b5936585a36e08703fd6c36a8d27bbb0c57165

Observation aec18274-6b2c-4ac8-8320-231110c0218b · outbound

This paper cites Multi-agent reinforcement learning: A review of challenges and applications.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Multi-agent reinforcement learning: A review of challenges and applications

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:31.628867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:31.628867Z digest=sha256:8c3c57aefb9b486979a62f5ac37d5ab3845494437a4a3aad40705bf0a60c7cef

Observation 5d9241b2-eeb8-4ea6-b3a8-f5ced24ab1b6 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:31.713985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:31.713985Z digest=sha256:038d3af5c7ff242e21ec5c758e58e0d41d4820b033b346e683bd338d66a457ac

Observation 78de714f-8f74-4c57-b429-3147a784ba24 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Process Reinforcement through Implicit Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:31.801869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:31.801869Z digest=sha256:1e0f6d55d2a2c49f7e7d34ec8b547be373eb9b505b7179356ce9bbccc6460d53

Observation 6d12bb68-7528-4280-9e78-e7c7d7d238ba · outbound

This paper cites Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:31.860107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:31.860107Z digest=sha256:8aa2238bc77c9a125278189d8558227f686022f6e9b2b2b6e90710c343ab0a9d

Observation f938d6e4-c9ab-470b-bbbe-fa16b4fd35a4 · outbound

This paper cites Multi-agent systems: A survey.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Multi-agent systems: A survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:31.976176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:31.976176Z digest=sha256:ce9adcc3db5b3f7d4257fd1973ff164dcf8a436b2d9926e4812fca7d6ab02a6c

Observation b2afc82b-2a1a-495a-a462-851c017dea1c · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.038564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.038564Z digest=sha256:c6a732457423747d27d2645762d69aeff39ecf9596402783d17394f94113ef5b

Observation d7c323eb-4fe7-4925-8e5c-cd209bf84eb6 · outbound

This paper cites AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.134517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.134517Z digest=sha256:fb59b84fab4b1982abe1efe47a48d918c88300aff794b99c70d0cc926a2f69b9

Observation 1ebec6f0-8c81-4910-8bcf-7fee5a73e4fa · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.221011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.221011Z digest=sha256:1947bb87b372725c0138b920e6d8bcf4ec32fccba47b974f070ca356d3ec74c0

Observation 446c71b1-1cc7-47ac-bd20-23c2c81e9da5 · outbound

This paper cites How we built our multi-agent research system.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL How we built our multi-agent research system

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.315489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.315489Z digest=sha256:77231c131d90e86e98a3434eee851b9377ee050f0a89f5b49c0915887f93dce4

Observation 94b0014d-7d06-4d5c-aa20-005f30600bee · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.437149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.437149Z digest=sha256:b6eec069cc3b6aef47426beafa5ea7060035bcc0ebe1dc3ebd44696c9875fb77

Observation d13b30e1-d8ff-4535-975a-5062af49d861 · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Skywork Open Reasoner 1 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.471543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.471543Z digest=sha256:7b8f4890887b51793ada218a1196718f32c4fdea78fa24f1ec67f74481687671

Observation b85af062-d1d0-499d-b050-0154e419afa7 · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.528248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.528248Z digest=sha256:a21201a6fb88768fcd847f51e22aedaf97a3c7095f581a8021f54e17a6f4c344

Observation c6f539eb-1a58-4bed-a208-5a9535f3dc05 · outbound

This paper cites OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.626007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.626007Z digest=sha256:530c731cf52eb39d0bcd1401556609700ef17ef14bd4b88227adf77a9abd9a97

Observation f0fcd6a6-61df-4133-afe5-f7e729f6796b · outbound

This paper cites AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.689715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.689715Z digest=sha256:3d4ca843460d0b423f3f50d68ae62617c6b6c2e243e2ba440904086ea12a50d2

Observation 569bcbb2-9b69-42af-8c16-e24c18f1e10a · outbound

This paper cites Qwen2.5-Coder Technical Report.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Qwen2.5-Coder Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.723395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.723395Z digest=sha256:9cadcbfdd5e041433298391ab2a429b27b86f08796e1313749bdb44b8f92ba12

Observation 5ba5575f-73c2-4c8c-b042-4dddfe5900d2 · outbound

This paper cites Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.791294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.791294Z digest=sha256:8ef6bd85054e37a151dca268f158185f439bb597b5d3367f9d279f776c9fb7ae

Observation ab7cd166-2d4f-424c-b04a-ff638f752448 · outbound

This paper cites Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.820793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.820793Z digest=sha256:7a4ffd7329ccf91acb88f47fec115acd451c346bd27fef290b66ff92e6f0b8a3

Observation 56debdf8-3f47-45cf-abac-8b59149eb863 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.881563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.881563Z digest=sha256:1545c42389375add8165d8605121ab085cf5f67b6c181d64a09e8fba696834f7

Observation 0e843709-2b3a-4e48-b8a1-38af653cd782 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.948738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.948738Z digest=sha256:6ebf6a93ae97e05aedfb657f7aad751a73e52615bf9f72e3ae35688d560a3c1a

Observation f375e221-b36a-49db-b778-fcbd77e91648 · outbound

This paper cites Reveal: Self-evolving code agents via iterative generation-verification, 2025.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Reveal: Self-evolving code agents via iterative generation-verification, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.984142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.984142Z digest=sha256:10c9acf5e06c0b4c99a5ba7e779f578f996a082ce1191d1e56e7e346f0e10cf0

Observation 8b88e872-4291-459d-b147-7b7809f15a36 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.011052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.011052Z digest=sha256:31b309e9bae689c2973bd86600f401538b42339f84a68a733df8ef5ffb90b7d6

Observation 491219c8-5731-4cf3-9ead-c526dd628817 · outbound

This paper cites Sequence-level knowledge distillation.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Sequence-level knowledge distillation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.095474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.095474Z digest=sha256:0d00cb4f3496475d04de10e588e1cbb5a0bff15799f6c0563322ce9b4e584102

Observation 4d480276-2117-48db-bed9-e2b3747df3f8 · outbound

This paper cites Natural questions: a benchmark for question answering research.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Natural questions: a benchmark for question answering research

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.161310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.161310Z digest=sha256:894389d4396396361a895ce3c8d2c589c8c1f0014bb96b64a5f43d30e6417f31

Observation 1a6955b5-0238-40d6-8a10-fabf57a2607e · outbound

This paper cites Camel: Communicative agents for "mind" exploration of large language model society.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Camel: Communicative agents for "mind" exploration of large language model society

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.222898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.222898Z digest=sha256:98a17407f8cf52bae89c4656003bb937102670496e51d49e80e365cff73019e1

Observation 35eff444-9688-4119-b2fc-d4120c69eaff · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.317703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.317703Z digest=sha256:9f595f08e2213a4f9af392909a484bebdf0caa697e84248b9e6ac0e45b0f2bf9

Observation d45b9d34-84d0-4829-ab88-b378f61b7e81 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.383273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.383273Z digest=sha256:4929ea5ec394585cd8dea84885546057f77f99f8e8d900a90d184325f0abb8d7

Observation b7497d46-9b24-4489-a3a3-658154dbf6b9 · outbound

This paper cites WebThinker: Empowering Large Reasoning Models with Deep Research Capability.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.450894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.450894Z digest=sha256:bae4d92b78e1c7f4a44b9931ae8d104644a052d9d57f1e390f451b4e2a593a64

Observation 0824de45-a673-4ba2-a2a3-3a113de407a4 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.485908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.485908Z digest=sha256:277d40df74de7ec9b00cb2b8440d63f1f4e6888401c3ca22ed0ea6bbf7f5d3a8

Observation d780e1ec-75b0-4e6b-a57a-888b514102d9 · outbound

This paper cites Competition-level code generation with alphacode.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Competition-level code generation with alphacode

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.551858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.551858Z digest=sha256:709cf32bec2fc4cfb3a9ebc33588ac54c848c917f0da6d23ea0e61d038f1d5b7

Observation 9af51948-67fd-470a-bb5f-24d2555930fa · outbound

This paper cites Let's Verify Step by Step.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Let's Verify Step by Step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.617991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.617991Z digest=sha256:db88765a242007af5bbcbd0756fc0cd5b6781fac2fa575ab3641ed66a9166e73

Observation 56168054-a414-4018-b0d5-23cf642948e2 · outbound

This paper cites Inference-time scaling for generalist reward modeling.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Inference-time scaling for generalist reward modeling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.681349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.681349Z digest=sha256:56fd1d2d4a2f6358c81af1816f4759fe7f068471e91838b5ba21afeb37fc4a4a

Observation daeccf5a-4e54-4b1d-99e2-96006c7d481b · outbound

This paper cites Decoupled Weight Decay Regularization.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Decoupled Weight Decay Regularization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.741732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.741732Z digest=sha256:57efb61cde057f50ce05a34df0d2520e28a473c2c156f9763d91d0820d20d0c6

Observation 82b74efe-b908-4be7-8a9d-f20ce2219763 · outbound

This paper cites Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.803769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.803769Z digest=sha256:8dcd54f4c8b8f29fdc07fd87b1ca672c3bb5def31e79e5bde4760747ccdae8a7

Observation c9526066-f0ec-4fa6-9041-eb7abad91742 · outbound

This paper cites When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.871288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.871288Z digest=sha256:63029f0178756c25a911e0d30d909a129a2eac7574c0ae16398f7af6b87aba51

Observation c3dda23f-b55c-45f0-a2ee-e71e1aa333ff · outbound

This paper cites Gaia: a benchmark for general ai assistants.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Gaia: a benchmark for general ai assistants

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.961254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.961254Z digest=sha256:1baa41ad126634e0cc3e3e6ab1708a7ef33cb5a24d1ba7d84365b45d8eee7d9f

Observation 12f83a6e-a7c5-4102-801f-ee5284d61ecb · outbound

This paper cites American invitational mathematics examination (aime) 2024.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL American invitational mathematics examination (aime) 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.994732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.994732Z digest=sha256:326f54943deea9fe8437d961ff7dab5577e745c7db6d8d4e0e58b255ce330532

Observation 71e45e29-ac31-4ee9-a15a-c1cf165fd34a · outbound

This paper cites American invitational mathematics examination (aime) 2025.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL American invitational mathematics examination (aime) 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.028802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.028802Z digest=sha256:3897139cbaa62537c6dbfacc5d6e50111ed62e5734e58cdf03c57ef528ded230

Observation dd56ec39-9d2e-4dd0-9d70-fd0e6f2ab48b · outbound

This paper cites Codeforces.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Codeforces

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.061528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.061528Z digest=sha256:1e1c5f112b5770e7c1c6bf0d0f3bf1f35a7eaec6eb8da756fb1583790d1a25be

Observation 57664f53-66ae-40ee-bf40-62956e47ff95 · outbound

This paper cites Humanity's Last Exam.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Humanity's Last Exam

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.124997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.124997Z digest=sha256:6199f2eabe499be80f43bd97f6d9192a31eefc39e07adda1f63d1022cbd6c15a

Observation ce665dea-d3f5-4d73-860c-e3fdff5debb1 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Measuring and Narrowing the Compositionality Gap in Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.183382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.183382Z digest=sha256:1f11320d6b94161ed1bc1d469cc3c4e292732038cbee7304caa18e3945257b60

Observation 9f7c568a-a605-4617-984f-dbfdab6eaa81 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ToolRL: Reward is All Tool Learning Needs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.215415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.215415Z digest=sha256:fe1f0a02daa6ec98c7f702d5541a3d46d04673745b4968ec08bd40cd120bc677

Observation bc2447ea-9700-4597-8a65-008dda813d42 · outbound

This paper cites Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.272995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.272995Z digest=sha256:09959cbc681bd8f5931759fbfdad3d9dec73f5c8558c2b4ec650290d4b9e72b1

Observation 05afe5d6-a392-4423-9300-045722a72b91 · outbound

This paper cites Qwen2.5 Technical Report.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.308026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.308026Z digest=sha256:5d7c211fd43fc450bf29e5b72cb4a9b06dbdd5c624c600c5b1d9b14cef34f931

Observation 4f0b6d47-7d03-47a1-b67e-6530e49c7431 · outbound

This paper cites ‘smolagents‘: a smol library to build great agentic systems.https://github.com/huggingface/smolagents, 2025.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ‘smolagents‘: a smol library to build great agentic systems.https://github.com/huggingface/smolagents, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.367351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.367351Z digest=sha256:a2a6b95df8b97c690c281e21f07ccb65dc6c7ad993896fc6256dfd8ea99c5068

Observation 3067b16b-3f58-4d30-8537-c6cba9e32084 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.486689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.486689Z digest=sha256:ec690fcdc761c197d4b06ca215011c9b66f25806612854d489231d4cd31d3174

Observation d21dd24e-8f78-4877-99fa-aac28be7692b · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL HybridFlow: A Flexible and Efficient RLHF Framework

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.579121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.579121Z digest=sha256:8e9116705ec008b0d3f4b7dbd9cf2f9a364eea3d51a60376b44aa70ebc79927f

Observation 6813e7ef-1462-4ad8-84b5-999c92740970 · outbound

This paper cites TaskCraft: Automated Generation of Agentic Tasks.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL TaskCraft: Automated Generation of Agentic Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.611121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.611121Z digest=sha256:beddfc474246f8ffa26b6fbbf4392075d6bd3e9b3ddf9a1891a35fcd5655c12b

Observation ad93aa0b-f9c9-45de-b47f-7ae4d9c8160f · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.684171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.684171Z digest=sha256:b01581c549e9e648c52ed7b6682358a6cc5224b1e096bc29739acc76ea9dbca3

Observation 5d5138e2-3aec-4a06-852f-36d9df1b15c8 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.744588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.744588Z digest=sha256:b62c6b42c03b7abb2220c5213b85576084347c64ba684a02edddd8a81cd70718

Observation b0e531ab-4bc4-429e-996f-fb502cc01dd8 · outbound

This paper cites Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.819776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.819776Z digest=sha256:ee6f657ebddb1ca20c99e3103d05fdd547c2baf8dad9df51a81facf7b31bfa54

Observation 2b63bb7a-bc4c-493c-a671-6eb18354746d · outbound

This paper cites Agent kb: Leveraging cross-domain experience for agentic problem solving.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Agent kb: Leveraging cross-domain experience for agentic problem solving

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:34.891354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:34.891354Z digest=sha256:9a03c786d98c91cc2ae19d18d2b56f01e78532bb683d8872c4aaab72e7c73b4f

Observation a0fc01ef-b2d6-40a4-bf48-752f4cd7f1bf · outbound

This paper cites WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.000344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.000344Z digest=sha256:ac0167795f2cda80b89dcfdc3ff266ad90d38c200d04434ef2b15963daa157ac

Observation 997d1e1f-332b-40b3-8ef6-f4b208246bd0 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.101608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.101608Z digest=sha256:427709cd5cd61b50a68963f31671e1284b28763759f401631ba314ce3ad99a93

Observation 7e68af04-f22e-4ed1-9c4a-d3c81c4ded6d · outbound

This paper cites Verl-tool: A version of verl to support tool use, 2025.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Verl-tool: A version of verl to support tool use, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.157650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.157650Z digest=sha256:7a295e8219fdf1f7ba8df9969ffbe182728cab71664d169a5bba42c6714d47e7

Observation dd42928b-5df2-4c13-9dca-a8e7060461e5 · outbound

This paper cites Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.272023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.272023Z digest=sha256:126aab1e9d085a1ae2466a5a7b806cc5bf450b783beefed155b9275618f08451

Observation 978d58b2-0f67-45da-9586-f4a534701f94 · outbound

This paper cites Musique: Multihop questions via single-hop question composition.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Musique: Multihop questions via single-hop question composition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.312642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.312642Z digest=sha256:24369c3b488ba2b201d4f355850cfc2be2d22f114e6eb4d9ec7eede3fc31e470

Observation d91c3373-7184-4626-9b61-c92451bd1110 · outbound

This paper cites Otc: Optimal tool calls via reinforcement learning.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Otc: Optimal tool calls via reinforcement learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.336906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.336906Z digest=sha256:ef8fb14df7beb9a5e832c88d9f7dbc7008642f3a24cfa0facb7b9a42acd23792

Observation 47a1e444-233c-42af-bd0b-a92e31b51857 · outbound

This paper cites StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.446112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.446112Z digest=sha256:c7f267ba783f0e3f4d499fc40b949fadd853223e5345a0793b27bafbb2e336a5

Observation ff451199-ebf8-41dd-a9cf-d0eae2555210 · outbound

This paper cites Corag: A cost-constrained retrieval optimization system for retrieval-augmented generation.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Corag: A cost-constrained retrieval optimization system for retrieval-augmented generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.521369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.521369Z digest=sha256:ca62509598fdf5ba538bd8c9f0902f702f45e53c0f995b290c2fead0da8ac505

Observation 81e9dadf-6486-4031-b452-4f976bd78665 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Chain-of-thought prompting elicits reasoning in large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.594243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.594243Z digest=sha256:049fa4fda64eb54345284b742bfc45dfcc232be12cb7ef392755ecbce54284ff

Observation c4d45d17-24ac-4d7b-9ab3-b79be03868be · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.713967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.713967Z digest=sha256:e6160897d1641b53b3030cc2f1187a3d2ae3241f063705d70bcdbf21d41d80ce

Observation be2f730c-502f-49ba-afd5-8c9677e23336 · outbound

This paper cites AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.839876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.839876Z digest=sha256:54da19906f7a84f023db003493ca30a778ff980d29b1e0967b0de4a74c8fdeaa

Observation e5141165-c921-4a66-8436-6ff4afbefcc0 · outbound

This paper cites WebDancer: Towards Autonomous Information Seeking Agency.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL WebDancer: Towards Autonomous Information Seeking Agency

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.940452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.940452Z digest=sha256:948c997e9e9b3fbb5da25b73f46e7e060307ddeb7d75cf60d47185523a6e3d75

Observation ad09509b-22a5-476b-9da2-a04b43cfa779 · outbound

This paper cites Simpletir: End-to-end reinforcement learning for multi-turn tool-integrated reasoning.https://simpletir.notion.site/report, 2025.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Simpletir: End-to-end reinforcement learning for multi-turn tool-integrated reasoning.https://simpletir.notion.site/report, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:36.051067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:36.051067Z digest=sha256:9ac2bfd74c3181b4ba32620b8e97f5b1bd064c61e50572a24e5b01e32cd4ea0c

Observation 265c1ce6-2293-4363-80f6-774048b00e63 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:36.137902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:36.137902Z digest=sha256:2f44b2da49195dcd42b275de030f58f7137e7633618f512e42a5565b3319f7ec

Observation b4dd1e80-a7f9-410c-8c91-6a50e0b59175 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL React: Synergizing reasoning and acting in language models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:36.262430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:36.262430Z digest=sha256:7468cc6b0bd53b380329bc73a0176ba2da23de25c864b80a483ac93166643a22

Observation a7267780-6bf5-4d8c-954b-797bca6902e2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:36.346541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:36.346541Z digest=sha256:f0b4e00ae9f4405bacfc199aaaaaaf28b6b8c170decf5520aa792f33a658e974

Observation 55570e4e-d987-4a9d-890f-b5f8c7a3f39c · outbound

This paper cites Auto-RAG: Autonomous Retrieval-Augmented Generation for Large Language Models.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Auto-RAG: Autonomous Retrieval-Augmented Generation for Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:36.480706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:36.480706Z digest=sha256:d2dbf76bc36799e975f5dbc4d59d90ca6d996d58a0c32fbebe3ad7522885cae8

Observation 5acf09dd-76c4-4cb9-8a9a-2a43bec7bf88 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:36.557480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:36.557480Z digest=sha256:b8b5608d4ab20bddf771e93072b6ab6ef59f4642617d078309c51132e64ad659

Observation 976deefd-64fb-4a82-a025-15911957bbc9 · outbound

This paper cites Flowmind: automatic workflow generation with llms.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Flowmind: automatic workflow generation with llms

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:36.657481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:36.657481Z digest=sha256:383db77bff7bc26610fe974dd4c185cb158dc7404aae245e728a81426438d204

Observation 5cecdfa7-9a73-4d04-8df6-fbb414618969 · outbound

This paper cites EvolveSearch: An Iterative Self-Evolving Search Agent.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL EvolveSearch: An Iterative Self-Evolving Search Agent

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:36.739651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:36.739651Z digest=sha256:194dd472f1ca5a2461cec2ba239ba198d5747bdf101d045e35d6fd91f4899534

Observation 2bdecd40-5aca-4a54-9bbc-091346685d9a · outbound

This paper cites AFlow: Automating Agentic Workflow Generation.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL AFlow: Automating Agentic Workflow Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:36.870589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:36.870589Z digest=sha256:39d3b6e7d291e47e5f7a2911538812c2275005f26e1993bcae9b962a98d27c48

Observation 3e140002-0657-44be-bb0f-5d16e054be44 · outbound

This paper cites Process vs.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Process vs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.004168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.004168Z digest=sha256:f113e35740668e9feefdbd6ea8feafb20a47055d12a850d893bb81c2dd40b492

Observation aab162f6-2a18-4c07-9bd7-a58fc8dc8939 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.103918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.103918Z digest=sha256:7f38c710398090f202a0c3981194180e59ce22d53aaab64ba3b7c0209a04c41d

Observation 594b156b-bb91-45a3-a408-e3da399ca22e · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.248925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.248925Z digest=sha256:bb6dd891c5ff217416ed9f793fe8a21d38402c12d5bafdeaaef0566c8d1947e3

Observation f0758f8d-b297-46e9-8cc1-d30b5fcd370d · outbound

This paper cites OpenResearcher: Unleashing AI for Accelerated Scientific Research.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL OpenResearcher: Unleashing AI for Accelerated Scientific Research

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.318790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.318790Z digest=sha256:b864bf1a9063818942f166e4ec9339861673f6cb611b37a2555d66efaf94af50

Observation a1e83de9-b3f6-4826-8813-6dc0726be362 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.396632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.396632Z digest=sha256:ab809cf319dac3d41900e8479a6d4e125c2f60e78634149e763a9f7bcdfeb9ab

Observation 9e66a4e0-1668-4f1f-8d1c-035c6775baf3 · outbound

This paper cites Agents: An Open-source Framework for Autonomous Language Agents.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Agents: An Open-source Framework for Autonomous Language Agents

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.471097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.471097Z digest=sha256:e2d8c492b7b1415e06dd6c402508de3f39e893ff8f1a18932e7382bedd4dca13

Observation 3d81235d-624b-42cd-af34-ed281e30b754 · outbound

This paper cites Symbolic Learning Enables Self-Evolving Agents.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Symbolic Learning Enables Self-Evolving Agents

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.538262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.538262Z digest=sha256:aee91532258e8b3c1ea5d7fa2e1bf37bf7e2f214dc84949ac22ab049e34f5a8b

Observation bea47f24-507f-4a7c-aacb-ab5965079b03 · outbound

This paper cites OAgents: An Empirical Study of Building Effective Agents.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL OAgents: An Empirical Study of Building Effective Agents

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.634454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.634454Z digest=sha256:f0e5bd2d9f95ef518ff7e0130ac3c90b803542ee445c58ef80520f21b9e36d80

Observation 4bc80da3-53d8-4761-b32d-2fd4cf67c913 · outbound

This paper cites Scaling Test-time Compute for LLM Agents.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Scaling Test-time Compute for LLM Agents

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.738674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.738674Z digest=sha256:b29ad5799b39e27ee0f1c5faba953e3fa56ece4842ac4cf60ccabd425337c9ff

Observation 692098ca-827b-401b-8f0c-3d73f7049b76 · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.828596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.828596Z digest=sha256:4e5af9f4da342c053aeaec0a2ca473bae505eb8ff6387093a4d945591825cadf

Observation 249683ff-7b58-4757-983b-819eff462a8a · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.840843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.840843Z digest=sha256:bc7fad6a42aa97b12483b80ae93432aa147fc10bda7baa4115db9c19c4726d53

Observation 1cd27c36-f15d-4f28-8e14-c1db3961abfa · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.886039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.886039Z digest=sha256:86abc73afcdd01db784e6496b71a20f20a649426db10eac2012702cc35f726b2

Observation a60e1ba7-42e8-4d3f-a083-84864fc73b46 · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:37.914911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:37.914911Z digest=sha256:f57ba2509057aee55c9b02af50e40471bc363888fea300a5bad69146ead1e98c

Observation 873d6a4c-b7d7-498f-ba4f-5dfc30521fd8 · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.027448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.027448Z digest=sha256:0bc4a10a8f4b6eb1b8ff68bf81b680ff4d309a912fea2add3caa84f30d8a699e

Observation 905295ae-3d75-4b20-9f22-39e4d96e6773 · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.141052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.141052Z digest=sha256:3fe5ba44d0c0c9f2c87e223f46c7365b2efeb39df0fcdfb5042dda48bbf8f2e8

Observation 0919bf46-5b23-41f7-90d4-74524349be45 · outbound

This paper cites An archive of all existing APOD pages (current date through.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL An archive of all existing APOD pages (current date through

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.174023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.174023Z digest=sha256:fee659a23a88d26d63ffa68d019eb50c9b638e6acd5b68b627d12c8c75d53116

Observation 7e8e6feb-dc51-4b6c-aead-2bb9fbc372e6 · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.207382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.207382Z digest=sha256:67e453d348dc9d0749b837756587d48ba10320acbc0405d8e4b4e97ccddc71da

Observation 06be36e4-0a7b-4720-932d-6f6bc3e8dac4 · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.234371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.234371Z digest=sha256:04920f9c33ae3b36af84e70f9335bbea12bd4744860c53d21d0be60e330d6d3a

Observation 2b93abf2-062e-4341-a327-adacede45e6c · outbound

This paper cites 1, 2015 (Credit: NASA/Bill Ingalls).

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL 1, 2015 (Credit: NASA/Bill Ingalls)

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.271828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.271828Z digest=sha256:ec9e0ca1f51ae5d474e8fe595f0159ba95f911e0bcdad21f77acb0b3139aa04c

Observation e796f871-e906-41de-aff2-01168f04413e · outbound

This paper cites </observation> Step 3 <think> Step 1 of the task is to identify the NASA Astronomy Picture of the Day (APOD) from the first week of August 2015 showing city lights on the horizon.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL </observation> Step 3 <think> Step 1 of the task is to identify the NASA Astronomy Picture of the Day (APOD) from the first week of August 2015 showing city lights on the horizon

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.294891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.294891Z digest=sha256:aecbb794da261818e5b0df88a78bc2fc55cac8bb7337a25d2bb8d848732c6ad0

Observation 36753cd6-d66e-4f7a-bae2-1df7277ade9b · outbound

This paper cites Marquette had a population of 20,629 at the.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Marquette had a population of 20,629 at the

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.340917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.340917Z digest=sha256:98ce1fc27f3d9e82b66ac29399abfa27ebe5dd1d0d622b22edd4a4501d756264

Observation 43afb520-cf77-4d5f-8a29-315a88ecbeff · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.385756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.385756Z digest=sha256:b54e77dc429f4a68fc1a7cccaf34aaccd050273e67d9eb4ecc78809348bcb30e

Observation 856fc924-0d4e-4a7a-a21b-72606be3a7a2 · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.424799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.424799Z digest=sha256:eb9437e4c94d8c738af43f800789276e37eb641f5589a940b67163268875fed6

Observation 6c50ade7-1dde-42f3-8be2-eaef972b8c75 · outbound

This paper cites “Back in the 1600’s he set up several missions, including.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL “Back in the 1600’s he set up several missions, including

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.472605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.472605Z digest=sha256:c1c1f2134f25339e5470439f35e5a92082e327f45f5fc310621534f226f342c5

Observation 51893399-7f40-4f42-8504-10392a842094 · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.516192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.516192Z digest=sha256:ab0bfd4b49b054e867e8cc8683b2e5fd172730a6ef6e46c948baf8840cdd876c

Observation 06e14bd8-408b-471c-ba65-a08968c95961 · outbound

This paper cites Completed in 1894, the Marquette Building brings Chicago’s early history to life in an artistic and elegant setting.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Completed in 1894, the Marquette Building brings Chicago’s early history to life in an artistic and elegant setting

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.557553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.557553Z digest=sha256:fe397120076303133c8a544bfbd1aa92197c3acdf5c4598ed27c684b1a0a88c3

Observation 80c25442-ad78-4c32-9ada-b5f4885910b9 · outbound

This paper cites an unresolved cited work.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:38.586633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:38.586633Z digest=sha256:b81b2e570069450dbdc8fe0f8470f523b7121aeb6e2b53ab380ed7e3037606c5

Pith citing papers

Observation dcdc812e-5f22-4df7-b1b8-81b57d114aa0 · inbound

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents cites this paper.

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:37.907213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:37.907213Z digest=sha256:59d2ebda1f252013a127753160b058f09ade2c472837c4776ccb9d3f42e2036e

Observation 1a7eaf6c-332a-4b96-a4a8-af5f03ecb7d2 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 283

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.747920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:b1acc71496aa4dc3526adb1efb5bb0088dafc6c0af0492ef8aca2dbf99587cba

Observation f01437bd-012f-4fca-8e25-4f829e0cd832 · inbound

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs cites this paper.

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:52.383322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:52.383322Z digest=sha256:c33e3b1a91120498f8a4ca3c4423bd4e1110fe196327abddd41dc2787928c0da

Observation e54dbf78-fc94-44f5-8dbc-82d62a9e7f40 · inbound

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling cites this paper.

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:45:17.898438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T21:44:18.744201Z digest=sha256:797dddab972777ee3818b05e7374e634e6f53c82c5736f59f60d6d1bbfcd5299

Observation fcd7cc45-b888-4cd7-b8ac-6651321df2f3 · inbound

Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework cites this paper.

Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:20:53.524753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:35:29.920327Z digest=sha256:f12e902d46497989ad9ad97d50c26a64ebbf8c83c982366e8ebff9085b524b4b

Observation 45d13e58-7e23-4823-89be-325b8f3bb308 · inbound

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent cites this paper.

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-05T15:11:10.916430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-05T15:03:50.420072Z digest=sha256:f885377d6672a16d7121e89cfa5fed3f64941e1b0513986c194656de8f3ddcf6

Observation 3c6f6425-b17e-4b45-aa4c-911cbb85099b · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.213509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:e9ada31058e7537ec66d1051933e74f0ecb4b9662446a1b29fc10b14c9ae56f5

Observation 61c22ce0-a9eb-4916-b90d-779314c5c6a1 · inbound

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate cites this paper.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:20.540725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:da390565f36a990beb904e8358d20fcd02673a4e84b393fa6ed380a5316af610

Observation 0c4865a5-0294-4186-b4c5-e2124955073d · inbound

SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States cites this paper.

SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:41:06.769144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T17:20:53.540362Z digest=sha256:fe589352b8fac5ed8108f22d9127995300b27c4939dfe55ba6e668af62c22bde

Observation 1e23cc51-8da2-4836-a10b-219e68569a74 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.388059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:40071627400f3a20dccaff8269e5f1e38c8f0c7e914422505f9771451db1478a

Observation ee513b98-3b27-4c01-8302-c504f50af3bd · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.013251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.013251Z digest=sha256:e5da1d87a9eece1c7775ab3ae7cebebf73339ca8ad764fc0fea4f04ee3dcec28

Observation d7b11d48-501a-4101-8a2e-0be63c69dd0f · inbound

Scaling Mobile Agent Systems: From Capability Density to Collective Intelligence cites this paper.

Scaling Mobile Agent Systems: From Capability Density to Collective Intelligence Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:24.624276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T00:48:57.036609Z digest=sha256:9f5c748662ec91610fcd1c1b02263493fb0e7de289e971e9ebcf31c0d3da1b0e

Observation f1c0f03c-60f7-42d7-9d7c-12631d1e739a · inbound

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning cites this paper.

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.426375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T17:58:05.817581Z digest=sha256:3eb8e6cf5b7c51c6b56b2bddae5734409ddb31ee43d5124377f4cb09e0626fd7

Observation d80bed98-4296-410e-bd21-44b6724818ba · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:34:41.023546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:00b3a4f78129ca2e2a471fbcdd54bdfbfc67c8b265cb92c82a03641d09b52ed5

Observation 7dae01a7-445c-4944-b479-069ec03d4324 · inbound

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning cites this paper.

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:28.234364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T17:28:58.574865Z digest=sha256:f613d0f3eff3cb0db73347a4bec59410503bc257a415792158992c40eed726a5

Observation 9aa5ba37-f362-4f0e-8428-9d85470c07af · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 262

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:50:48.446656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:99998998fd854cf4b494b807f8dc34b13852c7aa9bd390f6465610c01da16f71

Observation b1b9892c-b075-46dc-8e03-96c7a9adf5cc · inbound

Mathematical methods of reinforcement learning cites this paper.

Mathematical methods of reinforcement learning Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 121

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.743625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:ed1adfbbbc3c9010f5c966109d25e552bc4fd26102df595bcc284674c3b758c4