Pith. sign in

Paper Citation Record · LEDGER

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

As of 21 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 27 inbound Pith citation observations for arXiv:2501.11425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11425 v3

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:22:02.778783Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:31:12.889246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T13:56:19.181721Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved39
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc30c406-cd27-445d-a861-b0f1b83d3663 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training The claude 3 model family: Opus, sonnet, haiku

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.432865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.432865Z digest=sha256:478cc25ac5dd157cfcf0bad10166f67e10ebb781dc5188ecc557862597f3b0b0

Observation de5b2e6b-ec7d-4f59-9342-ee5d6bcc9aa9 · outbound

This paper cites A survey of monte carlo tree search methods.IEEE Transactions on Computational Intelligence and AI in games, 4 (1):1–43, 2012.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training A survey of monte carlo tree search methods.IEEE Transactions on Computational Intelligence and AI in games, 4 (1):1–43, 2012

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.971502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.439192Z digest=sha256:f57a04f4dd8f1eb102688dca23255c3b0f4a9780a27b22c28996e6485e373f3a

Observation 8056b4fb-a666-4045-a1bb-097288ad86ce · outbound

This paper cites AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.444592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.444592Z digest=sha256:a7df3c75b2babbe7672ca256703289895a25a23750c8b54bec6317c98ade21f0

Observation c2bc9fb5-a5b4-4ef4-8c9d-68613b0e36d1 · outbound

This paper cites Teaching large language models to self-debug.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Teaching large language models to self-debug

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.953505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.450516Z digest=sha256:afc846edac1c4959d0a27f94a9e3fc72a57656703ca1ab9ebaa94f62181dd4f5

Observation 9787c36e-e909-4162-b804-af0ca607f57a · outbound

This paper cites Mindsearch: Mimicking human minds elicits deep ai searcher.arXiv preprint arXiv:2407.20183, 2024.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Mindsearch: Mimicking human minds elicits deep ai searcher.arXiv preprint arXiv:2407.20183, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.455638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.455638Z digest=sha256:88d7cb9a0172cba8c1b20d6fe728f57503c7d4af576080c3069447524f733ce5

Observation 6210d6f9-f99d-4d3f-81e6-2725fdfa4412 · outbound

This paper cites Agent-FLAN: Designing data and methods of effective agent tuning for large language models.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Agent-FLAN: Designing data and methods of effective agent tuning for large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.460876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.460876Z digest=sha256:c16d7039064592bd571541a84cd113e3fc06e05add89a79162f762a70ee41b75

Observation dd7fcaad-0590-42bf-ad33-d4d5966e6ef4 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Mind2web: Towards a generalist agent for the web

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.936605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.467259Z digest=sha256:b553fe9c26db107330bb868b2af8f38b7962a1980212a615252002aef067ddb8

Observation 26f30bf4-4a45-49d9-a9fa-2141e6728f8e · outbound

This paper cites Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:22:03.435529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.472300Z digest=sha256:ed439e43bc512cf79af580bc57a55201a003db33461debab5c7a303fb09b11b1

Observation c0765e97-8f2c-47cc-b15e-79b0dccb2673 · outbound

This paper cites The Llama 3 Herd of Models.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.477330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.477330Z digest=sha256:b17a84a1ebeefc771e3f062ab885e101230ddf3d88e418ffb2e64c949bf04bbd

Observation b23979a3-e773-43ec-88df-a1913136c8e2 · outbound

This paper cites CRITIC: Large language models can self-correct with tool-interactive critiquing.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training CRITIC: Large language models can self-correct with tool-interactive critiquing

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.917830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.482619Z digest=sha256:a533f48245ed5582c3d70da2593ad0b1305c559f9cf3d359e910f21fa674a755

Observation da80a68a-9e6f-4df9-b7a2-8b56d5762239 · outbound

This paper cites Reasoning with language model is planning with world model.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Reasoning with language model is planning with world model

Reference 11

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T18:22:03.901045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.488670Z digest=sha256:3f6b42d7cfc828fcac1e9a2edec4a0340b3a8b57ce2943d1de7dfb858ec3afe8

Observation ce9379a9-42a0-4f69-a7e5-203032654627 · outbound

This paper cites GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.493706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.493706Z digest=sha256:fbbd65851dcd1a96ba3f8b6927477c01686b2384a27c017190e951785560846b

Observation a00a9c5c-f345-46d2-90a7-f5e9191715aa · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Large Language Models Cannot Self-Correct Reasoning Yet

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.499194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.499194Z digest=sha256:f49ed07640c15f54df6107a0fd5bd8292dc7f164b012c6e1247028ed9df1e680

Observation bc36c2b5-9564-4f71-baa2-68aab2079569 · outbound

This paper cites When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.504569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.504569Z digest=sha256:e5967ac8860d2091defcaa232b32b49de12387292bea14e7c6b081f432336ae7

Observation 8c0470dd-b146-4bbc-87a6-5b3811586b23 · outbound

This paper cites Language models can solve computer tasks.Advances in Neural Information Processing Systems, 36, 2024.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Language models can solve computer tasks.Advances in Neural Information Processing Systems, 36, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.882033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.509272Z digest=sha256:ed4dcb94c81cfc8589d2e3c156d9d81e28b19884f0781e867a925b38df984ba6

Observation d1bbc986-49d0-4ca2-896b-78f13456cb4c · outbound

This paper cites Bandit based monte-carlo planning.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Bandit based monte-carlo planning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.513954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.513954Z digest=sha256:41d85c946a3499175c741291924687ed5f4da08309e41da51f7829d8345a82fb

Observation 81a94272-d4cb-4793-83b2-cab28001cc60 · outbound

This paper cites Tree search for language model agents.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Tree search for language model agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.518668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.518668Z digest=sha256:3dca2d0fed56d5cb3d40929244c6cbfe30ad267b49c213883e8e664ce811f87e

Observation 3ea87d23-bed6-47eb-bf51-54565b8895b2 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Training Language Models to Self-Correct via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.523043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.523043Z digest=sha256:d26934f9fa7f499001d26a08ebc910109ea7146443ebbb7119e51a0652ff4cfe

Observation dbe50344-23bc-4dd6-b5b5-afd588893e4e · outbound

This paper cites Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:22:03.251016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.528542Z digest=sha256:60482f56af519ab75a464c58c5f7ad254b20d934e27516650a28f8dcef2bb296

Observation e4b7a4b8-e042-4425-be3a-058db3c31911 · outbound

This paper cites CriticBench: Benchmarking LLMs for Critique-Correct Reasoning.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.533704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.533704Z digest=sha256:2f696b2985223ed1ed3ece080d4c6ec15be54a3b4ce8f46509dd73e4b96eb6b1

Observation d3775be2-59b5-40c0-89c5-ac32b53d0c83 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.539124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.539124Z digest=sha256:8e2f6c7a07dd8c9a6dd2f82fc4ad0105db62af5057977d3d41fcfe15cd44b82f

Observation 65759579-653f-4f03-b936-60b3d1dc65e8 · outbound

This paper cites Agentbench: Evaluating LLMs as agents.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Agentbench: Evaluating LLMs as agents

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.850307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.545155Z digest=sha256:30df0dae84b954067b59bab24d26c64543482c2231a99de4a05f3c0e07d2b6c3

Observation 50791d4d-44c9-4e19-bc7f-943e920f4d32 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Self-refine: Iterative refinement with self-feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.833741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.550138Z digest=sha256:f6db1cf259ccb34b60f94d746e1e955c990cce596427014b393e712d29ffc076

Observation 81f6ec6e-87d6-4675-9b33-216e3f464499 · outbound

This paper cites CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.554860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.554860Z digest=sha256:a10fa3c702735b8dd7b8256c271a1311a729cfe94ea23c8865514e8cdabadb84

Observation 2ba0a2d0-9a25-41a7-b62b-6bab1660447b · outbound

This paper cites Skill set optimization: Reinforcing language model behavior via transferable skills.arXiv,.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Skill set optimization: Reinforcing language model behavior via transferable skills.arXiv,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.818445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.560049Z digest=sha256:1078f0ba6bf663305724624489e7dfc43ae86c60817c9dd49111011002d94442

Observation 1aab42fc-a660-47e3-8191-1d5f5235d877 · outbound

This paper cites Is self-repair a silver bullet for code generation? InThe TwelfthInternational Conference on Learning Representations, 2023.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Is self-repair a silver bullet for code generation? InThe TwelfthInternational Conference on Learning Representations, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.800890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.569970Z digest=sha256:f69ff06f18cf8c1a360c1e7d41403024f798ef93b2a03adb51740615b19e451c

Observation 4ff89f03-d10d-41b3-a00f-8d650f6bfa51 · outbound

This paper cites Chatgpt, 2022.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Chatgpt, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.785588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.574882Z digest=sha256:98428786296c18ebeb67844d9fbc916880e073d8010ed9e6035bd172d3032e01

Observation a3d8e3b5-4dfe-4198-a0c2-ab44ed9e8891 · outbound

This paper cites GPT-4 Technical Report.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.579427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.579427Z digest=sha256:9e3ca6183ad66bb5d31736b4415ccdb947db1207a4c1906272622112280d22e8

Observation 017dd6c6-f96a-49f6-b60e-fbf466411766 · outbound

This paper cites Automatically correcting large language models: Surveying the landscape of diverse automated correction 16 strategies.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Automatically correcting large language models: Surveying the landscape of diverse automated correction 16 strategies

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.770225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.584365Z digest=sha256:a2e56895e07243ae7d23ca6bd75a44a0fb7fe2c10d21619dfef783ed2176eb7e

Observation 9b16a689-438d-40d6-bb91-282ac9963ae4 · outbound

This paper cites Large Language Models Can Self-Improve At Web Agent Tasks.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Large Language Models Can Self-Improve At Web Agent Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.589106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.589106Z digest=sha256:e971db38fc858762842ddd6017c9b6b5353753811d976163b07ada2b42ac7ea3

Observation cf92a693-a6f9-40d5-bcd9-91eab11af435 · outbound

This paper cites ADaPT: As-needed decomposition and planning with language models.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training ADaPT: As-needed decomposition and planning with language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.593979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.593979Z digest=sha256:5cb3422d3ca68c2c1c8066c4dd452c7fc0345ae107a64833a0fc8359093c94a5

Observation 2ade1b4b-ae86-431f-a876-ce7674bababd · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.598640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.598640Z digest=sha256:b1b5b8c2dee899c93e86c4af78ffdbee33815aeae382084277c1185ca13e2e09

Observation 9bfb9819-8c81-4aef-87d9-67fdbb1116d2 · outbound

This paper cites Agent planning with world knowledge model.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Agent planning with world knowledge model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.752891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.604585Z digest=sha256:8c989cd52be579c14639d3db912d5998832d0ddcf88e51a5d2bf278599584744

Observation 5bf36f80-3d4f-4ec9-b0a5-af1cdd27e2a3 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Direct preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.735679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.609858Z digest=sha256:a84c89292aa226c3ed8238fa9c7e80feafa92573413b4ac8ec1598b5161153f8

Observation 4aabfb35-7752-428e-8b4a-bf5326e186d0 · outbound

This paper cites Tarr, William W.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Tarr, William W

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.719042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.615781Z digest=sha256:edae39ca94e3f1ba42526a686d9d3d9b0457eb2b0f62232659f45611bcbd85fe

Observation 5809b1a9-2387-48c7-95d6-9c2197943fd2 · outbound

This paper cites Training Language Models with Language Feedback at Scale.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Training Language Models with Language Feedback at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.620898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.620898Z digest=sha256:8ec3a72b70018f52b6194fbc444dde4e611f04cfa9d3507815809fb510082880

Observation 7a7a265f-2706-4240-b3d2-97d936605bf7 · outbound

This paper cites Direct multi-turn preference optimization for language agents.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Direct multi-turn preference optimization for language agents

Reference 37

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T18:22:03.702618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.626075Z digest=sha256:c1ed960bffb8a9637ea389f36541411239de6e70fe573369f4ae02d9ab0e1774

Observation 0969a0b0-779e-48d3-8434-8c62e6ceed97 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.630904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.630904Z digest=sha256:d5fa45f73873b013eb5818c53aca91ac681d47d4621a2fdc81b7a655af7b9846

Observation 074fb72d-495b-4f9c-9f98-5387ed6e8771 · outbound

This paper cites AgentBank: Towards generalized LLM agents via fine-tuning on 50000+ interaction trajectories.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training AgentBank: Towards generalized LLM agents via fine-tuning on 50000+ interaction trajectories

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.636678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.636678Z digest=sha256:ef15bcb03f6189d5189a2f56bb0c574f2346bb035c18ca5194cb17f75df8c005

Observation 80d42ed4-d03e-4cec-97a7-a1ae631c99e8 · outbound

This paper cites Trial and error: Exploration- based trajectory optimization of LLM agents.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Trial and error: Exploration- based trajectory optimization of LLM agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.642438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.642438Z digest=sha256:91b963475ee86075c3bf11d90ba4cf208b8b2d9182f9536f651a3c09f330dbb6

Observation 14980220-a53a-42f9-8068-999085826d5a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.647644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.647644Z digest=sha256:ba41032bf998ff7d980508a67d6e983d438243331753a153ee7315def80f9edd

Observation 0c7b7ecc-f1b2-475e-9124-50a1bffb1152 · outbound

This paper cites LLMs cannot find reasoning errors, but can correct them given the error location.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training LLMs cannot find reasoning errors, but can correct them given the error location

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.654533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.654533Z digest=sha256:28d9f100ed3a6ecf8a44311ba0b3fac644a6f124e191fa9024e70dd615f05766

Observation 2f57b7be-685c-4589-a5f1-d83132e06c45 · outbound

This paper cites E2CL: Exploration-based error correction learning for embodied agents.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training E2CL: Exploration-based error correction learning for embodied agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.660044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.660044Z digest=sha256:2702f82224eb2d5e2645cce55eb0bc12c1a95af0a2ac9faab2c587634847def8

Observation 52a95e2d-0e36-4fe4-88d8-5ed75631bc11 · outbound

This paper cites Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.665226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.665226Z digest=sha256:c435375f1c6009c2320723bcab78e8795ff634f4cbab1da01f146e933452bcd5

Observation 1191ddaf-e89b-4d2a-890d-0857e25a1068 · outbound

This paper cites an unresolved cited work.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.669844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.669844Z digest=sha256:68e88737b36e20070def2b473b0495ad99f849c2957cd68f9efa0406a692afaf

Observation 04a7e4b1-7a16-494b-9010-50aa380defec · outbound

This paper cites Generating sequences by learning to self-correct.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Generating sequences by learning to self-correct

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.686466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.674460Z digest=sha256:54ee976e4628d676ef24c432bb344f62e7ebd66080d7a0d3b665c1b332f4efd3

Observation 10a268ff-9cd9-4e05-b738-7e099d664208 · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.679923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.679923Z digest=sha256:011487e1c2374813d3b17a99542612103ba33022f447385e4bac2e8101c4b567

Observation cf3f601e-e429-4d2c-9ab2-ec4c26db51de · outbound

This paper cites AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.685354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.685354Z digest=sha256:d89b2ab14b122e82c4f73638cbd143d4e34a0bb36e18fd9b084667bf3181ff37

Observation c206319f-92cb-444b-9adf-3a15c3694d30 · outbound

This paper cites Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.691028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.691028Z digest=sha256:e8a486702108c77619fa00772222dfefe576a47556cca0de299ac870464f65ea

Observation 4cd919ec-6f6e-450d-a938-e4d10ae54cd9 · outbound

This paper cites Revealing the Barriers of Language Agents in Planning.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Revealing the Barriers of Language Agents in Planning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.696782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.696782Z digest=sha256:09e8478cade1c6575a5fb2854e70c83a3d6cd4ba1ab5174d9b53bed5c7d4a737

Observation c075135a-30b0-4840-a78f-ac5030ce8299 · outbound

This paper cites Watch every step! LLM agent learning via iterative step-level process refinement.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Watch every step! LLM agent learning via iterative step-level process refinement

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.668662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.702547Z digest=sha256:1d924f10c731491cc3111aeed8dcde2b8863ff8a4d4034928571cb6e7396f0da

Observation e74debc2-423f-4817-9ebd-665e86ea25dc · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.650796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.714040Z digest=sha256:0dd22150a0ad3fed97e21a8dd5ed425c2d72650881ef842aa3f8bf16512995c8

Observation dc5e646f-feee-4270-b631-4d500d6a64c6 · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.93.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training doi: 10.18653/v1/2024.emnlp-main.93

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.708296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.708296Z digest=sha256:2a9cce0b3fa43c3f4c0c10b0a0c8dbe0a549538aa8953ce2e2dccd1cc25c998a

Observation 93c63a08-8297-46d0-a87d-5d035a6249d8 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training React: Synergizing reasoning and acting in language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.617617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.725688Z digest=sha256:96063fb71945764202f7b4fc9051ed03085f88365601d0ac9cc8830769383bda

Observation cf35c683-d0ab-423a-882c-f5201de890fb · outbound

This paper cites Griffiths, Yuan Cao, and Karthik R Narasimhan.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Griffiths, Yuan Cao, and Karthik R Narasimhan

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.634254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.719972Z digest=sha256:42691e18b7253b5d961ddf8877abeaea0508c2267146e39cae15e6910b29fc78

Observation 92a6e8a6-28df-4483-836a-1c185e033216 · outbound

This paper cites AgentTuning: Enabling generalized agent abilities for LLMs.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training AgentTuning: Enabling generalized agent abilities for LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.737096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.737096Z digest=sha256:3f2015e27722cf32c81f67271293fe2cf18002d9a990f8c08a57452a96a748d7

Observation 24abe105-cf0b-434e-9722-75430a43dce3 · outbound

This paper cites EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.731402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.731402Z digest=sha256:cf303c510ebcae1560c5f566a550e0e40f733d05c59e6625cd050ca45a6dce96

Observation 055d0052-2503-4156-b4c9-94ca80654ecb · outbound

This paper cites Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.748089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.748089Z digest=sha256:f10b29f250927ce022c31123e974eb07a87e008f2cec5e1e1cfd303802c9a524

Observation 2c684ccf-89ed-4073-8934-75df5e4e3385 · outbound

This paper cites Mr-ben: A meta-reasoning benchmark for evaluating system-2 thinking in llms.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Mr-ben: A meta-reasoning benchmark for evaluating system-2 thinking in llms

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.599236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.742711Z digest=sha256:203384acd9f9c48647beb28aeaf318b4e9f0f36bd64bd2cd6da65d2b8736f8ac

Observation a562d886-1e95-4db8-aedd-b9fac972a046 · outbound

This paper cites TimeArena: Shaping efficient multitasking language agents in a time-aware simulation.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training TimeArena: Shaping efficient multitasking language agents in a time-aware simulation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.581994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.759929Z digest=sha256:5db478ac7c1710c824be8ba11b1bb93f99b126e5ad00a0955b53f897c3597bed

Observation f1a7ad6e-ff5a-4b07-a3ee-63a1deb9b859 · outbound

This paper cites Understanding the Dark Side of LLMs' Intrinsic Self-Correction.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Understanding the Dark Side of LLMs' Intrinsic Self-Correction

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.753853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.753853Z digest=sha256:12ca39ef1500a8b74c5e41d888ecd00e046eec6ccae478297c004b357b7c215d

Observation bb57cd0c-894e-404e-ad6c-ce3f6c9de42d · outbound

This paper cites Large language models as commonsense knowledge for large-scale task planning.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Large language models as commonsense knowledge for large-scale task planning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:22:03.553059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.774196Z digest=sha256:f500e30ad40768b8512e84379117b292489fcd811c781ee9ac7aa66ea97c4e3b

Observation c61f8360-c129-4225-bc9b-1b7963c77796 · outbound

This paper cites doi: 10.18653/v1/2024.acl-long.215.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training doi: 10.18653/v1/2024.acl-long.215

Reference 63

Resolution
verified exact
doi, observed 2026-08-10T18:22:02.819846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T18:22:02.765103Z digest=sha256:8de0f5272fbcf6faea98a92c8f6feae420d4b8fa62bb51b5eb4e04a441727c16

Observation 862c5247-0abc-48b1-a4bf-12cf69eea0eb · outbound

This paper cites Expel: Llm agents are experiential learners.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Expel: Llm agents are experiential learners

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.769678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.769678Z digest=sha256:a18171f6cf8f7d2676e1c8679ed76ee2b6c5ca8d3be20f86e848d1314dc86c74

Observation ad1a84b3-d726-436e-a107-0b84beaf7b08 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.778783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.778783Z digest=sha256:24d54c4dddf31532463d9117b85e39d64a8deb162801b2582ec1efc290235ad8

Observation f3c65df3-ac3a-477f-bea1-81b2af148280 · outbound

This paper cites Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.564724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.564724Z digest=sha256:01b20f0b268288e9ecfa24a2f495b6f2d675a2ac676d23e27f05fef7d06e7d35

Pith citing papers

Observation 6783f560-53d4-472e-a0b7-7c0b9d4f7635 · inbound

To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization cites this paper.

To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T18:07:52.239881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:07:52.239881Z digest=sha256:6214a139667e40bf0109b8fa9fec9afe8a9ffdf3b026242c1617aa225c03444b

Observation 8e371134-90ec-4c96-b54d-11d522178d0e · inbound

Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience cites this paper.

Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 262

Resolution
unresolved
no resolver link, observed 2026-08-15T23:31:12.889246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:31:12.889246Z digest=sha256:b207183a07cd2afd6c4fcf9e65db3f302e2040cd3519de5c52efc1ebfc3b0cbf

Observation 07ef6a90-fbf8-41b6-b362-2ba868b5ac26 · inbound

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution cites this paper.

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:03.357476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:03.357476Z digest=sha256:073724d1b42043fb76debecf27cd0f0d3d091b9d8a2ebabb454ce4f4f4ff9cf5

Observation d54b30e9-74f0-491c-911b-2e4fe225f983 · inbound

Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary cites this paper.

Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.878565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T11:34:44.319579Z digest=sha256:295478b4164b3d07d94118961bc9a528e5351c1ce28b3901b3ca19d5bf8728b0

Observation 43fc11a3-37d9-4aed-a4af-3d44e94031d9 · inbound

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning cites this paper.

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:07.875213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:07.875213Z digest=sha256:d1261e8a406fb968be9bf2cbb71f55e6ad4f2097742aaa7eb5b84ffee03fabb2

Observation cc3e3565-2606-4351-a445-408950042b7d · inbound

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents cites this paper.

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:27:37.449450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T00:27:37.360221Z digest=sha256:5e15a4f06f94c33d44b18cddcaea0826497c99ed3645b54fb2c0d0f464f8cbf6

Observation 60ac31f0-aca5-4678-9ad2-e093e0cffafa · inbound

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning cites this paper.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:05.817610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:05.817610Z digest=sha256:b83eef93242b67b125dfcb00203f1c29d66bc4ad75888e5add29d25167fa4465

Observation 5e784283-0055-4fc5-9ae7-06df5910fdd3 · inbound

SAND: Boosting LLM Agents with Self-Taught Action Deliberation cites this paper.

SAND: Boosting LLM Agents with Self-Taught Action Deliberation Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:23.023790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:47:23.023790Z digest=sha256:b7ba7417e4eb01dd8034a77f68df07abc0b9aa2e362f2a1170d8e6a6f207fdc1

Observation 11389ede-4427-4e8b-9513-85c8adb70359 · inbound

Agent Safety Alignment via Reinforcement Learning cites this paper.

Agent Safety Alignment via Reinforcement Learning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:35.842058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:35.842058Z digest=sha256:d71872b0215b14bd39fdfbb0dd2ad30e823875e43788477c51b1c65669733ef5

Observation 51204792-88b2-4000-8f53-bc77a7647bbc · inbound

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents cites this paper.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.676834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.676834Z digest=sha256:0a046014fb4ffccddb59d4253099eaa69271f8a79a029cf385c3e3fc5c1479aa

Observation e6063efe-068b-4e87-b5ff-a213379dd495 · inbound

ReQuestNet: A Foundational Learning model for Channel Estimation cites this paper.

ReQuestNet: A Foundational Learning model for Channel Estimation Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.124905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.124905Z digest=sha256:29401f3cd5ab1c5b291d70e214148bf9c860a13859a815724096a3545e27547b

Observation 954732c8-7d21-451e-80ed-82a9d6e7e726 · inbound

Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments cites this paper.

Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 58

Resolution
malformed identifier
arxiv_id, observed 2026-05-18T23:56:55.227957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T23:53:41.256273Z digest=sha256:e8ce3c94002dfd86d5f1e0bb1b1979613a5defe0b4ce9471db0b478f3c025738

Observation de9b7b6a-e106-4af2-bfbf-effc413786b6 · inbound

TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving cites this paper.

TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:54.216137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:52:09.041721Z digest=sha256:6df7ff1473c9a23edc953d72b67429d71009582ec549dd1d08bdc14ee4b7ab1b

Observation bcc60619-bd81-4b0f-b25f-4c9e98bba6b2 · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:56.874791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:36f20490fdfbb6eecf818eb1191699796db712970818d474e6b936b1238b1a8a

Observation 16135e9c-96b2-45d1-a29a-6a40c1678038 · inbound

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents cites this paper.

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:31:00.981125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T15:28:07.981488Z digest=sha256:c46386a0cf530724ca8e3e3e804fca111fbddb1f635421d75eac1c87f68a25a2

Observation 564c48f5-03ab-4bf5-bec0-99c7ad73ab0c · inbound

SEAL: Synergistic Co-Evolution of Agents and Learning Environments cites this paper.

SEAL: Synergistic Co-Evolution of Agents and Learning Environments Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:40.927696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T13:38:10.466713Z digest=sha256:a641272ce04ba2446846a9926cf51e8b9a656dd8b9bd37db594df9122227a3ba

Observation ab2e037c-6af1-4670-ad6a-2f084ca6cb25 · inbound

COMAP: Co-Evolving World Models and Agent Policies for LLM Agents cites this paper.

COMAP: Co-Evolving World Models and Agent Policies for LLM Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:24.851089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T14:27:50.260308Z digest=sha256:2c6a24f0b81fbd58abe54ef96c26f033b1640ce68e906dca4add040938a98b7e

Observation 1636119b-fa91-4868-b79f-a072ea86b6a0 · inbound

Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection cites this paper.

Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:22.687885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T20:55:02.866723Z digest=sha256:45c1c841a8fc085b297a74d53a0a1df55a4173d0cc389389ee5f3808b5b16ebf

Observation a797b419-0d5b-43d1-8e7a-6a9debd70232 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 105

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:a81d07ccade306c354a36c29b6da999b31f2ee6682b63c95d0c38dbb870d7f83

Observation 382ab218-744c-4c8e-b635-2ac60247a19e · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T13:56:19.183080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-09T13:51:49.149342Z digest=sha256:55ba0232f2114b3ae81dac9bd8e086fa9a06e38abfe1cec1905706c46464419c

Observation 92210d39-2e0c-402e-acb5-0ec6fe9095ca · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-13T06:45:27.857034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:45:27.857034Z digest=sha256:ee0c43d95ecbb2306329b052619429d9f7a1cd346fad06efa096efba2797da43

Observation e96571b2-8592-4b12-97e6-b2a69fdcc141 · inbound

Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool Orchestration cites this paper.

Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool Orchestration Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T22:37:39.546572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:37:39.546572Z digest=sha256:dd1e319da73e63ff373bde80468cabfb638894a17519a40e8c2d2babf664a012

Observation 1af4d129-b140-4e06-ad80-fd63bfe22bdc · inbound

Agents in the Wild: Where Research Meets Deployment cites this paper.

Agents in the Wild: Where Research Meets Deployment Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T12:44:32.128287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:44:32.128287Z digest=sha256:68c1eceb53272dc4e549d407fca165e9232342264da3fe2109ca86a68d45832f

Observation b173eeee-323c-4da1-baa0-0e874dc85a6f · inbound

Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems cites this paper.

Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-31T00:46:12.934662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:46:12.934662Z digest=sha256:dd804eb03f3711203c3b6a7d2b3abaea360e3ec859c01a678f5ca47a9f2708fd

Observation fad88add-183b-4f1f-bc9b-07f6cf689383 · inbound

From Failures to Supervision: DynamicEnvPlan for Robust Long-Horizon Embodied Planning cites this paper.

From Failures to Supervision: DynamicEnvPlan for Robust Long-Horizon Embodied Planning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T00:37:57.352425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:37:57.352425Z digest=sha256:888b6c052d54477713820640e7071964b2640642542a46743ad6cf8c902a9702

Observation 17822924-6ec9-4ab3-b904-8366e7604baf · inbound

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale cites this paper.

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 12

Resolution
malformed identifier
no resolver link, observed 2026-08-15T14:26:48.754960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:26:48.754960Z digest=sha256:a45cc82c5f1a4cfc73f775362d51872f4bc306b990495af548c8922efbc706e6

Observation 20bb84e0-bdfb-42d2-a03e-4b3ccbd56b01 · inbound

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale cites this paper.

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:26:48.759983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:26:48.759983Z digest=sha256:81e8d3553b75a71ebe90567c52d6b60e5374999c8fff3665b78682f94fe6c2ce