Pith. sign in

Paper Citation Record · LEDGER

Learning Agentic Policy from Action Guidance

As of 1 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2605.12004.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12004 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:02:49.206053Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:34:39.055539Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-01T12:16:17.934468Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact56
  • verified fuzzy21
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1639d43e-7ea4-4e77-9b48-1f80e58be624 · outbound

This paper cites Claude Opus 4.6 model card.

Learning Agentic Policy from Action Guidance Claude Opus 4.6 model card

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.831018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:ad67c3ab7aa8ea15e2650db7183fc667a63a60e3e0f0d89c77df01f5441f219e

Observation 5691e81c-db21-469e-914f-a15cb41526d4 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Learning Agentic Policy from Action Guidance $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.802510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:82125d4d0d943f025f812ea70bd7c57cc7de26514797dac274cfcc25578ff662

Observation cf10f969-fe6e-4803-be0b-9cd962e0eda4 · outbound

This paper cites Fine- tuning web agents: It works, but it’s trickier than you think.

Learning Agentic Policy from Action Guidance Fine- tuning web agents: It works, but it’s trickier than you think

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.816125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:5eb03f813598a01a5be8e359b9661ba759708428e3d1eb732294675772465bb5

Observation f2176436-10eb-4ea0-bbf8-da2b115a9287 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Learning Agentic Policy from Action Guidance SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:25c6da17003630b187b4428990806d2ed9a2e39a734d5d6fb1313dd5d07d1af1

Observation a2d792c1-5c8e-444f-8ae2-27ecb873b73a · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

Learning Agentic Policy from Action Guidance xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.799791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f34329f63793d15e21f9fabdca990bc3328046f907c54f5a5c49c62e52d3e65c

Observation 357d6141-dc4b-4d87-8a14-a127566deef6 · outbound

This paper cites Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning.

Learning Agentic Policy from Action Guidance Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:33.109892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1e17f7878945f85b75bcd91600da6c3e478a15c24c67993e263401f1ad8dbf38

Observation 987e40b8-e036-4fc2-85eb-758838d08a28 · outbound

This paper cites GPG: A simple and strong reinforcement learning baseline for model reasoning.

Learning Agentic Policy from Action Guidance GPG: A simple and strong reinforcement learning baseline for model reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.825455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:8574787c317ca63d8b3c9a482242644574b53b7eff9c827e7740dca32df3b0dd

Observation fbe351f8-dc2b-4f91-8b4d-b0e3a4584bae · outbound

This paper cites Redsearcher: A scalable and cost-efficient framework for long-horizon search agents.

Learning Agentic Policy from Action Guidance Redsearcher: A scalable and cost-efficient framework for long-horizon search agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.587389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:4cd46c7774e82200b08a90642d01b5799d5c702b35b5f63c48f57f39a135e886

Observation b3e77c83-4a51-4118-9c6d-fc14d1fed60d · outbound

This paper cites Harder is better: Boosting mathematical reasoning via difficulty-aware GRPO and multi-aspect question reformulation.

Learning Agentic Policy from Action Guidance Harder is better: Boosting mathematical reasoning via difficulty-aware GRPO and multi-aspect question reformulation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.776771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:b614b9eadc73672098af7044b79009524517e90b12648af97ad547d2c74c6387

Observation de9932e6-8f5c-4a75-86a9-c4c8d28d60d8 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114.

Learning Agentic Policy from Action Guidance Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.818141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:00e4c31d21a60d01084771b9317a5b9737bf0ddeeb47b23e4c6d81f1bc2b3b70

Observation f5550734-41ba-41e7-80b5-8341adc77647 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Learning Agentic Policy from Action Guidance OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.519833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f6d71cd47a9fb7580519f4b8b8c6612e445b957c8986711a41985080b6f9001d

Observation a0d8fdd6-9f20-4899-a0f7-7fd7a4f750e6 · outbound

This paper cites Wildclawbench.

Learning Agentic Policy from Action Guidance Wildclawbench

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.823623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:14c7d7feffa1beabd6b530fadc4475041138b5b5239a6f7a177885f18418f348

Observation 751d4fcb-e620-4ec5-abcc-dc6947058c96 · outbound

This paper cites Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning.

Learning Agentic Policy from Action Guidance Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.727198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f9d07c19bebb87c4f3cfe54987450933b1f262da39e842b82493c6c426c1a7de

Observation 37eda508-fa87-42b7-be1d-f85b7e9c116e · outbound

This paper cites Agentic Reinforced Policy Optimization.

Learning Agentic Policy from Action Guidance Agentic Reinforced Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:57:12.173630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:6c5eb0b066eb9e68074cbf7fa7216179c7332e7c80b5db72dd279220dd3b9280

Observation bec250b6-7d34-4c30-bf01-d732a8d11b3a · outbound

This paper cites Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence.

Learning Agentic Policy from Action Guidance Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.621909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1cf0e49c5afc8c39dfcc9bd74b0192c7805650ca5acbb1b045b0795d445f8349

Observation 5efedcd7-945a-4fa4-8eda-6e88d8374c97 · outbound

This paper cites Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks.

Learning Agentic Policy from Action Guidance Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:32:18.834213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:021c6cbdeb939a4584adb8fb9688add004bc4277afdccef59f9ebd620a1f7222

Observation 4a61f7ee-90f2-460f-9b23-03f66a981999 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Learning Agentic Policy from Action Guidance Group-in-Group Policy Optimization for LLM Agent Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.632913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:67be521452a36299780cc5336db272f3406e4685b3fe5e70a95b60a44e539ed5

Observation 859ffaa4-7137-4c9d-bc2d-2b543d12052e · outbound

This paper cites SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning.

Learning Agentic Policy from Action Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.645623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:9d182d7830be3c38c3ceb71d97621dda3580078affc26b3009e9a3fb0b108d15

Observation 1e817b8d-464e-40b4-b648-316d0bf06b9a · outbound

This paper cites Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl.

Learning Agentic Policy from Action Guidance Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.649435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:5b00d30199d338e9f3bb663383049ff5a3aaf95124a5e67efb13971d104d4716

Observation 009f5b83-d135-430d-a75b-484f7a33c050 · outbound

This paper cites Actor-curator: Co-adaptive curriculum learning via policy-improvement bandits for rl post-training.arXiv preprint arXiv:2602.20532.

Learning Agentic Policy from Action Guidance Actor-curator: Co-adaptive curriculum learning via policy-improvement bandits for rl post-training.arXiv preprint arXiv:2602.20532

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.590870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:341f2c5ebb10edd0675ea737ae97c61b13a90f51f733fbf220b735ae8c2e9975

Observation 3af81f34-30cf-47d1-bbd8-d1e391518386 · outbound

This paper cites Deep q-learning from demonstrations.

Learning Agentic Policy from Action Guidance Deep q-learning from demonstrations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.778716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:3b56c59513207bd00c4c7be47bac75b6298e3fbd7766675e1bbd787602fa8234

Observation 50fbd06d-8e5a-424c-bb98-d95aa186f427 · outbound

This paper cites Boosting mllm reasoning with text-debiased hint-grpo.

Learning Agentic Policy from Action Guidance Boosting mllm reasoning with text-debiased hint-grpo

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.798467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:8e6213e0ed5b4ca0292d7b49cae2c9a1d5a8490fa79bc75b8e321f21ce07de42

Observation b82336af-dcd4-47ec-bc91-ffc20befc04b · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Learning Agentic Policy from Action Guidance Reinforcement Learning via Self-Distillation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.597502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:239d5bf2850f3cd306a9e42c4d3e8345cf0b46313c6ba5c105a84bd32cd5bdd1

Observation 08144029-b394-4e13-8a6d-176e06855d6c · outbound

This paper cites Tree search for LLM agent reinforcement learning.

Learning Agentic Policy from Action Guidance Tree search for LLM agent reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.780576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:2b48b565f25a7ce809398e5b2728f32522d6231436b25b2c8c7234434e4e8864

Observation 0be9105e-bea4-4158-beac-d4f44700b3c5 · outbound

This paper cites Thinking with map: Reinforced parallel map-augmented agent for geolocalization.ACL.

Learning Agentic Policy from Action Guidance Thinking with map: Reinforced parallel map-augmented agent for geolocalization.ACL

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.812409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:c6cc4adfea0c06bc660010cb08e632991fdc9b261bb8d06cb989633d3245f660

Observation 8635f681-43ae-46dd-94c9-dfaef4cb5faa · outbound

This paper cites Vcrl: Variance-based curriculum reinforcement learning for large language models.

Learning Agentic Policy from Action Guidance Vcrl: Variance-based curriculum reinforcement learning for large language models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.676707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:3da4d045087a0d4eb738d4c4d062d9694043e39ead3fd78c11b38261a1e8baa3

Observation 80061224-480d-4b83-b19b-6846e08e04c9 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Learning Agentic Policy from Action Guidance SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.688917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:d07ffadd434051207383ede43f9af03e0102790cc8a410910083a8b3f73bfd70

Observation afda1e21-8088-47ae-9ae2-b7a6effdc9d4 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Learning Agentic Policy from Action Guidance Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.606250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:ee8d7f2cbd60952fa184d17d20943cd8b407ebb4c2cfa60909b1601259831425

Observation 52a638ee-6fb7-4688-97ab-09df64d5fd0f · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

Learning Agentic Policy from Action Guidance WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:37:09.773663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f1f40b51321e4bec79f6c2190efcda1852841800ff3853828b4b9787ef52f1f8

Observation b970073b-9c08-4745-9f20-11aae7c9e7b9 · outbound

This paper cites Adacurl: Adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting.

Learning Agentic Policy from Action Guidance Adacurl: Adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.814168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:de9017bdc6d6e350098aee339ab88d097acfb430080f8f773427753616095f24

Observation 55358957-318a-4682-941b-47aa643d0dac · outbound

This paper cites WebThinker: Empowering Large Reasoning Models with Deep Research Capability.

Learning Agentic Policy from Action Guidance WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:14:25.573630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:8a28818885d634cc87ee744d688765f31b080e399328ef1d68fbf2b120d74378

Observation 92c5b202-9bfe-4c5c-b64a-08617c072a60 · outbound

This paper cites Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model.

Learning Agentic Policy from Action Guidance Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.618672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f0ad86183e16d26451eb24c99f7537ea3322477dc84dd58ab925579a0c1e353f

Observation cfd20f2a-199e-4d14-ae46-5d9da119b7bb · outbound

This paper cites Guided exploration with proximal policy optimization using a single demonstration.

Learning Agentic Policy from Action Guidance Guided exploration with proximal policy optimization using a single demonstration

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.821791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:17288c7920b55e3439e8892c13e6c52153b020e4c15c56268d484d592332f2d1

Observation 66b3e568-7ed0-4777-bb46-04d01342c9e7 · outbound

This paper cites Truthfulqa: Measuring how models mimic hu- man falsehoods.

Learning Agentic Policy from Action Guidance Truthfulqa: Measuring how models mimic hu- man falsehoods

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.820085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:01846c8c14b34b1205ff92a71ec5d0cdfb6e6faf72afd6838711771010815810

Observation 8caa542f-b769-4914-8565-be8a51ce2d1a · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning Agentic Policy from Action Guidance DeepSeek-V3 Technical Report

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.796733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:6e7bde249ade78b5bed1ec9ce04708e2e9155ff2d3b28646716838dd243eb7ab

Observation 2361523b-ecd7-42c7-a7aa-dac6d962aa89 · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

Learning Agentic Policy from Action Guidance Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.594178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:c9983fcc9973f1aa2e7a99b173b9cbc6904d16bc44ceeed0d14d413e9b588dca

Observation b737ab17-d520-460c-a169-a0d582078b9c · outbound

This paper cites Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions.

Learning Agentic Policy from Action Guidance Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.787952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:d122df9e3d9212e6c54ee734285df4464b603acce4dcc67f7ff43a4b4023798c

Observation 826ac4a2-1c2d-4ade-9115-ee3388b75676 · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

Learning Agentic Policy from Action Guidance SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.776747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:cb844cf8bfc7cf269376b1fe21ef52523bbdddd1d3dbec7eb06be3266d4ae0c9

Observation cf36beac-9f5d-45f3-9e96-e84907002074 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Learning Agentic Policy from Action Guidance GAIA: a benchmark for General AI Assistants

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.603371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:2e7c784caff423c39721e8539919c2ba3229c06ecd416385c652d16f484fab35

Observation 14e1b130-4e90-4896-aea3-d3393945bed6 · outbound

This paper cites Minimax m2.1 system card.

Learning Agentic Policy from Action Guidance Minimax m2.1 system card

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.829342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:8bc24847ea2557f81e441b7f0fe74e29ae2b60cb58e89ed12ce268b1aa75b7d7

Observation bdfa9d7a-d43d-4f0b-b3dc-39f5de834860 · outbound

This paper cites Over- coming exploration in reinforcement learning with demonstrations.

Learning Agentic Policy from Action Guidance Over- coming exploration in reinforcement learning with demonstrations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.806867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:94fad949333038804ffe16998ff8d4b104dd52f4f0f64e901c4128df303894ab

Observation 2cf22fff-26b7-4055-817f-97f6907687bc · outbound

This paper cites Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models.

Learning Agentic Policy from Action Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:07:17.751601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:b786d9d5499c9b9201d69b844b558e5a2b72262526b920c6fc3afb4c059c3fcd

Observation a00052b6-468d-4553-a72a-21ff3ba94931 · outbound

This paper cites Gpt-5.4 thinking system card.

Learning Agentic Policy from Action Guidance Gpt-5.4 thinking system card

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.810613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:cd054e35205bd382f83d7b70bb9697320064cae4a3332780c4c2fcda311aa64d

Observation b097e1fa-7b21-481a-acf6-bde3b98e240a · outbound

This paper cites Iterative reasoning preference optimization.Advances in Neural Information Processing Systems, 37:116617–116637.

Learning Agentic Policy from Action Guidance Iterative reasoning preference optimization.Advances in Neural Information Processing Systems, 37:116617–116637

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.802870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:6d2e557ab9d65e0c78b98755fafc916cb73dfeb6e781a4886cb1c71f87c9ff55

Observation 2fba3040-3165-4db1-8b1a-54dc1ca17d4e · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Learning Agentic Policy from Action Guidance UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.699644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:2aadb017a4bb38702c1761152d55842cf5fa22db1a4380a1a0f081b84373a87c

Observation a57effd7-b998-4217-8826-62f5937f4fe7 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Learning Agentic Policy from Action Guidance Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.755910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1e93a3d6c750f845998de688bd676abcd46b16c2f1ed233f994c1d8b89086a75

Observation 9c44e73f-ff6c-4474-9312-9bd559a419c2 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Learning Agentic Policy from Action Guidance GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.773989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:fefc1713a935bb9614e7460ac489d6656434569ff99603995ed89d0ea4808bb1

Observation d467159e-4082-404f-8d31-199b819351da · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning Agentic Policy from Action Guidance Proximal Policy Optimization Algorithms

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.600157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:3345ae642ce4db18a75ea5c7f8fb4613243fae2e04e9cc9cbc39f632886d8202

Observation 1633b736-430c-4d9c-862d-6bc615c3575a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learning Agentic Policy from Action Guidance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.696982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:5dbc07dd4c1f41654f6a4cf121d9c26df28f685f5a9595f2eae4de9529cc626c

Observation fc9c0f91-b266-4d95-b705-4562cdd3e576 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Learning Agentic Policy from Action Guidance Self-Distillation Enables Continual Learning

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.685825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:c503c5789c9ae9ce3c25277279687964077a85ca5aae80899d1a467ab37b2f8d

Observation fa15cef6-ae82-438f-9a32-97dd9aa78c15 · outbound

This paper cites OpenAI GPT-5 System Card.

Learning Agentic Policy from Action Guidance OpenAI GPT-5 System Card

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.679643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:37223991bc4bb5576c74f581fb98a32fc66de4c756fee118bc952572178dfd91

Observation 95c47958-1b89-4fc5-ad6c-b886c217c52a · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Learning Agentic Policy from Action Guidance Kimi K2.5: Visual Agentic Intelligence

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.768488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:9085a7c14b909f51adbe58eeda0ea0a02488d16dec2f94533dea19ee7cd03367

Observation 485864f6-187f-404a-a5ab-3f6c1c9529ab · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February.

Learning Agentic Policy from Action Guidance Qwen3.5: Accelerating productivity with native multimodal agents, February

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.805037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:62ed918fa7a384e20f10ae5c6276001f536446188bff3325d74a2eaa933f07d3

Observation e2912aa8-4b67-4e68-b97c-f70d1e74a105 · outbound

This paper cites an unresolved cited work.

Learning Agentic Policy from Action Guidance Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-13T11:07:39.800591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:6dee3774e9f42a7d9159b23415b15198cab66cd334ae15e119547b857cbd8006

Observation 68746824-b279-4f01-862c-ff4041c7e0d0 · outbound

This paper cites Tongyi DeepResearch Technical Report.

Learning Agentic Policy from Action Guidance Tongyi DeepResearch Technical Report

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:56:57.433092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:b6d14371b850424a3fdde213ce2c1f57754e99b24018d1a3466f9bce7187c3ad

Observation f4e2c5f8-93f7-4bc0-b219-a26eaf80f776 · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965.

Learning Agentic Policy from Action Guidance Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.774687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T07:38:14.455455+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:55c830ac920d501b8378bb2f7ef84c280a004098db139b9affe0f7fbaab7a465

Observation 8dc39b33-8e3c-4204-ba86-c86d6618adae · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Learning Agentic Policy from Action Guidance Deep Reinforcement Learning and the Deadly Triad

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.711777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:4c4272001eca6781c803ba6f6640d28eb79e50a442e1df388c097d5e7c7eadf2

Observation 87baeeb7-af4a-4cee-b02e-f8467ffb4008 · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Learning Agentic Policy from Action Guidance Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.758867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:4005a7978a300d8ba5b8e9408ea4e84801e4db42a684f68bacba929d74e7db20

Observation 74e5c648-4a6d-4977-9805-bacabdda8f9f · outbound

This paper cites Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem.

Learning Agentic Policy from Action Guidance Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.761647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:ddf2c47a0428caf00529278ea77d8ae2864b51f58018e36cbc298adc974e08b1

Observation db34b9fc-5d71-432d-993e-62773e8d8907 · outbound

This paper cites OpenClaw-RL: Train Any Agent Simply by Talking.

Learning Agentic Policy from Action Guidance OpenClaw-RL: Train Any Agent Simply by Talking

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.765019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f33e7c7ef47a959e35de4f1782414970f0806a1e4c2e5e8e4a3b08ce9a177eeb

Observation 7711b20f-5074-4cc0-b9e5-82086d47af10 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Learning Agentic Policy from Action Guidance RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:06528a3730f0f2593c4aacb630e4f0674d4087c634c4af3c878f366c086fcc5d

Observation f9d25609-d38b-42b2-8332-d678f8004c54 · outbound

This paper cites Agentic Reasoning for Large Language Models.

Learning Agentic Policy from Action Guidance Agentic Reasoning for Large Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.657956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T10:08:10.217033+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:e205efce3c3921caa19b710ad7cb1918a29f27dc1a05488eff323e34419872a3

Observation f4c1fd2f-5f2e-44e6-af5b-fb0dce07b088 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Learning Agentic Policy from Action Guidance WebWalker: Benchmarking LLMs in Web Traversal

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:07:17.791143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1058a54ef14faca80041716eb290694b4f7e8a1b6239774ee63e70013d1b1865

Observation 82846ed5-0a12-479c-a3b3-b13b3b5eddd7 · outbound

This paper cites Learn hard problems during rl with reference guided fine-tuning.

Learning Agentic Policy from Action Guidance Learn hard problems during rl with reference guided fine-tuning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.702708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:04ca2c7cce28f563509b4b194643decf6bd84770a78ceeb62d75a6ccc11f56a1

Observation a51275f2-6cbc-4d04-9b51-8f2c2dfe705d · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

Learning Agentic Policy from Action Guidance Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.808900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:bdf58b388079b5335e3166e2d069c37280b5b027e380c26a098fc8f2eb77be93

Observation 1d69ddf5-36fb-4f9f-ae7f-c240932e4ed3 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Learning Agentic Policy from Action Guidance Learning to Reason under Off-Policy Guidance

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:17:03.075876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:206b610b85c2a00fc59a7c4744502072e65183decfeb42b167d1857dcc93d49c

Observation 64b3bdc3-82c7-4ed2-8889-7e699ce8bea5 · outbound

This paper cites GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL.

Learning Agentic Policy from Action Guidance GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:04:34.683381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:bfd9a200dca8ed9e8ddb6e70196ac8c15874ee4438910c14ab6d544b80fc2738

Observation 9f6a887b-9d47-4c88-ae0f-10733a8b4686 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Learning Agentic Policy from Action Guidance React: Synergizing reasoning and acting in language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.827410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:e71e3e7fdc4e905d1da6333b3bbc2d763f33fc07c23c743452cc2e4944aa03f9

Observation afad9ae9-cc3b-4580-a25f-ea603716bab3 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Learning Agentic Policy from Action Guidance $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.666389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:ad2bbd3fefa1040ebdda521094ebf51ceabd18baf16cbaf226cd892ad11a4c9f

Observation 69d3835c-60f0-442b-9165-e06fad3beb15 · outbound

This paper cites Coba-rl: Capability-oriented budget allocation for reinforcement learning in llms.

Learning Agentic Policy from Action Guidance Coba-rl: Capability-oriented budget allocation for reinforcement learning in llms

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.669790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:dddab5f5d7c79c534ab38639286d87c75ae966d1dab4279537df86ec30da9990

Observation 7f6c0cb2-ad96-4afd-8d81-644cb60a574a · outbound

This paper cites arXiv preprint arXiv:2603.21383 , year=.

Learning Agentic Policy from Action Guidance arXiv preprint arXiv:2603.21383 , year=

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.663444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:3efaeb34573863548950d03cd1e588e4eab2c8296703f24a52a03c4ff1dc1b5e

Observation 41b76d43-58f0-4bec-a003-97980b44661c · outbound

This paper cites MedResearcher-R1: Expert-Level Medical Deep Researcher via A Knowledge-Informed Trajectory Synthesis Framework.

Learning Agentic Policy from Action Guidance MedResearcher-R1: Expert-Level Medical Deep Researcher via A Knowledge-Informed Trajectory Synthesis Framework

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.656657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:d9c0ebadecb9201e70df250466bd418ca620f890b7200b5b0773befb4c7f4331

Observation 1c777d36-c31c-4b72-bea6-5de65d9332bc · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Learning Agentic Policy from Action Guidance DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.694521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:245ffcda86fea8fb826c0a6291ede6205cc2543df432630a044e97d6f710557b

Observation 66a6d182-aa4a-4d38-b0c8-a96694d6572c · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Learning Agentic Policy from Action Guidance Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.659881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:37a64394db1ca837c4f57d9a7c2a720484a2ad205716ee48e3efda66fd629b55

Observation a850c3b3-a2dc-4ad3-8e8c-acf481ef8173 · outbound

This paper cites Agentevolver: Towards efficient self-evolving agent system.

Learning Agentic Policy from Action Guidance Agentevolver: Towards efficient self-evolving agent system

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.785140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:572a2fdcd746710c9b9778901c2b3ecac46aab27ef44d22a9e627f5b99c2d4a8

Observation ff787be5-ad14-4556-bcbd-6f07b151209e · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for LLMs: A Survey.

Learning Agentic Policy from Action Guidance The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.608879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:e70fb7c4babf289678394d600a29640b86252de57094394c0ab8adde473c1d4d

Observation 5eddc1e9-79b6-4ecb-998d-34d290444f99 · outbound

This paper cites arXiv preprint arXiv:2508.11408 , year=.

Learning Agentic Policy from Action Guidance arXiv preprint arXiv:2508.11408 , year=

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:07:17.705737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:bef38348fec72ed7c26cdb4ffd0f48f9ca9f698b209e4e09ce84563e18c902be

Observation d45a8540-4c29-4f0d-8030-aef3cc060ce8 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Learning Agentic Policy from Action Guidance Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.779421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:ace9aa4da2963be9a6719f945118a63d6633f6d71523a4e9509d655a94e253d3

Observation 4bd1baa4-5550-45eb-b3a4-96e166e18099 · outbound

This paper cites Prosperity before collapse: How far can off-policy rl reach with stale data on llms?.

Learning Agentic Policy from Action Guidance Prosperity before collapse: How far can off-policy rl reach with stale data on llms?

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.708884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:c4316b021c3599193bb4aeb421fc3f4216ef66c4ddfb61e79d1e9a28f5ee57a5

Observation 379ba4ab-3a65-469a-bd7e-d9cdcdfc50cf · outbound

This paper cites Code2world: A gui world model via renderable code generation.

Learning Agentic Policy from Action Guidance Code2world: A gui world model via renderable code generation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.639387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:6908407ca6b1fb80427dd2ee19f907fc7a006559f46f3fbdf66503bc91bbc2e8

Observation 66043e16-8a57-43b6-9729-a26b0da6908c · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Learning Agentic Policy from Action Guidance Instruction-Following Evaluation for Large Language Models

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.642066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:4fc2bdc4435f063bdaa705368690b154cd7eab84ba6184baa58c8e3b0d0a4bb2

Observation 3bb80b65-9f3d-4e17-a552-07be46bb0b8d · outbound

This paper cites BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese.

Learning Agentic Policy from Action Guidance BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:04:50.009871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1409fac0552fe6f62e45a7cc2147b4eb3eaf331410ade6c87bb9f3d71c9b15e0

Pith citing papers

Observation 1e1e43b4-438e-4481-bbd8-f32e36544441 · inbound

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning cites this paper.

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning Learning Agentic Policy from Action Guidance

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-08-01T07:39:17.666069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T10:08:10.546089+00:00.

source=arxiv_source observed=2026-08-01T07:34:39.055539Z digest=sha256:a23de9ba8222c74a528227e835eaae5753f8b5a0e996d3248dcc63ad32928aa2