Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:02:49.206053Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2605.12004.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:02:49.206053Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:34:39.055539Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-01T12:16:17.934468Z
82 of 82 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1639d43e-7ea4-4e77-9b48-1f80e58be624 · outbound
Learning Agentic Policy from Action Guidance Claude Opus 4.6 model card
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5691e81c-db21-469e-914f-a15cb41526d4 · outbound
Learning Agentic Policy from Action Guidance $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation cf10f969-fe6e-4803-be0b-9cd962e0eda4 · outbound
Learning Agentic Policy from Action Guidance Fine- tuning web agents: It works, but it’s trickier than you think
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f2176436-10eb-4ea0-bbf8-da2b115a9287 · outbound
Learning Agentic Policy from Action Guidance SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a2d792c1-5c8e-444f-8ae2-27ecb873b73a · outbound
Learning Agentic Policy from Action Guidance xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 357d6141-dc4b-4d87-8a14-a127566deef6 · outbound
Learning Agentic Policy from Action Guidance Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 987e40b8-e036-4fc2-85eb-758838d08a28 · outbound
Learning Agentic Policy from Action Guidance GPG: A simple and strong reinforcement learning baseline for model reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation fbe351f8-dc2b-4f91-8b4d-b0e3a4584bae · outbound
Learning Agentic Policy from Action Guidance Redsearcher: A scalable and cost-efficient framework for long-horizon search agents
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b3e77c83-4a51-4118-9c6d-fc14d1fed60d · outbound
Learning Agentic Policy from Action Guidance Harder is better: Boosting mathematical reasoning via difficulty-aware GRPO and multi-aspect question reformulation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation de9932e6-8f5c-4a75-86a9-c4c8d28d60d8 · outbound
Learning Agentic Policy from Action Guidance Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f5550734-41ba-41e7-80b5-8341adc77647 · outbound
Learning Agentic Policy from Action Guidance OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a0d8fdd6-9f20-4899-a0f7-7fd7a4f750e6 · outbound
Learning Agentic Policy from Action Guidance Wildclawbench
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 751d4fcb-e620-4ec5-abcc-dc6947058c96 · outbound
Learning Agentic Policy from Action Guidance Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 37eda508-fa87-42b7-be1d-f85b7e9c116e · outbound
Learning Agentic Policy from Action Guidance Agentic Reinforced Policy Optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation bec250b6-7d34-4c30-bf01-d732a8d11b3a · outbound
Learning Agentic Policy from Action Guidance Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5efedcd7-945a-4fa4-8eda-6e88d8374c97 · outbound
Learning Agentic Policy from Action Guidance Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4a61f7ee-90f2-460f-9b23-03f66a981999 · outbound
Learning Agentic Policy from Action Guidance Group-in-Group Policy Optimization for LLM Agent Training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 859ffaa4-7137-4c9d-bc2d-2b543d12052e · outbound
Learning Agentic Policy from Action Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1e817b8d-464e-40b4-b648-316d0bf06b9a · outbound
Learning Agentic Policy from Action Guidance Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 009f5b83-d135-430d-a75b-484f7a33c050 · outbound
Learning Agentic Policy from Action Guidance Actor-curator: Co-adaptive curriculum learning via policy-improvement bandits for rl post-training.arXiv preprint arXiv:2602.20532
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3af81f34-30cf-47d1-bbd8-d1e391518386 · outbound
Learning Agentic Policy from Action Guidance Deep q-learning from demonstrations
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 50fbd06d-8e5a-424c-bb98-d95aa186f427 · outbound
Learning Agentic Policy from Action Guidance Boosting mllm reasoning with text-debiased hint-grpo
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b82336af-dcd4-47ec-bc91-ffc20befc04b · outbound
Learning Agentic Policy from Action Guidance Reinforcement Learning via Self-Distillation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 08144029-b394-4e13-8a6d-176e06855d6c · outbound
Learning Agentic Policy from Action Guidance Tree search for LLM agent reinforcement learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0be9105e-bea4-4158-beac-d4f44700b3c5 · outbound
Learning Agentic Policy from Action Guidance Thinking with map: Reinforced parallel map-augmented agent for geolocalization.ACL
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8635f681-43ae-46dd-94c9-dfaef4cb5faa · outbound
Learning Agentic Policy from Action Guidance Vcrl: Variance-based curriculum reinforcement learning for large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 80061224-480d-4b83-b19b-6846e08e04c9 · outbound
Learning Agentic Policy from Action Guidance SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation afda1e21-8088-47ae-9ae2-b7a6effdc9d4 · outbound
Learning Agentic Policy from Action Guidance Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 52a638ee-6fb7-4688-97ab-09df64d5fd0f · outbound
Learning Agentic Policy from Action Guidance WebSailor: Navigating Super-human Reasoning for Web Agent
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b970073b-9c08-4745-9f20-11aae7c9e7b9 · outbound
Learning Agentic Policy from Action Guidance Adacurl: Adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 55358957-318a-4682-941b-47aa643d0dac · outbound
Learning Agentic Policy from Action Guidance WebThinker: Empowering Large Reasoning Models with Deep Research Capability
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 92c5b202-9bfe-4c5c-b64a-08617c072a60 · outbound
Learning Agentic Policy from Action Guidance Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation cfd20f2a-199e-4d14-ae46-5d9da119b7bb · outbound
Learning Agentic Policy from Action Guidance Guided exploration with proximal policy optimization using a single demonstration
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 66b3e568-7ed0-4777-bb46-04d01342c9e7 · outbound
Learning Agentic Policy from Action Guidance Truthfulqa: Measuring how models mimic hu- man falsehoods
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8caa542f-b769-4914-8565-be8a51ce2d1a · outbound
Learning Agentic Policy from Action Guidance DeepSeek-V3 Technical Report
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2361523b-ecd7-42c7-a7aa-dac6d962aa89 · outbound
Learning Agentic Policy from Action Guidance Large Language Model Agent: A Survey on Methodology, Applications and Challenges
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b737ab17-d520-460c-a169-a0d582078b9c · outbound
Learning Agentic Policy from Action Guidance Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 826ac4a2-1c2d-4ade-9115-ee3388b75676 · outbound
Learning Agentic Policy from Action Guidance SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation cf36beac-9f5d-45f3-9e96-e84907002074 · outbound
Learning Agentic Policy from Action Guidance GAIA: a benchmark for General AI Assistants
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 14e1b130-4e90-4896-aea3-d3393945bed6 · outbound
Learning Agentic Policy from Action Guidance Minimax m2.1 system card
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation bdfa9d7a-d43d-4f0b-b3dc-39f5de834860 · outbound
Learning Agentic Policy from Action Guidance Over- coming exploration in reinforcement learning with demonstrations
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2cf22fff-26b7-4055-817f-97f6907687bc · outbound
Learning Agentic Policy from Action Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a00052b6-468d-4553-a72a-21ff3ba94931 · outbound
Learning Agentic Policy from Action Guidance Gpt-5.4 thinking system card
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b097e1fa-7b21-481a-acf6-bde3b98e240a · outbound
Learning Agentic Policy from Action Guidance Iterative reasoning preference optimization.Advances in Neural Information Processing Systems, 37:116617–116637
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2fba3040-3165-4db1-8b1a-54dc1ca17d4e · outbound
Learning Agentic Policy from Action Guidance UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a57effd7-b998-4217-8826-62f5937f4fe7 · outbound
Learning Agentic Policy from Action Guidance Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9c44e73f-ff6c-4474-9312-9bd559a419c2 · outbound
Learning Agentic Policy from Action Guidance GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d467159e-4082-404f-8d31-199b819351da · outbound
Learning Agentic Policy from Action Guidance Proximal Policy Optimization Algorithms
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1633b736-430c-4d9c-862d-6bc615c3575a · outbound
Learning Agentic Policy from Action Guidance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation fc9c0f91-b266-4d95-b705-4562cdd3e576 · outbound
Learning Agentic Policy from Action Guidance Self-Distillation Enables Continual Learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation fa15cef6-ae82-438f-9a32-97dd9aa78c15 · outbound
Learning Agentic Policy from Action Guidance OpenAI GPT-5 System Card
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 95c47958-1b89-4fc5-ad6c-b886c217c52a · outbound
Learning Agentic Policy from Action Guidance Kimi K2.5: Visual Agentic Intelligence
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 485864f6-187f-404a-a5ab-3f6c1c9529ab · outbound
Learning Agentic Policy from Action Guidance Qwen3.5: Accelerating productivity with native multimodal agents, February
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e2912aa8-4b67-4e68-b97c-f70d1e74a105 · outbound
Learning Agentic Policy from Action Guidance Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 68746824-b279-4f01-862c-ff4041c7e0d0 · outbound
Learning Agentic Policy from Action Guidance Tongyi DeepResearch Technical Report
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f4e2c5f8-93f7-4bc0-b219-a26eaf80f776 · outbound
Learning Agentic Policy from Action Guidance Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T07:38:14.455455+00:00.
Observation 8dc39b33-8e3c-4204-ba86-c86d6618adae · outbound
Learning Agentic Policy from Action Guidance Deep Reinforcement Learning and the Deadly Triad
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 87baeeb7-af4a-4cee-b02e-f8467ffb4008 · outbound
Learning Agentic Policy from Action Guidance Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 74e5c648-4a6d-4977-9805-bacabdda8f9f · outbound
Learning Agentic Policy from Action Guidance Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation db34b9fc-5d71-432d-993e-62773e8d8907 · outbound
Learning Agentic Policy from Action Guidance OpenClaw-RL: Train Any Agent Simply by Talking
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7711b20f-5074-4cc0-b9e5-82086d47af10 · outbound
Learning Agentic Policy from Action Guidance RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f9d25609-d38b-42b2-8332-d678f8004c54 · outbound
Learning Agentic Policy from Action Guidance Agentic Reasoning for Large Language Models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T10:08:10.217033+00:00.
Observation f4c1fd2f-5f2e-44e6-af5b-fb0dce07b088 · outbound
Learning Agentic Policy from Action Guidance WebWalker: Benchmarking LLMs in Web Traversal
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 82846ed5-0a12-479c-a3b3-b13b3b5eddd7 · outbound
Learning Agentic Policy from Action Guidance Learn hard problems during rl with reference guided fine-tuning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a51275f2-6cbc-4d04-9b51-8f2c2dfe705d · outbound
Learning Agentic Policy from Action Guidance Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1d69ddf5-36fb-4f9f-ae7f-c240932e4ed3 · outbound
Learning Agentic Policy from Action Guidance Learning to Reason under Off-Policy Guidance
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 64b3bdc3-82c7-4ed2-8889-7e699ce8bea5 · outbound
Learning Agentic Policy from Action Guidance GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9f6a887b-9d47-4c88-ae0f-10733a8b4686 · outbound
Learning Agentic Policy from Action Guidance React: Synergizing reasoning and acting in language models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation afad9ae9-cc3b-4580-a25f-ea603716bab3 · outbound
Learning Agentic Policy from Action Guidance $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 69d3835c-60f0-442b-9165-e06fad3beb15 · outbound
Learning Agentic Policy from Action Guidance Coba-rl: Capability-oriented budget allocation for reinforcement learning in llms
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7f6c0cb2-ad96-4afd-8d81-644cb60a574a · outbound
Learning Agentic Policy from Action Guidance arXiv preprint arXiv:2603.21383 , year=
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 41b76d43-58f0-4bec-a003-97980b44661c · outbound
Learning Agentic Policy from Action Guidance MedResearcher-R1: Expert-Level Medical Deep Researcher via A Knowledge-Informed Trajectory Synthesis Framework
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1c777d36-c31c-4b72-bea6-5de65d9332bc · outbound
Learning Agentic Policy from Action Guidance DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 66a6d182-aa4a-4d38-b0c8-a96694d6572c · outbound
Learning Agentic Policy from Action Guidance Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a850c3b3-a2dc-4ad3-8e8c-acf481ef8173 · outbound
Learning Agentic Policy from Action Guidance Agentevolver: Towards efficient self-evolving agent system
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ff787be5-ad14-4556-bcbd-6f07b151209e · outbound
Learning Agentic Policy from Action Guidance The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5eddc1e9-79b6-4ecb-998d-34d290444f99 · outbound
Learning Agentic Policy from Action Guidance arXiv preprint arXiv:2508.11408 , year=
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d45a8540-4c29-4f0d-8030-aef3cc060ce8 · outbound
Learning Agentic Policy from Action Guidance Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4bd1baa4-5550-45eb-b3a4-96e166e18099 · outbound
Learning Agentic Policy from Action Guidance Prosperity before collapse: How far can off-policy rl reach with stale data on llms?
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 379ba4ab-3a65-469a-bd7e-d9cdcdfc50cf · outbound
Learning Agentic Policy from Action Guidance Code2world: A gui world model via renderable code generation
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 66043e16-8a57-43b6-9729-a26b0da6908c · outbound
Learning Agentic Policy from Action Guidance Instruction-Following Evaluation for Large Language Models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3bb80b65-9f3d-4e17-a552-07be46bb0b8d · outbound
Learning Agentic Policy from Action Guidance BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1e1e43b4-438e-4481-bbd8-f32e36544441 · inbound
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning Learning Agentic Policy from Action Guidance
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T10:08:10.546089+00:00.