Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T10:33:54.851493Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2607.10601.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T10:33:54.851493Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:44:31.301557Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-04T12:16:13.904932Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4bd8000a-17c5-497e-991b-75af1d98781d · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories On-policy distillation of language models: Learning from self-generated mistakes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bed126f-49a0-4b34-8771-8c9097274e48 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories A general theoretical paradigm to understand learning from human preferences
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a742f75-8987-4797-96ca-d4981f57b7da · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fadb3a24-6dea-449d-8890-a67f0292a037 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories FireAct: Toward Language Agent Fine-tuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa1023e0-2103-4140-8bcf-38a7d6edef7f · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2032254-9900-4d09-a479-56845208ee84 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ATLaS: Agent Tuning via Learning Critical Steps
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 073b40c5-dbb9-486f-9b6b-5a46a17d6e3f · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Self-play fine-tuning converts weak language models to strong language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b629639-8ed1-4636-8dfb-7d7e727e1127 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daced041-f2dc-4f63-8d6b-cc3dd8ae344b · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aabe613-5e49-464b-bcd6-d37bf6f05257 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Mind2Web: Towards a generalist agent for the web
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eea484d1-9b2d-40cf-947d-18c3272a465f · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories KTO: Model Alignment as Prospect Theoretic Optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddddbc5f-1585-4c36-8a98-841c2807c6af · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories AgentRefine: Enhancing Agent Generalization through Refinement Tuning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea5c0a9f-0bad-41e1-848d-3d75d73784c8 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Solving the granularity mismatch: Hierarchical preference learning for long-horizon llm agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecceda03-3c18-4ebb-bdfb-f024a19b2c26 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Gemma 3 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b04911-4930-4a14-b996-e942cbefe43b · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1323640d-a83f-4a73-99fe-49b9b9ef3352 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c6e5d8-ecde-4168-a827-3ea7abf5ff01 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b2b2cf-224f-4ae4-9ad6-dde87fe12866 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39980b5c-efff-4ecf-817c-a9bd5c061323 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f9fcd0f-ff5f-4d3d-905f-9611cafbb6d2 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Hammer: Robust function- calling for on-device language models via function masking.arXiv preprint arXiv:2410.04587, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd85973-f39f-4e1c-88f6-0945f4952fda · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ToolACE: Winning the Points of LLM Function Calling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eea86bb-7289-4a91-95de-a03bb0a7515f · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79571b25-a819-4894-a3e3-928ce0f60229 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Gui agents: A survey
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce7beb51-7129-4ea0-883d-91ca30119bf8 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Training language models to follow instructions with human feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e5ed80b-ce8e-440d-ac78-7c1869bd53f5 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb841ce-f1d5-482e-a2b8-b447e2059dbb · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18440b04-b9a5-4dea-b32d-f6f4e07bcc55 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac12049-04b6-44bc-9e63-d089bc29f8c2 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a548b2e-ffed-4753-9dc6-204b79192997 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Direct preference optimization: Your language model is secretly a reward model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8dbf53-6548-4bfb-970e-1b80295858fd · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Proximal Policy Optimization Algorithms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7144fbd5-b679-453c-8af3-4f5464f345c6 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61d49fa2-1d15-422e-80b3-39416bc77069 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dbc19d7-538f-4172-8aea-d9e62cf42b6d · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Direct Multi-Turn Preference Optimization for Language Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a648e3d0-d204-426e-9c40-b3c0df0fe317 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd8d7a3-0a9a-4b3a-8383-414f080a6a23 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e11918f-c40e-4fcf-9e2c-4f56c9e8a377 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc48e6f6-ed47-47ba-a876-ac2299bc4867 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving.arXiv preprint arXiv:2601.01426, 2026
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eddeb439-8fc2-4d3c-bff0-b692f58a41f7 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Triplets better than pairs: Towards stable and effective self-play fine-tuning for LLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072b71ac-60e3-40ac-b0ad-f30414dc02fe · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d66964fa-2af2-4a5d-a84c-76b7f7c114ef · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbcf4ede-7dad-48e6-b29b-dd7363a87829 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories On the generalization of SFT: A reinforcement learning perspective with reward rectification.arXiv preprint arXiv:2508.05629, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e80a9c3c-7e51-4714-8bcb-d02e0649a8c2 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1e739c-bb86-4664-9402-992f048bad50 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f322476f-585f-4aa2-acac-e6eca0458269 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Qwen3 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eae69f32-f3d9-40a1-b647-63633bc230ea · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938b0a07-ceca-4086-8f3c-6be037a98c7f · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories React: Synergizing reasoning and acting in language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab2e065b-db38-425d-a90c-00716a677089 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4763284b-6737-4384-8075-b7db468ec3b4 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Pivotrl: High accuracy agentic post-training at low compute cost.arXiv preprint arXiv:2603.21383, 2026
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c1d6fa-76c3-40ae-a3d0-e44b3ee16705 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Self-Rewarding Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb0ce6a-0dd2-45e4-bfe3-044c7ba7c105 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agenttuning: Enabling generalized agent abilities for llms
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa68c9e-1438-475f-b7bf-7c4778a8b393 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ToolACE-R: Model-aware iterative training and adaptive refinement for tool learning.arXiv preprint arXiv:2504.01400, 2025
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c67c93-b0a2-4325-b8c2-92037ed6bb82 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories first_name
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3388738-55db-44b5-bc57-f5edc284d166 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories user_id":
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec77e9ca-df12-4c1e-9d2e-954a088d3cab · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories name": "Alice
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12432ffd-2b1f-45ec-9999-c5d43c59e399 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories user_id":
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e87c705e-95b9-402d-9331-7386dabd061e · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories user_id":
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbbb2dd-43d8-461c-a4a0-dc1205f589a9 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories error":
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312805d7-2730-439b-a138-166929038dd1 · outbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories user_id":
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875642c6-38ee-4344-9c6f-05ebcba0861e · inbound
Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2f3f8ba-9fab-439d-9502-a0dcb14cd4ea · inbound
Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories
Reference 2026
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.