Pith. sign in

Paper Citation Record · LEDGER

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories

As of 21 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2607.10601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10601 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T10:33:54.851493Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:44:31.301557Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T12:16:13.904932Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4bd8000a-17c5-497e-991b-75af1d98781d · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories On-policy distillation of language models: Learning from self-generated mistakes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:9ee804e550e0ed82a35fbbd11e57651e0c30fe597357ee742f58386255189448

Observation 7bed126f-49a0-4b34-8771-8c9097274e48 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories A general theoretical paradigm to understand learning from human preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:7e1117ba26ea140d7ca61319b5464a459321fb9cd74794e9b3e694df0b5a3cd8

Observation 2a742f75-8987-4797-96ca-d4981f57b7da · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:c7f764a2860261459d43cdb88bd3f45e713253f4ae3bc227d0172cb45fc61faa

Observation fadb3a24-6dea-449d-8890-a67f0292a037 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories FireAct: Toward Language Agent Fine-tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:44687cc8aee5de44675bf64224583ef97315f949534af7e77aa757f190a52c05

Observation aa1023e0-2103-4140-8bcf-38a7d6edef7f · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:85607a5998b4573b58268f2ff507bb0188ce01016595097f1c5a618c68fb4f16

Observation a2032254-9900-4d09-a479-56845208ee84 · outbound

This paper cites ATLaS: Agent Tuning via Learning Critical Steps.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ATLaS: Agent Tuning via Learning Critical Steps

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:2dc4585f2b0f26f7b81a3a8249cfefb31c5234ac7dd50f0c8f259801ae8b02e6

Observation 073b40c5-dbb9-486f-9b6b-5a46a17d6e3f · outbound

This paper cites Self-play fine-tuning converts weak language models to strong language models.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Self-play fine-tuning converts weak language models to strong language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:61d307faa378e2b224e2122745a5b6b810f9604be30be7c419bfc17af5572874

Observation 1b629639-8ed1-4636-8dfb-7d7e727e1127 · outbound

This paper cites Rethinking DPO: The Role of Rejected Responses in Preference Misalignment.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Rethinking DPO: The Role of Rejected Responses in Preference Misalignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:2584f949e35396c3c8c552113f3783a2f98e76c3fdfeae26543a14fb2df5807c

Observation daced041-f2dc-4f63-8d6b-cc3dd8ae344b · outbound

This paper cites Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:7a7c7ecae5d1ad99fe4c9179d7329b75414f0a42f5ae7135f6fd1dbaa367b413

Observation 0aabe613-5e49-464b-bcd6-d37bf6f05257 · outbound

This paper cites Mind2Web: Towards a generalist agent for the web.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Mind2Web: Towards a generalist agent for the web

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:b4f3979cf6a23f390dc2f0e59c98692df1590a5fb01a812586d10a69ebb951f0

Observation eea484d1-9b2d-40cf-947d-18c3272a465f · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories KTO: Model Alignment as Prospect Theoretic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:06684b9517e8ace2ad5edbfd04c492b122cbdc5d52991a782eaf7a3297872468

Observation ddddbc5f-1585-4c36-8a98-841c2807c6af · outbound

This paper cites AgentRefine: Enhancing Agent Generalization through Refinement Tuning.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories AgentRefine: Enhancing Agent Generalization through Refinement Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:9648b4bd6bbf506979f77796636730755db522d2c93b3b9f772523fc14135243

Observation ea5c0a9f-0bad-41e1-848d-3d75d73784c8 · outbound

This paper cites Solving the granularity mismatch: Hierarchical preference learning for long-horizon llm agents.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Solving the granularity mismatch: Hierarchical preference learning for long-horizon llm agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:5070b43defddb7b7ea66567dcb5e7f1088c65d3c3f78812cb9049306e48946a6

Observation ecceda03-3c18-4ebb-bdfb-f024a19b2c26 · outbound

This paper cites Gemma 3 Technical Report.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Gemma 3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:4c440069d5c4cf19042ddd039786793d34d6d64405dd60a3f36a88c11d8ab3f6

Observation 13b04911-4930-4a14-b996-e942cbefe43b · outbound

This paper cites StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:d65ab645bd60b2d2ab401b11c56649affb95dfefa7749a8d8865b97e906e2ade

Observation 1323640d-a83f-4a73-99fe-49b9b9ef3352 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:29b3deccf461758372b2fcf3e9bd6f5e22b53a2ab162f778b31bef1258dc72c5

Observation 95c6e5d8-ecde-4168-a827-3ea7abf5ff01 · outbound

This paper cites DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:f5140e76ab9c5d51b69bb098e0cb1d52cf1b5a7e628f40715e8ae504132be6df

Observation 87b2b2cf-224f-4ae4-9ad6-dde87fe12866 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:8e2c88e1041fc735e4d560dccde9af602958ff0db2a670b28b6685276bbefaa1

Observation 39980b5c-efff-4ecf-817c-a9bd5c061323 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:a16ff1304414ae59ba6b862d35a768881bbb09b1b0fe2f16a6300fc0652e4eeb

Observation 7f9fcd0f-ff5f-4d3d-905f-9611cafbb6d2 · outbound

This paper cites Hammer: Robust function- calling for on-device language models via function masking.arXiv preprint arXiv:2410.04587, 2024.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Hammer: Robust function- calling for on-device language models via function masking.arXiv preprint arXiv:2410.04587, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:8a41db9b481fbf03f5e31ce4990b2ff5fdd1f5d16eb1d9f10f77eef8c9879278

Observation 5cd85973-f39f-4e1c-88f6-0945f4952fda · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ToolACE: Winning the Points of LLM Function Calling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:af8211d7a0c36013546201b77cdb84dff0ba9d0f42f45b307cf2da22ddabd69e

Observation 5eea86bb-7289-4a91-95de-a03bb0a7515f · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:49257392423768cb2921bc6a384ba577aa6136a6d4d41237bc6305e89a589292

Observation 79571b25-a819-4894-a3e3-928ce0f60229 · outbound

This paper cites Gui agents: A survey.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Gui agents: A survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:b245be18d6eb69747f72573b352cdcbae91fdb394371ba1d4358285259fc4613

Observation ce7beb51-7129-4ea0-883d-91ca30119bf8 · outbound

This paper cites Training language models to follow instructions with human feedback.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Training language models to follow instructions with human feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:b296dd39a0a235f00e2dabc1282a182bd26fbe5d05a86dce2c114150e5c55815

Observation 2e5ed80b-ce8e-440d-ac78-7c1869bd53f5 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:d428789599403190c239b413bb63f6deb903ba7e846f4bf7098b9cffe757eb6d

Observation bfb841ce-f1d5-482e-a2b8-b447e2059dbb · outbound

This paper cites Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:00b188540ca6b72ff7b7a6f0e89bd3a53e4e5206f843c312af2e8c5ff3ea4e99

Observation 18440b04-b9a5-4dea-b32d-f6f4e07bcc55 · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:09817f4623715b1127635e9d73247e13fc37be7ec3e93a7b1e4e89069018dc6f

Observation cac12049-04b6-44bc-9e63-d089bc29f8c2 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:926ec7e3b6c055afcd3f29c3ddff4f60b24296d49e04ee531ee249dfaea95296

Observation 8a548b2e-ffed-4753-9dc6-204b79192997 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:e2b12404f1668c9eac8a7f4bed098ca000ab9688bda64556e04bdc5470af8260

Observation 0c8dbf53-6548-4bfb-970e-1b80295858fd · outbound

This paper cites Proximal Policy Optimization Algorithms.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:bbb9bb3a4f115736a35e3c71976e159c6407ca1971f45a8115d2ad6e9727ec9f

Observation 7144fbd5-b679-453c-8af3-4f5464f345c6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:761df847957162841c99f12be83ecd2aaac2471b6be7b9f85b3e2b2b0e80dd3b

Observation 61d49fa2-1d15-422e-80b3-39416bc77069 · outbound

This paper cites ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:858ed6fed469899abc97c938befa15f8f15f3be1817761ed87977d0bdc85f664

Observation 2dbc19d7-538f-4172-8aea-d9e62cf42b6d · outbound

This paper cites Direct Multi-Turn Preference Optimization for Language Agents.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Direct Multi-Turn Preference Optimization for Language Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:81fddfa1bc04d54a984fdabbc6e3bd2ec7db4d119af37383a7302cc32ec1fdf7

Observation a648e3d0-d204-426e-9c40-b3c0df0fe317 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:a7013e3ed0914364ac64f856d37710d02918c1d780134fabe4356078ce9a3c4c

Observation 7fd8d7a3-0a9a-4b3a-8383-414f080a6a23 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:59054f72ac7a7f31ab2f30b6c5e236d651d5b7f74845793f0904f669e1acdd75

Observation 9e11918f-c40e-4fcf-9e2c-4f56c9e8a377 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:60b3e43729deee179f704df1ba7c1727f62368728d2245bcbbb37c37b6d46d43

Observation dc48e6f6-ed47-47ba-a876-ac2299bc4867 · outbound

This paper cites Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving.arXiv preprint arXiv:2601.01426, 2026.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving.arXiv preprint arXiv:2601.01426, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:e01a75914371abbbce437b95ecd6daa90dd069816262e3db248253719ab8cdfe

Observation eddeb439-8fc2-4d3c-bff0-b692f58a41f7 · outbound

This paper cites Triplets better than pairs: Towards stable and effective self-play fine-tuning for LLMs.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Triplets better than pairs: Towards stable and effective self-play fine-tuning for LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:af27d7b322dd7ee6ed654fbf7b3f524b3b9ce0dfb50587729a464484bae027a7

Observation 072b71ac-60e3-40ac-b0ad-f30414dc02fe · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:3bd2affd962eb8aae8fb7678c272d26fe01546deb4c589711ae1b59bdfcc21fd

Observation d66964fa-2af2-4a5d-a84c-76b7f7c114ef · outbound

This paper cites Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:edaafea425852ea3729757023ddbc2db7572ab3a4798d47d9ceb9bc79edd7550

Observation bbcf4ede-7dad-48e6-b29b-dd7363a87829 · outbound

This paper cites On the generalization of SFT: A reinforcement learning perspective with reward rectification.arXiv preprint arXiv:2508.05629, 2025.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories On the generalization of SFT: A reinforcement learning perspective with reward rectification.arXiv preprint arXiv:2508.05629, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:8871f7d2feaad31a16d7e17fbea278586e6a64ce4eeb000565c75e78e101fd58

Observation e80a9c3c-7e51-4714-8bcb-d02e0649a8c2 · outbound

This paper cites an unresolved cited work.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:6a7768d5cde5b5b53a057416ed6835ba677fa2796c506282b3a20b054f58f8f7

Observation 6b1e739c-bb86-4664-9402-992f048bad50 · outbound

This paper cites AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:5742a3eb4ae5fb8effd3e8362522956e2b10ef6896b859ab022e7eee1fc4a8a6

Observation f322476f-585f-4aa2-acac-e6eca0458269 · outbound

This paper cites Qwen3 Technical Report.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Qwen3 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:dbbe05766b8d0ea97eb3abe7dfa99510310adff6d9bf789f4f6c41df1c359d58

Observation eae69f32-f3d9-40a1-b647-63633bc230ea · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:e8a9449847fe3b48cab33c43a490d61f6d8bae3c754e3814f4477c9c4b55b728

Observation 938b0a07-ceca-4086-8f3c-6be037a98c7f · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories React: Synergizing reasoning and acting in language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:057e786f0f0aa684439c45b0bc7ac773610fc48d4af7ef51388dce30cadc40bc

Observation ab2e065b-db38-425d-a90c-00716a677089 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:98c7a8b91802353daadf3d8506bb83ee31248c5a81e26a5e7571e5711c22defd

Observation 4763284b-6737-4384-8075-b7db468ec3b4 · outbound

This paper cites Pivotrl: High accuracy agentic post-training at low compute cost.arXiv preprint arXiv:2603.21383, 2026.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Pivotrl: High accuracy agentic post-training at low compute cost.arXiv preprint arXiv:2603.21383, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:759d5d93e07a8d124fc265dd7393fffbeaffafdb6fe5217282991c466511ea28

Observation e7c1d6fa-76c3-40ae-a3d0-e44b3ee16705 · outbound

This paper cites Self-Rewarding Language Models.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Self-Rewarding Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:edd413e31f509606995dc7ef99fbed2efcae341e3da75cc36ce56530329923a0

Observation 6cb0ce6a-0dd2-45e4-bfe3-044c7ba7c105 · outbound

This paper cites Agenttuning: Enabling generalized agent abilities for llms.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agenttuning: Enabling generalized agent abilities for llms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:d89a6a986dbca0ba7875d854270187b3b37bfa675f0c24cd1b998e99ab53554d

Observation 8fa68c9e-1438-475f-b7bf-7c4778a8b393 · outbound

This paper cites ToolACE-R: Model-aware iterative training and adaptive refinement for tool learning.arXiv preprint arXiv:2504.01400, 2025.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ToolACE-R: Model-aware iterative training and adaptive refinement for tool learning.arXiv preprint arXiv:2504.01400, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:debc4dc1a09d91fa6c3e9f64014a3387610598f032897bc88719ea12bea3ac00

Observation 39c67c93-b0a2-4325-b8c2-92037ed6bb82 · outbound

This paper cites first_name.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories first_name

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:4a7a7096dec4c56ff1e7f054e42f0679ac827d8154ecf35fb8c4ec37b213d766

Observation a3388738-55db-44b5-bc57-f5edc284d166 · outbound

This paper cites user_id":.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories user_id":

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:9e97907cff313413851f414565472467768e93cd8fdb2c3ec2cde38a3de3a391

Observation ec77e9ca-df12-4c1e-9d2e-954a088d3cab · outbound

This paper cites name": "Alice.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories name": "Alice

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:5eeae3e29a08640fe93ce2b136d6ea210e7aebe16fab47849f53ab8c51ce1573

Observation 12432ffd-2b1f-45ec-9999-c5d43c59e399 · outbound

This paper cites user_id":.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories user_id":

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:2e77a436e2a78dfe1772027c35eee5a7e1eeba96c78accb1deedb98219a4a7ad

Observation e87c705e-95b9-402d-9331-7386dabd061e · outbound

This paper cites user_id":.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories user_id":

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:c2e93f7e7eac1fbc476a7eafa029b5106fd81d42ecdc9e4556cac4abe4c02b77

Observation dfbbb2dd-43d8-461c-a4a0-dc1205f589a9 · outbound

This paper cites error":.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories error":

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:a925ef68ea75738e82c4bc6682ab3f8ec2a1bd91549e78428f2a4655a1771e1d

Observation 312805d7-2730-439b-a138-166929038dd1 · outbound

This paper cites user_id":.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories user_id":

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:5ba0ad0b380cad5e0eab0c22e6057bb5c35ba7e844b691ed585ecd84ee3a7a08

Pith citing papers

Observation 875642c6-38ee-4344-9c6f-05ebcba0861e · inbound

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit cites this paper.

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:31.189084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:31.189084Z digest=sha256:4c9abc967c9f10166942e36e96100c3c9e43e5415c650049bacb60744713b1e0

Observation c2f3f8ba-9fab-439d-9502-a0dcb14cd4ea · inbound

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit cites this paper.

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories

Reference 2026

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T09:49:29.305770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T09:44:31.301557Z digest=sha256:9c9776842a0d6dd379f19ea6c6203ac140d3f372047585020879313457593713