Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T21:44:39.699017Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.27973.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T21:44:39.699017Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4b4aa3bc-6628-42d1-a393-165418b0f174 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5831a163-a6c7-44db-8028-17f4af99b395 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Gemini: A Family of Highly Capable Multimodal Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e22867-69a9-4e42-88da-edfa1627f947 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Qwen2.5 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9fa85a6-9ae5-46c7-a736-b8f3626439dc · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents DeepSeek-V3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f02e486-dab4-461f-8114-afd539c2ae60 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Multimodal Web Navigation with Instruction-Finetuned Foundation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 634df6b6-7e15-4fa5-8211-c2c498cd8216 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents GPT-4V(ision) is a Generalist Web Agent, if Grounded
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 903a6313-e57f-489c-88f1-55f281659901 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Embodied agent interface: Benchmarking llms for embodied decision making.Advances in Neural Information Processing Systems, 37:100428–100534, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6556022e-ee80-4c45-81db-d214d24cd033 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents MIT press Cambridge, 1998
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 385bcef2-03b7-4696-b0e5-0e3a6d3d3b84 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71c2bca8-b13a-487f-9e18-37234ffc5b1b · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents OpenAI o1 System Card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 411df13f-8af4-443c-8036-881b431581ac · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 888a4484-cdc8-4d06-85cc-3db0d89068bb · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8735516-36a0-43e9-89f9-5203bb40c61a · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Group-in-Group Policy Optimization for LLM Agent Training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8ae0c1-cf39-4b32-98a2-c000c4af9b68 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents General agents need world models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e4a8bd-41a0-4f7a-87c1-794798b2141c · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Reinforcement Learning with Unsupervised Auxiliary Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6db1d2c-16ab-4d3a-9085-5c7787abe26a · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Deep- mdp: Learning continuous latent space models for representation learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d86f2f0-9407-47eb-9b2e-6975c1723600 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Vlms- guided representation distillation for efficient vision-based reinforcement learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5328c020-c40e-4471-a32f-1bef29fd3f08 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Data-Efficient Reinforcement Learning with Self-Predictive Representations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d95d0b2-43f3-403f-98bf-6a6ec15cb62e · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Agent Learning via Early Experience
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb474bf4-dcbe-46be-bb53-1dfdf73bab9c · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Reinforcement world model learning for llm-based agents.arXiv preprint arXiv:2602.05842, 2026
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdac9022-35f5-449c-8d0b-a75e148c70f1 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c015d4a-6083-4f83-9e76-5284edfb809e · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aac405ef-6282-48e6-be71-c8d28ba29d35 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65cc9ddf-b3c2-4104-9aed-8f694fcc11c9 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d413b60-1d57-46da-87c7-279b4f935e39 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c0496fa-01bb-4c21-8b8d-35eb7d750d84 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents WebDancer: Towards Autonomous Information Seeking Agency
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d1f9e4-8cea-4cb8-a25d-12a5c6f16966 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents WebSailor: Navigating Super-human Reasoning for Web Agent
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e21fec-cc70-4b66-83e2-2944416c30cd · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Rt-2: Vision-language-action models transfer web knowledge to robotic control
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1bbfc75-3b80-4683-8353-be3af30bad12 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents React: Synergizing reasoning and acting in language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774269c8-0ef8-4ddf-8c5e-e370ef54b1d2 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05423b04-eb26-4c2c-a480-7126191203d3 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Fine-Tuning Language Models from Human Preferences
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 189ff840-67d9-4843-8c0e-de4eae2f4613 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Proximal Policy Optimization Algorithms
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b915ecb9-cb16-42a9-beb5-851cc140193c · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 858ec9ad-6c65-497f-9cb6-c0e3ea8e723b · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Understanding R1-Zero-Like Training: A Critical Perspective
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f06992f8-47c7-4e56-8a08-c7341f06a8c4 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca8784e-c1de-493f-8d81-83c2274d1d71 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Agentic Reinforced Policy Optimization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d929e877-9682-493b-9c17-ef7bfa0dd1f8 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Dream to Control: Learning Behaviors by Latent Imagination
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93875012-557a-42a8-90df-798d46b23199 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Mastering Atari with Discrete World Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c8da77b-8f99-4c4b-8fb3-d0898bc32cc7 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Mastering Diverse Domains through World Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d205fcf-7e3c-482f-8f60-5dbe3578bfcc · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Reason- ing with language model is planning with world model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa4ef01-dc0d-47cc-a9e7-253c083f8ffb · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f202d84d-b982-4228-9ab8-288546579a13 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c918da-4385-4142-8181-dcf78676b305 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6198f19b-ce98-49eb-b616-8555e4fa7ad7 · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff20e295-7b85-4f67-b39b-369ce0b7a07f · outbound
TAPO: Transition-Aware Policy Optimization for LLM Agents Training Verifiers to Solve Math Word Problems
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.