Pith. sign in

Paper Citation Record · LEDGER

TAPO: Transition-Aware Policy Optimization for LLM Agents

As of 4 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.27973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27973 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T21:44:39.699017Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b4aa3bc-6628-42d1-a393-165418b0f174 · outbound

This paper cites GPT-4 Technical Report.

TAPO: Transition-Aware Policy Optimization for LLM Agents GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.585532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.585532Z digest=sha256:ace1159b4fd6d49d3469cdb749d5d234fb39451c70c0093ebaffe694ae896ddf

Observation 5831a163-a6c7-44db-8028-17f4af99b395 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.588812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.588812Z digest=sha256:49339193b0ba4e9b75f2b6cf1805015a28843ad8c899290c6e297d169b760b1d

Observation b8e22867-69a9-4e42-88da-edfa1627f947 · outbound

This paper cites Qwen2.5 Technical Report.

TAPO: Transition-Aware Policy Optimization for LLM Agents Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.592129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.592129Z digest=sha256:0ae2b7ae3c7220b3d6a312159ff2fb93dc7f631fe8320ea9fe490de4a6f99929

Observation a9fa85a6-9ae5-46c7-a736-b8f3626439dc · outbound

This paper cites DeepSeek-V3 Technical Report.

TAPO: Transition-Aware Policy Optimization for LLM Agents DeepSeek-V3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.595092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.595092Z digest=sha256:2b0c944b57bfd79f5f96394bf61e383bd3b12ea901b783db7634c8be93d1f500

Observation 1f02e486-dab4-461f-8114-afd539c2ae60 · outbound

This paper cites Multimodal Web Navigation with Instruction-Finetuned Foundation Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.598107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.598107Z digest=sha256:dc8c9232e7be66bf34b9368da48645fabebabfc07f7e3613a83fd847ee38e8cc

Observation 634df6b6-7e15-4fa5-8211-c2c498cd8216 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

TAPO: Transition-Aware Policy Optimization for LLM Agents GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.601537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.601537Z digest=sha256:724e21d081f3f47a538907c7e4004339e34e011f42b3e8442cca09bd7e771945

Observation 903a6313-e57f-489c-88f1-55f281659901 · outbound

This paper cites Embodied agent interface: Benchmarking llms for embodied decision making.Advances in Neural Information Processing Systems, 37:100428–100534, 2024.

TAPO: Transition-Aware Policy Optimization for LLM Agents Embodied agent interface: Benchmarking llms for embodied decision making.Advances in Neural Information Processing Systems, 37:100428–100534, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.604589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.604589Z digest=sha256:6ffa1c619b6ce7f0f257dc6c6951fd6764c6661509d09865e585794c22b0791b

Observation 6556022e-ee80-4c45-81db-d214d24cd033 · outbound

This paper cites MIT press Cambridge, 1998.

TAPO: Transition-Aware Policy Optimization for LLM Agents MIT press Cambridge, 1998

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.606756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.606756Z digest=sha256:e39a0151a7d3aa5a9931a3118f8b864f21758329e0bc7e01038a551a9b7d2e12

Observation 385bcef2-03b7-4696-b0e5-0e3a6d3d3b84 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

TAPO: Transition-Aware Policy Optimization for LLM Agents Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.609034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.609034Z digest=sha256:30a1c6d6c81fe052a924e1f552135b1ea491ae96ec007ae80008f00b4df1471c

Observation 71c2bca8-b13a-487f-9e18-37234ffc5b1b · outbound

This paper cites OpenAI o1 System Card.

TAPO: Transition-Aware Policy Optimization for LLM Agents OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.611311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.611311Z digest=sha256:4bd5c6c73747474e9aa73eb7eaf087dda128e10e0b6fd2fa8b247d7f20614ef9

Observation 411df13f-8af4-443c-8036-881b431581ac · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.614561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.614561Z digest=sha256:dc9856d1461dcfc4dbf878dda165923a6139d14ca049c934c9995bf529b4a4bc

Observation 888a4484-cdc8-4d06-85cc-3db0d89068bb · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.617527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.617527Z digest=sha256:967393ec0d8965e6bde0169be9eeb638bc37e3432cc33cc10e225c119f3162e7

Observation c8735516-36a0-43e9-89f9-5203bb40c61a · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

TAPO: Transition-Aware Policy Optimization for LLM Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.620076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.620076Z digest=sha256:9dd8d0ee7ae67d086bcea63fc4b1f5d5371971ba04748aebc853ad16f76b5c03

Observation 1c8ae0c1-cf39-4b32-98a2-c000c4af9b68 · outbound

This paper cites General agents need world models.

TAPO: Transition-Aware Policy Optimization for LLM Agents General agents need world models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.622957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.622957Z digest=sha256:33319d3a02eeb6f76fd53bf9cdc6b382be14ae21ba233eb4c1318ccdbb1abfaa

Observation c7e4a8bd-41a0-4f7a-87c1-794798b2141c · outbound

This paper cites Reinforcement Learning with Unsupervised Auxiliary Tasks.

TAPO: Transition-Aware Policy Optimization for LLM Agents Reinforcement Learning with Unsupervised Auxiliary Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.625477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.625477Z digest=sha256:3b661faada64a6a205ec070bee605348ff19a7dbcd8ba5fb765c0adcf68e45e6

Observation e6db1d2c-16ab-4d3a-9085-5c7787abe26a · outbound

This paper cites Deep- mdp: Learning continuous latent space models for representation learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents Deep- mdp: Learning continuous latent space models for representation learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.628079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.628079Z digest=sha256:75c6c1ddc5ce1ccd7756dd859787372fd932205dbadb87c0859814a520066db1

Observation 4d86f2f0-9407-47eb-9b2e-6975c1723600 · outbound

This paper cites Vlms- guided representation distillation for efficient vision-based reinforcement learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents Vlms- guided representation distillation for efficient vision-based reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.630204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.630204Z digest=sha256:0a928d0cebe1acd675ac03933d31266814a90225a7a9d2fde1d105c21de0b943

Observation 5328c020-c40e-4471-a32f-1bef29fd3f08 · outbound

This paper cites Data-Efficient Reinforcement Learning with Self-Predictive Representations.

TAPO: Transition-Aware Policy Optimization for LLM Agents Data-Efficient Reinforcement Learning with Self-Predictive Representations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.632400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.632400Z digest=sha256:adaec2a1a5e0cb40de47f09a52f427509e98b19ab4c9d2bf9845367f0d6d3ce7

Observation 5d95d0b2-43f3-403f-98bf-6a6ec15cb62e · outbound

This paper cites Agent Learning via Early Experience.

TAPO: Transition-Aware Policy Optimization for LLM Agents Agent Learning via Early Experience

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.634912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.634912Z digest=sha256:da11b2c978bb17623b5998d9f40ce9a7c1f8b8fb0b6049e43bdbc98b8905ce23

Observation fb474bf4-dcbe-46be-bb53-1dfdf73bab9c · outbound

This paper cites Reinforcement world model learning for llm-based agents.arXiv preprint arXiv:2602.05842, 2026.

TAPO: Transition-Aware Policy Optimization for LLM Agents Reinforcement world model learning for llm-based agents.arXiv preprint arXiv:2602.05842, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.637487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.637487Z digest=sha256:9c040f5947e68e467b666ee7244f02b89ac871c1c6e8d29fe22090f396bca48c

Observation bdac9022-35f5-449c-8d0b-a75e148c70f1 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022.

TAPO: Transition-Aware Policy Optimization for LLM Agents Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.639844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.639844Z digest=sha256:9361147fcf100248d8a9c6495bdcd2c8340b3a71ec226ad48e1f5bf464d6ccf4

Observation 4c015d4a-6083-4f83-9e76-5284edfb809e · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.642114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.642114Z digest=sha256:96233995a7d2eb8d06024ec8e25479ba83bc28c5862e187c458ca8f4e3a854f7

Observation aac405ef-6282-48e6-be71-c8d28ba29d35 · outbound

This paper cites CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges.

TAPO: Transition-Aware Policy Optimization for LLM Agents CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.644614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.644614Z digest=sha256:99ceae4f580e02f54526693a2c9345b3d53bf8316fb292688a3329625dd7c859

Observation 65cc9ddf-b3c2-4104-9aed-8f694fcc11c9 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

TAPO: Transition-Aware Policy Optimization for LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.647018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.647018Z digest=sha256:c1a6d4dd33bf32647b050e77803e34ba56eaad4cb0e54b66dc948087a862ac53

Observation 2d413b60-1d57-46da-87c7-279b4f935e39 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.649499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.649499Z digest=sha256:c9f313f73ec3cd505ec9768d9539aa596bf5a396d1ca5719c1839b37b3277078

Observation 6c0496fa-01bb-4c21-8b8d-35eb7d750d84 · outbound

This paper cites WebDancer: Towards Autonomous Information Seeking Agency.

TAPO: Transition-Aware Policy Optimization for LLM Agents WebDancer: Towards Autonomous Information Seeking Agency

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.651881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.651881Z digest=sha256:5eabe1431226eb734671f37d0c9965f284fb952075be7bf48c8244e5c958a7ff

Observation d3d1f9e4-8cea-4cb8-a25d-12a5c6f16966 · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

TAPO: Transition-Aware Policy Optimization for LLM Agents WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.654314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.654314Z digest=sha256:9f5ca64c2d35e83fa9126d3742fad380d1bd2dcd36433b785fd757d0f07ee297

Observation 32e21fec-cc70-4b66-83e2-2944416c30cd · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

TAPO: Transition-Aware Policy Optimization for LLM Agents Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.656683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.656683Z digest=sha256:668ed1fe2eff932e0647e9d6fd2ea11b96a710763677406e37d30e1cb5a97651

Observation e1bbfc75-3b80-4683-8353-be3af30bad12 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

TAPO: Transition-Aware Policy Optimization for LLM Agents React: Synergizing reasoning and acting in language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.658791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.658791Z digest=sha256:06cea3820f5c2a270bfaacc531c3ef4ad4e72d5289827bdb91b667c3d568a34d

Observation 774269c8-0ef8-4ddf-8c5e-e370ef54b1d2 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

TAPO: Transition-Aware Policy Optimization for LLM Agents Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.661240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.661240Z digest=sha256:2415632eb10f02bdce38b28c413444277355b9eebe31a7248f56b3fcc3ac66c2

Observation 05423b04-eb26-4c2c-a480-7126191203d3 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

TAPO: Transition-Aware Policy Optimization for LLM Agents Fine-Tuning Language Models from Human Preferences

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.663478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.663478Z digest=sha256:5f70d71f447eb5855d787cd6394a8ec90ff95f6e7f1d42e14ffd70d6981c3474

Observation 189ff840-67d9-4843-8c0e-de4eae2f4613 · outbound

This paper cites Proximal Policy Optimization Algorithms.

TAPO: Transition-Aware Policy Optimization for LLM Agents Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.666212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.666212Z digest=sha256:f693b056c1ec269da11b0f0e9b35acc09ff5776b4bbbce89faa2534f785295d9

Observation b915ecb9-cb16-42a9-beb5-851cc140193c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

TAPO: Transition-Aware Policy Optimization for LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.668576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.668576Z digest=sha256:481c4a7573c1986e26555097e8d16624a306aa8715cf93a3b9cf90203a0db1de

Observation 858ec9ad-6c65-497f-9cb6-c0e3ea8e723b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

TAPO: Transition-Aware Policy Optimization for LLM Agents Understanding R1-Zero-Like Training: A Critical Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.670928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.670928Z digest=sha256:ce6bc1d767ca152eaa19aa4bc782d937f7224bb7e4177fdf08314c856cc9696b

Observation f06992f8-47c7-4e56-8a08-c7341f06a8c4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.673802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.673802Z digest=sha256:3d811ba38a6c95068f0b7d01cdf76cf36bc87f86e4ade8adee3648fbf20c20f8

Observation 0ca8784e-c1de-493f-8d81-83c2274d1d71 · outbound

This paper cites Agentic Reinforced Policy Optimization.

TAPO: Transition-Aware Policy Optimization for LLM Agents Agentic Reinforced Policy Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.676348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.676348Z digest=sha256:658f2605e8c075987e5eb5732c6b735c771b43909ee7a86e9687f446d55780ba

Observation d929e877-9682-493b-9c17-ef7bfa0dd1f8 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

TAPO: Transition-Aware Policy Optimization for LLM Agents Dream to Control: Learning Behaviors by Latent Imagination

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.678831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.678831Z digest=sha256:22bfd6f4a950f85e8e48c3af03f2bf24d927b4f7dd812a9438555b14f9726d7c

Observation 93875012-557a-42a8-90df-798d46b23199 · outbound

This paper cites Mastering Atari with Discrete World Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Mastering Atari with Discrete World Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.681305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.681305Z digest=sha256:3e6c31f206b5ef65551838f794a9e26fd979d4c3ad8175eab89d00b2a5b1d256

Observation 6c8da77b-8f99-4c4b-8fb3-d0898bc32cc7 · outbound

This paper cites Mastering Diverse Domains through World Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Mastering Diverse Domains through World Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.683897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.683897Z digest=sha256:ba023e6a4f459369481b8cb87d907d271c43d1e26b56c1cab057b97dc2d7e33c

Observation 9d205fcf-7e3c-482f-8f60-5dbe3578bfcc · outbound

This paper cites Reason- ing with language model is planning with world model.

TAPO: Transition-Aware Policy Optimization for LLM Agents Reason- ing with language model is planning with world model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.686699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.686699Z digest=sha256:d8ece91884a670750a7138d419ed02ee7bf02192cced9a50f1ce3adbdc74fecf

Observation 5aa4ef01-dc0d-47cc-a9e7-253c083f8ffb · outbound

This paper cites Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation.

TAPO: Transition-Aware Policy Optimization for LLM Agents Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.688810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.688810Z digest=sha256:17ca8f64199fe3ffc382efc5c6ec54a722bc2fe92f456ab453065ef60d537965

Observation f202d84d-b982-4228-9ab8-288546579a13 · outbound

This paper cites Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents.

TAPO: Transition-Aware Policy Optimization for LLM Agents Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.691367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.691367Z digest=sha256:7c8a82d17cbe3b0c07b088da8e8af6fdbe80d7a20ecb98d05ddaee09986454a2

Observation 53c918da-4385-4142-8181-dcf78676b305 · outbound

This paper cites WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model.

TAPO: Transition-Aware Policy Optimization for LLM Agents WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.693932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.693932Z digest=sha256:963a937599bb565adf31f27d5bf7dd10e248efa1e52e998760e6a60cb56ea8c9

Observation 6198f19b-ce98-49eb-b616-8555e4fa7ad7 · outbound

This paper cites RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents.

TAPO: Transition-Aware Policy Optimization for LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.696164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.696164Z digest=sha256:cd9aa8b7366bea7f8f670a703fe746766e8ccae16118df9a6694e784b284a85d

Observation ff20e295-7b85-4f67-b39b-369ce0b7a07f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TAPO: Transition-Aware Policy Optimization for LLM Agents Training Verifiers to Solve Math Word Problems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.699017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.699017Z digest=sha256:21e34601ffe4c2cd92f3e924927cdde69557eb92cd91284ca9b6f8bd3ea73f84

Pith citing papers

No inbound Pith citation observations are available.