Pith. sign in

Paper Citation Record · LEDGER

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.22008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22008 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:18:42.180146Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d7dec4f-6150-4ed1-994b-60799a054b25 · outbound

This paper cites It is an open-world city simulation originally proposed by Sestini et al.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning It is an open-world city simulation originally proposed by Sestini et al

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:43.932432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:41.702652Z digest=sha256:da1e3142f793ae67bac212048253331263ad7908b39d5b673d9f392d84dc23ee

Observation 8fd18dfc-2dfc-47e6-bd2e-4db9c7b22d91 · outbound

This paper cites Semi-supervised reward learning for offline reinforcement learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Semi-supervised reward learning for offline reinforcement learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:18:42.627088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:40.147376Z digest=sha256:e017d01cb68967d8f82dd7717a73cc587c4a3cb7024bbf452198fa69f427be60

Observation 712750b4-bcd7-4a39-903d-6e8105fc559e · outbound

This paper cites Imitating Human Behaviour with Diffusion Models.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Imitating Human Behaviour with Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.482865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.482865Z digest=sha256:4546218cb6d8bd3ba363d4a225d3746d3fcfa59adb31831dde7bc799c9192b4f

Observation db243e78-9b73-4234-934d-fe74ed65f3cc · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.659299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.659299Z digest=sha256:16f96dfbc9004d5f49fa5fd853f51f2b140fe20954a6c59e419a648d0fcda99b

Observation 1b91c065-8c05-4f75-97eb-340212b64dd2 · outbound

This paper cites Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-Ensembles.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-Ensembles

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:18:42.411597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:41.072225Z digest=sha256:023f01e61876c3dd62f47243dd4df238b0c1957bbffc3265788891e7af8271fd

Observation bfbfab27-e9ab-46be-8d03-b18b17121187 · outbound

This paper cites Offline Learning from Demonstrations and Unlabeled Experience.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Offline Learning from Demonstrations and Unlabeled Experience

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:41.211864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:41.211864Z digest=sha256:3be0faeb940a5a61cb0924b639881f5e2cc5eb51fcd64c4aceb5093c6a69eeb7

Observation d0fc60be-e3a7-473b-a2d5-a12719025707 · outbound

This paper cites Offline Reinforcement Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Offline Reinforcement Learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:44.829576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:41.318621Z digest=sha256:7f24e90610bad917f4f75333731ef7288d2badff4d60c50cc212adcbe96600ae

Observation bf7933e9-9537-4676-b198-4e8b16331d02 · outbound

This paper cites As we describe in Section 3, it is especially suitable for offline settings.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning As we describe in Section 3, it is especially suitable for offline settings

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:44.521428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:41.444412Z digest=sha256:91dee631089ecbade24938c155b5d77b2bce15f5411823edddd9742150dc3062

Observation 42e8ad3a-8bb7-4dbd-a2bf-00e9e752eedc · outbound

This paper cites All of these approaches require optimal expert demonstrations – that is, demonstrations generated by an optimal policy.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning All of these approaches require optimal expert demonstrations – that is, demonstrations generated by an optimal policy

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:44.250384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:41.581591Z digest=sha256:eee960f696851ac5a78a4aa5f2ca0fd502af8f60639af13fa9f9f6c8346547e6

Observation 23f83ca3-7541-496e-b363-2e8a34dc3744 · outbound

This paper cites An episode is marked a success if the agent reaches the goal before the timeout.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning An episode is marked a success if the agent reaches the goal before the timeout

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:43.646139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:41.855809Z digest=sha256:87f7b8a595adaa53debd78767e5fbf702f21dabc4f95b65aa0cb6b1fd6d282f4

Observation 59acd6bf-03da-4eaa-96a4-29a66e2f455c · outbound

This paper cites We provide more details about the baselines in Section 4.2 of the main paper.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning We provide more details about the baselines in Section 4.2 of the main paper

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:43.400993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:42.053775Z digest=sha256:bd4d5e2bac9d7c50a97c986180a7b15ffa96aeffa985a3fec40210e4fd0ec989

Observation f204ddd1-e672-44b6-95c5-8bafe6d88323 · outbound

This paper cites an unresolved cited work.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:18:43.133188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:42.180146Z digest=sha256:5750d2f17f2aeb7a4bf36866afe200680e95802f20e16003a27fec856744c92e

Observation d4c60b3a-dd2e-4220-9dd3-04619ddab4d2 · outbound

This paper cites Efficient Active Imitation Learning with Random Network Distillation.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Efficient Active Imitation Learning with Random Network Distillation

Reference 1995

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:18:42.862234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:39.654517Z digest=sha256:a51c72a757d374f48f2c44c704b00f09c832152f6cd1deee3cbd57599ce73ab6

Observation 0468167d-bdc4-4e3a-8b6a-eb989e5aefb8 · outbound

This paper cites The Provable Benefits of Unsupervised Data Sharing for Offline Reinforcement Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning The Provable Benefits of Unsupervised Data Sharing for Offline Reinforcement Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:39.898864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:39.898864Z digest=sha256:362259570f4cd88a3b4a56b603112117c7321ce4ba239f44ea91b2a290fb78fa

Observation 31bad089-20d2-4f3c-b523-8c9ce563d6b3 · outbound

This paper cites Technical challenges of deploying reinforcement learning agents for game testing in aaa games.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Technical challenges of deploying reinforcement learning agents for game testing in aaa games

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:45.697343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:39.752822Z digest=sha256:4383b429418448aae415abf144632c6396096d9bdd07e94ee310627d7ee430e3

Observation d0f97daf-17e3-4e84-9b5d-5f5a97969c03 · outbound

This paper cites Demonstration-efficient inverse reinforcement learning in procedurally generated environments.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Demonstration-efficient inverse reinforcement learning in procedurally generated environments

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:45.468020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:40.791397Z digest=sha256:52650da0797fc9d0534fdb0b53b9cf11c63fbb9af178a6eb383e46d10ae31a97

Observation 213c6f6d-44c1-4820-bc60-a440fd8fe231 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.309764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.309764Z digest=sha256:09a1dd1567f6d7f193a72fb34285fe50fb06d9dd50e78ee00506b63b8c00c00f

Observation 2209042a-0467-4f47-87c4-06c0de714503 · outbound

This paper cites Towards informed design and validation assistance in computer games using imitation learn- ing.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Towards informed design and validation assistance in computer games using imitation learn- ing

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:45.099069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:40.929974Z digest=sha256:7da810161c6f8b87433842a960be1e33e0a54ebc565f2679d2f6a568e2833f97

Observation 3ebf8c51-d315-40b4-bddf-d33506d93911 · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Beyond Reward: Offline Preference-guided Policy Optimization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.007912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.007912Z digest=sha256:37102cafa04a233606b846e4ef7fae0ef5bb1ef18866d9ba82edb28d088f1d78

Pith citing papers

No inbound Pith citation observations are available.