Pith. sign in

Paper Citation Record · LEDGER

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations

As of 23 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2501.12668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12668 v3

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:00:39.404564Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 356a74c4-edf7-4426-b32a-c996890c4e7d · outbound

This paper cites Variational Option Discovery Algorithms.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Variational Option Discovery Algorithms

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.284426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.284426Z digest=sha256:a7a0e2a4d5450df9fcf152d68e0f7622ce3dc67671f66a0dea46d855e439d2b3

Observation 1a5f856b-89a9-4d64-aace-8ff73ecf8c79 · outbound

This paper cites Reinforcement Learning with NBDI In downstream learning, our objective is to learn a skill policy πθ(z|st) that maximizes the expected sum of discounted rewards, parameterized by θ.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Reinforcement Learning with NBDI In downstream learning, our objective is to learn a skill policy πθ(z|st) that maximizes the expected sum of discounted rewards, parameterized by θ

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:00:39.815397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.391734Z digest=sha256:34bca79ae442af1a781b2b1c6603e018f12fd43ebd9a0c071b1c3e34b073a916

Observation 0a825fa3-7dc9-4f25-83fd-7223025ca041 · outbound

This paper cites Hierarchical Few-Shot Imitation with Skill Transition Models.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Hierarchical Few-Shot Imitation with Skill Transition Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.309382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.309382Z digest=sha256:1904fe6cdd4bf120df2e4737e177b8b36e874adf89fcb8989d6e1e72568ad405

Observation e98c0197-42c4-4961-822b-aa4f939e66bb · outbound

This paper cites Deep Successor Reinforcement Learning.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Deep Successor Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.319295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.319295Z digest=sha256:a4f5b26ab12d7e6543422e49aaa8ea3f71621f96a74fd13f5e05bc38aad97859

Observation 7ae88fbb-def1-4a3c-8136-190aae601a21 · outbound

This paper cites Parrot: Data-Driven Behavioral Priors for Reinforcement Learning.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Parrot: Data-Driven Behavioral Priors for Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.360466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.360466Z digest=sha256:f77964f46aaad03ec52ab2c1b82ccb6e95b53d2990249188299fa941837e713a

Observation c5c6dbcf-e155-416a-bfb6-54226364f196 · outbound

This paper cites The agent starts at the white point and aims to reach the red point.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations The agent starts at the white point and aims to reach the red point

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:00:39.796725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.396128Z digest=sha256:cdfa6a76f5da6b756a00412d1d7a17378e93acd33829a28de5140b09ad1244b1

Observation 3aa4090e-ae0a-429a-8780-0928e906c805 · outbound

This paper cites Note that task and environment setup differ between the training data and the downstream task, demonstrating the model’s capacity to handle unseen downstream tasks.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Note that task and environment setup differ between the training data and the downstream task, demonstrating the model’s capacity to handle unseen downstream tasks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:00:39.780037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.400432Z digest=sha256:f9abe8696d54f2750d86216862516ae5b83daed3ae9b116979637c8827bffdd1

Observation e096401d-a79c-4272-b851-7e3688af8321 · outbound

This paper cites In the maze and sparse block stacking environment, a single fully-connected layer with hidden dimension 52 and 70 have been used, respectively.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations In the maze and sparse block stacking environment, a single fully-connected layer with hidden dimension 52 and 70 have been used, respectively

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:00:39.765484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.404564Z digest=sha256:4c3e5fa39e7a74a461e158487abc2edaceb2aa5d370092845763303b66e1cd89

Observation 26ba5c93-b7d0-4a1b-96c8-7a4988efb30f · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Offline Reinforcement Learning with Implicit Q-Learning

Reference 1993

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.314060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.314060Z digest=sha256:6f252ea0bdf0b7d35eca97544c6277ab67b8ccb821e218e91340ea33d80abc7c

Observation 3777762c-810c-4abc-9f7d-332f75ec639a · outbound

This paper cites Exploration by Random Network Distillation.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Exploration by Random Network Distillation

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.295533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.295533Z digest=sha256:0b4b00bb23989c1f126b1c46b2c29e42241ee0de08a0c672e3d41b30f0c25ba4

Observation 73f2092a-90a9-4955-ba5f-8eebce436c02 · outbound

This paper cites an unresolved cited work.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Unresolved cited work

Reference 1997

Resolution
unresolved
raw_fallback, observed 2026-08-10T17:00:39.835376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.385758Z digest=sha256:df3d0b059ba9da5d7037e53103aab9bab61e38ad50c0e109317708fd8631998e

Observation 131ef7c3-d7a5-457e-9e57-d18f5fe42b64 · outbound

This paper cites Mujoco: A physics engine for model-based control.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Mujoco: A physics engine for model-based control

Reference 1998

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:00:39.856047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.370836Z digest=sha256:a53a99cfa0914ddc9bdb37c74c6d23192d8e36cadeb8826c44345a0cb97f04ea

Observation 30ba7568-bd4b-4387-b138-9e2b5a9d5430 · outbound

This paper cites Q-cut—dynamic discovery of sub-goals in reinforcement learning.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Q-cut—dynamic discovery of sub-goals in reinforcement learning

Reference 2001

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:00:39.890122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.329938Z digest=sha256:d89a9b46975e83dabe5a3b0535fa32bb91d08c5b4002a7d42b08bf3ca451ce7f

Observation 984b62c6-35aa-430a-b697-02ce6b3b832e · outbound

This paper cites Skill-based Meta-Reinforcement Learning.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Skill-based Meta-Reinforcement Learning

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.334889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.334889Z digest=sha256:ee20c737249a5857ead2d8568e2e8bcb7dd6616548ff920a8ecb504a7ab49b77

Observation acdc3fa3-1cc9-4ea2-991f-9fa2ffed6108 · outbound

This paper cites Variational Intrinsic Control.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Variational Intrinsic Control

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.304456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.304456Z digest=sha256:d266e97b30e9666bdc8404edf1b60893381fa1eea34927d8995c12d6d9725711

Observation 33e7d08d-7ea8-422b-a1d8-7557cdafc3f7 · outbound

This paper cites P., and Barto, A.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations P., and Barto, A

Reference 2004

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:00:39.872478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.352330Z digest=sha256:69944d1c184dff4f6eff4f0d4cef130943f20cd362a6227baad163a2dd94e39e

Observation 50cf7c16-3959-491a-ae21-4eec5534cd17 · outbound

This paper cites URL https://doi.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations URL https://doi

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.356602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.356602Z digest=sha256:6741b4f466ccc13e552cb27f1e87380fd7259a8cf6739baa0f68f0ecac6d5c42

Observation 044f9e51-79e9-4135-a1a5-745a1f0ce861 · outbound

This paper cites TRAIL: Near-Optimal Imitation Learning with Suboptimal Data.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations TRAIL: Near-Optimal Imitation Learning with Suboptimal Data

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.375264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.375264Z digest=sha256:fc9b618f4acdef8f0967a7495cf80a6885c55a585535e0728daa920aaa23cb4a

Observation ee194b79-695a-4891-b1f6-e642eb3daf26 · outbound

This paper cites Eigenoption Discovery through the Deep Successor Representation.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Eigenoption Discovery through the Deep Successor Representation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.325081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.325081Z digest=sha256:65e88436731e94a7faa6a4be080699d3141a1050d912ccf55df03bca37b5a8be

Observation 9ebb9f36-1bc3-48b6-bb20-4d7db3969973 · outbound

This paper cites OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.289737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.289737Z digest=sha256:816fd46c2da28f6aa17e3698dc67d6add6781bbbf317abf4379a0345487f2e0b

Observation 66e98dcc-5346-4e4a-85a7-c15980ea3dc4 · outbound

This paper cites Demonstration-Guided Reinforcement Learning with Learned Skills.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Demonstration-Guided Reinforcement Learning with Learned Skills

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.340097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.340097Z digest=sha256:181959fe2526c69252f50fd6fd6ea5972cd9fda3cdbf29f600d9c6ae6899b667

Observation 5fa293b8-4b3c-4717-9d21-9f2de7321e1e · outbound

This paper cites Option Discovery in Hierarchical Reinforcement Learning using Spatio-Temporal Clustering.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Option Discovery in Hierarchical Reinforcement Learning using Spatio-Temporal Clustering

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:39.364979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:39.364979Z digest=sha256:89ed922be0502a82d803938da2f5b64fc40329388b40b20330d71119e410f79a

Observation b6eb3ddb-7587-4d19-a166-2aef55ca556d · outbound

This paper cites Investigating the effectiveness of bpe: The power of shorter sequences.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations Investigating the effectiveness of bpe: The power of shorter sequences

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:00:39.905139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.300158Z digest=sha256:cb807a977b4b09037a5f7148d83401a980745167c9c20c6875301d1cf8a57f6f

Observation 6ffefef2-e886-4bfb-adc1-3e990a54a59e · outbound

This paper cites MO2: Model-Based Offline Options.

NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations MO2: Model-Based Offline Options

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T17:00:39.595870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T17:00:39.347876Z digest=sha256:ba41bba0d645ee3b18cccbca9f7b5147699665b436c70c9f5481f25d453802a0

Pith citing papers

No inbound Pith citation observations are available.