Pith. sign in

Paper Citation Record · LEDGER

Continual Task Learning through Adaptive Policy Self-Composition

As of 17 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2411.11364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11364 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:43:02.792410Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 492b9325-5d1f-4314-88f1-23a4ce261f3d · outbound

This paper cites Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal.

Continual Task Learning through Adaptive Policy Self-Composition Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:43:02.942633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:43:02.743725Z digest=sha256:1ffa7472c31d0dd43bb771f92685af781eb1ad0d8d098bb6105a111502cd9e76

Observation 9e214f58-064d-4f9d-a688-313a39ce7ce6 · outbound

This paper cites Merging decision transformers: Weight averaging for form- ing multi-task policies.

Continual Task Learning through Adaptive Policy Self-Composition Merging decision transformers: Weight averaging for form- ing multi-task policies

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:43:03.066575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:43:02.752679Z digest=sha256:d95ae71b0dda2831e001483cf2202ef0a1b96d3a4cdae6031810e8df90df6c67

Observation fb4209c1-50ce-4b10-9444-46a4b35a1433 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Continual Task Learning through Adaptive Policy Self-Composition Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.757009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.757009Z digest=sha256:e0deaccd048882aa2657b21164288bc431704377e3c9e79d1ea25870e7a9fd19

Observation fce135ad-4685-4a96-9c9b-dbfae5d5279e · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Continual Task Learning through Adaptive Policy Self-Composition Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.765627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.765627Z digest=sha256:c5e57227a3e1a7d59b78993cbfbb71a3b7c0d94e81c81e9aa5380059941a6a5d

Observation 0a4f1e8c-0551-40c0-b356-3208de4efe06 · outbound

This paper cites Progressive Neural Networks.

Continual Task Learning through Adaptive Policy Self-Composition Progressive Neural Networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.770151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.770151Z digest=sha256:d7ff20975678b4892a45e9f387c21b8f055bbf9002097ad05214d7b1118df589

Observation 31a53479-bd7f-4433-9596-05097faefc12 · outbound

This paper cites t-DGR: A Trajectory-Based Deep Generative Replay Method for Continual Learning in Decision Making.

Continual Task Learning through Adaptive Policy Self-Composition t-DGR: A Trajectory-Based Deep Generative Replay Method for Continual Learning in Decision Making

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.774690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.774690Z digest=sha256:7d0c0d5f360a4a28c018f87ee3c3e2adc4d58b02b95d34c3729d047a0e6664b8

Observation 1c49ffab-c00d-41c2-8a79-759697d104bc · outbound

This paper cites Balanced Destruction-Reconstruction Dynamics for Memory-replay Class Incremental Learning.

Continual Task Learning through Adaptive Policy Self-Composition Balanced Destruction-Reconstruction Dynamics for Memory-replay Class Incremental Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:43:02.840699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:43:02.779039Z digest=sha256:845282238785fa6b8d7e37a02b73c3291c2d487b2802aa9b7cb864ce5bbf396d

Observation e1e270dc-1b3a-413a-8d50-130b122187de · outbound

This paper cites an unresolved cited work.

Continual Task Learning through Adaptive Policy Self-Composition Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:43:03.052836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:43:02.783658Z digest=sha256:1c7a222ea52ea95d78d0c1fb71c93c2563cc65849cc676c55fbef6b8011c4d10

Observation 23dc4be9-2448-4cea-a043-4816697601cf · outbound

This paper cites an unresolved cited work.

Continual Task Learning through Adaptive Policy Self-Composition Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:43:03.033737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:43:02.788298Z digest=sha256:4f847c6992326a0c7fcd8e103adf8eee308cf9effa9ee8b2790f477319a8246e

Observation ef5243dd-cdf0-4791-87dc-a0d844c903e0 · outbound

This paper cites In contrast, regularization-based and rehearsal-based methods introduce additional loss terms to mitigate catastrophic forgetting.

Continual Task Learning through Adaptive Policy Self-Composition In contrast, regularization-based and rehearsal-based methods introduce additional loss terms to mitigate catastrophic forgetting

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:43:03.017312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:43:02.792410Z digest=sha256:3bf42f2b739d9e214da0a6d035d85bd3e60040a85559719cb829bbde4ab90277

Observation 8fa348ff-459a-421f-8677-502e70672e88 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Continual Task Learning through Adaptive Policy Self-Composition Offline Reinforcement Learning with Implicit Q-Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.748298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.748298Z digest=sha256:d7ef456d5ce92ab17b66dd170dc3800353c9a3f08353162d942cd7a6ed5a66b7

Observation f1596077-3fed-4702-b872-d3586f90ee8c · outbound

This paper cites Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?.

Continual Task Learning through Adaptive Policy Self-Composition Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.726068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.726068Z digest=sha256:cd869563a14322a054b14e60a75020fd84e7531eebdeffac84c7255f66174fc0

Observation 31607587-facf-4d87-aa1b-4756d7114e95 · outbound

This paper cites OER: Offline Experience Replay for Continual Offline Reinforcement Learning.

Continual Task Learning through Adaptive Policy Self-Composition OER: Offline Experience Replay for Continual Offline Reinforcement Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.734624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.734624Z digest=sha256:4559137aa43f728be7b29ff748bcd31c5b5acebc2e0c0d237307311c2665b9f8

Observation c432601c-beec-4caa-9361-be3d2d71009e · outbound

This paper cites Efficient Lifelong Learning with A-GEM.

Continual Task Learning through Adaptive Policy Self-Composition Efficient Lifelong Learning with A-GEM

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.720966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.720966Z digest=sha256:ee0b2c4f34186c7517806ff64b5cd64b64023f62d6d5b10b3119af439ff94e50

Observation 374a8992-ed8a-43da-80e1-9aa8c588381b · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Continual Task Learning through Adaptive Policy Self-Composition Off-policy deep reinforcement learning without exploration

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.730607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.730607Z digest=sha256:9a019c6e2dd95921fe69f43f017d53c87e70c551b623a15cc5b86a90ac7e4b8c

Observation 71f4c87a-8d58-44b1-ae0d-a2c67a171689 · outbound

This paper cites Variational Continual Learning.

Continual Task Learning through Adaptive Policy Self-Composition Variational Continual Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.761308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.761308Z digest=sha256:4381d75098ee68c4b85c0460a7e47ecf664817dc6215236c5d0a27e2cbd1913f

Observation 3666ea64-1760-4a4e-8f54-422824ee58d2 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Continual Task Learning through Adaptive Policy Self-Composition LoRA: Low-Rank Adaptation of Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T18:43:02.739411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:43:02.739411Z digest=sha256:065c1fe2c9e2ebea065ecc7351403dde49c41390fee5f31ef3b1600e1728d51b

Pith citing papers

No inbound Pith citation observations are available.