Pith. sign in

Paper Citation Record · LEDGER

Self-Composing Policies for Scalable Continual Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2506.14811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14811 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:39.186629Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:20.198877Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:58:57.542814Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4be8b203-e0fb-4c71-8697-77f2f8a2ce5d · outbound

This paper cites P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z.

Self-Composing Policies for Scalable Continual Reinforcement Learning P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.844779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.026047Z digest=sha256:965ef00a736ae5e846726cf9a5b3c351b4ed6b2023a27eccb2e162f7da954846

Observation 29d1ae43-8179-4f64-84fb-f5ad099f7a0e · outbound

This paper cites Therefore, the computational cost of the internal policy is constant and independent of the number of modules (i.e., number of tasks), Tint(n) = O(1).9 Total Complexity.

Self-Composing Policies for Scalable Continual Reinforcement Learning Therefore, the computational cost of the internal policy is constant and independent of the number of modules (i.e., number of tasks), Tint(n) = O(1).9 Total Complexity

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.530805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.153365Z digest=sha256:bfe4813f3001c5a7ff75c77c02f68950f540322b78dba81354510cc2716be6de

Observation 21cc310c-e4f7-4b09-a6c0-1410f23e4681 · outbound

This paper cites 9For the sake of simplicity, this definition of the internal policy ignores possible activation and normalization layers.

Self-Composing Policies for Scalable Continual Reinforcement Learning 9For the sake of simplicity, this definition of the internal policy ignores possible activation and normalization layers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.518424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.157501Z digest=sha256:95fb8d062c4b65ba6b5ff503956a5cfcc70f920e8eaaaf9574089cc3aca7df52

Observation 7d9e6bba-2c93-477c-9eaf-78918cfc303b · outbound

This paper cites an unresolved cited work.

Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:39.318128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.186629Z digest=sha256:b009b08e22c0c40a79691040274112003512f636e0f6d345b069838baddfe84d

Observation 43f60adf-e129-40b9-b512-02b32f0f0990 · outbound

This paper cites an unresolved cited work.

Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:39.478802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.168332Z digest=sha256:9d96cb2c45691db4d0d73dddadd57ed761ff7ae409dfbfdf3c2e85cafa717d83

Observation 92e3bb62-861f-408c-b619-21bf06002c47 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Self-Composing Policies for Scalable Continual Reinforcement Learning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.788091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.051687Z digest=sha256:b215ec4b7b25e9ac5951bd9a527f8b411097617f0d1eb73a4fe3ab720cfa4eb2

Observation f5af5f3d-0b1a-4818-a396-6ef6ff47680b · outbound

This paper cites Finally, note that the diagonal of the matrix has no especially positive transfer values.

Self-Composing Policies for Scalable Continual Reinforcement Learning Finally, note that the diagonal of the matrix has no especially positive transfer values

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.429837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.171801Z digest=sha256:40c5417180d0ee94482064f2ef4fd8e1ffad420cb296d9c4f5107f2a61f60703

Observation 317d1efb-d88a-4b17-a8bc-4acfd852c627 · outbound

This paper cites DARLA: Improving zero-shot transfer in reinforcement learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning DARLA: Improving zero-shot transfer in reinforcement learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.742282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.065896Z digest=sha256:0feb0106beb00b29fa8883f51e666fea9e11a37b6a8f994c813b713342090410

Observation 34695499-3440-48c5-af70-84f4bfedb10b · outbound

This paper cites A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al.

Self-Composing Policies for Scalable Continual Reinforcement Learning A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.718636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.073608Z digest=sha256:e1b660dba1673704b3838605c1e80709d08d15f080852fe5224a9062a161e1a0

Observation 1d659b91-6234-4ad0-a7b6-199a0bf6b0c5 · outbound

This paper cites and Lazebnik, S.

Self-Composing Policies for Scalable Continual Reinforcement Learning and Lazebnik, S

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.694872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.080755Z digest=sha256:f0d80e5459777695e45f4e1da6f5478bd8f85787c6ebc31bc255b0226b5cae8d

Observation b1880b98-c986-479a-807b-d8a5e7f7e7b4 · outbound

This paper cites and Cohen, N.

Self-Composing Policies for Scalable Continual Reinforcement Learning and Cohen, N

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.683418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.084153Z digest=sha256:26511efa941db10c33df2fb5ca4e479bf2f1baaf07036dd3ac2ba69a36a083be

Observation 667d9567-5c7f-4e7d-99c0-c3f727a7ed7b · outbound

This paper cites A., van Seijen, H., and EATON, E.

Self-Composing Policies for Scalable Continual Reinforcement Learning A., van Seijen, H., and EATON, E

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.671874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.087500Z digest=sha256:f8f7c401b327fcab8426e12b2c10f394a5c67307635e0d592ae09d715b990cc0

Observation abf8d190-968e-4e5e-bdc5-332a66797733 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Self-Composing Policies for Scalable Continual Reinforcement Learning DINOv2: Learning Robust Visual Features without Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:39.090901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:39.090901Z digest=sha256:1d3b1ca7abfacda750013244a5f624559c493d172725f05a6e0818377d15e723

Observation b373c6e9-68e2-495b-9dca-23d67dc39192 · outbound

This paper cites Routing net- works: Adaptive selection of non-linear functions for multi-task learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning Routing net- works: Adaptive selection of non-linear functions for multi-task learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.660530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.094846Z digest=sha256:75b76974bc73d4323a9b8b69f5985f53cdc22cbf5659e2a78cabc2b116534b51

Observation cfa1ce0d-f908-4a27-975a-5f9289193564 · outbound

This paper cites Routing Networks and the Challenges of Modular and Compositional Computation.

Self-Composing Policies for Scalable Continual Reinforcement Learning Routing Networks and the Challenges of Modular and Compositional Computation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:39.098469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:39.098469Z digest=sha256:c67acd7d11ecb3f12660b2e52c74b4d6377c9fd4b106e6fc8f8a05ff506cb148

Observation 0746266a-c85f-4f8b-9142-dda724f93ded · outbound

This paper cites Policy Distillation.

Self-Composing Policies for Scalable Continual Reinforcement Learning Policy Distillation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:39.102242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:39.102242Z digest=sha256:d4acaeef40eed9fe1c634cd4ba0523adc50a0add8aa4d920d2efc73d535ef0d5

Observation da3f200e-103e-4462-a5b4-ac87ac9a71cb · outbound

This paper cites Proximal Policy Optimization Algorithms.

Self-Composing Policies for Scalable Continual Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:39.113642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:39.113642Z digest=sha256:86da23439f01ffb850be1e390f5b55a9738dbb99bc96dfb158580f8c224f6c71

Observation c7818e6b-4ab2-4712-af14-45ad343e455b · outbound

This paper cites V ., Montone, G., and O’Regan, J.

Self-Composing Policies for Scalable Continual Reinforcement Learning V ., Montone, G., and O’Regan, J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.636813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.117234Z digest=sha256:523815ec291bf1916716a21b1627472a3c8aac99ff604b42e9e3b38000b4f090

Observation 0d88f8cf-b4f8-4749-b2d1-37c7a231577c · outbound

This paper cites Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions.

Self-Composing Policies for Scalable Continual Reinforcement Learning Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:39.124267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:39.124267Z digest=sha256:ae9976f414cdeeb7a05049f73cc261c3d621fb8b14b7c3c992e98a54fcf9d738

Observation 2b2bef20-85a1-4998-9847-c983e7a125c1 · outbound

This paper cites Continual World: A robotic benchmark for continual reinforcement learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning Continual World: A robotic benchmark for continual reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.613338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.128218Z digest=sha256:e3184a343a56634aa2be6355f00901837e8ecd600ffa6b15b8c806f2f22e8464

Observation db873ac4-15e5-46c8-b2ee-299fc137e60c · outbound

This paper cites Disentangling transfer in continual reinforce- ment learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning Disentangling transfer in continual reinforce- ment learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.600999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.131580Z digest=sha256:193a17597ff2836ecd2adef163c906ef5619bcaf93f08acb15486f05136ea6d3

Observation 0be4bb6a-e033-4289-8d30-9dfa8b22ba9c · outbound

This paper cites Supermasks in superposition.

Self-Composing Policies for Scalable Continual Reinforcement Learning Supermasks in superposition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.589259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.135082Z digest=sha256:e3aab9e6532f06bb9200aa1981c03b320b60e79f77b4337d67b2e938825da058

Observation 36588c15-1180-4e78-8b3f-0c2f3a36d725 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

Self-Composing Policies for Scalable Continual Reinforcement Learning Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.577244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.138636Z digest=sha256:364c2d337331913e00232b28dcc091ce809852fe5ad1f7258d86b9d0aa1d54a2

Observation 15547404-9f2a-4fe1-ad86-8cce9832ae89 · outbound

This paper cites an unresolved cited work.

Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:39.565574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.142435Z digest=sha256:5c10a2205f4a7c22abc9839261af375695c9edfae785af5b4b1a393f7d0b901b

Observation 869064b6-bead-4ec4-9e81-efe63dee22ae · outbound

This paper cites Gradient surgery for multi-task learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning Gradient surgery for multi-task learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.554336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.146103Z digest=sha256:374998dbc8222e523dbd912cee9dede2720eac772b8e9aac25a53e28a16a2db8

Observation 6c868104-f8c8-4cbd-b35c-f38e7c2a5702 · outbound

This paper cites an unresolved cited work.

Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:39.542556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.149785Z digest=sha256:6e3c9ffd47d9b894a333ac0b843918828e4a156748edfb192e91f9f0e0d85aa1

Observation cdbc379b-2fa6-4fc8-8f39-c6290181e1db · outbound

This paper cites an unresolved cited work.

Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:39.506391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.161145Z digest=sha256:c6273089d5bfcf8079aa3b9c138f7a30d39b7cfb224d2402b8b06d86c73fcdb1

Observation 0345a13f-de38-45e8-b284-cea6ba45f785 · outbound

This paper cites Trucks are longer vehicles, and thus, more difficult to avoid.

Self-Composing Policies for Scalable Continual Reinforcement Learning Trucks are longer vehicles, and thus, more difficult to avoid

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.494181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.164840Z digest=sha256:4e35177b9fbc3653ac53da2ed3283537de935e0338d6c952d1c443af27168496

Observation 0b094072-7127-4a85-94a4-20fc1d88beb7 · outbound

This paper cites Figure D.2c provides the FTr matrix of the last sequence, Freeway.

Self-Composing Policies for Scalable Continual Reinforcement Learning Figure D.2c provides the FTr matrix of the last sequence, Freeway

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.378668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.175216Z digest=sha256:f44ff925a7f146a681d069a9489e6b91709fca0141088baf92550441581beceb

Observation 08374901-c3a0-459e-865a-8deda8b0d0ee · outbound

This paper cites All methods share the same common hyperparameters in every task sequence.

Self-Composing Policies for Scalable Continual Reinforcement Learning All methods share the same common hyperparameters in every task sequence

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.347211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.178681Z digest=sha256:1830e4f9927edcb6a55198774f085024c97b5021f20e0fef8d25e059ad69fe22

Observation e28708de-c8ff-486f-8755-3ffb3c06e354 · outbound

This paper cites an unresolved cited work.

Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work

Reference 512

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:39.333468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.182401Z digest=sha256:eb19cec1fcb4bebcc0247e467492ad38e17d2a60bcfc52bc75f1c5bd35ac02e4

Observation 03d52f89-da45-4ca9-aa3e-968d08590255 · outbound

This paper cites Conflict- averse gradient descent for multi-task learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning Conflict- averse gradient descent for multi-task learning

Reference 1952

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.706409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.077162Z digest=sha256:e87a862c363b12723a0408d760790adbcc1655204c0db3651f0ca9831d6ee6e8

Observation 993a1326-8f2f-417a-84b5-566c30dfb297 · outbound

This paper cites Building a subspace of policies for scalable continual learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning Building a subspace of policies for scalable continual learning

Reference 1999

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.776482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.055459Z digest=sha256:47c5e29794c0c32f01529fd94c2c115b7f46c47b36e01e80a4f7063075e1f0ae

Observation a94b8a4a-42ed-40ac-8340-94c1eac3f064 · outbound

This paper cites D., Juraf- sky, D., et al.

Self-Composing Policies for Scalable Continual Reinforcement Learning D., Juraf- sky, D., et al

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.811117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.039394Z digest=sha256:f6e6248399dea08c00dd8a9ac48cbcc6063ca47e0dcc876cb7d158eb32e29f04

Observation a8b89b48-f42c-401b-8d0b-1156bdeaed33 · outbound

This paper cites Curriculum learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning Curriculum learning

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.822258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.034501Z digest=sha256:fc84654d8bd2b87084c5156e106df6b8f540c8f5f56ca42374ed167317849f34

Observation a011d90e-6011-427b-a4aa-4868b87dc841 · outbound

This paper cites Progressive Neural Networks.

Self-Composing Policies for Scalable Continual Reinforcement Learning Progressive Neural Networks

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:39.106161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:39.106161Z digest=sha256:32fd8502ed894f16dcdef91527b56e2763ef0bcf4a4e1b30926b7fbe8a45acf0

Observation 7744fc9b-bf68-445d-95e7-a0aef631af76 · outbound

This paper cites A., Veˇcer´ık, M., Roth¨orl, T., Heess, N., Pascanu, R., and Hadsell, R.

Self-Composing Policies for Scalable Continual Reinforcement Learning A., Veˇcer´ık, M., Roth¨orl, T., Heess, N., Pascanu, R., and Hadsell, R

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.648531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.110024Z digest=sha256:3ff451c6723ec80815e3c1e8af7c6c072cbcc7e1f722f85026a08f598ed8b042

Observation 208f7a10-c9d9-4c64-9bd9-f8a3d831883a · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforce- ment learning with a stochastic actor.

Self-Composing Policies for Scalable Continual Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforce- ment learning with a stochastic actor

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.753568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.062617Z digest=sha256:fde2031a19911afa43cac512b762c7e481a575a2a778beba0e94fde7aa256ea7

Observation 626cc428-a2d3-44ef-8a06-2ebd839fbf4a · outbound

This paper cites Don't forget, there is more than forgetting: new metrics for Continual Learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning Don't forget, there is more than forgetting: new metrics for Continual Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:39.047610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:39.047610Z digest=sha256:bb1a16f81cc04509777fd84abd89d360cb36b879877c617c77f200d60bbbf2f2

Observation 12c29cdb-6be0-44a5-9ffa-3db05fe64d27 · outbound

This paper cites W., Heess, N., Osindero, S., and Pascanu, R.

Self-Composing Policies for Scalable Continual Reinforcement Learning W., Heess, N., Osindero, S., and Pascanu, R

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.799539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.043751Z digest=sha256:d74f3f7a25c5600e04d091cb7cc56a26a124f1ff4121cf693fc1e52e6e42b26c

Observation eaa7c3c6-0804-454c-8758-75ec424bab0f · outbound

This paper cites an unresolved cited work.

Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:39.833522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.030531Z digest=sha256:9208f567e49769625acab5962c5d90af40238d3beba8e99172ce2ae2ab6a062e

Observation 8e1bfe15-7c74-46a1-bc8b-e363508d15d1 · outbound

This paper cites N., Kaiser,Ł., and Polosukhin, I.

Self-Composing Policies for Scalable Continual Reinforcement Learning N., Kaiser,Ł., and Polosukhin, I

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.625070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.120540Z digest=sha256:0dcce163c255e07d1dd2ffcfceee950706ba95d12fad449062883d153a2afb11

Observation a3fdf65c-1d57-4c0c-99c8-c73b6c679a50 · outbound

This paper cites Compacting, picking and growing for unforgetting continual learning.

Self-Composing Policies for Scalable Continual Reinforcement Learning Compacting, picking and growing for unforgetting continual learning

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.730466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.069719Z digest=sha256:c84e5992a7a93f1a7fe83bb1ce13bd5e6c27292c506efe765e9cdb9dc8aa3c30

Observation 40921fb1-adaa-4519-bde1-bd44063d3215 · outbound

This paper cites G., Menick, J., Munos, R., and Kavukcuoglu, K.

Self-Composing Policies for Scalable Continual Reinforcement Learning G., Menick, J., Munos, R., and Kavukcuoglu, K

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:39.764931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:55:39.059141Z digest=sha256:a0109ee9749b16d601493ae760fbae933b110ee57edd25e5f88dcb8d3d164349

Pith citing papers

Observation 6ae13a21-2126-4b58-a649-085c98d14515 · inbound

Curriculum-Adapted Robust Reinforcement Learning for UAV Deconfliction in Adversarial Environments cites this paper.

Curriculum-Adapted Robust Reinforcement Learning for UAV Deconfliction in Adversarial Environments Self-Composing Policies for Scalable Continual Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:20.198877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:20.198877Z digest=sha256:1b8d77778ce0f4a41f467bc56a40c5bc18daaadb2e07be6f802f6221f7516191

Observation cd4eb217-c347-4179-bb53-fdc3cfd694b9 · inbound

When Robots Sleep: Offline Skill Consolidation for Shared-Policy Robot Learning cites this paper.

When Robots Sleep: Offline Skill Consolidation for Shared-Policy Robot Learning Self-Composing Policies for Scalable Continual Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.544492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T01:04:01.317563Z digest=sha256:82c93c62888ae8b21e79899d05a175e2e201e4a0c8511c0365be82b31e2f7502