Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:39.186629Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2506.14811.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:39.186629Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:20.198877Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:58:57.542814Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4be8b203-e0fb-4c71-8697-77f2f8a2ce5d · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 29d1ae43-8179-4f64-84fb-f5ad099f7a0e · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Therefore, the computational cost of the internal policy is constant and independent of the number of modules (i.e., number of tasks), Tint(n) = O(1).9 Total Complexity
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 21cc310c-e4f7-4b09-a6c0-1410f23e4681 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning 9For the sake of simplicity, this definition of the internal policy ignores possible activation and normalization layers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7d9e6bba-2c93-477c-9eaf-78918cfc303b · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 43f60adf-e129-40b9-b512-02b32f0f0990 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 92e3bb62-861f-408c-b619-21bf06002c47 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning An image is worth 16x16 words: Transformers for image recognition at scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f5af5f3d-0b1a-4818-a396-6ef6ff47680b · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Finally, note that the diagonal of the matrix has no especially positive transfer values
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 317d1efb-d88a-4b17-a8bc-4acfd852c627 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning DARLA: Improving zero-shot transfer in reinforcement learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 34695499-3440-48c5-af70-84f4bfedb10b · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1d659b91-6234-4ad0-a7b6-199a0bf6b0c5 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning and Lazebnik, S
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b1880b98-c986-479a-807b-d8a5e7f7e7b4 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning and Cohen, N
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 667d9567-5c7f-4e7d-99c0-c3f727a7ed7b · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning A., van Seijen, H., and EATON, E
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation abf8d190-968e-4e5e-bdc5-332a66797733 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning DINOv2: Learning Robust Visual Features without Supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b373c6e9-68e2-495b-9dca-23d67dc39192 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Routing net- works: Adaptive selection of non-linear functions for multi-task learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cfa1ce0d-f908-4a27-975a-5f9289193564 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Routing Networks and the Challenges of Modular and Compositional Computation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0746266a-c85f-4f8b-9142-dda724f93ded · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Policy Distillation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3f200e-103e-4462-a5b4-ac87ac9a71cb · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7818e6b-4ab2-4712-af14-45ad343e455b · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning V ., Montone, G., and O’Regan, J
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0d88f8cf-b4f8-4749-b2d1-37c7a231577c · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2bef20-85a1-4998-9847-c983e7a125c1 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Continual World: A robotic benchmark for continual reinforcement learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation db873ac4-15e5-46c8-b2ee-299fc137e60c · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Disentangling transfer in continual reinforce- ment learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0be4bb6a-e033-4289-8d30-9dfa8b22ba9c · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Supermasks in superposition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 36588c15-1180-4e78-8b3f-0c2f3a36d725 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Smoothquant: Accurate and efficient post-training quantization for large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 15547404-9f2a-4fe1-ad86-8cce9832ae89 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 869064b6-bead-4ec4-9e81-efe63dee22ae · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Gradient surgery for multi-task learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6c868104-f8c8-4cbd-b35c-f38e7c2a5702 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cdbc379b-2fa6-4fc8-8f39-c6290181e1db · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0345a13f-de38-45e8-b284-cea6ba45f785 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Trucks are longer vehicles, and thus, more difficult to avoid
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0b094072-7127-4a85-94a4-20fc1d88beb7 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Figure D.2c provides the FTr matrix of the last sequence, Freeway
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 08374901-c3a0-459e-865a-8deda8b0d0ee · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning All methods share the same common hyperparameters in every task sequence
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e28708de-c8ff-486f-8755-3ffb3c06e354 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work
Reference 512
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 03d52f89-da45-4ca9-aa3e-968d08590255 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Conflict- averse gradient descent for multi-task learning
Reference 1952
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 993a1326-8f2f-417a-84b5-566c30dfb297 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Building a subspace of policies for scalable continual learning
Reference 1999
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a94b8a4a-42ed-40ac-8340-94c1eac3f064 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning D., Juraf- sky, D., et al
Reference 2009
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a8b89b48-f42c-401b-8d0b-1156bdeaed33 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Curriculum learning
Reference 2013
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a011d90e-6011-427b-a4aa-4868b87dc841 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Progressive Neural Networks
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7744fc9b-bf68-445d-95e7-a0aef631af76 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning A., Veˇcer´ık, M., Roth¨orl, T., Heess, N., Pascanu, R., and Hadsell, R
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 208f7a10-c9d9-4c64-9bd9-f8a3d831883a · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforce- ment learning with a stochastic actor
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 626cc428-a2d3-44ef-8a06-2ebd839fbf4a · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Don't forget, there is more than forgetting: new metrics for Continual Learning
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12c29cdb-6be0-44a5-9ffa-3db05fe64d27 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning W., Heess, N., Osindero, S., and Pascanu, R
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eaa7c3c6-0804-454c-8758-75ec424bab0f · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Unresolved cited work
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8e1bfe15-7c74-46a1-bc8b-e363508d15d1 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning N., Kaiser,Ł., and Polosukhin, I
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a3fdf65c-1d57-4c0c-99c8-c73b6c679a50 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning Compacting, picking and growing for unforgetting continual learning
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 40921fb1-adaa-4519-bde1-bd44063d3215 · outbound
Self-Composing Policies for Scalable Continual Reinforcement Learning G., Menick, J., Munos, R., and Kavukcuoglu, K
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6ae13a21-2126-4b58-a649-085c98d14515 · inbound
Curriculum-Adapted Robust Reinforcement Learning for UAV Deconfliction in Adversarial Environments Self-Composing Policies for Scalable Continual Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd4eb217-c347-4179-bb53-fdc3cfd694b9 · inbound
When Robots Sleep: Offline Skill Consolidation for Shared-Policy Robot Learning Self-Composing Policies for Scalable Continual Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.