Pith. sign in

Paper Citation Record · LEDGER

Adaptive Policy Backbone via Shared Network

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2509.22310.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.22310 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:59.794278Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a9ca6ba-bd01-49f2-80b8-778104f2140b · outbound

This paper cites degenerate.

Adaptive Policy Backbone via Shared Network degenerate

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:50:00.153254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:59.794278Z digest=sha256:1e726b251339ce4999345a85c60c236ba17098d10f0f6cae40a2655a67b3fff3

Observation 4e85b819-979a-406e-9ec7-9ea4bb07a619 · outbound

This paper cites On First-Order Meta-Learning Algorithms.

Adaptive Policy Backbone via Shared Network On First-Order Meta-Learning Algorithms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.695466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.695466Z digest=sha256:9b189a022b59a0dfc1bdb2a971162b08df2f9cdf379dcc57ca8da0060371a2c4

Observation 1deb65c4-9b8e-40c6-8485-f3d51ecd5df9 · outbound

This paper cites Solving Rubik's Cube with a Robot Hand.

Adaptive Policy Backbone via Shared Network Solving Rubik's Cube with a Robot Hand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.711210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.711210Z digest=sha256:05febffc2acb6e4bddc442245e0385d8e65074189e0db3e706f49056bed89c5a

Observation b28df3d7-f8de-4139-9a13-3594f750127c · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Adaptive Policy Backbone via Shared Network Behavior Regularized Offline Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.718704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.718704Z digest=sha256:93462cf21f8993e7cabe03ba11e69b66ea2a647bdfd2edf1e865747655fe9a0f

Observation 1ceecf3a-9e06-4f0c-85de-24865542deba · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Adaptive Policy Backbone via Shared Network Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.722543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.722543Z digest=sha256:b709a8076db5a03aa8a99d5e37e34ef71b7d176bf460bd1f8ad3cf87b2326c7d

Observation 3a83e20b-e3e1-4619-b4f5-4167e284c1ce · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Adaptive Policy Backbone via Shared Network Offline Reinforcement Learning with Implicit Q-Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.726441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.726441Z digest=sha256:121916a54817b60d4c9d542eae5bbccab05d437f389f0f895720f9a47a9389f1

Observation 05c43a16-7500-433b-adfa-5bdbf776c883 · outbound

This paper cites Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML.

Adaptive Policy Backbone via Shared Network Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.730352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.730352Z digest=sha256:a78a2f27934e338f6701ebd7e1cf64ec3a20c2bed3ac5b0cb09ea435e4e8c156

Observation 414ac2dd-1cdd-4af4-b619-c4e96c4c8cc5 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Adaptive Policy Backbone via Shared Network LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.742716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.742716Z digest=sha256:8bb10bca32f59553d48fa3b18b1b373fb9636fa22b5e2fcba019dc50e075a145

Observation 292373e9-f99b-4da4-810e-9247c6fa5f60 · outbound

This paper cites BitFit: Simple parameter-efficient fine- tuning for transformer-based masked language-models.arXiv preprint arXiv:2106.10199,.

Adaptive Policy Backbone via Shared Network BitFit: Simple parameter-efficient fine- tuning for transformer-based masked language-models.arXiv preprint arXiv:2106.10199,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.746322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.746322Z digest=sha256:287455e4ecc9e2d14aefb9326144a3ea670505d442b889bdfd43dc74a6207011

Observation 2ac2e632-91cf-4f6d-87b4-7fadc3e7afd5 · outbound

This paper cites Lossless Adaptation of Pretrained Vision Models For Robotic Manipulation.

Adaptive Policy Backbone via Shared Network Lossless Adaptation of Pretrained Vision Models For Robotic Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.749675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.749675Z digest=sha256:261f2931c091fbd28e2f433e9188b6b34ceda6fb12e0132475139e324db2c76d

Observation 3086fa2c-6aa3-46e0-b15f-0a51c63f3e23 · outbound

This paper cites Universal Successor Features Approximators.

Adaptive Policy Backbone via Shared Network Universal Successor Features Approximators

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.753524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.753524Z digest=sha256:6042341ae6b0668e92f335dc1199c107d54ad6dabe37b5ce15945817719fa940

Observation 741d5ac3-a158-4028-83ef-8c6e7ea260c1 · outbound

This paper cites Advantages and Limitations of using Successor Features for Transfer in Reinforcement Learning.

Adaptive Policy Backbone via Shared Network Advantages and Limitations of using Successor Features for Transfer in Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.757366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.757366Z digest=sha256:ce6725a246a9d67705b4a9bbcb1c2eb99a94707874124f213ddf95100ada9bb6

Observation 9477da09-850c-4b91-bf38-89168baef5ea · outbound

This paper cites Some Considerations on Learning to Explore via Meta-Reinforcement Learning.

Adaptive Policy Backbone via Shared Network Some Considerations on Learning to Explore via Meta-Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.761012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.761012Z digest=sha256:d3d10dabc97c0bca6e0f41d712be6a35f08f13f81d457c0bbdfe198b244d85dc

Observation 26847317-8382-4273-971e-fb3937ac9729 · outbound

This paper cites ES-MAML: Simple Hessian-Free Meta Learning.

Adaptive Policy Backbone via Shared Network ES-MAML: Simple Hessian-Free Meta Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.764887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.764887Z digest=sha256:9132ab8ca1d9d176ad8fa8f4f5e7678be6ae0db9e8ca69a2c5240f4ebaad9001

Observation c9e7e083-7a17-48c6-a939-29950eedd8a7 · outbound

This paper cites VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning.

Adaptive Policy Backbone via Shared Network VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.768452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.768452Z digest=sha256:dbd56355f1e072119d1d965edb13beaf9d7cb2caabaf1ab577b55fbb523080c0

Observation 828391e7-c6ee-418c-8edf-340adefe9786 · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

Adaptive Policy Backbone via Shared Network A Tutorial on Meta-Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.772066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.772066Z digest=sha256:8435331c1a19980c99d0926b79ed798c6d3cfde85ff2a4a13d40fe81cff16e9f

Observation b23d4ff5-a654-4fd2-92da-040e0ae7a8fe · outbound

This paper cites Learning to Learn: Meta-Critic Networks for Sample Efficient Learning.

Adaptive Policy Backbone via Shared Network Learning to Learn: Meta-Critic Networks for Sample Efficient Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.775920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.775920Z digest=sha256:e0202f78ba44eb22e95502b8079cfd0225bad278431f46e4aa375f256398669d

Observation 95864adc-3203-4c64-8dba-bbcdb68dde2d · outbound

This paper cites Parameter Space Noise for Exploration.

Adaptive Policy Backbone via Shared Network Parameter Space Noise for Exploration

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.783118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.783118Z digest=sha256:bdf0c524036836879bb3a38f32fb481019e83edaedc2e8e1a8a820586577383b

Observation da34e2a3-9d08-40aa-b1df-f3980800a6b0 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Adaptive Policy Backbone via Shared Network Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.786714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.786714Z digest=sha256:eaf6fb1af9da30fd67320c1b4e4a8f7c1e57cb2cd752546d4793ae09a3421327

Observation ac7fa0b2-8f80-474d-90df-718b9516b643 · outbound

This paper cites In meta-test, the agent is evaluated atθ= 1.5π, that is,x= 3 cosθ,y= 3 sinθ.

Adaptive Policy Backbone via Shared Network In meta-test, the agent is evaluated atθ= 1.5π, that is,x= 3 cosθ,y= 3 sinθ

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:50:00.166748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:59.790252Z digest=sha256:324aa8d3c61fd68798c9949b960a72a825c93b6903106514cdf5a6006c7a5094

Observation 973ae3f2-e8bb-4640-aa38-2310ccc672c9 · outbound

This paper cites Revisit Policy Optimization in Matrix Form.

Adaptive Policy Backbone via Shared Network Revisit Policy Optimization in Matrix Form

Reference 2007

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:50:00.138816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:59.686659Z digest=sha256:0aa627fda1822eea05574740006086925699af5056806b2c69135c9e9b7a0759

Observation 1aef2f06-c341-481c-9fe6-06b30aa21ddc · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

Adaptive Policy Backbone via Shared Network RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.703482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.703482Z digest=sha256:9cf24b1167f439d323975734d661207438c6d7eaf5d701d5b9783a074ac5774d

Observation 136b4f67-5e1c-4156-9b51-c756204b629b · outbound

This paper cites A Simple Neural Attentive Meta-Learner.

Adaptive Policy Backbone via Shared Network A Simple Neural Attentive Meta-Learner

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.691118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.691118Z digest=sha256:4de710f2fdebffd00c479c1cb98b7e290fbe82bac4e2d01b9acfe2c7a2c47150

Observation 74f6ec7a-7cbf-47d2-90d4-83ad4c2c0aba · outbound

This paper cites Learning to reinforcement learn.

Adaptive Policy Backbone via Shared Network Learning to reinforcement learn

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.699119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.699119Z digest=sha256:3b147d2816332dc678fcbec5fe3b3c2278eaea9cad226cd9323b2d159023dd48

Observation 21437c3d-b1cd-4bcc-8ecb-08a9d841156a · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Adaptive Policy Backbone via Shared Network Dota 2 with Large Scale Deep Reinforcement Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.707291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.707291Z digest=sha256:14be88c7eda82ac329c266049d00f68d7607b111197f67ec24edc35db1f5ddfc

Observation 01db3086-9d60-49cb-b969-5ed2172b8e29 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Adaptive Policy Backbone via Shared Network Off-policy deep reinforcement learning without exploration

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:50:00.180795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:59.714968Z digest=sha256:e7af69f14ab84e44c999a76a37422c117ca9c6d0335c2ce4c3856953d6b1490f

Observation a0314d63-cc19-4a94-818a-629cd1cd223c · outbound

This paper cites Surgical Fine-Tuning Improves Adaptation to Distribution Shifts.

Adaptive Policy Backbone via Shared Network Surgical Fine-Tuning Improves Adaptation to Distribution Shifts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.738570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.738570Z digest=sha256:f5a372bcb4dc09b91fd208fd3a88b3c9ca5aba81a0f4d529efa46e11cc65fdd8

Observation a21779bf-7180-49c6-9252-5ea53c32fa28 · outbound

This paper cites Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution.

Adaptive Policy Backbone via Shared Network Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.734533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.734533Z digest=sha256:48127bc0b7994454ccf0c7094a829de9aec2f6016f8f6b416668d4354ad65543

Observation 6c9b0a5a-8041-40a8-a9c6-d9eb14e3526f · outbound

This paper cites Sharing Knowledge in Multi-Task Deep Reinforcement Learning.

Adaptive Policy Backbone via Shared Network Sharing Knowledge in Multi-Task Deep Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:59.779619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:59.779619Z digest=sha256:9255ee81a01c5097cce4624326c59cbe6af263a4ce778fcb6d721a20bfe7708f

Pith citing papers

No inbound Pith citation observations are available.