Pith. sign in

Paper Citation Record · LEDGER

Expert Behavior Prior Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2607.21302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21302 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:56:36.541257Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 341f92dd-3769-4d31-8a59-b7bf920ffaab · outbound

This paper cites Diverse imitation learning via self-organizing generative models,.

Expert Behavior Prior Reinforcement Learning Diverse imitation learning via self-organizing generative models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:29.715147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:29.715147Z digest=sha256:3f966a261ae6d6758734e79be1d0d1b31181420544a05d173ec2c2955d1f20a3

Observation b6a8681b-336e-43ca-80ee-7433cefb0314 · outbound

This paper cites Markov balance satisfaction improves performance in strictly batch offline imitation learning,.

Expert Behavior Prior Reinforcement Learning Markov balance satisfaction improves performance in strictly batch offline imitation learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:29.767460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:29.767460Z digest=sha256:416ceb1bf2a215ef18f70f939b3ff300e16e9a625d7e9b9cf2063b0b2d07598e

Observation dbcf8448-4dcf-4eda-8814-79ce693b9ed4 · outbound

This paper cites X-IL: Exploring the Design Space of Imitation Learning Policies.

Expert Behavior Prior Reinforcement Learning X-IL: Exploring the Design Space of Imitation Learning Policies

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:29.832819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:29.832819Z digest=sha256:be315685b56bd0fb35a48543ac5b53c45521ef39a1dfbaf1549386010e54cd02

Observation c0c9f0f7-1bf0-4375-86bb-5fa0567a7e39 · outbound

This paper cites Augmenting decision with hypothesis in reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Augmenting decision with hypothesis in reinforcement learning,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:29.872831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:29.872831Z digest=sha256:d3cbfedb2fe5d63aa40dd3561a33454b8c547d44218cabb71a2d3d98b4ba68f6

Observation 83e7b3f0-9c47-41bc-8384-a76252bf6064 · outbound

This paper cites Why so pessimistic? estimating uncertainties for offline rl through ensembles, and why their independence matters,.

Expert Behavior Prior Reinforcement Learning Why so pessimistic? estimating uncertainties for offline rl through ensembles, and why their independence matters,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:29.974702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:29.974702Z digest=sha256:23931e4f886b16dfecb92e254c2556e75f32557578830de1da42858c5cdc6cdd

Observation cee8cfde-a681-4a64-b018-10afd8a8aa72 · outbound

This paper cites Epistemic bellman operators,.

Expert Behavior Prior Reinforcement Learning Epistemic bellman operators,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.142701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.142701Z digest=sha256:7a931a7fe29009f7c20874cdc8dfa07d503efde9db2e4d80b39fef9f076f0abd

Observation eb72262a-016a-474f-a607-9083bebed730 · outbound

This paper cites Behavior priors for efficient reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Behavior priors for efficient reinforcement learning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.302455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.302455Z digest=sha256:66a5d758037bf64714a0e0172b5282c37ee06a00f3b85f4cf428c714502e0c7b

Observation 3b766335-3a59-4364-ac77-73b6d6037151 · outbound

This paper cites Pre-training goal-based models for sample-efficient reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Pre-training goal-based models for sample-efficient reinforcement learning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.470370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.470370Z digest=sha256:c03dd7fa164e870b7e3e96504cb05dff95634e3caf467ac7d9dabe6a96c9e638

Observation 2f28a8c9-f018-48b8-ab9b-19bfaa08277a · outbound

This paper cites BLEND: Behavior-guided Neural Population Dynamics Modeling via Privileged Knowledge Distillation.

Expert Behavior Prior Reinforcement Learning BLEND: Behavior-guided Neural Population Dynamics Modeling via Privileged Knowledge Distillation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.530259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.530259Z digest=sha256:f6631cecb3f4dc863999e512c963ea0eff78e988aaa2ad670925d1c2dfdec12d

Observation b915a079-3c54-4e29-9cef-2a701c37ffc7 · outbound

This paper cites Jump-start reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Jump-start reinforcement learning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.638326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.638326Z digest=sha256:9f57c9abffd3afb3d32c2827ce63b0fa7fc658c3806cfe41e1ddd363e7cc5ed8

Observation 6081f832-9100-423c-bd19-aca97dfae6e3 · outbound

This paper cites Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers,.

Expert Behavior Prior Reinforcement Learning Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.730465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.730465Z digest=sha256:f7c250458c00f0cc8f82434ec20f74abf16705bf27705abd132dfa572cca45fb

Observation 346620d3-e6f3-4c03-b532-fb9a0a79556e · outbound

This paper cites Learn to supervise: Deep reinforcement learning-based prototype refinement for few-shot motor fault diagnosis,.

Expert Behavior Prior Reinforcement Learning Learn to supervise: Deep reinforcement learning-based prototype refinement for few-shot motor fault diagnosis,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.799416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.799416Z digest=sha256:debd411925d934c30edaf3c6ba5881c1b13db1c88fdfd97800236481319566a5

Observation 94b59805-fd35-4a10-bf2b-3d3929510011 · outbound

This paper cites Efficient online reinforcement learning with offline data,.

Expert Behavior Prior Reinforcement Learning Efficient online reinforcement learning with offline data,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.858137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.858137Z digest=sha256:6c545d69204ed4e7339256cb3a5ab6dfdcff4a341f4848c1cc0da6dd52a1985a

Observation a6d2af2e-c156-48bf-8659-1371ea1dd623 · outbound

This paper cites Leveraging offline data in online reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Leveraging offline data in online reinforcement learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.925627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.925627Z digest=sha256:e51b4a6abab03960cdfddfdd709b14876f6f5abae4edcc2164e595ddd5256b34

Observation afb0555d-6cf0-413f-8b5c-7a1f828849d6 · outbound

This paper cites Enhancing Reinforcement Learning Agents with Local Guides.

Expert Behavior Prior Reinforcement Learning Enhancing Reinforcement Learning Agents with Local Guides

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:30.986206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:30.986206Z digest=sha256:3e893901ac059d93217e8f0cf110ce11c20d47842d26b9323be0ef8861c04252

Observation 8506d4f5-30cb-49ef-9af4-cfcda1e306c8 · outbound

This paper cites Leveraging demonstrations to improve online learning: Quality matters,.

Expert Behavior Prior Reinforcement Learning Leveraging demonstrations to improve online learning: Quality matters,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.059219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.059219Z digest=sha256:d0e037e09820ac46a980cf2d2ed13e16c25af63b4340215f3e5785f6bdcafee8

Observation 32933b84-abbe-4abc-a2d2-22c5d7e74678 · outbound

This paper cites Iterative regularized policy optimization with imperfect demonstrations,.

Expert Behavior Prior Reinforcement Learning Iterative regularized policy optimization with imperfect demonstrations,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.071799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.071799Z digest=sha256:e9c3f86cab9fd7bcd5b670ed96ea5a9c99d8ae2eb582b38c071d71d4e93b2182

Observation 91d4652f-f4bf-4688-9b00-7c6f2ee3783f · outbound

This paper cites Constraint- adaptive policy switching for offline safe reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Constraint- adaptive policy switching for offline safe reinforcement learning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.074838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.074838Z digest=sha256:febb36d30c4c05bdd371c586ccfda043e95eff23a148753da1d2e1d560872e67

Observation 3e5e2604-fdd2-4c8f-90cb-bba73b3c3757 · outbound

This paper cites Residual skill policies: Learning an adaptable skill-based action space for rein- forcement learning for robotics,.

Expert Behavior Prior Reinforcement Learning Residual skill policies: Learning an adaptable skill-based action space for rein- forcement learning for robotics,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.095430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.095430Z digest=sha256:180094ff3459ac5d6b603f2a6f01bb0fe5e1eb9d924ab5deba210a6d74d070b9

Observation a8653a43-5ca3-403c-9606-da4b949fe954 · outbound

This paper cites Leveraging Skills from Unlabeled Prior Data for Efficient Online Exploration.

Expert Behavior Prior Reinforcement Learning Leveraging Skills from Unlabeled Prior Data for Efficient Online Exploration

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.151219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.151219Z digest=sha256:45c86126873650992cfb0fda3c3b13fbd91ac20e3b38f6a56ab7f074f6d261ab

Observation 7e6a71d8-ecec-40a5-83f5-548cde8187e0 · outbound

This paper cites Policy regularization with dataset constraint for offline reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Policy regularization with dataset constraint for offline reinforcement learning,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.237029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.237029Z digest=sha256:d3a9f86e68661ace3333da26456ebd891be36663e40166957b9dfdfc231a0c62

Observation 81e955d9-8414-4ee5-a57c-db274c0f9a2d · outbound

This paper cites Accelerating exploration with unlabeled prior data,.

Expert Behavior Prior Reinforcement Learning Accelerating exploration with unlabeled prior data,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.388210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.388210Z digest=sha256:44a638f34da752e18e81687eeb3312188858de90a48adeb471bc3916dde956e3

Observation 5d2a733e-aa25-4c28-b5b3-f7949b75cba8 · outbound

This paper cites Cross-domain offline policy adaptation with optimal transport and dataset constraint,.

Expert Behavior Prior Reinforcement Learning Cross-domain offline policy adaptation with optimal transport and dataset constraint,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.578131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.578131Z digest=sha256:596b0344127c0651b6e9499378ff72ad2046e18438b7211e7c480d3b7ef6f0d8

Observation 6dc0356f-c43f-41cb-87fc-0962c26c76f9 · outbound

This paper cites Policy gradient for rectangular robust markov decision processes,.

Expert Behavior Prior Reinforcement Learning Policy gradient for rectangular robust markov decision processes,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.722979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.722979Z digest=sha256:e54e2ea30083c683a98f6c024751a47e3f580b4cbda712cb11bd89da1500efed

Observation 0ff1a6f6-c2e5-4349-98f8-215a34f9e7fb · outbound

This paper cites Reinforcement learning: An introduction,.

Expert Behavior Prior Reinforcement Learning Reinforcement learning: An introduction,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:31.884835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:31.884835Z digest=sha256:d06955cc6debb971949372689dc1c8309d8bc7a14dc5ee48bff8cd4b87f57ae5

Observation b62b83be-8784-46e8-a09a-9f9566b69423 · outbound

This paper cites Is q-learning provably efficient?.

Expert Behavior Prior Reinforcement Learning Is q-learning provably efficient?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:32.058259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:32.058259Z digest=sha256:9fd2eccf66f5baa6e97202e90f90c7f45836e51951aab4bea79e02d16907f187

Observation ea73e4e2-ee0a-4120-ac1e-bbf5940092a4 · outbound

This paper cites Actor-critic alignment for offline-to-online re- inforcement learning,.

Expert Behavior Prior Reinforcement Learning Actor-critic alignment for offline-to-online re- inforcement learning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:32.275765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:32.275765Z digest=sha256:2f0ae0e3bda08a5536d6a4fbcedbeb46b1a48e583402cc42d5f2bde93b225dd8

Observation 5feb362e-7a47-4d7a-991e-4671991e57bf · outbound

This paper cites Adaptive policy learning for offline-to-online reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Adaptive policy learning for offline-to-online reinforcement learning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:32.412266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:32.412266Z digest=sha256:d9d5503fd79389ff82778f80898f2903dd27798b609b87e553dcfbfa37418939

Observation 21f7a3da-bb6e-4f61-94fc-e2a21133bdbc · outbound

This paper cites Optimistic critic reconstruction and constrained fine-tuning for general offline-to-online rl,.

Expert Behavior Prior Reinforcement Learning Optimistic critic reconstruction and constrained fine-tuning for general offline-to-online rl,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:32.531640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:32.531640Z digest=sha256:c2e366bfa98d96402a7e7d21f67b20f22917cff9818ef491e2b10dadd970a2cf

Observation 23ab0009-a461-439e-b9bf-0698c090ced3 · outbound

This paper cites Tree-based batch mode rein- forcement learning,.

Expert Behavior Prior Reinforcement Learning Tree-based batch mode rein- forcement learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:32.675158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:32.675158Z digest=sha256:f5b533652c89849244090956284d513c9ef7473e3bb7fdc74a00f821e12c34fa

Observation 7fa71875-f4e4-4297-a91c-5c32be4a3538 · outbound

This paper cites Mildly conservative q-learning for offline reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Mildly conservative q-learning for offline reinforcement learning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:32.794597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:32.794597Z digest=sha256:b94bceb42247886979ec6380181a843d3dc1e44a9a2c1c815ab9dd024bb56894

Observation a1447621-db8d-4328-ba4e-95b364d4eae0 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Expert Behavior Prior Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:32.835740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:32.835740Z digest=sha256:82480c4ae021b742c959932f96ed8f91f60af2a577e4118bc156585a0c70f2a7

Observation f7f0e50a-4191-49fe-be71-a2cb55e7e547 · outbound

This paper cites De-pessimism offline reinforcement learning via value compensation,.

Expert Behavior Prior Reinforcement Learning De-pessimism offline reinforcement learning via value compensation,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:32.899188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:32.899188Z digest=sha256:2646cc5860b1589171098d081f21c51a7cc8d2bb48099898848c57f1ca1daeb9

Observation 09c13e59-946a-4bba-947d-eaffdbddc620 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,.

Expert Behavior Prior Reinforcement Learning Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:32.963742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:32.963742Z digest=sha256:ec4d65317e2b0f0c9e9c19468ec488b5a800b431aa85eb85413570cee57746a6

Observation 822d88c3-11b9-4f6d-954b-86655039d45c · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Expert Behavior Prior Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:33.064948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:33.064948Z digest=sha256:27f0b19c530753ed0915d92a7788aa1c6705c35ee31ac23f415b9fd96c170a85

Observation 3f20a930-d0d0-4395-8ffa-a72e6d36787b · outbound

This paper cites Don’t start from scratch: Leveraging prior data to automate robotic reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Don’t start from scratch: Leveraging prior data to automate robotic reinforcement learning,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:33.240341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:33.240341Z digest=sha256:557d46cb87276774abf547b02f90e9033b693a76f36fda652608a66d9e6bb5dc

Observation c5072f68-74b8-47aa-8502-26451274812d · outbound

This paper cites Behavior prior representation learning for offline reinforce- ment learning,.

Expert Behavior Prior Reinforcement Learning Behavior prior representation learning for offline reinforce- ment learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:33.362398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:33.362398Z digest=sha256:0ab729ce5ff567e51165c02bc43f13e90d14f6dc86960b40b1c5b9950a4365df

Observation 03b83487-1f98-4c71-9d77-7bbfcd9df768 · outbound

This paper cites One ACT Play: Single Demonstration Behavior Cloning with Action Chunking Transformers.

Expert Behavior Prior Reinforcement Learning One ACT Play: Single Demonstration Behavior Cloning with Action Chunking Transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:33.450588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:33.450588Z digest=sha256:90a49e80ec19e6a62d79d3c1d434ee31768569e59cdbd2618908a35155e824de

Observation 163ac0ef-2bef-4e25-9473-990c566512c1 · outbound

This paper cites Policy optimization with demonstrations,.

Expert Behavior Prior Reinforcement Learning Policy optimization with demonstrations,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:33.531132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:33.531132Z digest=sha256:93fbf66e9a8ff14d8d33c5b4169ee038e67a3a9a5715f45d97234939e13f2516

Observation dba72cb6-62df-4dd0-96f0-98af582b032c · outbound

This paper cites Reinforcement Learning with Sparse Rewards using Guidance from Offline Demonstration.

Expert Behavior Prior Reinforcement Learning Reinforcement Learning with Sparse Rewards using Guidance from Offline Demonstration

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:33.684912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:33.684912Z digest=sha256:4a7d1430bc5fa6e05c2926c340ad9a589d1b4159624663fcceb496b314f58011

Observation 5893c168-8d3b-444e-9354-fbbfb894a289 · outbound

This paper cites Goal-conditioned on-policy reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Goal-conditioned on-policy reinforcement learning,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:33.799415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:33.799415Z digest=sha256:a00fce2d102d9fdd79626b6efec8b9c168238f008dc599b8a1925b79cbc86f68

Observation 6fa04518-1d3f-411f-92cd-b92d001e8d69 · outbound

This paper cites Recurrent experience replay in distributed reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Recurrent experience replay in distributed reinforcement learning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:33.866798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:33.866798Z digest=sha256:cc8a571577c900759e82d91bd6c6a2f22e57e221333fa4aa30fd4ace1ff2d733

Observation cdd59388-1c54-4fdd-aff1-630c621322fb · outbound

This paper cites Learning Sparse Control Tasks from Pixels by Latent Nearest-Neighbor-Guided Explorations.

Expert Behavior Prior Reinforcement Learning Learning Sparse Control Tasks from Pixels by Latent Nearest-Neighbor-Guided Explorations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:33.967615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:33.967615Z digest=sha256:66825f449cc15dbbd916f117b30a01ab584b8314d4c3d16306a807a47ca86766

Observation 0a641db0-4e66-40b2-b67b-68049ba7f9b3 · outbound

This paper cites Theoretically principled deep rl acceleration via nearest neighbor function approximation,.

Expert Behavior Prior Reinforcement Learning Theoretically principled deep rl acceleration via nearest neighbor function approximation,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:34.218354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:34.218354Z digest=sha256:b51cf86231c0f32da30a038a07c212407d5527aadd8ebe7b63f41023d97a9352

Observation 62b46839-b198-4750-b1c5-8e89c68393cc · outbound

This paper cites A review of recurrent neural net- works: Lstm cells and network architectures,.

Expert Behavior Prior Reinforcement Learning A review of recurrent neural net- works: Lstm cells and network architectures,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:34.293904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:34.293904Z digest=sha256:5cb7dcc33c6b4984b17d5eca0457fb080cb56f4e6ff2ddd25606fa48d8702057

Observation 504f9b9c-e716-4a8d-a92f-ab32a0141259 · outbound

This paper cites Frustratingly easy regularization on representation can boost deep reinforcement learning,.

Expert Behavior Prior Reinforcement Learning Frustratingly easy regularization on representation can boost deep reinforcement learning,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:34.393332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:34.393332Z digest=sha256:dceb991ed965bbd1424777cc2fce70553020269024fd7d1f92c2008077bae33d

Observation 6f4f9eb3-2fd6-44e7-8dd6-80ec3b3fde8e · outbound

This paper cites Q-learning with nearest neighbors,.

Expert Behavior Prior Reinforcement Learning Q-learning with nearest neighbors,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:34.574737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:34.574737Z digest=sha256:63ea5c20b5fdf2db5d61e8dc0adcc50a192e7aed08207409c34abbb459610788

Observation 44512d60-f245-4f66-b622-b0dcfd475e60 · outbound

This paper cites Improving policy exploitation in online reinforcement learning with instant retrospect action,.

Expert Behavior Prior Reinforcement Learning Improving policy exploitation in online reinforcement learning with instant retrospect action,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:34.698097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:34.698097Z digest=sha256:66f478276b3d463adc9ef08bc521d1c4b92f0b16c5da1fe53b25412b95d69506

Observation aa4d02f3-c43d-474d-a816-8729a6e9bd6c · outbound

This paper cites Seizing serendipity: exploiting the value of past success in off-policy actor-critic,.

Expert Behavior Prior Reinforcement Learning Seizing serendipity: exploiting the value of past success in off-policy actor-critic,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:34.845854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:34.845854Z digest=sha256:bc71de3ef1d76fe02bf66c3723fac81ad58c2acc3beeceef86ac0da9e6c83e35

Observation e8476f44-e366-4a04-8c2b-2f01e5b160b9 · outbound

This paper cites Offline-boosted actor-critic: Adaptively blending optimal historical behaviors in deep off-policy rl,.

Expert Behavior Prior Reinforcement Learning Offline-boosted actor-critic: Adaptively blending optimal historical behaviors in deep off-policy rl,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:34.950027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:34.950027Z digest=sha256:b5e9c4c8f130bc71280efd8a62d5957cd6398d7639ccc3bced8d8737cd02ffcb

Observation 8ea73369-3cd0-4d66-9a5c-adc939e45cd0 · outbound

This paper cites Off-policy deep reinforcement learning without exploration,.

Expert Behavior Prior Reinforcement Learning Off-policy deep reinforcement learning without exploration,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.055000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.055000Z digest=sha256:0dfdcaa4a84f4c5259605c57ecab6c47a3389e97c25360edd7577e9b8a8e36c9

Observation b3fa0b96-1174-4cda-a03e-5c0168346b18 · outbound

This paper cites Auto-Encoding Variational Bayes.

Expert Behavior Prior Reinforcement Learning Auto-Encoding Variational Bayes

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.179046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.179046Z digest=sha256:269f5977655ca649627c52aab03c57612332ac3624310b749903c681ec64c950

Observation b910d697-74b6-4a3b-b22e-3b59658c6b1d · outbound

This paper cites Reparameterized policy learning for multimodal trajectory optimization,.

Expert Behavior Prior Reinforcement Learning Reparameterized policy learning for multimodal trajectory optimization,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.302638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.302638Z digest=sha256:8e2c591443e6fd2e2d7d8c4ba70fe88e00592b8899886c9acd8dd89de298eb00

Observation ce2529a7-64a0-4bbf-9ea1-37190d9bb340 · outbound

This paper cites DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation.

Expert Behavior Prior Reinforcement Learning DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.396183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.396183Z digest=sha256:c7bb5ce758edbcb7ba0e5f0499604c3929c95f68760b4146058bece28c955eba

Observation ee6eb6eb-f9d9-4fcd-8a93-026744b08b85 · outbound

This paper cites Addressing function approxima- tion error in actor-critic methods,.

Expert Behavior Prior Reinforcement Learning Addressing function approxima- tion error in actor-critic methods,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.482128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.482128Z digest=sha256:0f98fa11267f1acbe1a9883b6aa8204d3a1c1caca424dc5276694b2e8e2ce6ef

Observation a77a6591-4287-4227-93d6-34a572ae9783 · outbound

This paper cites an unresolved cited work.

Expert Behavior Prior Reinforcement Learning Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.589963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.589963Z digest=sha256:5a76428abc901b7685acc856fe9f3c8e84b0065ffc8960042a7faa7f90d0261a

Observation 80fa20c8-a962-4b8a-a3e5-a9eca32fc7bc · outbound

This paper cites Continuous control with deep reinforcement learning.

Expert Behavior Prior Reinforcement Learning Continuous control with deep reinforcement learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.699676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.699676Z digest=sha256:cdb677417937e2b9fbec08578a7cc70dbda018f2e26b6e8fccf1a4b43e0d5020

Observation 9301c955-d440-4711-82d7-a5e3f9ba95c9 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

Expert Behavior Prior Reinforcement Learning Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.807665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.807665Z digest=sha256:b162005aa14be2bac3add06ef059ebbe758475ecae18a748430b96ed3304eecd

Observation c158b7c2-32ec-4edf-9475-07c3e811f996 · outbound

This paper cites an unresolved cited work.

Expert Behavior Prior Reinforcement Learning Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.918222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.918222Z digest=sha256:e4bd6b83916f00d57eff69ef4f485809b43b5cf8741eb1e3859cfd5c5b9ddb1c

Observation c7e0ee36-8abc-45c2-a41f-3faedef10911 · outbound

This paper cites Reinforcement learning with stochastic reward machines,.

Expert Behavior Prior Reinforcement Learning Reinforcement learning with stochastic reward machines,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:36.031209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:36.031209Z digest=sha256:b279d7284c26d772ca567388c339e3b22eee8f6e9abe53bd1def90758cfc25d3

Observation ef020d41-844a-4952-b14b-0a9e874f033b · outbound

This paper cites Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors.

Expert Behavior Prior Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:36.138710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:36.138710Z digest=sha256:6503dd8b73e18481ef7261fa0267fe293aff72f47e0b0ceed802683f63bb923f

Observation a1fd98d0-4f57-49cb-b69a-1f3c4f1e8399 · outbound

This paper cites Softmax deep double deterministic policy gradients,.

Expert Behavior Prior Reinforcement Learning Softmax deep double deterministic policy gradients,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:36.248100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:36.248100Z digest=sha256:2fc243bf51f24bd3700998ebee895649b542d48639b7802f78460ba4c03790c9

Observation 2b60f18e-79ec-4e30-9aa6-98c606b6ff08 · outbound

This paper cites Logit standardization in knowledge distillation,.

Expert Behavior Prior Reinforcement Learning Logit standardization in knowledge distillation,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:36.289956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:36.289956Z digest=sha256:bbb4fe41b960d7bb2c0d4852125fb7e7ed09f8ff9cebc6a8dd72646b6a9e677f

Observation fa6be7b5-bd2b-4c1e-8ea0-157f77c92cb7 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Expert Behavior Prior Reinforcement Learning Distilling the Knowledge in a Neural Network

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:36.434176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:36.434176Z digest=sha256:29ae7f025cc5a8dfda1a1cc86dc95d012ebbaa690f420cbe143123bf24d541dd

Observation 88022178-65f0-4374-bf61-c4e6ad1621bd · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice,.

Expert Behavior Prior Reinforcement Learning Deep reinforcement learning at the edge of the statistical precipice,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:36.541257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:36.541257Z digest=sha256:cf2d0afa097e3856d220431008ca1f5c3b2a3c824689a5b1f9f49392356e742c

Pith citing papers

No inbound Pith citation observations are available.