Pith. sign in

Paper Citation Record · LEDGER

How Should We Meta-Learn Reinforcement Learning Algorithms?

As of 8 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 1 inbound Pith citation observation for arXiv:2507.17668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17668 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:48:52.758471Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:39:32.117730Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact4
  • verified fuzzy26
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a3d7ccb0-92bc-4646-b308-e5a76539886e · outbound

This paper cites Loss of Plasticity in Continual Deep Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Loss of Plasticity in Continual Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.616239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.616239Z digest=sha256:3e5f9d3425885e6301186969bab8816753ec7e4999c6128be9f456f26b9f7838

Observation 24da469e-b54b-4146-bb26-ec474352ed58 · outbound

This paper cites Towards Characterizing Divergence in Deep Q-Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Towards Characterizing Divergence in Deep Q-Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.681243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.681243Z digest=sha256:9061fc253369e38b7297a7533cd1f62879796fd1643e7a177cb5b85d737e7ef5

Observation f48faf5f-8f10-4095-b923-e3fe5bf3f8f4 · outbound

This paper cites A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.755232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.755232Z digest=sha256:73cdb38c8888e01623433a4eb4091ace111dbcb704170259ef272625e1c31600

Observation b19f019a-93e9-4b0d-b221-c531836e6370 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

How Should We Meta-Learn Reinforcement Learning Algorithms? Deep reinforcement learning at the edge of the statistical precipice

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.809562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.809562Z digest=sha256:25b3f6bbf33bd1c72aca40720abb99b7269183d40853b93bb00d670d77e2a3f9

Observation 44b1374f-aa2b-4bff-a13f-84907f107a90 · outbound

This paper cites A Generalizable Approach to Learning Optimizers.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Generalizable Approach to Learning Optimizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.948249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.948249Z digest=sha256:86295d81db386f655f08a71ffd24736b793ce182e0045f91edaa2df656846945

Observation 34b95b01-ac5e-4eaf-afc5-9dd2e849b3f3 · outbound

This paper cites Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas.

How Should We Meta-Learn Reinforcement Learning Algorithms? Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.089106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.089106Z digest=sha256:bf7f78ce4677a2619ee8a640e53fa5ace44dee4bec2e7d6f50c51b7080769fc4

Observation 2fafa30b-8079-487a-bd81-65d3d1db55d3 · outbound

This paper cites An information-theoretic perspective on intrinsic motivation in reinforcement learning: A survey.

How Should We Meta-Learn Reinforcement Learning Algorithms? An information-theoretic perspective on intrinsic motivation in reinforcement learning: A survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.201906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.201906Z digest=sha256:824de9a1cc0bf36d0b3bfa4bfdfa8447d4094fda38324892ecd554a5d565a2b9

Observation 59369bea-951d-4879-922e-ac071fdf0ea3 · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Tutorial on Meta-Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.339004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.339004Z digest=sha256:c62edc5dd53d1cacc528ea0d12af7043fd51df44883246d50c4457f138cde9e2

Observation 8fd95d5d-5e2d-46b6-84db-12a0aa2f778f · outbound

This paper cites OpenAI gym, 2016.

How Should We Meta-Learn Reinforcement Learning Algorithms? OpenAI gym, 2016

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.485426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.485426Z digest=sha256:24f5c737d6307502f569fe9219cb4f9cac48ad1d0b3c6cf7da91932e6d4e6f43

Observation 379f9e2e-23a4-4b86-8f11-d59bb3c045d6 · outbound

This paper cites Exploration by Random Network Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Exploration by Random Network Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.624540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.624540Z digest=sha256:91389a982ca66c01cb83a1e2b61fa29bebea4994d047e03796360ff6770d919e

Observation 5dc19526-a0f3-4c51-bbe0-c15aaf14490d · outbound

This paper cites Boltzmann exploration done right.

How Should We Meta-Learn Reinforcement Learning Algorithms? Boltzmann exploration done right

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.793806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.793806Z digest=sha256:bb992b9fb2fd62cb13c36c8b657187e0de2f6fac325779c671dae581ae8bfe38

Observation 2049f298-757a-4752-a563-7d78d809c499 · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Symbolic Discovery of Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.939586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.939586Z digest=sha256:dc32c82f5d8ea1714d7ea9dc5549593ecf22ac90d0e9716bdc84914bcbbbef4d

Observation 1e634b23-c977-4fb1-adee-0abef0b42e18 · outbound

This paper cites Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl.

How Should We Meta-Learn Reinforcement Learning Algorithms? Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.030210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.030210Z digest=sha256:0b9ff246d7c97eac69363ca4bb6e40bfe6937f1cebffe89747910a63c9ede918

Observation 1aea3f56-9912-422e-8722-f2befb70c15f · outbound

This paper cites Discovering Symbolic Models from Deep Learning with Inductive Biases.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering Symbolic Models from Deep Learning with Inductive Biases

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.105628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.105628Z digest=sha256:bb75643c18f7d88ebf44c282109fbc19e921ef842a39f68cd01eadc66063225e

Observation 707eb909-41c4-4c9a-bf29-55fb711e05c8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.158711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.158711Z digest=sha256:759adfae2a4a740f9e4781cafb435dc94f9362959950a106aa7f2917f055bba2

Observation 96506966-4351-4c0c-80b6-4511d5ffe9e5 · outbound

This paper cites Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.239937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.239937Z digest=sha256:b1dea7b4b1c3206cfc92d08529b10a4e70ec615492648dd8fb799cecd65cd2ff

Observation 82b12f95-c64e-4c2d-849b-84e96f4da4d9 · outbound

This paper cites Loss of plasticity in deep continual learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Loss of plasticity in deep continual learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.309720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.309720Z digest=sha256:d4de5a0b9249ab4e38a830f4b7a151c2e372d818cba690c2df8241dbcd440743

Observation bad298e5-b312-469a-9370-21d35555a317 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.395863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.395863Z digest=sha256:c983f80d2edd0f22b81b5235afef853be384ee1482b5d038de57f97f04a30d34

Observation f70732c3-e78c-447d-b9f7-793c88ad0058 · outbound

This paper cites Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps.

How Should We Meta-Learn Reinforcement Learning Algorithms? Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:48:54.113207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:46.513473Z digest=sha256:4dd93beeae446c6e03ef78b6e750e0b9faff014acbff2132d20cedde4a24793a

Observation e97cd32c-805c-4709-bced-cce321896e67 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

How Should We Meta-Learn Reinforcement Learning Algorithms? OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.585185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.585185Z digest=sha256:286416b45037dd4811fd7de2f266fb8baf6947d523abb2d38e189377ae07175d

Observation 6454b23b-eafe-4b95-bb70-59f7f5a77502 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Model-agnostic meta-learning for fast adaptation of deep networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.675019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.675019Z digest=sha256:9e2d03c8da0c65e3138e6a24142060eec82c501197ad48ba58559f7ffc66d0f4

Observation 23895565-0168-4277-9fd2-a917f64ddf2e · outbound

This paper cites Noisy Networks for Exploration.

How Should We Meta-Learn Reinforcement Learning Algorithms? Noisy Networks for Exploration

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.738150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.738150Z digest=sha256:27f741486226bb9057b3079b90d4e30d8ac9f3c4e6d103c5c80095360304e2c0

Observation faaafd20-4489-409b-b34e-4e5bd93cbcb1 · outbound

This paper cites Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem.

How Should We Meta-Learn Reinforcement Learning Algorithms? Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.822940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.822940Z digest=sha256:c00eafc18a1800310e607786e40d40b2b614a6dfa2febf217a76337e03e6ca8d

Observation 0648e23b-f653-4be6-b3c1-e4613c3559fc · outbound

This paper cites Born Again Neural Networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Born Again Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.904574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.904574Z digest=sha256:b996ab3d9afb965d9261c6483d75f5acad550a76d891cc1c441276eb61f9dd21

Observation c4e7164c-4f0c-4ab6-a9c1-9520d6a6f49a · outbound

This paper cites Goldie, Chris Lu, Matthew T.

How Should We Meta-Learn Reinforcement Learning Algorithms? Goldie, Chris Lu, Matthew T

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:59.178524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.000738Z digest=sha256:90a3d409db8cd879c0ab037bdfee6b614c064e706a548a8521b1b48fcb3e8f54

Observation 0d14a9d1-f1ad-4f76-9b59-997042574c7c · outbound

This paper cites Benchmarking the Spectrum of Agent Capabilities.

How Should We Meta-Learn Reinforcement Learning Algorithms? Benchmarking the Spectrum of Agent Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.078571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.078571Z digest=sha256:18fcb00996ba1fe182a4d3cfcf87fe120b8c681ba041eb999835dd493ad675ce

Observation 3d14ae1c-d8f4-47b8-a380-a881f6b7f07d · outbound

This paper cites Distilling the Knowledge in a Neural Network.

How Should We Meta-Learn Reinforcement Learning Algorithms? Distilling the Knowledge in a Neural Network

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.149052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.149052Z digest=sha256:9131add5f3ef80964ea173ceafa2ad7c166cec7f6620fc5cc7a4bfe4c2473f53

Observation 4080903f-56ca-481e-9526-3d647585a01f · outbound

This paper cites Long short-term memory.

How Should We Meta-Learn Reinforcement Learning Algorithms? Long short-term memory

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.222901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.222901Z digest=sha256:35a8bc555b1fe238738ff4a867c2f7db0dbbc82955c4912f3bab298c5a3358ba

Observation dc9c16ba-60c7-4c79-90e9-7a662088ee40 · outbound

This paper cites Automated Design of Agentic Systems.

How Should We Meta-Learn Reinforcement Learning Algorithms? Automated Design of Agentic Systems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.306436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.306436Z digest=sha256:50a8f76a943c2e880704d7aed6b186a0caa045115e0900f119f05a35adc385b7

Observation 41b311f6-d302-40b8-a5ea-f13008d342b1 · outbound

This paper cites Transient non-stationarity and generalisation in deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Transient non-stationarity and generalisation in deep reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.888824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.383576Z digest=sha256:236d5dd770786d3162295200192f90294f33d6ee037c40d104879dae48b1b70f

Observation 0edfac8e-8273-47a5-aa23-a12663df7d34 · outbound

This paper cites Transient non-stationarity and generalisation in deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Transient non-stationarity and generalisation in deep reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.610653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.471532Z digest=sha256:4bf49900a5aaf9a8a527296fe8f3c00f87c196f15c4c781ed3c0e633117babe5

Observation 7d42ccf3-7005-4fd5-8841-a55667aa143c · outbound

This paper cites Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.861384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.539950Z digest=sha256:360c9d10f1551f0ce0410e0db563d4a2fef04daa33c3aaac51bf650214902fed

Observation 820da51f-8024-4178-abbc-91a7e9405067 · outbound

This paper cites Discovering temporally-aware reinforcement learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering temporally-aware reinforcement learning algorithms

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.350351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.617075Z digest=sha256:8f42cbfd91e56caa2aea3ab997799404a5c66dde137541de77cad2fdb53f495c

Observation 6bf5b200-36ad-4e13-9410-c95c9baf1e10 · outbound

This paper cites Improving policy optimization with generalist-specialist learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving policy optimization with generalist-specialist learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.129243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.682731Z digest=sha256:469face48f4cadfbfb596359ecd89795c4597df16ac6b3e6500d2089a9aa64c6

Observation 3a1a5cb6-507c-49c1-adf6-d89d5bf2585e · outbound

This paper cites Meta Learning Backpropagation And Improving It.

How Should We Meta-Learn Reinforcement Learning Algorithms? Meta Learning Backpropagation And Improving It

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.690763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.763840Z digest=sha256:4bf41ae21aa70079c38dfb5ebc52a78e4ccc0e7091e2b8a40eac5224e71fe4ad

Observation 267ff5a4-13c3-4ac9-98bb-2b92691851ef · outbound

This paper cites Improving generalization in meta reinforcement learning using learned objectives.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving generalization in meta reinforcement learning using learned objectives

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.828958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.817050Z digest=sha256:63bb23aba05da58190cac76af36c89d1360f0abecd272df7df640dba33a8aca2

Observation 2517e4d0-afda-4d88-927b-a6d3a685f536 · outbound

This paper cites Mirror Learning: A Unifying Framework of Policy Optimisation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Mirror Learning: A Unifying Framework of Policy Optimisation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.878994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.878994Z digest=sha256:b29da412b2a4547ac987f760fbf681b7fbe377130b91b2a221af263a7921ef43

Observation 853ffeb4-20c3-4278-8903-a0cc889d3f0f · outbound

This paper cites Learning to Optimize for Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Learning to Optimize for Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.922233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.922233Z digest=sha256:c74186acd7d5d707d74cce75d59edb9e83ad0bf7c2f8c0c325b2d1bd9a585077

Observation 670b210f-af41-461e-9f14-69132d7b1c52 · outbound

This paper cites gymnax: A JAX -based reinforcement learning environment library, 2022 a.

How Should We Meta-Learn Reinforcement Learning Algorithms? gymnax: A JAX -based reinforcement learning environment library, 2022 a

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.659181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.972077Z digest=sha256:85bed9dc28ef575ca72f6297a5aa89d84dd18b9f8f02c451058b62ebf624de40

Observation 1c5b6137-edd7-4d39-98f1-7d410a52472f · outbound

This paper cites evosax: JAX-based Evolution Strategies.

How Should We Meta-Learn Reinforcement Learning Algorithms? evosax: JAX-based Evolution Strategies

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.110073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.110073Z digest=sha256:6fbd07ad6fe7fd66a6c024f223259671c08b9d38d86cab2080aa9da33238e008

Observation 2a2eb53a-4375-4fbb-9164-1db3ebb27504 · outbound

This paper cites In-context reinforcement learning with algorithm distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? In-context reinforcement learning with algorithm distillation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.496541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.255410Z digest=sha256:33aa894aed5ad4aeb8c46d750c0a109721e516c52af11824474a81e98701f095

Observation a65675a6-b2a8-42da-91e0-7c76d1806eac · outbound

This paper cites Evolution through Large Models.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolution through Large Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.383996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.383996Z digest=sha256:ec5f3ce3e710e2f67a3ffa9ccc882ac11bca2f4d58b3dd2b0275545eeb912f3f

Observation ea817af4-650b-4692-a29a-e8901cb12399 · outbound

This paper cites Rediscovering orbital mechanics with machine learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Rediscovering orbital mechanics with machine learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.338948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.486735Z digest=sha256:366152d24552ff3a44a6a3bb3b068adbe85c5baf7534aa719f5aba69826b5b71

Observation 35ca94b3-032a-4957-86d1-36418967adeb · outbound

This paper cites Discovered policy optimisation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovered policy optimisation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.186310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.575011Z digest=sha256:0aa92f030cb427b6e4bde4dd6139964b98c5b50ede8e2d10ae9c4c3edcdd0c63

Observation 4d9a5bb5-e31a-4381-89b2-6e54d990ab38 · outbound

This paper cites Discovering Preference Optimization Algorithms with and for Large Language Models.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering Preference Optimization Algorithms with and for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.754244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.754244Z digest=sha256:a4575608aa07537c2b5208f0ed9a17d13c15b895feed6ae630122bb5ec6e8501

Observation af015423-f7fa-41e0-9775-a370f6364081 · outbound

This paper cites Behaviour Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Behaviour Distillation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.015958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.891797Z digest=sha256:37c32956e7f941dce8e6c781254271ab36517c817505c2984a2c228c8f0cc165

Observation 0c5f61b8-a66c-4923-a626-748f3e6bdea9 · outbound

This paper cites Understanding plasticity in neural networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding plasticity in neural networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.023177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.023177Z digest=sha256:5e8cea4266b5ced0aab48ea2a1adf66f8bd13138b6ea60214a9daab51f07eeb0

Observation 0336ceb2-4e37-44af-9cc1-59d24cf54dfd · outbound

This paper cites Craftax: a lightning-fast benchmark for open-ended reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Craftax: a lightning-fast benchmark for open-ended reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.824630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.134832Z digest=sha256:3d3277b84ee51962df586790c9757fea907b11ed2ebab4654695e0d0d773be85

Observation 759aebf1-ca1f-4df1-8d37-a0074a1264b0 · outbound

This paper cites Interpretable machine learning methods applied to jet background subtraction in heavy-ion collisions.

How Should We Meta-Learn Reinforcement Learning Algorithms? Interpretable machine learning methods applied to jet background subtraction in heavy-ion collisions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.652401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.301257Z digest=sha256:ff237f4f644d6b95eb838bc79ef41fcabc86c418f2ed906f311239e3ae5ee32f

Observation 6ce3b857-d15e-46fb-98c2-40ec67b58731 · outbound

This paper cites Meta-Learning Update Rules for Unsupervised Representation Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Meta-Learning Update Rules for Unsupervised Representation Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.432448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.432448Z digest=sha256:b37d82051e8b5258aca975a8b4f03ef27c99d9a307ec457bd5db74294627f14d

Observation 0df60976-ee7b-48c8-8d27-81fe624486ab · outbound

This paper cites Understanding and correcting pathologies in the training of learned optimizers.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding and correcting pathologies in the training of learned optimizers

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:48:53.486491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.527570Z digest=sha256:4eba925da87f042b503cdf871fbff76fd93b01f85d48c057925895adccf56569

Observation 44f62159-29b7-427f-9199-f6b3536a3d0c · outbound

This paper cites Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves.

How Should We Meta-Learn Reinforcement Learning Algorithms? Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.605196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.605196Z digest=sha256:76b2358cffab4818d6d2d1c9f81d42077c7bff0ebad81ea53f71c96053518ad9

Observation 0610a14e-4f3a-47f2-9159-606f1641d390 · outbound

This paper cites Gradients are Not All You Need.

How Should We Meta-Learn Reinforcement Learning Algorithms? Gradients are Not All You Need

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.755270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.755270Z digest=sha256:f069eee77e54775756c9c393436fd152568cfaba133574e146a87b4cda0f60f5

Observation 76f8e25e-1c33-4720-a7f4-075bb6db140e · outbound

This paper cites VeLO: Training Versatile Learned Optimizers by Scaling Up.

How Should We Meta-Learn Reinforcement Learning Algorithms? VeLO: Training Versatile Learned Optimizers by Scaling Up

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.872182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.872182Z digest=sha256:b70463ef5a701cdd2f08b73808e207a1fb32995f50f1c4d949d2e97e44722555

Observation a99fa079-00a9-49d9-b6e0-52fbad2f45cb · outbound

This paper cites Self-distillation amplifies regularization in hilbert space.

How Should We Meta-Learn Reinforcement Learning Algorithms? Self-distillation amplifies regularization in hilbert space

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.459946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.019447Z digest=sha256:6bce6e4c67c0ec09dd94e6acbfc2176f9bd24d6244081b88969e946783c6a2e0

Observation 72858513-525e-4cb8-8404-30205921394e · outbound

This paper cites Small batch deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Small batch deep reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.295508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.139633Z digest=sha256:00139d3fbb59150a049f2584b9d07eb3207474a6e455f648560492ec85fddaad

Observation afcfcb2f-aebb-41cd-89c1-e2a2d7896257 · outbound

This paper cites Discovering reinforcement learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering reinforcement learning algorithms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.134460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.254660Z digest=sha256:e0a8fdbfe867171899ac2709f36ba1e5bb23c2b82e3cd9e0a77ebd6fc0931226

Observation 8075000c-14fd-47d8-9aed-dfdac0ce7168 · outbound

This paper cites Openai o3-mini, January 2025.

How Should We Meta-Learn Reinforcement Learning Algorithms? Openai o3-mini, January 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.965617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.382465Z digest=sha256:42f6c06b0efe72590d772b1d854f1b28a7c919dba96a0b706f021b29387b9406

Observation 836da2ff-865c-4390-b4c4-0b61b1be5c16 · outbound

This paper cites Stabilizing transformers for reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Stabilizing transformers for reinforcement learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.789331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.452623Z digest=sha256:97f756e74f82cf95ddba858bdeae228ac5a8df05703e661f12b818fe7b9a8759

Observation b3da679a-906b-4c9d-810a-8580885f52c3 · outbound

This paper cites Evolving Curricula with Regret - Based Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolving Curricula with Regret - Based Environment Design

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.589660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.495021Z digest=sha256:ef11da93c1cca04cb302bf2dad2feff84650b4ced4c99c6f5e4c56fa5596fc9d

Observation 8f69eaaf-88c3-44ef-a1dd-9c73bda00a80 · outbound

This paper cites Parameter Space Noise for Exploration.

How Should We Meta-Learn Reinforcement Learning Algorithms? Parameter Space Noise for Exploration

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.593230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.593230Z digest=sha256:1744c4fde520f36db087b7f29da564b26ba4823ab7e4f925faafb3dddc277b86

Observation d359abff-4acd-473c-8099-de350efa7649 · outbound

This paper cites Tunability: Importance of hyperparameters of machine learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Tunability: Importance of hyperparameters of machine learning algorithms

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.346301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.706376Z digest=sha256:fd9fd68c644aee57097a5ae6fb8173bda1733ec71d445e51886f29daa9663b8a

Observation a2e2d967-75c0-47e9-8a37-fb87210c6462 · outbound

This paper cites Evolutionsstrategie : Optimierung technischer systeme nach prinzipien der biologischen evolution.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolutionsstrategie : Optimierung technischer systeme nach prinzipien der biologischen evolution

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.153556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.800901Z digest=sha256:c008b559d256c3ec26dfa8caded770b114bfd2494e13a36fb94fc56541f37fac

Observation 9bb36e11-6709-4c14-9b4c-ea544ef8c666 · outbound

This paper cites Pawan Kumar, Emilien Dupont, Francisco J.

How Should We Meta-Learn Reinforcement Learning Algorithms? Pawan Kumar, Emilien Dupont, Francisco J

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.860051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.860051Z digest=sha256:1b1291b2ae62ec60392c3028a89b321b47e9fcf886a7841c999591b99bf79758

Observation f77d5312-72e7-4aeb-93d5-b885a82980a1 · outbound

This paper cites Policy Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Policy Distillation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.969246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.969246Z digest=sha256:85cb7f028a912278cc524c4a1af10ca247bf94d48cb9c4780a7e737934767641

Observation 1dfd04f3-40b7-4c24-ac86-7f19ddad9f85 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.080292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.080292Z digest=sha256:a5eacb28422ee70d709442746a4bbc644e125f97704edbbaf25cf2bd2900a0a5

Observation 21e3588a-9d4c-44ad-b16e-341ea7d14213 · outbound

This paper cites Proximal Policy Optimization Algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Proximal Policy Optimization Algorithms

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.130615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.130615Z digest=sha256:3ced2976c7a95ac4f25ce69afe5a40bf6d6e9eddfd7701d64a3471ffd30ff39a

Observation 30db874b-70dd-4448-bbc1-1cb7d6c5b9ef · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

How Should We Meta-Learn Reinforcement Learning Algorithms? High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.229277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.229277Z digest=sha256:10ce1b86c1b94745ea1ef39cb156bf59ab8864a6632423a35c07ec31dd782ad8

Observation 945e19ab-25ed-45ef-90d3-58f813481d85 · outbound

This paper cites The Dormant Neuron Phenomenon in Deep Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? The Dormant Neuron Phenomenon in Deep Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.326750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.326750Z digest=sha256:539d34e4c53bb02212cfa3968e7c4c94721a9aa3fac81262cd384871ce6bb3da

Observation d765ea3d-991b-4127-946d-725e108653ad · outbound

This paper cites Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.431138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.431138Z digest=sha256:e4da2d1a608ab87e12bd8f5f9bf4851b5c38fecee4356296d7f341047cbe8ac4

Observation 3f177879-b4e0-4a29-adab-9a732bd859f2 · outbound

This paper cites Generalizable Symbolic Optimizer Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Generalizable Symbolic Optimizer Learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.021535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.499138Z digest=sha256:0a333685045a4f6990a46eba479bc0ee06f3c1b29e0996e25ff777151361224a

Observation cf324fbd-f556-48f8-ba71-a16468de3b0e · outbound

This paper cites Position: Leverage Foundational Models for Black-Box Optimization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Position: Leverage Foundational Models for Black-Box Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.550816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.550816Z digest=sha256:6855b05f0d3b0dec2e9c77bc607ecae1e4a93ab1435c4f29c53fe2a42fb2d1bc

Observation 4d92d76f-80a5-4dc5-9d43-0693b7e3567e · outbound

This paper cites Maxinfo RL : Boosting exploration in reinforcement learning through information gain maximization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Maxinfo RL : Boosting exploration in reinforcement learning through information gain maximization

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.840199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.635219Z digest=sha256:51dabdba7ab1a3a6d2250bb4244a0294f16c57ed1bd010ef253c98679124ac92

Observation 9f2931ef-91ed-4c19-9777-88dfd6d5e8e6 · outbound

This paper cites Sutton and Andrew Barto.

How Should We Meta-Learn Reinforcement Learning Algorithms? Sutton and Andrew Barto

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.668222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.716178Z digest=sha256:3939dce2f16ec9f3b01bb2355ea6f710a6abd7177e13f2eb4bd3db94572e9cfa

Observation 3acf0046-0b42-41dd-8def-cac57e2a55b3 · outbound

This paper cites Improving deep reinforcement learning by reducing the chain effect of value and policy churn.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving deep reinforcement learning by reducing the chain effect of value and policy churn

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.798344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.798344Z digest=sha256:2b789b1ab1c20000a09eaa6c2db6f13cc72efb2e9b2787aca1c01855db4bd82e

Observation 7658bc7e-8099-4ad4-8afb-4a365f10451b · outbound

This paper cites MuJoCo : A physics engine for model-based control.

How Should We Meta-Learn Reinforcement Learning Algorithms? MuJoCo : A physics engine for model-based control

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.925909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.925909Z digest=sha256:26de816a77e584d2ae712b765a16a9c4442636b16db4e6d69a3f53bd412bbdfd

Observation 9d4d48f1-3bb7-4255-a172-039e321357a0 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

How Should We Meta-Learn Reinforcement Learning Algorithms? Deep Reinforcement Learning and the Deadly Triad

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.022023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.022023Z digest=sha256:4e318287b3816687d6b55411fc870db0f11f6ec64e294213465324acfcadac14

Observation b9178950-d006-48d7-84c0-46e344cce11d · outbound

This paper cites Attention Is All You Need.

How Should We Meta-Learn Reinforcement Learning Algorithms? Attention Is All You Need

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.137791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.137791Z digest=sha256:9c51ad4aff80488b19eedfd8771199efca2ba0a521ee872bca9cd1cd0e55b0bd

Observation 8f07cb6a-ce7b-45e4-8917-29771434a0ae · outbound

This paper cites Dataset Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Dataset Distillation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.204307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.204307Z digest=sha256:addb1113ab8053a380746e251b4c4b67dcd9281d05c01bfd2b9e8178e4c127d6

Observation 59339402-f08a-4723-bd2c-7434f8185b54 · outbound

This paper cites Natural Evolution Strategies.

How Should We Meta-Learn Reinforcement Learning Algorithms? Natural Evolution Strategies

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.127951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.287410Z digest=sha256:5195b1ed5a2df299bbf95af41b3f3608c19622463d7baad2a80892040f20d776

Observation 985ed171-fb21-495a-b2b0-84d329f3aa4b · outbound

This paper cites Understanding short-horizon bias in stochastic meta-optimization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding short-horizon bias in stochastic meta-optimization

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.523597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.381492Z digest=sha256:5949fac48039d969e600022becb5916c852ae5ff2ed2583e20e3102ca7fc3f04

Observation 1e7ab7c4-6b5d-4d2a-8ff5-7a5056aac1b8 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

How Should We Meta-Learn Reinforcement Learning Algorithms? MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.488734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.488734Z digest=sha256:5289e12bc34536225008816ec178d42e0abc45ff82dfb8a82e8630a3a3128b14

Observation cc47c1d0-9afe-4395-bec3-813cad431c96 · outbound

This paper cites Self-distillation as instance-specific label smoothing.

How Should We Meta-Learn Reinforcement Learning Algorithms? Self-distillation as instance-specific label smoothing

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.343900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.573001Z digest=sha256:dfd984389c7afbab06e1af5df85d1b6d8bec0132c01189018c3b53991d0c66ea

Observation 44945f70-98c5-4abf-b02c-65bc7fd38f80 · outbound

This paper cites Symbolic Learning to Optimize: Towards Interpretability and Scalability.

How Should We Meta-Learn Reinforcement Learning Algorithms? Symbolic Learning to Optimize: Towards Interpretability and Scalability

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:52.957731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.681440Z digest=sha256:d330ea4d538ae45dd934778cfbe4cc4fd320d267e6fce5ee018853050d964311

Observation 41439ad1-985a-4ebb-b5bf-0348a9646f81 · outbound

This paper cites write newline.

How Should We Meta-Learn Reinforcement Learning Algorithms? write newline

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.758471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.758471Z digest=sha256:ef6e55e4641d35ff2d5813953ef9f69b324cc260e75fc791de289ed082599528

Pith citing papers

Observation c32431f8-ff8d-4845-9333-b867b3e88878 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback How Should We Meta-Learn Reinforcement Learning Algorithms?

Reference 243

Resolution
verified exact
local_arxiv, observed 2026-08-03T04:44:19.150566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-03T04:39:32.117730Z digest=sha256:5f2f4812da7c9070e3a6dcc82e21f47b6add1f7749141b16a046f5ebac715ea1