Pith. sign in

Paper Citation Record · LEDGER

How Should We Meta-Learn Reinforcement Learning Algorithms?

As of 12 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 1 inbound Pith citation observation for arXiv:2507.17668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17668 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:48:52.758471Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:39:32.117730Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact4
  • verified fuzzy26
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a3d7ccb0-92bc-4646-b308-e5a76539886e · outbound

This paper cites Loss of Plasticity in Continual Deep Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Loss of Plasticity in Continual Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.616239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.616239Z digest=sha256:31bd74e5399b63e0b124b91500689459d189c4af558b230ed76e73289f9f6dc9

Observation 24da469e-b54b-4146-bb26-ec474352ed58 · outbound

This paper cites Towards Characterizing Divergence in Deep Q-Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Towards Characterizing Divergence in Deep Q-Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.681243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.681243Z digest=sha256:b7c9f4f324d7f5249571ada6d721ad525b57a9abeff4964524d33b119a59bacd

Observation f48faf5f-8f10-4095-b923-e3fe5bf3f8f4 · outbound

This paper cites A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.755232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.755232Z digest=sha256:cacc749c40520ce8cf28da2162866055e111ac420e609a4a07457fe579a633dc

Observation b19f019a-93e9-4b0d-b221-c531836e6370 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

How Should We Meta-Learn Reinforcement Learning Algorithms? Deep reinforcement learning at the edge of the statistical precipice

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.809562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.809562Z digest=sha256:429e0633957f6b52a410779042b46673f4dcca5e520f6184885cee444222942e

Observation 44b1374f-aa2b-4bff-a13f-84907f107a90 · outbound

This paper cites A Generalizable Approach to Learning Optimizers.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Generalizable Approach to Learning Optimizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.948249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.948249Z digest=sha256:00e5984cf9e73b2864e73091616383bc0c7de85dbc20d8f3aae7e0dee6174bec

Observation 34b95b01-ac5e-4eaf-afc5-9dd2e849b3f3 · outbound

This paper cites Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas.

How Should We Meta-Learn Reinforcement Learning Algorithms? Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.089106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.089106Z digest=sha256:b07518c250f709449b9a9149db1e1ed9f1545cff395d75881bb5a4c968a94c99

Observation 2fafa30b-8079-487a-bd81-65d3d1db55d3 · outbound

This paper cites An information-theoretic perspective on intrinsic motivation in reinforcement learning: A survey.

How Should We Meta-Learn Reinforcement Learning Algorithms? An information-theoretic perspective on intrinsic motivation in reinforcement learning: A survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.201906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.201906Z digest=sha256:169ea79de450cd4ec0febf8e06304d6112a6a6251316421d45cc8e021fe6641a

Observation 59369bea-951d-4879-922e-ac071fdf0ea3 · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Tutorial on Meta-Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.339004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.339004Z digest=sha256:b36783f5b6d6932fe75eec0b2e6b2ad2a6ef97fd807bfc46003aa25b15e48518

Observation 8fd95d5d-5e2d-46b6-84db-12a0aa2f778f · outbound

This paper cites OpenAI gym, 2016.

How Should We Meta-Learn Reinforcement Learning Algorithms? OpenAI gym, 2016

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.485426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.485426Z digest=sha256:796852b917093f1cca8f906e28d6e4e9c4e7a10add4c2eaa6ef69af258a88000

Observation 379f9e2e-23a4-4b86-8f11-d59bb3c045d6 · outbound

This paper cites Exploration by Random Network Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Exploration by Random Network Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.624540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.624540Z digest=sha256:eba4d7c0694faf57fa6cbb5e1d9f65d3a2de570eeefef54a580fd4e7b81f6130

Observation 5dc19526-a0f3-4c51-bbe0-c15aaf14490d · outbound

This paper cites Boltzmann exploration done right.

How Should We Meta-Learn Reinforcement Learning Algorithms? Boltzmann exploration done right

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.793806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.793806Z digest=sha256:94bf6fc81ee5beb5241e6d7aa31102a1c888af96464b0c83930b5191ca41f6bd

Observation 2049f298-757a-4752-a563-7d78d809c499 · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Symbolic Discovery of Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.939586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.939586Z digest=sha256:94e567ebb427ebc0ee5acb65656262135e84a62fe762a40069f3f7fae52c4d21

Observation 1e634b23-c977-4fb1-adee-0abef0b42e18 · outbound

This paper cites Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl.

How Should We Meta-Learn Reinforcement Learning Algorithms? Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.030210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.030210Z digest=sha256:b998f33d979800a1079877f0ac2eddeeb8d34a29e92dd2138311df9c464c360c

Observation 1aea3f56-9912-422e-8722-f2befb70c15f · outbound

This paper cites Discovering Symbolic Models from Deep Learning with Inductive Biases.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering Symbolic Models from Deep Learning with Inductive Biases

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.105628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.105628Z digest=sha256:dfc4d1ac152900e9f8bb931ce63b6ffe899268f4e2454f8da88a027bd4d38d83

Observation 707eb909-41c4-4c9a-bf29-55fb711e05c8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.158711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.158711Z digest=sha256:3a1e69aebe0f188da66556789fb7dbd5713fa8cd82e47b7e6714a099faf6c722

Observation 96506966-4351-4c0c-80b6-4511d5ffe9e5 · outbound

This paper cites Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.239937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.239937Z digest=sha256:ef65c840ec5cc909ed3516ad3cfab58bdcf8d396655f8bd8eae628f8bf4a2e22

Observation 82b12f95-c64e-4c2d-849b-84e96f4da4d9 · outbound

This paper cites Loss of plasticity in deep continual learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Loss of plasticity in deep continual learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.309720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.309720Z digest=sha256:98e532c78719e4b00e4ca803385a0d9627bfc8bcc1e70cf8f8d617d44104140d

Observation bad298e5-b312-469a-9370-21d35555a317 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.395863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.395863Z digest=sha256:b41e7e5c29707208810326ad78ee29b482055d344694f7a60cfa94ae12289f7b

Observation f70732c3-e78c-447d-b9f7-793c88ad0058 · outbound

This paper cites Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps.

How Should We Meta-Learn Reinforcement Learning Algorithms? Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:48:54.113207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:46.513473Z digest=sha256:7a1ca4eea16502d306c9c276ef0553025805a43dc5c2b0e4b28c0cb70275751a

Observation e97cd32c-805c-4709-bced-cce321896e67 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

How Should We Meta-Learn Reinforcement Learning Algorithms? OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.585185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.585185Z digest=sha256:6d5f921bc510f30f5db687b09d6902b9aa21b43c54a0cfafa1527ed13a235af1

Observation 6454b23b-eafe-4b95-bb70-59f7f5a77502 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Model-agnostic meta-learning for fast adaptation of deep networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.675019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.675019Z digest=sha256:ff8b1413c43856a2d25c63b84d5dfbaac18a27b4c9c960d8190d6957e11d310e

Observation 23895565-0168-4277-9fd2-a917f64ddf2e · outbound

This paper cites Noisy Networks for Exploration.

How Should We Meta-Learn Reinforcement Learning Algorithms? Noisy Networks for Exploration

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.738150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.738150Z digest=sha256:65688fef4dc73ba998a8c03c88ff16452d8a0674bc23cd75182265587836860c

Observation faaafd20-4489-409b-b34e-4e5bd93cbcb1 · outbound

This paper cites Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem.

How Should We Meta-Learn Reinforcement Learning Algorithms? Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.822940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.822940Z digest=sha256:46ee3c32678080314a7457aa375249d3a37f48f043ef86e77194e16a678d8cbf

Observation 0648e23b-f653-4be6-b3c1-e4613c3559fc · outbound

This paper cites Born Again Neural Networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Born Again Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.904574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.904574Z digest=sha256:d70991e161e1f680e494071b7aab9b71e69d008a69a844f68508021867a1e6e3

Observation c4e7164c-4f0c-4ab6-a9c1-9520d6a6f49a · outbound

This paper cites Goldie, Chris Lu, Matthew T.

How Should We Meta-Learn Reinforcement Learning Algorithms? Goldie, Chris Lu, Matthew T

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:59.178524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.000738Z digest=sha256:e2564533c4dd64e50ac800a838926670e077a826e3695cf69cfbe1e4ffbc46e8

Observation 0d14a9d1-f1ad-4f76-9b59-997042574c7c · outbound

This paper cites Benchmarking the Spectrum of Agent Capabilities.

How Should We Meta-Learn Reinforcement Learning Algorithms? Benchmarking the Spectrum of Agent Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.078571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.078571Z digest=sha256:03ddc9014fc361fbbd160d7d97f2f66935ff95f7e1ba5b5aa03c014a739e203e

Observation 3d14ae1c-d8f4-47b8-a380-a881f6b7f07d · outbound

This paper cites Distilling the Knowledge in a Neural Network.

How Should We Meta-Learn Reinforcement Learning Algorithms? Distilling the Knowledge in a Neural Network

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.149052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.149052Z digest=sha256:66688ef86fd6abf533e5c7ca61eb1e10424c8e9c465e322d85ac863b7407ea0e

Observation 4080903f-56ca-481e-9526-3d647585a01f · outbound

This paper cites Long short-term memory.

How Should We Meta-Learn Reinforcement Learning Algorithms? Long short-term memory

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.222901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.222901Z digest=sha256:790a279eabdd3041c2634400d7028fdc36db514dc152baab16fe1cda942bc1ce

Observation dc9c16ba-60c7-4c79-90e9-7a662088ee40 · outbound

This paper cites Automated Design of Agentic Systems.

How Should We Meta-Learn Reinforcement Learning Algorithms? Automated Design of Agentic Systems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.306436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.306436Z digest=sha256:13404068a42ab4bbe43465804edeec2c14052335ce2baa6d44e4b5ecd813e028

Observation 41b311f6-d302-40b8-a5ea-f13008d342b1 · outbound

This paper cites Transient non-stationarity and generalisation in deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Transient non-stationarity and generalisation in deep reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.888824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.383576Z digest=sha256:644c22d73e58195bebc86540854f441e115f4aedfdd89047a359a13a75d7b9f9

Observation 0edfac8e-8273-47a5-aa23-a12663df7d34 · outbound

This paper cites Transient non-stationarity and generalisation in deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Transient non-stationarity and generalisation in deep reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.610653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.471532Z digest=sha256:e77b71aa09e0215440ef74892ce0c968734fac05a83be360217c58cdebe13c37

Observation 7d42ccf3-7005-4fd5-8841-a55667aa143c · outbound

This paper cites Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.861384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.539950Z digest=sha256:2a13a77c6ba88488cbc3e7b467d2f4dbc5f2aec17841c5e96f625863dd43d729

Observation 820da51f-8024-4178-abbc-91a7e9405067 · outbound

This paper cites Discovering temporally-aware reinforcement learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering temporally-aware reinforcement learning algorithms

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.350351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.617075Z digest=sha256:306cd331749cce4f4bb5c2b847a763bdd8b7d9f40fc2fac307b256218748dc79

Observation 6bf5b200-36ad-4e13-9410-c95c9baf1e10 · outbound

This paper cites Improving policy optimization with generalist-specialist learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving policy optimization with generalist-specialist learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.129243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.682731Z digest=sha256:b084ded03e3a5183594f3127bfc85eb49616028a6b434f4b8057d84eada12f14

Observation 3a1a5cb6-507c-49c1-adf6-d89d5bf2585e · outbound

This paper cites Meta Learning Backpropagation And Improving It.

How Should We Meta-Learn Reinforcement Learning Algorithms? Meta Learning Backpropagation And Improving It

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.690763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.763840Z digest=sha256:9e48ee50c7538d8207430e8cb8d78a35d708e7a1632399d0db8930d562a6de2b

Observation 267ff5a4-13c3-4ac9-98bb-2b92691851ef · outbound

This paper cites Improving generalization in meta reinforcement learning using learned objectives.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving generalization in meta reinforcement learning using learned objectives

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.828958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.817050Z digest=sha256:2fb3a8cd3c0e68473d195fcacfed647fda28a05c339e1d462bb257f5f6b0779b

Observation 2517e4d0-afda-4d88-927b-a6d3a685f536 · outbound

This paper cites Mirror Learning: A Unifying Framework of Policy Optimisation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Mirror Learning: A Unifying Framework of Policy Optimisation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.878994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.878994Z digest=sha256:369038658d7168bb95e9c39dba21e156abe64d93bdaf2b8adf225ac253ae1e97

Observation 853ffeb4-20c3-4278-8903-a0cc889d3f0f · outbound

This paper cites Learning to Optimize for Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Learning to Optimize for Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.922233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.922233Z digest=sha256:2a1c853dedbd9466cf3a9e07911bc3c9c4be84bfa44d43ab0524511af9533e22

Observation 670b210f-af41-461e-9f14-69132d7b1c52 · outbound

This paper cites gymnax: A JAX -based reinforcement learning environment library, 2022 a.

How Should We Meta-Learn Reinforcement Learning Algorithms? gymnax: A JAX -based reinforcement learning environment library, 2022 a

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.659181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.972077Z digest=sha256:2be6f9f57ef098c36a7d17d3cd5ed4c96b74c46c0b2a70cc4ac4d262aab25f0e

Observation 1c5b6137-edd7-4d39-98f1-7d410a52472f · outbound

This paper cites evosax: JAX-based Evolution Strategies.

How Should We Meta-Learn Reinforcement Learning Algorithms? evosax: JAX-based Evolution Strategies

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.110073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.110073Z digest=sha256:4a75cb513467510e3c93f62d349dabab3b8a544a82870882d60716e2ee82ad02

Observation 2a2eb53a-4375-4fbb-9164-1db3ebb27504 · outbound

This paper cites In-context reinforcement learning with algorithm distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? In-context reinforcement learning with algorithm distillation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.496541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.255410Z digest=sha256:c3d04bf96ed6e0a577e2a5d05051095367eeb85c05d3b1231419a1860dcff7b3

Observation a65675a6-b2a8-42da-91e0-7c76d1806eac · outbound

This paper cites Evolution through Large Models.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolution through Large Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.383996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.383996Z digest=sha256:4161e0555be9beb48888eec1363b3d23646bc6dddb0e5f5332845be4978befd6

Observation ea817af4-650b-4692-a29a-e8901cb12399 · outbound

This paper cites Rediscovering orbital mechanics with machine learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Rediscovering orbital mechanics with machine learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.338948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.486735Z digest=sha256:2d99e6bf97ce166fbe9e5100d376f0316eda84a56dabcd668ed4702befa0814f

Observation 35ca94b3-032a-4957-86d1-36418967adeb · outbound

This paper cites Discovered policy optimisation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovered policy optimisation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.186310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.575011Z digest=sha256:19c458b5b5f1abc34d1127e11a3cc711435e606fb75f56ef41b836ce92794e63

Observation 4d9a5bb5-e31a-4381-89b2-6e54d990ab38 · outbound

This paper cites Discovering Preference Optimization Algorithms with and for Large Language Models.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering Preference Optimization Algorithms with and for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.754244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.754244Z digest=sha256:7860f549b52fe0396ff7baa869ddbc74b0d5398e7bff1db71bfefd0423158f72

Observation af015423-f7fa-41e0-9775-a370f6364081 · outbound

This paper cites Behaviour Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Behaviour Distillation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.015958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.891797Z digest=sha256:0c8618efda6310fa02a60135e53885b16bca523280ca8bcdbfef3e0a62ba3e1a

Observation 0c5f61b8-a66c-4923-a626-748f3e6bdea9 · outbound

This paper cites Understanding plasticity in neural networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding plasticity in neural networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.023177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.023177Z digest=sha256:0d1d8ed06bb6e1eec2bf4391c1eb37b2ae4b9be67a73c7da6ba7d02254f5a212

Observation 0336ceb2-4e37-44af-9cc1-59d24cf54dfd · outbound

This paper cites Craftax: a lightning-fast benchmark for open-ended reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Craftax: a lightning-fast benchmark for open-ended reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.824630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.134832Z digest=sha256:40f3f8dcda1d7cb1afbeaef59f3e23326c89567a9489459aecebfc4a376ea2b3

Observation 759aebf1-ca1f-4df1-8d37-a0074a1264b0 · outbound

This paper cites Interpretable machine learning methods applied to jet background subtraction in heavy-ion collisions.

How Should We Meta-Learn Reinforcement Learning Algorithms? Interpretable machine learning methods applied to jet background subtraction in heavy-ion collisions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.652401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.301257Z digest=sha256:bc4907e6693c95057672dd5fcb7dfa28e89067930e380b700f0c03c821fc5f40

Observation 6ce3b857-d15e-46fb-98c2-40ec67b58731 · outbound

This paper cites Meta-Learning Update Rules for Unsupervised Representation Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Meta-Learning Update Rules for Unsupervised Representation Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.432448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.432448Z digest=sha256:751e708cf605c3881826faa008342921ca0272dcce80c665654703fcfa3f2f7b

Observation 0df60976-ee7b-48c8-8d27-81fe624486ab · outbound

This paper cites Understanding and correcting pathologies in the training of learned optimizers.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding and correcting pathologies in the training of learned optimizers

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:48:53.486491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.527570Z digest=sha256:d4d313aad6d8c1685a275bfe98d6a45f9aee8fd510e3346c244d4e8a7a6beac9

Observation 44f62159-29b7-427f-9199-f6b3536a3d0c · outbound

This paper cites Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves.

How Should We Meta-Learn Reinforcement Learning Algorithms? Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.605196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.605196Z digest=sha256:96ba1c9248231c295195d9dbca65a53ac4f568428895532eaf21c73a9290e3bd

Observation 0610a14e-4f3a-47f2-9159-606f1641d390 · outbound

This paper cites Gradients are Not All You Need.

How Should We Meta-Learn Reinforcement Learning Algorithms? Gradients are Not All You Need

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.755270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.755270Z digest=sha256:19326faf1dfa78b3671aecd3d2c63d0b8c422743174fdb3a2e02d7324aa0be0c

Observation 76f8e25e-1c33-4720-a7f4-075bb6db140e · outbound

This paper cites VeLO: Training Versatile Learned Optimizers by Scaling Up.

How Should We Meta-Learn Reinforcement Learning Algorithms? VeLO: Training Versatile Learned Optimizers by Scaling Up

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.872182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.872182Z digest=sha256:b206a604b560755a541a8fe1dd23a7d6c44505944842fe43a531ac50224a4c83

Observation a99fa079-00a9-49d9-b6e0-52fbad2f45cb · outbound

This paper cites Self-distillation amplifies regularization in hilbert space.

How Should We Meta-Learn Reinforcement Learning Algorithms? Self-distillation amplifies regularization in hilbert space

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.459946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.019447Z digest=sha256:80988a250a797fc2630f6de38097073a4a614562adccb87ba844e463d1d82ab3

Observation 72858513-525e-4cb8-8404-30205921394e · outbound

This paper cites Small batch deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Small batch deep reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.295508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.139633Z digest=sha256:29b50bab97c1a38d9ba043a8aeaa5031b2732e5cfa8e483a301d9e8e1eb61807

Observation afcfcb2f-aebb-41cd-89c1-e2a2d7896257 · outbound

This paper cites Discovering reinforcement learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering reinforcement learning algorithms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.134460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.254660Z digest=sha256:a2975f563899ff68450d6d316058a4c34d6f98cf49649b0a0c6d918317e8f636

Observation 8075000c-14fd-47d8-9aed-dfdac0ce7168 · outbound

This paper cites Openai o3-mini, January 2025.

How Should We Meta-Learn Reinforcement Learning Algorithms? Openai o3-mini, January 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.965617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.382465Z digest=sha256:caf1fe925603c8f22f09286e31867534d195a677f620bed58cfc0f5b6ddf96f3

Observation 836da2ff-865c-4390-b4c4-0b61b1be5c16 · outbound

This paper cites Stabilizing transformers for reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Stabilizing transformers for reinforcement learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.789331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.452623Z digest=sha256:322d39eef408e6998df05baf758e18577ceee4a5b5e4cdeb2395dad8f93d99f8

Observation b3da679a-906b-4c9d-810a-8580885f52c3 · outbound

This paper cites Evolving Curricula with Regret - Based Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolving Curricula with Regret - Based Environment Design

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.589660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.495021Z digest=sha256:6d303dfd198c35f12c1ddcaae5f607968160f624bc653db7a5475acee063d132

Observation 8f69eaaf-88c3-44ef-a1dd-9c73bda00a80 · outbound

This paper cites Parameter Space Noise for Exploration.

How Should We Meta-Learn Reinforcement Learning Algorithms? Parameter Space Noise for Exploration

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.593230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.593230Z digest=sha256:23b4b175b727691400009bbbb1351143eb2cd3f1f4eab1a1b3c0ca2af783951d

Observation d359abff-4acd-473c-8099-de350efa7649 · outbound

This paper cites Tunability: Importance of hyperparameters of machine learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Tunability: Importance of hyperparameters of machine learning algorithms

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.346301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.706376Z digest=sha256:aced8c153e717902492267b137faa175e2af14002afd35cb5604ba14de660a90

Observation a2e2d967-75c0-47e9-8a37-fb87210c6462 · outbound

This paper cites Evolutionsstrategie : Optimierung technischer systeme nach prinzipien der biologischen evolution.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolutionsstrategie : Optimierung technischer systeme nach prinzipien der biologischen evolution

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.153556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.800901Z digest=sha256:7254acfb6318421ba87b112a16c9f8c58a8e13b5e6b00bd6ce1d00342e2bb15d

Observation 9bb36e11-6709-4c14-9b4c-ea544ef8c666 · outbound

This paper cites Pawan Kumar, Emilien Dupont, Francisco J.

How Should We Meta-Learn Reinforcement Learning Algorithms? Pawan Kumar, Emilien Dupont, Francisco J

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.860051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.860051Z digest=sha256:d98aadd85cacb1f4d3e787fed6f223109533141d797d2d16d986f258557ca58a

Observation f77d5312-72e7-4aeb-93d5-b885a82980a1 · outbound

This paper cites Policy Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Policy Distillation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.969246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.969246Z digest=sha256:4705561ddabe23032e41eb68c5b5ba9f7bdc64a0d138c0ccedd7e63b80bf2ad5

Observation 1dfd04f3-40b7-4c24-ac86-7f19ddad9f85 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.080292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.080292Z digest=sha256:324c3822d38d336bd1589069761090c6ec0ae8bf5367ae1396a73380bc9e8bfa

Observation 21e3588a-9d4c-44ad-b16e-341ea7d14213 · outbound

This paper cites Proximal Policy Optimization Algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Proximal Policy Optimization Algorithms

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.130615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.130615Z digest=sha256:cf1ab557233324a06e82081d62ca817647cbf333ff064740238724d71469ae5f

Observation 30db874b-70dd-4448-bbc1-1cb7d6c5b9ef · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

How Should We Meta-Learn Reinforcement Learning Algorithms? High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.229277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.229277Z digest=sha256:ee61f9136914378acea7162495a0938d2a9ede96482334527d794e6ea1f190fe

Observation 945e19ab-25ed-45ef-90d3-58f813481d85 · outbound

This paper cites The Dormant Neuron Phenomenon in Deep Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? The Dormant Neuron Phenomenon in Deep Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.326750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.326750Z digest=sha256:46071ca52ac06b270baf5aa33b9668ad00bd80c23a3051e7351c46003808452a

Observation d765ea3d-991b-4127-946d-725e108653ad · outbound

This paper cites Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.431138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.431138Z digest=sha256:f96198a96e2b2e94cb741745a108ff9993a7560b93377186fcd4f551f40c81ff

Observation 3f177879-b4e0-4a29-adab-9a732bd859f2 · outbound

This paper cites Generalizable Symbolic Optimizer Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Generalizable Symbolic Optimizer Learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.021535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.499138Z digest=sha256:74c28d21bdfa599c0cdcb4c024a88058d1dd7542c78894812fc36f332afd0d93

Observation cf324fbd-f556-48f8-ba71-a16468de3b0e · outbound

This paper cites Position: Leverage Foundational Models for Black-Box Optimization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Position: Leverage Foundational Models for Black-Box Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.550816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.550816Z digest=sha256:67024d01e75ce5685f4bc7a35c91420bd6e014e784cc671dad591b2677e183f3

Observation 4d92d76f-80a5-4dc5-9d43-0693b7e3567e · outbound

This paper cites Maxinfo RL : Boosting exploration in reinforcement learning through information gain maximization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Maxinfo RL : Boosting exploration in reinforcement learning through information gain maximization

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.840199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.635219Z digest=sha256:9bc5179bb1f13a3c9c06b17461bcf9c9eb9c0e2d1fe5dbe3c012b6c88738ce89

Observation 9f2931ef-91ed-4c19-9777-88dfd6d5e8e6 · outbound

This paper cites Sutton and Andrew Barto.

How Should We Meta-Learn Reinforcement Learning Algorithms? Sutton and Andrew Barto

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.668222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.716178Z digest=sha256:75c8c268ff197f8df8bee9345c83368a7bc0785262767a6358dd9d34c9c113ef

Observation 3acf0046-0b42-41dd-8def-cac57e2a55b3 · outbound

This paper cites Improving deep reinforcement learning by reducing the chain effect of value and policy churn.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving deep reinforcement learning by reducing the chain effect of value and policy churn

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.798344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.798344Z digest=sha256:608488d0a9c0dd696e64379a447f7dbd41beea06e67ce6ab9650b3cd04565193

Observation 7658bc7e-8099-4ad4-8afb-4a365f10451b · outbound

This paper cites MuJoCo : A physics engine for model-based control.

How Should We Meta-Learn Reinforcement Learning Algorithms? MuJoCo : A physics engine for model-based control

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.925909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.925909Z digest=sha256:8094de3005ea125ce48807b19c1d00fed6fe88557caed8c9c38a4691f78c983a

Observation 9d4d48f1-3bb7-4255-a172-039e321357a0 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

How Should We Meta-Learn Reinforcement Learning Algorithms? Deep Reinforcement Learning and the Deadly Triad

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.022023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.022023Z digest=sha256:d3116b9d2ce351a920d3191acfab59fcc58f2ced6723d74d87af1c2179f53996

Observation b9178950-d006-48d7-84c0-46e344cce11d · outbound

This paper cites Attention Is All You Need.

How Should We Meta-Learn Reinforcement Learning Algorithms? Attention Is All You Need

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.137791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.137791Z digest=sha256:97402603996035e99965f72729b1031c4769510f213e6b4964ce31441e30f264

Observation 8f07cb6a-ce7b-45e4-8917-29771434a0ae · outbound

This paper cites Dataset Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Dataset Distillation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.204307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.204307Z digest=sha256:bdda77f21cd16f16e1f5bf88e98c629a6607c3a76d671556def1a8812f85a274

Observation 59339402-f08a-4723-bd2c-7434f8185b54 · outbound

This paper cites Natural Evolution Strategies.

How Should We Meta-Learn Reinforcement Learning Algorithms? Natural Evolution Strategies

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.127951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.287410Z digest=sha256:3f3f8eb200c9f0c373a11611707d2cef5f460842c2f6a89ef99f8b0e5ff5e1ab

Observation 985ed171-fb21-495a-b2b0-84d329f3aa4b · outbound

This paper cites Understanding short-horizon bias in stochastic meta-optimization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding short-horizon bias in stochastic meta-optimization

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.523597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.381492Z digest=sha256:027cdbd3685f97b2d9554b66aa9531c47cfc3140d3f19fd0241b895feaaa0374

Observation 1e7ab7c4-6b5d-4d2a-8ff5-7a5056aac1b8 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

How Should We Meta-Learn Reinforcement Learning Algorithms? MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.488734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.488734Z digest=sha256:eebcbace42855134262ea4fca0018940cfb71b81ee1b9315f6d5460157bd1004

Observation cc47c1d0-9afe-4395-bec3-813cad431c96 · outbound

This paper cites Self-distillation as instance-specific label smoothing.

How Should We Meta-Learn Reinforcement Learning Algorithms? Self-distillation as instance-specific label smoothing

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.343900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.573001Z digest=sha256:0ce26b91b4e59ae271b37437bc530d94015e35c77467a5fd6f11d3b9cf565a22

Observation 44945f70-98c5-4abf-b02c-65bc7fd38f80 · outbound

This paper cites Symbolic Learning to Optimize: Towards Interpretability and Scalability.

How Should We Meta-Learn Reinforcement Learning Algorithms? Symbolic Learning to Optimize: Towards Interpretability and Scalability

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:52.957731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.681440Z digest=sha256:43f01b158d414dd25fe7006e0a19469ecb4b9d32c23ca21cdcf814e591cb7838

Observation 41439ad1-985a-4ebb-b5bf-0348a9646f81 · outbound

This paper cites write newline.

How Should We Meta-Learn Reinforcement Learning Algorithms? write newline

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.758471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.758471Z digest=sha256:0e5757043a01a9e9c8d0386bdb9baed0b3884de0e10b3bfd3eae3775b02df2d9

Pith citing papers

Observation c32431f8-ff8d-4845-9333-b867b3e88878 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback How Should We Meta-Learn Reinforcement Learning Algorithms?

Reference 243

Resolution
verified exact
local_arxiv, observed 2026-08-03T04:44:19.150566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-03T04:39:32.117730Z digest=sha256:b018035750ae455de0fe30c38cc62a211678036ea24d4cea6eb83e2d588f0914