Pith. sign in

Paper Citation Record · LEDGER

Offline Reinforcement Learning with Penalized Action Noise Injection

As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.02356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02356 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:38:35.725134Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1da2545c-de84-4043-9793-2d18e1ffe030 · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Offline Reinforcement Learning with Penalized Action Noise Injection Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:38.434048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:31.737091Z digest=sha256:d22c9fa409c321bb4b138c5aee2bd4a8cbe6aff91cbec265ebf1686063560ba0

Observation 3e5d8264-1aa1-4027-9faf-c087bdbf7e9a · outbound

This paper cites Layer Normalization.

Offline Reinforcement Learning with Penalized Action Noise Injection Layer Normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.842375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.842375Z digest=sha256:cc48a0ac96942a6428ab55e67be838921acec72a459a9cb22d362edad17caf7a

Observation dede0fb2-7726-4036-bc49-c3a9da87a142 · outbound

This paper cites Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling.

Offline Reinforcement Learning with Penalized Action Noise Injection Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.961205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.961205Z digest=sha256:15aef92020125e252c4a158aa8e9ec8f8c47afaf485830b2e0651de04377d163

Observation f7d77ab3-7008-4262-a7e8-eb0341d5621b · outbound

This paper cites Score Regularized Policy Optimization through Diffusion Behavior.

Offline Reinforcement Learning with Penalized Action Noise Injection Score Regularized Policy Optimization through Diffusion Behavior

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.067287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.067287Z digest=sha256:cb2f6a791985caef6ac2baf836bfb20ebb9e4922d3c48d948d43ac446588f70d

Observation 30bc8fe0-d178-4f77-ac42-81ce3fe4e4d5 · outbound

This paper cites Diffusion Policies creating a Trust Region for Offline Reinforcement Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Diffusion Policies creating a Trust Region for Offline Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.141102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.141102Z digest=sha256:1b70b46a5ddc93617bea695d9d8aa3a1de9e48a62de72ea19e2f958ba9f64720

Observation 270e3281-3349-405a-be63-20c824aff9ab · outbound

This paper cites Heavy-tailed denoising score matching.

Offline Reinforcement Learning with Penalized Action Noise Injection Heavy-tailed denoising score matching

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:38:36.017010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:32.216742Z digest=sha256:1d1a272d6e630abeecdcb137226051a074b9f422d1f725348c3017b665af5770

Observation cfd8ec7d-c3f3-4b95-a8e1-9eed79c10ea8 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.343671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.343671Z digest=sha256:10123c65941473cb3d8bc6058338a7b9a04483adfdf06b6ea9e90910ef3efa18

Observation 2458abd8-0a18-49dc-8f1e-74484648aefc · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection A minimalist approach to offline reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.427027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.427027Z digest=sha256:c37edd34a49e2dddaf9618077cc23cb0dc3ba483eabc78a0d52f302ab92a53d9

Observation fe4cf0f3-adb7-45c5-9de0-b0f5b52e3330 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Offline Reinforcement Learning with Penalized Action Noise Injection Addressing function approximation error in actor-critic methods

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.494366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.494366Z digest=sha256:4ff4415ae9d9d6c2129776b9c20c85bb0c0e2d28b26f86380ce2ae5064881c5b

Observation 3b8f0a35-73b2-4bb6-91b2-d81f6d6e16cc · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Offline Reinforcement Learning with Penalized Action Noise Injection Off-policy deep reinforcement learning without exploration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.621106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.621106Z digest=sha256:0dacf84f8c76733f18917ccc07a70e81bf39fd2b181421bc647313aaa3efc830

Observation 0bdfee5c-e361-4565-82e3-156ebaf484f5 · outbound

This paper cites Calculus of variations.

Offline Reinforcement Learning with Penalized Action Noise Injection Calculus of variations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.743577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.743577Z digest=sha256:4404dc377dcf05e46bbae5895d912a9566c8c76a7e5dda1770389423bfcd3d7f

Observation 44a5caa0-2ef4-435c-bf9d-d4675b8f5884 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Offline Reinforcement Learning with Penalized Action Noise Injection IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.861799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.861799Z digest=sha256:0910d43b4f82e29214e7462f6a35e4ec4d6e44c510bca945f3f2c032c3c9a4ab

Observation 3fd83fcf-0547-4a0b-9dc1-ecabbe7bc4a1 · outbound

This paper cites Estimation of non-normalized statistical models by score matching.

Offline Reinforcement Learning with Penalized Action Noise Injection Estimation of non-normalized statistical models by score matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.978532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.978532Z digest=sha256:2fbacbc0f3e15f5c53797f60300af16f1d8258012e250297440a2fe062ac2307

Observation 23d6c552-6feb-46af-97a5-0fbb323cb07f · outbound

This paper cites Understanding diffusion objectives as the elbo with simple data augmentation.

Offline Reinforcement Learning with Penalized Action Noise Injection Understanding diffusion objectives as the elbo with simple data augmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:38.232042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.074290Z digest=sha256:0f45f6321eae5207ed7d1ff6334e47d24ff0d77975f0b5351be2d5fd543cdb62

Observation 0601121a-7f5e-46fb-ab8b-600e70a417b3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Offline Reinforcement Learning with Penalized Action Noise Injection Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:33.201984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:33.201984Z digest=sha256:4b8beed47bcbc6299827906c1011e90138acd314b0a8ccffb86b8fc62fddb2d9

Observation 8dcf4a96-1f10-4f7d-8c6b-853a671b92f4 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Offline Reinforcement Learning with Implicit Q-Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:33.288918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:33.288918Z digest=sha256:5c3aeb38420c1e31b3119835b6a335c8cbd9b3cdfee67826c172a80b4bddcd3f

Observation 09e7a1eb-a7a6-4c39-b6e4-c53cf57c1574 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Offline Reinforcement Learning with Penalized Action Noise Injection Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:33.405392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:33.405392Z digest=sha256:7c4c26c3a0da480cdbf92ea66b3146f9b4a475f77ced8c3f6c45c7bd411dde70

Observation bcf8ceda-bcaa-428f-bb39-bc5937405986 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Conservative q-learning for offline reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:38.059986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.524554Z digest=sha256:7501b625e66d7ef5bab9507d1a3214ae21acd0b388ddbb3a08780804d82df6ee

Observation 41b9557a-32c3-4a56-ba75-368c9b45e136 · outbound

This paper cites Reinforcement learning with augmented data.

Offline Reinforcement Learning with Penalized Action Noise Injection Reinforcement learning with augmented data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.879872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.627961Z digest=sha256:d6454af5281ed632c33c9e071e468be600a04e1849b809ae991a7d965ca26109

Observation 55ec212d-70ea-418b-a4dd-af6133addaaa · outbound

This paper cites Batch reinforcement learning with hyperparameter gradients.

Offline Reinforcement Learning with Penalized Action Noise Injection Batch reinforcement learning with hyperparameter gradients

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.675965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.739155Z digest=sha256:bf4cb5b4bcfc16c30591741581511dcc7e74b6e79d73ae874761a26205d33b0d

Observation ec0ef95b-8d50-4a57-abaf-d3cabbd5cc84 · outbound

This paper cites Learning energy-based models in high-dimensional spaces with multiscale denoising-score matching.

Offline Reinforcement Learning with Penalized Action Noise Injection Learning energy-based models in high-dimensional spaces with multiscale denoising-score matching

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.521833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.848855Z digest=sha256:19b7369d3fd5df8e5e0d3cf38e4ac2cd23eb6d9bf4564bde607859c76fe66a72

Observation d3ce1b0b-1a87-4a90-96ad-568a994fe870 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.258953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:33.985104Z digest=sha256:c7085d9cd13f33a21e8aa12dc23e236f2d2cb0765d37f6d3f30ae828cd0ec8af

Observation 51d92faf-a390-4ba8-9e6e-a0c6d0d6d8af · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Offline Reinforcement Learning with Penalized Action Noise Injection AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.078884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.078884Z digest=sha256:7c8a8244e79a21c3ee93c0e2df104cbd2474ee3f88f35b50be7d70416af11509

Observation baf425d1-58da-4a6f-bb21-d574dbcd88a6 · outbound

This paper cites Anti-exploration by random network distillation.

Offline Reinforcement Learning with Penalized Action Noise Injection Anti-exploration by random network distillation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:37.079058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.197168Z digest=sha256:5a9db27c6c949d4029f96fda985b63672d688f8949c649f786e9e96a448121c2

Observation b65795e4-72c4-4282-bfeb-70a1979aaa5b · outbound

This paper cites Heavy-Tailed Diffusion Models.

Offline Reinforcement Learning with Penalized Action Noise Injection Heavy-Tailed Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.311608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.311608Z digest=sha256:11722bcbcedd9abcc5998a6a699161f0a64267ceeb9f3c4f34d90693bd6082be

Observation 78e68a9c-3574-42a0-8af4-e66fd739dad3 · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

Offline Reinforcement Learning with Penalized Action Noise Injection DreamFusion: Text-to-3D using 2D Diffusion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.429681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.429681Z digest=sha256:93bacbb188bfd207e04700ec3d493abbae37263634a5a94c36415cd78b9b69e6

Observation c922708f-0235-4672-9a6f-680c9919918d · outbound

This paper cites Efficient differentiable simulation of articulated bodies.

Offline Reinforcement Learning with Penalized Action Noise Injection Efficient differentiable simulation of articulated bodies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.915029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.550267Z digest=sha256:87cb965dc7652b16eac062f8e62dfecb4d84a9f4a3c7545d5faae81bbb5f506d

Observation 995cbc0a-a10c-42ba-8e3e-cbb70fa0c5ac · outbound

This paper cites Offline reinforcement learning as anti-exploration.

Offline Reinforcement Learning with Penalized Action Noise Injection Offline reinforcement learning as anti-exploration

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.743353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.671019Z digest=sha256:3c2962ead49666abe21aa2b3357c57ce4c2892a967c6f82871ac0f2b9c0437ee

Observation de5bffb6-547d-4f2c-a9f5-e3254fbdfef1 · outbound

This paper cites S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics.

Offline Reinforcement Learning with Penalized Action Noise Injection S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.568452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:34.772132Z digest=sha256:e60829e5817feaf00c05fab0d8a83a39a42c5ed7f90b6873c9fece64de2bff71

Observation bb38a4bf-2c58-4375-80ea-5a4aa65b7b1e · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Offline Reinforcement Learning with Penalized Action Noise Injection Generative modeling by estimating gradients of the data distribution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.862757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.862757Z digest=sha256:23d5d91f86b42c28c910cf779589ff75adb439f66853312b4b947812b6f32ea5

Observation 8a2cce95-1e29-4969-ab39-8b4baee2e0c2 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Offline Reinforcement Learning with Penalized Action Noise Injection Score-Based Generative Modeling through Stochastic Differential Equations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:34.981272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:34.981272Z digest=sha256:5a5933795686906f102e96abef0671dc7525e12d41dd5b9a24536468ca530c63

Observation af1d8a34-686b-4d5f-8006-e3459a95c306 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Revisiting the minimalist approach to offline reinforcement learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.087356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.087356Z digest=sha256:26f25bbfc195c1e6afe53e636591b78fd9d55503ab98d6fa9f59d0910f4ed9cd

Observation 78e6b11b-7441-4184-b72d-1864ae3d9c6f · outbound

This paper cites A connection between score matching and denoising autoencoders.

Offline Reinforcement Learning with Penalized Action Noise Injection A connection between score matching and denoising autoencoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.208587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.208587Z digest=sha256:31cae1119d44079b5c80e93800e9cb55cac4c1bd685b189c82e921a5c734c4e2

Observation 95dc4f19-86d8-4674-8e21-89a853191561 · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Offline Reinforcement Learning with Penalized Action Noise Injection Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.366704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.366704Z digest=sha256:cc5a9ca045c3087592ddbe074d71584eac6f100bd98263a99e3f6c9bab6f9c9d

Observation 0cfadee7-393b-42d9-8e70-a045f47d28c7 · outbound

This paper cites On scale mixtures of normal distributions.

Offline Reinforcement Learning with Penalized Action Noise Injection On scale mixtures of normal distributions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.366574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:35.479265Z digest=sha256:caa5a3926f5e7a31b82a7fd31d4a6ccad6a44cf71c70038e56164eccf31012ed

Observation 993d3e30-6b2c-4ac3-b75d-ea23486c921c · outbound

This paper cites Exploration and Anti-Exploration with Distributional Random Network Distillation.

Offline Reinforcement Learning with Penalized Action Noise Injection Exploration and Anti-Exploration with Distributional Random Network Distillation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:35.625180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:35.625180Z digest=sha256:3ba80ee4990b42bf2c2fadc5dec0e43ad4d5b7417791e60c5c3c19dc4267319e

Observation bea9b73e-a826-4506-ab75-db014269a869 · outbound

This paper cites Rorl: Robust offline reinforcement learning via conservative smoothing.

Offline Reinforcement Learning with Penalized Action Noise Injection Rorl: Robust offline reinforcement learning via conservative smoothing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.183894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:38:35.725134Z digest=sha256:1a6cc83f325dd5c6db8c1f52a8b63a22c0ca17359ca1c9c2405bc2543bb88ff8

Pith citing papers

No inbound Pith citation observations are available.