Pith. sign in

Paper Citation Record · LEDGER

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2502.00288.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00288 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:39:58.748802Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact5
  • verified fuzzy24
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 91374a6e-fc4f-4e94-bfff-27c6b7243ffc · outbound

This paper cites Layer Normalization.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.569489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.569489Z digest=sha256:ebafd8d2d0f2691fc22b142294f212ef9adf03b961635172f5d29c96400dfd41

Observation ef8aae6c-ceae-4080-bf79-253b2d321ea1 · outbound

This paper cites J., Smith, L., Kostrikov, I., and Levine, S.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network J., Smith, L., Kostrikov, I., and Levine, S

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.974310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.574141Z digest=sha256:dcc60313d76df3b3061d0d4da61ab29fdf3bcaae46fd4b656882049cff54b29c

Observation af8f9d16-85ec-47ec-8e05-7cc9d5413d18 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Dota 2 with Large Scale Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.577776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.577776Z digest=sha256:3e4bb0f50859a54eb5744f65872e03dc4927c0bbd5c2edb1a239a1a6f6cdd3e7

Observation b2f8dd92-cde7-4946-9ec5-23b8f5711018 · outbound

This paper cites Offline rl without off-policy evaluation.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Offline rl without off-policy evaluation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.963751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.582381Z digest=sha256:405e0b0b63faf0cfa3cbd5a96facc583d00154225bfb32f4df3df48ed7811f80

Observation a4fae73a-bb69-4f9e-a2e2-7d7a2fae07cf · outbound

This paper cites an unresolved cited work.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.586110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.586110Z digest=sha256:f8eaa1cd1997af340f37eb9889fca86ca09109115e76a7c354b1ffbe5329f921

Observation 0150fed3-b684-4b76-8ff9-00776e071b32 · outbound

This paper cites A., Salazar, G., Tran, H.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network A., Salazar, G., Tran, H

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.945981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.589949Z digest=sha256:032eb6250fbccd7608b748cced802c2ae4a5f0d17ee247a60bbb31d93a859002

Observation a4c28b55-7c16-4d0e-bbf5-c61969531f0d · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Decision transformer: Reinforcement learning via sequence modeling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.935274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.593890Z digest=sha256:812adb3b294f65ee1488a78ad360b7d646f661b90b9678d2a3c9858e2b3e8bba

Observation b0d109ad-e96a-4b00-994e-d6727e0c5b54 · outbound

This paper cites Rvs: What is essential for offline RL via supervised learning? In International Conference on Learning Representations, 2022.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Rvs: What is essential for offline RL via supervised learning? In International Conference on Learning Representations, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.924194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.597241Z digest=sha256:afd76b04497c9743758cfd48ecffa1d2832352909e9e18794c710ec6f3ad3c5b

Observation 1fecd2b1-1114-40f1-a0e2-6ef574376969 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Counterfactual multi-agent policy gradients

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.600622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.600622Z digest=sha256:794b632f7f3896cd2fdce0ac051a71e1b5f0828dcaee5bde4463903960a0f2d0

Observation db3b8a5c-55dd-40ad-a00f-948b388fe3ef · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.604186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.604186Z digest=sha256:d3d2b665248f43524eab56a1ca0a299a48f78c85c11ba0397966c0d9fc4dc050

Observation 7cf1abf2-b44b-437b-90f8-0fe4bb516db4 · outbound

This paper cites and Gu, S.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network and Gu, S

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.913496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.608251Z digest=sha256:871584bcd010fa6a9bdec679fe25711b994ff9a30ea7f4d8fc8430910c7327be

Observation 2e974eba-5512-4ef2-9190-856862d12ee9 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Addressing function approximation error in actor-critic methods

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.612097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.612097Z digest=sha256:114130bfb1b0b724de01b42df662694376eee748d54b4f7d41b077144b8a75e9

Observation 26ad1416-bb4e-4828-8072-f452aba5f1ce · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Reinforcement learning with deep energy-based policies

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.615756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.615756Z digest=sha256:1dd47b276b97970fe6d45049eb25563923d2f130c68b3670dda6c8c666184c43

Observation 2e7cc4ae-b972-48a4-be71-8365288c2187 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.619235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.619235Z digest=sha256:e0d1171a4f4e9278676304a546ca75dad0334280df3701140d2522dbcc70bf9d

Observation 762c91c2-2744-4237-b726-5ae2ce80c82c · outbound

This paper cites Modem: Accelerating visual model-based reinforcement learning with demonstrations.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Modem: Accelerating visual model-based reinforcement learning with demonstrations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.882950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.622632Z digest=sha256:29775ac45010a388f56bca55eaa2cf29bf31389b53238876091d98fb9b1e6ba7

Observation bd92b01e-1c3a-4daf-b9b1-0e7201050ce0 · outbound

This paper cites Neural networks: a comprehensive foundation.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Neural networks: a comprehensive foundation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.871795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.626098Z digest=sha256:cdc45cbd1243c2fa29105f62a1ded0e2278622b8d844ac76f8bc6b5b62d9d5f6

Observation 12eb4e6e-2e22-44b3-ad2c-b7be9411f06e · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Gaussian Error Linear Units (GELUs)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.629593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.629593Z digest=sha256:4632de5035f8d75afb79c013d15fc0ec218ed22029ffdead9a8299a0b941263d

Observation 7a2b416e-2a8f-4220-b613-e2dec9408142 · outbound

This paper cites Deep q-learning from demonstrations.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Deep q-learning from demonstrations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.633495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.633495Z digest=sha256:f18844c47e788034c47b173e5e3a2a881f2a7029c9fd40967232a0bc813d10b7

Observation 6c80aa3c-ffd2-4295-be7a-5301295dc6a5 · outbound

This paper cites B ayesian design principles for offline-to-online reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network B ayesian design principles for offline-to-online reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.860365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.637064Z digest=sha256:77b211f61b945c210fdbbbe50e1489b2d13dfd07c09bc2c4a7dd3fb823b496d2

Observation 493f95b0-9b97-4569-bfea-0995ca08aa92 · outbound

This paper cites R., and Davison, A.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network R., and Davison, A

Reference 20

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T19:39:59.593966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.642052Z digest=sha256:99e8dd5c2b55c748d8d6998e299100ebc1a7b673fdd971d59a357c50d531df82

Observation b47e30cf-10fb-42c2-b51d-1a33e754ab05 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Offline reinforcement learning with implicit q-learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.645779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.645779Z digest=sha256:8760adc711184ae5cf7eb8f34b95a9288c4316be0db6f2de79cb260065fd02f4

Observation ccf644af-9f28-4a5f-b182-1954113cb865 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Conservative q-learning for offline reinforcement learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.649212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.649212Z digest=sha256:0c44e888bf565a5b72296e78d5e8184931e6af29c91bf6916dee962242758c8e

Observation 538adca6-c19f-489d-9158-65b072b067f2 · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.837241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.652437Z digest=sha256:ef45d364e79ba4d1dec29527f86f49aba433c2568c3649d02ebdddac5ab42aab

Observation aaf53e1a-99d2-490a-aa35-c56a737c90fb · outbound

This paper cites Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.826694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.655705Z digest=sha256:3f3b67eb2a4d905e4e619d6fe99de24744d4be62b1ab9b55e0eddd0744c56605

Observation aaf30073-6ddb-40f9-9145-f0178adef183 · outbound

This paper cites A survey of convolutional neural networks: Analysis, applications, and prospects.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network A survey of convolutional neural networks: Analysis, applications, and prospects

Reference 25

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T19:39:59.411897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.658938Z digest=sha256:ef6069c04563cda461f4899ac8382208f2157b13c3304e5c71d6641b3cb4d2b0

Observation e884384c-00bc-4e93-aeb3-71aef5333199 · outbound

This paper cites Continuous control with deep reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Continuous control with deep reinforcement learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.662257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.662257Z digest=sha256:db0621e93710b7c6c3d064229615693265602cc9258a20a16bf4b0f16201098d

Observation 824cb68b-7c81-4977-99f0-f85c9e5c024a · outbound

This paper cites Discrete Sequential Prediction of Continuous Actions for Deep RL.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Discrete Sequential Prediction of Continuous Actions for Deep RL

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.665938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.665938Z digest=sha256:5ea14474460ef9a33fcd90497dcc4f65e1d1b3d44a6b38481be909046ff2dca4

Observation d4425530-19bc-4373-a556-f48c66a2efe0 · outbound

This paper cites A., Veness, J., Bellemare, M.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network A., Veness, J., Bellemare, M

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.669816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.669816Z digest=sha256:99a7a37526b42642ebd11b51eb08df9c988ecbea46314b5616c7f82ec07f47bd

Observation 02cadaba-4b3f-4c95-bc9b-b71a628de3d4 · outbound

This paper cites Overcoming exploration in reinforcement learning with demonstrations.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Overcoming exploration in reinforcement learning with demonstrations

Reference 29

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T19:39:59.176348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.673153Z digest=sha256:7413acdb40425c637f15f51fdb6fbb87067211f9d0f82b0fa05dcdeeea1b8f9d

Observation f3d12f85-0098-4afe-b502-e03eb612d060 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.809768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.676302Z digest=sha256:bebce8494048d2bf34af07c1fcb7dc6180d5cb104d085a483bdafa6a5799338c

Observation 52cea845-30ab-41f4-a614-e4c62b127473 · outbound

This paper cites Learning complex dexterous manipulation with deep reinforcement learning and demonstrations, 2018.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Learning complex dexterous manipulation with deep reinforcement learning and demonstrations, 2018

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.798314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.679741Z digest=sha256:8f8e3d377a72c92553b92e0cd8cf226f2646e81016af35154e6a55c7567dcef3

Observation 991c977d-4f6b-4758-b19a-9f829615dd29 · outbound

This paper cites an unresolved cited work.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:39:59.787378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.682953Z digest=sha256:d09dc2f466a7f15a36191201edbd0298bd026711c14ac43dafb1cd917a982281

Observation 358af754-ac01-460a-8232-e75feb34d301 · outbound

This paper cites Mastering atari, go, chess and shogi by planning with a learned model.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Mastering atari, go, chess and shogi by planning with a learned model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.686288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.686288Z digest=sha256:c69016ad96fee115dbae97e7ae80f079b514aad6204b1dc5ba9635649cf9a5b5

Observation ceb203cc-861f-4e3c-80ee-a5fddc944378 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Proximal Policy Optimization Algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.689604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.689604Z digest=sha256:bce9ebe211b06fbc0cde1bd1d1d3dba961367ea0edc96adf0d245ff11d328e6e

Observation b3b41d11-1f2b-412f-bcda-c5f7599ee9d2 · outbound

This paper cites and Abbeel, P.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network and Abbeel, P

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.693154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.693154Z digest=sha256:f8804e5fe47c10f7bef0ff14ffe1f89d11f96825b0e1e6c23bbea41c69d3585b

Observation c10b52f3-7744-4a07-94d5-28a36404e590 · outbound

This paper cites Continuous control with coarse-to-fine reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Continuous control with coarse-to-fine reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.769503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.696684Z digest=sha256:f7e3c8accb5944356bd0c6d1dd7e490e8b09de57c13aa9cfcd5029b08de4161f

Observation 26766bd2-613b-4f3f-bc8a-b5711008786f · outbound

This paper cites Solving continuous control via q-learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Solving continuous control via q-learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.758043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.700226Z digest=sha256:d50c2d0721deed769d7739c9725270ce8a9808ad47733c3c1fafac8788d26d98

Observation 04f49af1-4824-4fed-b2dd-03e662db1c07 · outbound

This paper cites Growing Q -networks: S olving continuous control tasks with adaptive control resolution.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Growing Q -networks: S olving continuous control tasks with adaptive control resolution

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.746815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.703922Z digest=sha256:c460a347a1f1416c12389aa6354f24516e0dba6933c7b52553dc255ba59d3fa7

Observation d5ef7617-735e-416b-94eb-1fa96bc507a2 · outbound

This paper cites Mastering the game of go without human knowledge.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Mastering the game of go without human knowledge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.707331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.707331Z digest=sha256:28b2668a37f30d1bdf1697286ea2f9a17961e99f2de10f8e2e71efe1510d37cd

Observation 590cead1-f946-4069-816b-55fde7943841 · outbound

This paper cites and Agrawal, S.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network and Agrawal, S

Reference 40

Resolution
verified exact
doi, observed 2026-08-09T19:39:58.802667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.710807Z digest=sha256:3a20b1985ac154f5f93989db60c0739627986d58fafd6985430120b5f40ad76a

Observation b2aca0e3-7fa3-4625-9fa3-72622036c302 · outbound

This paper cites Action branching architectures for deep reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Action branching architectures for deep reinforcement learning

Reference 41

Resolution
verified exact
doi, observed 2026-08-09T19:39:58.790610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.714256Z digest=sha256:20c02ad1389a0f2ce27813352dfd9672c790809065992a1ae145bf3c8e087d1b

Observation 76bb7886-7220-4fcb-b150-e0503c4d87ab · outbound

This paper cites Learning to represent action values as a hypergraph on the action vertices.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Learning to represent action values as a hypergraph on the action vertices

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.729926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.717906Z digest=sha256:a475fa41fbb633700ee8731d9c28b01c5c0f88c5398a1bd8b4927aeac7af4cd7

Observation c3230589-da4a-48a3-a5e1-f6b7a96708c4 · outbound

This paper cites Deep reinforcement learning with double q-learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Deep reinforcement learning with double q-learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.721491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.721491Z digest=sha256:e422c63aea58e13a127a9d4c85f9c5dbf205ea2505387e453773611236b65d64

Observation 909c3e06-0f3d-4fa0-a2cd-0c5d93d6deb8 · outbound

This paper cites Discriminator-weighted offline imitation learning from suboptimal demonstrations.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Discriminator-weighted offline imitation learning from suboptimal demonstrations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.719141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.725163Z digest=sha256:823fb631eb6ad9692db5b10c1ac94801c71a7a3a8b858ef00c4b5d956fbbb4f2

Observation f2673bbe-0e28-409c-acef-3c4fb7a9d297 · outbound

This paper cites Hd-cnn: Hierarchical deep convolutional neural networks for large scale visual recognition.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Hd-cnn: Hierarchical deep convolutional neural networks for large scale visual recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.708249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.728435Z digest=sha256:7ca6bdd42465de822edea510624f91bbe4bc3d4cd76d8f0da6470137129d70b3

Observation b7835a39-4772-4a97-9a2c-832d712639db · outbound

This paper cites Mastering visual continuous control: Improved data-augmented reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Mastering visual continuous control: Improved data-augmented reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.695917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.731948Z digest=sha256:944f4ee22727155bd1adfbaf959d5b4b403a05729a7aa7b5d3481b630d19d54e

Observation 51b5b9c5-0637-44b1-8770-9a8c6aa53510 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network The surprising effectiveness of ppo in cooperative multi-agent games

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.684908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.735090Z digest=sha256:676db758f4f47e8aeef629bd72066a7eb3e99f0fa3338105376a86bae9bfa8fe

Observation 0094e1ab-6b8c-48c7-92db-8cf0d2792c8b · outbound

This paper cites Policy expansion for bridging offline-to-online reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Policy expansion for bridging offline-to-online reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.674731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.738490Z digest=sha256:d551997d8688750c5bb30de2f711c527980fe77ddabfa237804ecfdb35c7394d

Observation 87a41440-f206-42de-8260-e1bae4ec3c3c · outbound

This paper cites Z., Kumar, V., Levine, S., and Finn, C.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Z., Kumar, V., Levine, S., and Finn, C

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.664137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.741871Z digest=sha256:fb139791b4a6bcc2d7e79917803e17022d1734316602ca5f0775774d7fc4e709

Observation c7add5a0-d9ee-428d-8d03-a620df032e05 · outbound

This paper cites D., Maas, A., Bagnell, J.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network D., Maas, A., Bagnell, J

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.653624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.745240Z digest=sha256:32b3cd8d54142324cb0d8af33538625f2ff686ea1b0c80e2fbc5b481a722f18d

Observation 4a2aa577-c749-4a7a-bd3d-e5e44f236f26 · outbound

This paper cites write newline.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.748802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.748802Z digest=sha256:237aaa7123ba66d9c521723b15deabb7908bfedcf3b1ee18bdc0517206d9a9ca

Pith citing papers

No inbound Pith citation observations are available.