Pith. sign in

Paper Citation Record · LEDGER

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving

As of 22 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2509.04712.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04712 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T06:00:50.704572Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee52a3a2-f5b2-4319-a44f-364d299d5565 · outbound

This paper cites A survey of autonomous driving: Common practices and emerging tech- nologies,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving A survey of autonomous driving: Common practices and emerging tech- nologies,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.509243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.507442Z digest=sha256:59b11dacabb24eadc96f4507b040a15ce92e480de529ee01aa275f84ad344fa0

Observation 0cb7e546-170d-4ba8-8088-1d89d2e43943 · outbound

This paper cites Lane change and merge maneuvers for connected and automated vehicles: A survey,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Lane change and merge maneuvers for connected and automated vehicles: A survey,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.485506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.513003Z digest=sha256:00b8d3d4c28e6a8d58aa7c78d07923d34149c480686e15ac9e7105be62a6ab57

Observation c93f8b65-980f-45c8-a51a-d3fb682a2c58 · outbound

This paper cites Automated lane change controller design,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Automated lane change controller design,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.463003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.518589Z digest=sha256:8a76c9f05814e055cd58aa1c001cd0ce814be38a3ea01b474f04620765288a31

Observation 0f9284b2-f756-4e47-8754-778eb9f38f79 · outbound

This paper cites Traffic dynam- ics: studies in car following,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Traffic dynam- ics: studies in car following,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.434134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.523478Z digest=sha256:902a6c4a220d607d074a852639420c01c9df371bc6f332e23e93a2fa32234a39

Observation f33213f1-b4a5-466d-837b-0d1dfb435c13 · outbound

This paper cites A behavioural car-following model for computer simulation,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving A behavioural car-following model for computer simulation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.409275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.529563Z digest=sha256:a1acfc8e84dd06ac14f745e4c1dcc36a3b17246f7f62c8a7135695cd6e094d18

Observation 8b341eb3-75a1-4475-bca7-c7d6930636bc · outbound

This paper cites Congested traffic states in empirical observations and microscopic simulations,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Congested traffic states in empirical observations and microscopic simulations,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.535468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.535468Z digest=sha256:84b913d591c6e09ec7a0cd479eee75c55b99c136f9558fd67f3e39b33621c7b5

Observation 20015de6-1c5c-4f10-aa39-1b2aa8b98703 · outbound

This paper cites General lane-changing model mobil for car-following models,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving General lane-changing model mobil for car-following models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.541460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.541460Z digest=sha256:2a0d63be8ef1d3f5f9a90953a8de5f6c7f7a0cbdb69f58de51d75e98f058b5e4

Observation 6715f95c-e45f-4db3-8ba8-878f389d535c · outbound

This paper cites Driving intention recognition and lane change prediction on the highway,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Driving intention recognition and lane change prediction on the highway,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.349165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.546470Z digest=sha256:89f823276c4b166451b08d58b0e5dcc78577af19a94dfaba0116d19c60156891

Observation c24b7b92-da87-4042-8347-13feec93eedd · outbound

This paper cites End to End Learning for Self-Driving Cars.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving End to End Learning for Self-Driving Cars

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.551475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.551475Z digest=sha256:f2245f760e206d50d7eab585133e758bb44ac84aa6ea14069c8b514ff4375bc7

Observation 375a326f-43e9-4b16-9a36-85937ffbcd2e · outbound

This paper cites Explaining How a Deep Neural Network Trained with End-to-End Learning Steers a Car.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Explaining How a Deep Neural Network Trained with End-to-End Learning Steers a Car

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.557642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.557642Z digest=sha256:6dcda9d00154609f52a5955b3c3826cf7fd23944ccac416971e3ee7230465e7d

Observation 55b9b7e8-12ba-4701-a5b2-05f268b3e490 · outbound

This paper cites End-to-end driving via conditional imitation learning,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving End-to-end driving via conditional imitation learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.319836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.563235Z digest=sha256:85bfdb5a253b3d9038e552963942d00a6c39ad007825d8be876e2673ff2872a7

Observation 36b2f5bc-7278-495e-8fb1-8b47a0162173 · outbound

This paper cites Urban driving with conditional imitation learning,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Urban driving with conditional imitation learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.292253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.567984Z digest=sha256:42a3fe556889cc407cee9ac58a627a4ad928849999887fdbf8223b2ce59f8ef5

Observation 55972cc3-481a-460e-a29a-3e839dd90330 · outbound

This paper cites an unresolved cited work.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.573228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.573228Z digest=sha256:ffef313d7a9afd6c1ae65942f1ad7f01e51bab4c2279f970a4cb00b73492f7c2

Observation d9bb2ed2-0075-4f50-a6b6-bf88234ca509 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.578999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.578999Z digest=sha256:71d9ba592e23f7539d770dbc89760678111b60a8628e57ae2a7a0a378c2b6a1c

Observation 0beb63b8-ef8f-4cca-a6c2-26501b0a1cd7 · outbound

This paper cites Soft actor- critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Soft actor- critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.584294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.584294Z digest=sha256:1c966578376ec4b184859ba736ec48edb07a64c3d037e000c8362de5e0eaa17d

Observation c6c14dc3-1631-44ad-a571-876a55ac214f · outbound

This paper cites Extensive Exploration in Complex Traffic Scenarios using Hierarchical Reinforcement Learning.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Extensive Exploration in Complex Traffic Scenarios using Hierarchical Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-05T06:00:50.800160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.589688Z digest=sha256:61c82d9fa983bfb20d0df1b352e4ccd7d2887b318a7141ef92a9867a9e5b98be

Observation f30b5aa9-2645-4da6-b0cb-3c763ee0f7fa · outbound

This paper cites Exploiting hier- archy for scalable decision making in autonomous driving,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Exploiting hier- archy for scalable decision making in autonomous driving,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.241149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.595194Z digest=sha256:09a0f922596caeda8107166d6035530ceb105ec3fbf1905ed7daacf3600e3254

Observation 82cfd841-8194-4b4f-a592-e118644d4a2f · outbound

This paper cites Learning hierarchical behavior and motion planning for autonomous driv- ing,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Learning hierarchical behavior and motion planning for autonomous driv- ing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.214963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.600298Z digest=sha256:1016d42bdbba2efdb897636203f0e61309b0040ec859c87ba118d1c4ad0b8dea

Observation 442e7056-f63b-42d2-9cec-8e1fe35220f9 · outbound

This paper cites Deep hierarchical rein- forcement learning for autonomous driving with distinct behav- iors,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Deep hierarchical rein- forcement learning for autonomous driving with distinct behav- iors,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.195558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.605747Z digest=sha256:fcf637c2f75e931fa388ffbed5e9a8f5922fb7b7b31b2f6e7cf87ff7ac98420e

Observation 9530e4af-1363-4bc5-8a88-e4d940d0f112 · outbound

This paper cites A rein- forcement learning approach to autonomous decision making of intelligent vehicles on highways,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving A rein- forcement learning approach to autonomous decision making of intelligent vehicles on highways,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.173646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.611238Z digest=sha256:efb1de24f8fdae1d8df9bb43edd7f1671385403e34250d62da53bdd09eab54f9

Observation 9b9c98ce-ed79-45fd-a6dc-46550f4f3f75 · outbound

This paper cites Integrating deep reinforcement learning with model-based path planners for automated driving,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Integrating deep reinforcement learning with model-based path planners for automated driving,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.153412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.616865Z digest=sha256:b5dc9b1156be1bc50919e79125b4b324b417162f74391986e1221bc6a2058fd9

Observation 3a1a448e-32d0-4c06-8cad-5fbcedbb70ff · outbound

This paper cites Lane change decision-making through deep reinforcement learning with rule- based constraints,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Lane change decision-making through deep reinforcement learning with rule- based constraints,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.131439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.622546Z digest=sha256:8fe9bfccf406622d2988d4ba0b5d12abf78041a35952b051f60d3cec36efeeaf

Observation 202b78e6-34ee-4f55-b85d-e941086fc858 · outbound

This paper cites Combining reinforcement learning with rule- based controllers for transparent and general decision-making in autonomous driving,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Combining reinforcement learning with rule- based controllers for transparent and general decision-making in autonomous driving,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.110195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.627359Z digest=sha256:a19d0e78358d7020334365c4d8deb17c2fc30e3f64b302a5197fb5fcc827130a

Observation 6948b72f-e7b5-4b3e-a656-133b06883507 · outbound

This paper cites A combined reinforcement learning and model predictive control for car- following maneuver of autonomous vehicles,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving A combined reinforcement learning and model predictive control for car- following maneuver of autonomous vehicles,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.087262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.632114Z digest=sha256:c484fcb6a50c9c8fdd8685f0924f29ed9f664f99c24a19943cc7b5bf1da69952

Observation 29d92c44-477d-48f7-b5f3-3535411ff1ed · outbound

This paper cites Combining reinforcement learning with model predic- tive control for on-ramp merging,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Combining reinforcement learning with model predic- tive control for on-ramp merging,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.064224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.637756Z digest=sha256:eef820617044412d99f5decd0958b9f73a1b529b964c8f7c4d9ad208a40a4ccb

Observation fb03a919-7a64-4c50-8a76-44a78b66fd6c · outbound

This paper cites A Hierarchical Architecture for Sequential Decision-Making in Autonomous Driving using Deep Reinforcement Learning.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving A Hierarchical Architecture for Sequential Decision-Making in Autonomous Driving using Deep Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-05T06:00:50.774633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.642885Z digest=sha256:ca75a689ec090445f594d152ed0cff6cc9cee78d106b8fcc8bbef37963cdf132

Observation 94b4fe8e-3b93-44d9-8ac6-7d5ae5304ca8 · outbound

This paper cites Driving decision and control for automated lane change behavior based on deep reinforcement learning,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Driving decision and control for automated lane change behavior based on deep reinforcement learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.043596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.648314Z digest=sha256:6a42a81bbdf8bd0d65eaa69e3d2bdb884484d29760a8cbf34eb3f47dd9af1e93

Observation 2ad53c79-efa3-406a-94ca-e017a3af85d7 · outbound

This paper cites Prioritized experience- based reinforcement learning with human guidance for au- tonomous driving,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Prioritized experience- based reinforcement learning with human guidance for au- tonomous driving,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.024417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.653383Z digest=sha256:33a30f3acc95e659b7f1b5edd41f0c8b8dab36772743721f7a0b6f3bfbdfb520

Observation 27b91336-1d58-4653-afa1-78e13170692f · outbound

This paper cites Learning to drive in a day,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Learning to drive in a day,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:51.002152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.658528Z digest=sha256:6f55a394be162e3c6b071a3f971a7edec0a6263d8de2dd1bdef4affd65a5c812

Observation 3b472d54-e28a-4949-b19b-de8fe36f6896 · outbound

This paper cites Efficient deep reinforcement learning with imitative expert priors for autonomous driving,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Efficient deep reinforcement learning with imitative expert priors for autonomous driving,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:50.978928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.663239Z digest=sha256:9f0950baa7ce08d6fb573707fc711d5c3a61195cc253841345f56b6047fa35e8

Observation 6bda6f56-dce9-46bf-9fa5-8cbe3c6a03d7 · outbound

This paper cites Imitation is not enough: Robustifying imitation with reinforcement learning for challenging driving scenarios,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Imitation is not enough: Robustifying imitation with reinforcement learning for challenging driving scenarios,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.668241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.668241Z digest=sha256:82140eee4fed2f9a4c5bc8e39d7fe37799c239dc9f323ff32022aecf77a66bba

Observation 19369424-5080-40a3-b3f4-36d24e2bd34e · outbound

This paper cites Boosted bellman residual minimization handling expert demonstrations,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Boosted bellman residual minimization handling expert demonstrations,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:50.949614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.672879Z digest=sha256:b6a9bb96b3171fac68a47e0fd065df7a9f2b3b5248831634fde7505471da657e

Observation 46d406f3-6250-4d75-a145-16ba275627c3 · outbound

This paper cites Deep q-learning from demonstrations,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Deep q-learning from demonstrations,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.678081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.678081Z digest=sha256:6b7f1891c1d981ff23f7b808b4f03e4c51e6f15ffb87d949b28631e87a96a16f

Observation e453e3c0-12a8-4260-b396-4acb29cf25ab · outbound

This paper cites Reward learning from human preferences and demonstrations in atari,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Reward learning from human preferences and demonstrations in atari,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.683143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.683143Z digest=sha256:a8cf780bc944e2b9c80170dc2df7f00214a65539f418d2b2778f0d63db5db2c6

Observation 10074819-6656-4277-8120-92136c49b0d5 · outbound

This paper cites SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.687992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.687992Z digest=sha256:bf1b24cefd19d522a5c027bcda2e35e045142892d551d04a6155c15292a6f33c

Observation 55434cac-0246-4a85-b7ec-c8d4ce512d71 · outbound

This paper cites An environment for autonomous driving decision- making,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving An environment for autonomous driving decision- making,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:50.907365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.694491Z digest=sha256:3ff6d10f1b118b403cde1a347279a64df9a8def979f46a35458e0b2146ae681e

Observation 319d0981-1e94-4a81-a857-42912ccd75b6 · outbound

This paper cites Conservative q- learning for offline reinforcement learning,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Conservative q- learning for offline reinforcement learning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:50.887294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T06:00:50.699611Z digest=sha256:3e467da1081480326960277abe016937fc674f33ec7a09524b3ae702019d73c9

Observation fc2ea35e-942f-4094-a434-f69b9e4cf02c · outbound

This paper cites Generative adversarial imitation learning,.

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving Generative adversarial imitation learning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:50.704572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:50.704572Z digest=sha256:d75fed7a4ee14179b6041fbffe191d6ec9641fd0612cc654a25f0fba47167a0e

Pith citing papers

No inbound Pith citation observations are available.