Pith. sign in

Paper Citation Record · LEDGER

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2506.03568.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03568 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:47.517274Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b953d159-f33d-4b89-859a-3537b2c5e1ab · outbound

This paper cites Learning to drive in a day,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning to drive in a day,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:49.081032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:46.996668Z digest=sha256:595b01ca7ba5dc9eb46f2dda777929d586f0d5e65a191d98eab6db5da666ce60

Observation ea909278-58b2-43ac-9010-e526a626c62c · outbound

This paper cites End to End Learning for Self-Driving Cars.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving End to End Learning for Self-Driving Cars

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.054690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.054690Z digest=sha256:42253f8cd5b9aec6eb8723cee0c5606e4ef4def5d88357d26e9a54be98881b82

Observation 26b284c1-6c32-4c92-bd57-e8d5f8d7c391 · outbound

This paper cites Dense reinforcement learning for safety validation of autonomous vehicles,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Dense reinforcement learning for safety validation of autonomous vehicles,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.062703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.062703Z digest=sha256:fe8644510f0f06a35f51abe13c5d42cb2aa5dee21a4dba75db47828ef385fae5

Observation b9599d1f-fd6d-4551-a9b6-ad054c3bd309 · outbound

This paper cites A survey of deep RL and IL for autonomous driving policy learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A survey of deep RL and IL for autonomous driving policy learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:49.016724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.074645Z digest=sha256:a10f1e944a387f5fda706da4fe155ab49382f31ab72d701304fe1ddc37f52060

Observation a10452b6-967f-4279-b900-5e0790312f10 · outbound

This paper cites A survey on autonomous vehicle control in the era of mixed-autonomy: From physics-based to ai-guided driving policy learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A survey on autonomous vehicle control in the era of mixed-autonomy: From physics-based to ai-guided driving policy learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.994054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.080690Z digest=sha256:81dab856cab1fabb89428d706c77d1297e99ef32a8de0a53771e1157a1b17364

Observation d6c01a77-4c2a-4f80-91d0-67b85257cbb5 · outbound

This paper cites Deep learning for safe autonomous driving: Current chal- lenges and future directions,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Deep learning for safe autonomous driving: Current chal- lenges and future directions,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.087129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.087129Z digest=sha256:77ffa8600bc957e1ad44a31f21fc0f6ccd60fbaca647cffe627bb58dc676d8e8

Observation 70e94c2b-9b6d-496f-9258-84821eaebf43 · outbound

This paper cites Survey of deep reinforcement learning for motion planning of autonomous vehicles,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Survey of deep reinforcement learning for motion planning of autonomous vehicles,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.092742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.092742Z digest=sha256:2f3a143f98ce509092d9daccbf95de3177856a9154fe1306acb8b1a557b807ca

Observation 33fc5d1f-f2a8-4f43-b19f-8a3c79b377b0 · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.099754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.099754Z digest=sha256:a024ea6e40c70e7364e400487b4416967e307913b93705df86a6adaa1d0b2014

Observation fbac7cfb-67d7-4cd9-9adb-34ae7100f6ce · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Deep reinforcement learning for autonomous driving: A survey,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.938561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.108862Z digest=sha256:906fc3fa9ba29c41fb81996d4f2f8f6544abbd9b3da0dd1a6d843572f383ce56

Observation 2a382cf7-ee65-43e4-a623-303b4ec4c4fa · outbound

This paper cites End-to-end urban driving by imitating a reinforcement learning coach,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving End-to-end urban driving by imitating a reinforcement learning coach,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.913663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.119541Z digest=sha256:3eabf71c5c49a92d7fff7e552b12be4f5f80c1fff8bb388a2305a05726cb1146

Observation 69b0d92e-7f70-4794-a84a-fc54ae364c6a · outbound

This paper cites Reward misdesign for autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Reward misdesign for autonomous driving,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.887744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.129878Z digest=sha256:f09867a5689b22cc798a92bdab992fe1bbda18d6736758327aa7a1520ec4949e

Observation 5d2bb129-4ecb-4225-92d3-85b487ae0178 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Scalable agent alignment via reward modeling: a research direction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.141529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.141529Z digest=sha256:981ab0bd5577f696ea675c1fb992531c7f620e80653887e5cb020580ebd9fc8f

Observation c5789285-6345-42fd-8fa0-d698171dc58a · outbound

This paper cites Demonstrating specification gaming in reasoning models.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Demonstrating specification gaming in reasoning models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.152726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.152726Z digest=sha256:9766ed90430bcf455ac75c3a1085c20dbc244291dca1a35f6f47c3809ca06e6e

Observation 3ecdb816-979a-4115-8595-9a98158cc2fd · outbound

This paper cites Toward human-in-the-loop AI: Enhancing deep reinforcement learning via real-time human guidance for autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Toward human-in-the-loop AI: Enhancing deep reinforcement learning via real-time human guidance for autonomous driving,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.861260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.159984Z digest=sha256:4b2fb3394d0a14d5bb1c29931632d6ff6fddc28e9dfb044d38868d1d396e73c2

Observation e200d637-9a49-4c9b-bcf8-f7d3425a8012 · outbound

This paper cites Hindsight credit assignment,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Hindsight credit assignment,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.841270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.176862Z digest=sha256:3c1fcba5b20c4a44ecc715e79dd737196949bd10c53f9655830117dcf0dbdd98

Observation 37afddd2-4f33-41ff-8d6a-2707ab0e72bc · outbound

This paper cites Trial without Error: Towards Safe Reinforcement Learning via Human Intervention.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Trial without Error: Towards Safe Reinforcement Learning via Human Intervention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.203825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.203825Z digest=sha256:7350fe4f9d073f2a7430937f89ac1c09374816bd9e30def9f81703a61f63eba7

Observation 7435c22e-26b8-487a-b920-9a468755d7a3 · outbound

This paper cites A survey on imitation learning techniques for end-to-end autonomous vehicles,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A survey on imitation learning techniques for end-to-end autonomous vehicles,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.820149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.211989Z digest=sha256:5b84c842122664e1784e253918bf2d4b376d8962404a7801bd8445167a0bdded

Observation 43b0205f-a0b0-4a5c-84ba-28463da293fb · outbound

This paper cites Conditional predictive behavior planning with inverse reinforcement learning for human-like autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Conditional predictive behavior planning with inverse reinforcement learning for human-like autonomous driving,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.790007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.224555Z digest=sha256:334c6a6e7eac012fcf6a63b93de4a600f2d47651e3a965d419604fb3ee4adeaf

Observation 52d91390-1107-4988-9158-bdb23e1d15f6 · outbound

This paper cites Pattern recognition and adaptive control,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Pattern recognition and adaptive control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.747424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.230363Z digest=sha256:9dedb5d9ac246e544ba9cf26932ac02b595cf212f555876308268b45ba23e1e9

Observation a4696d27-913a-475b-a557-d457f26f83c7 · outbound

This paper cites An Algorithmic Perspective on Imitation Learning.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving An Algorithmic Perspective on Imitation Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.235938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.235938Z digest=sha256:e1956c20f1f78fa241e5c93c9f7ad6dd06a3f2ceb65fe903cf48557aacdfd3d7

Observation 2f971b72-1e49-456d-8b9a-ff91feb8d2ef · outbound

This paper cites Learning a decision module by imitating driver’s control behaviors,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning a decision module by imitating driver’s control behaviors,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.726411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.242273Z digest=sha256:8a90aeb3387ad73289dd4b1b25b45e756e4f3f241e77c19c54edec3236aa6a14

Observation 33693b59-16e0-45ff-8b2c-7ff49c725d71 · outbound

This paper cites Conservative Safety Critics for Exploration.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Conservative Safety Critics for Exploration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.251327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.251327Z digest=sha256:43aa063a8ae7acc6b4c1f62a256df6435f70f3fd297d9c844af55de415f4d7bd

Observation 9ab2d0dd-d2b1-4721-b75a-7bfb2264ad49 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Behavior Regularized Offline Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.259634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.259634Z digest=sha256:dfa289f8febdea6166addc0be301130e86c2a3bf1f5dbc71c86628eb4a841be8

Observation 7e335325-1715-4ac0-a391-11f2cc86de03 · outbound

This paper cites Off-policy deep reinforcement learning without exploration,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Off-policy deep reinforcement learning without exploration,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.703985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.267759Z digest=sha256:a8328f76c122a2d91e22d34a1051155972f5a05be7cc1eca3baa5ba9e827dd87

Observation 0098b27e-2e4b-42a2-a36d-324d9d0fd4e7 · outbound

This paper cites Adversarial inverse rein- forcement learning with self-attention dynamics model,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Adversarial inverse rein- forcement learning with self-attention dynamics model,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.675772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.274511Z digest=sha256:da6cbbb34da2a98ebe3ae4c7c5a4781577c6e55d492378a550cb313adb439855

Observation 9434015b-8421-4144-b37f-18da759de9d6 · outbound

This paper cites Efficient reductions for imitation learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Efficient reductions for imitation learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.648437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.301201Z digest=sha256:9890c3c9649209b666475982534ccf999cdf26a0deddd95523a03d0502557028

Observation 4442ef43-f0c1-432b-9059-456c2821fa2c · outbound

This paper cites Exploring the limi- tations of behavior cloning for autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Exploring the limi- tations of behavior cloning for autonomous driving,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.625527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.314152Z digest=sha256:3167ca9fe3dde4a62453690748d20a1ca1303ac711928b8c4697e3683ce429c1

Observation 6bf27769-625a-48d1-ab28-945048db412f · outbound

This paper cites Behavioral cloning a correction,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Behavioral cloning a correction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.590978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.320883Z digest=sha256:dd29b64c5804ff3ab020d21dfdc189d594e25cd0277d8f13abdea9bf9bffb9d6

Observation ec8175b5-24a6-42d4-bdcc-c042ea77d078 · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.561472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.328632Z digest=sha256:e4a87a63fd5808a7774c327e1e1935acc61b03a8a1bd3cc8433c88f8c1cd7ebb

Observation 024cbaec-61db-4023-908b-e53548811411 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A reduction of imitation learning and structured prediction to no-regret online learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.533279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.335160Z digest=sha256:7511158c5499f64c485ba07e654ad4173879b227123b0b2c1cd0c4cb127ee09a

Observation f1763acf-d4a0-4831-a456-b69e1506e204 · outbound

This paper cites Query-Efficient Imitation Learning for End-to-End Autonomous Driving.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Query-Efficient Imitation Learning for End-to-End Autonomous Driving

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.345392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.345392Z digest=sha256:5680731aa5df1507e5f168da2c2aad25150f8997e96abdb9b147589313503005

Observation 70cfac0c-e7e9-4dfc-a259-3032037c6108 · outbound

This paper cites Hg-dagger: Interactive imitation learning with human experts,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Hg-dagger: Interactive imitation learning with human experts,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.354950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.354950Z digest=sha256:b8bf3c4cc8c76a5928783d81f5646c835b2a68042fb7cea5335402e313eb3855

Observation 8d0258af-51c5-4dba-a66c-2a63b26941c0 · outbound

This paper cites Thriftydagger: Budget-aware novelty and risk gating for interactive imitation learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Thriftydagger: Budget-aware novelty and risk gating for interactive imitation learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.489799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.379002Z digest=sha256:9c0d33f195d6f3bd6de31fd6b4863edcae9579378eb241d8ea4ea132d0efbda1

Observation 9a20d808-d509-43e5-b925-2926ac9b8c0d · outbound

This paper cites Expert intervention learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Expert intervention learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.467741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.384324Z digest=sha256:4d584b06bc19d7c537c22dd5a648eaa7a5a694021cbbf1c76141b6212b68c8a4

Observation 93b1d758-f991-4c18-92d1-157a51a2a22b · outbound

This paper cites Human-in-the-Loop Imitation Learning using Remote Teleoperation.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Human-in-the-Loop Imitation Learning using Remote Teleoperation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.390201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.390201Z digest=sha256:10f286d26cf77e90075cca80b695644a4b71800546b2156c3947d0d3e617c768

Observation 5cc54a30-2cf5-4e27-bca1-c6b64468a705 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Deep reinforcement learning from human preferences,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.434427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.398761Z digest=sha256:09a0dfd5e0c17bf729294e6443f90d21352b39cba25b0774a7f1620329bc9bd1

Observation 558d2604-d8a2-45cd-96b4-1858083f53f1 · outbound

This paper cites Batch active preference-based learning of reward functions,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Batch active preference-based learning of reward functions,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.407579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.404330Z digest=sha256:dea5f7e7c2b7409a659621f7642d39e59e1f75eae7051ca947c0a16b44146fa7

Observation 64f9f942-7044-4d1b-b85d-3446d11595d6 · outbound

This paper cites Learning Reward Functions by Integrating Human Demonstrations and Preferences.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning Reward Functions by Integrating Human Demonstrations and Preferences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.410249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.410249Z digest=sha256:2f2651c7ddad505165218967b0d1cd6a7ea95a7afb78c11203584cfb7bb6e9e1

Observation 67e09e0c-5b81-46f2-8a1d-49f877c648b0 · outbound

This paper cites Efficient learning of safe driving policy via human-AI copilot optimization,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Efficient learning of safe driving policy via human-AI copilot optimization,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.378955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.416144Z digest=sha256:82e20b3180bff54c8e89b17e0593cb4af0c3fe7daaca4f2494139430722b6bdd

Observation 67936181-90d0-495a-866a-28f330be8d8c · outbound

This paper cites Learning from active human involvement through proxy value propagation,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning from active human involvement through proxy value propagation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.353963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.421979Z digest=sha256:a83c16f5ea3667d43c71cfa28a7b565a901e7facef2800a44a78b30073a770b5

Observation ace12819-7be9-4dbb-8251-53a121e283b2 · outbound

This paper cites Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.427864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.427864Z digest=sha256:2bb2155742b69536032fab6ef45472553951ba4fe13bba58c1b5f2d95d371ad4

Observation 355d23c3-b0f1-4007-80e5-e2862f3e8c3b · outbound

This paper cites Socially situated artificial intelligence enables learning from human interaction,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Socially situated artificial intelligence enables learning from human interaction,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.332758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.433857Z digest=sha256:b0c5932ca882dec9a66cc973b54d143561dd72dfc37ed18964af14034c90b4c5

Observation 268b589d-5665-465a-8013-6846a2e03e11 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.301583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.439626Z digest=sha256:acef85412a7d5652dc9b1c95acdd5f477300094209dc5911801bb418363ac451

Observation 09d7c262-872f-4a76-b8bb-ab18800494a0 · outbound

This paper cites Human as AI mentor: Enhanced human-in-the-loop reinforcement learning for safe and effi- cient autonomous driving,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Human as AI mentor: Enhanced human-in-the-loop reinforcement learning for safe and effi- cient autonomous driving,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.271114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.446612Z digest=sha256:4f116bf91d8b79d74adaea2e665f8c5347f88b66700d59d7934b5aeceef93838

Observation dd045378-b146-4e55-9cf1-b5559197a214 · outbound

This paper cites Guarded policy optimization with imperfect online demonstrations,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Guarded policy optimization with imperfect online demonstrations,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.234155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.451753Z digest=sha256:3fee1ad3a10209e66cc7da0a65799ef92ab861bdf01ed3aa80c78ce8e0e153af

Observation d56724d3-6b63-4773-84a8-f20f074ef99b · outbound

This paper cites Trust region policy optimization,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Trust region policy optimization,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.209925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.457418Z digest=sha256:bc0a916089b8844ef293bc4d533e6de3b2c4a1e688f47982c9145791aa12539c

Observation fbc2e929-024c-40ff-8841-12f2dfc434be · outbound

This paper cites Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.176999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.462594Z digest=sha256:9aefd5ffdea504a540fb6fadb99802c0f932c395cf32bf34399f2f4bcaf34ec1

Observation b46c8c8a-366e-447a-bee1-d91116c5d1d4 · outbound

This paper cites Responsive safety in reinforce- ment learning by pid lagrangian methods,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Responsive safety in reinforce- ment learning by pid lagrangian methods,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.150462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.468686Z digest=sha256:e3842d1982e4354fca678cb897a5567a3ea0725462dd9c3b7b4e849221315883

Observation 8d55fd2c-4562-4271-8430-f18d391150b2 · outbound

This paper cites Learning to Walk in the Real World with Minimal Human Effort.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning to Walk in the Real World with Minimal Human Effort

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.474001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.474001Z digest=sha256:7f420791a23e2ae084282392853d7c1d86456754a92bc2f91acca0cf125ff5fd

Observation f771110b-c8bf-4dda-8076-11d34b2454a5 · outbound

This paper cites Conservative Q- learning for offline reinforcement learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Conservative Q- learning for offline reinforcement learning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.125073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.479725Z digest=sha256:b7e705ef9821d6db00a3fe91f012479dde089c1cc9e78ae7585a6e2f0609b72a

Observation 2b598292-45f2-4e28-9d94-55dfe16f77f7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Proximal Policy Optimization Algorithms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.485274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.485274Z digest=sha256:1a8e625e0b5c9c207605d0e7891ccee3220fc9f2c528aab348854902b28be9c4

Observation 389de93b-200b-4f6b-aa22-ce32ef49726e · outbound

This paper cites Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.085564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.497444Z digest=sha256:cdd0501928505a9330bf74580597f395b6835119246cf8ddff024480dbe252f4

Observation 0b2a3f9b-c5c4-4ba3-8bb6-24d538ca433f · outbound

This paper cites A framework for behavioural cloning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A framework for behavioural cloning,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.059247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.504779Z digest=sha256:25f4a71a74ae340969cc36c42415a278b21ad0c67489ccfae7a7f224f3552974

Observation ee91778d-321a-480a-a9dd-e402f4b66a43 · outbound

This paper cites Generative adversarial imitation learning,.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Generative adversarial imitation learning,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:48.037949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.511187Z digest=sha256:18e024b461fd042c1f13ecf2e116c0fa450b4db568545b19bc287f7fd7a96c34

Observation ec87f5b9-e218-4ec9-9066-043b308dffc2 · outbound

This paper cites an unresolved cited work.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:47.999906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:47.517274Z digest=sha256:47404f40d41451b10f3af3f6e852cb6dfa8c71484e4f24e43f1237aa597eb64f

Pith citing papers

No inbound Pith citation observations are available.