Pith. sign in

Paper Citation Record · LEDGER

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

As of 20 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2606.26527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.26527 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T05:24:53.844059Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf04b379-a558-485b-9187-494c8668fe6d · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Deep reinforcement learning for autonomous driving: A survey,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:3bc1134c6448ea795bae0add2561939c004378d09c441caa7213764e6f04f37a

Observation a2e17af3-df6a-4c55-a763-0e71219471f5 · outbound

This paper cites Safe reinforcement learning for autonomous lane changing using set-based prediction,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Safe reinforcement learning for autonomous lane changing using set-based prediction,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:15407966fa032365db191cffec81a4d45b6f5118f27309fb27072b40c6976770

Observation fce97bd5-f466-400d-af0c-9d9953b00104 · outbound

This paper cites Unsupervised reinforcement learning for multi-task autonomous driving: Expanding skills and cultivating curiosity,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Unsupervised reinforcement learning for multi-task autonomous driving: Expanding skills and cultivating curiosity,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:62428698d75774dd5d9b1d7e6ac4fcd2ae2efb9476c14c18240b3beed6a35947

Observation 5c14c1ad-31b6-4e3d-92ee-058642c07588 · outbound

This paper cites Driving tasks transfer using deep reinforcement learning for decision-making of autonomous vehicles in unsignalized intersection,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Driving tasks transfer using deep reinforcement learning for decision-making of autonomous vehicles in unsignalized intersection,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:068e480041ff61afd4b77b5ba0df679b5cdbcd184688f18c6c9a4d05e42c04a2

Observation e3d9fff5-edaa-4ff9-9209-7d806b0a3493 · outbound

This paper cites A perspective of q-value estimation on offline-to-online reinforcement learning,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy A perspective of q-value estimation on offline-to-online reinforcement learning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:05231a6af0c9df519977f4f87ab6f32f2cdf0a6b5aef033dd890595553818571

Observation c4d5d787-db8f-43ce-b0e2-7e1a84d496e7 · outbound

This paper cites Sim-to-lab-to-real: Safe reinforcement learning with shielding and generalization guarantees,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Sim-to-lab-to-real: Safe reinforcement learning with shielding and generalization guarantees,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:5175428932594a49409f254901dc35fa721ce629d550ce765a7bbfd83f02f92e

Observation 591fc16c-fcbc-446b-93fd-438e2e3138f4 · outbound

This paper cites Knowledge transfer from simple to complex: A safe and efficient reinforcement learning framework for autonomous driving decision-making,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Knowledge transfer from simple to complex: A safe and efficient reinforcement learning framework for autonomous driving decision-making,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:b9e7f007317e0c9dd6517d4f34a687ddae37893914d98130f0572a3453eadcf4

Observation b59c97a9-9156-47a6-a5bd-e86894a34787 · outbound

This paper cites Zero-shot deep reinforcement learning driving policy transfer for autonomous vehicles based on robust control,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Zero-shot deep reinforcement learning driving policy transfer for autonomous vehicles based on robust control,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:f03d82fd2566df66dd2787da788f02d1ec0297b3dd2ee6e7899a1a850db0744b

Observation 51924c67-6b7c-436b-b71b-447134fbb78f · outbound

This paper cites Safety reinforcement learning control via transfer learning,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Safety reinforcement learning control via transfer learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:c46223338dcc6d5e0f6547cfb078721d8deba7990f71f2fcc51cd4315b65d3fc

Observation 2b113a94-afcf-4c4a-9f7a-6b229e3bceb3 · outbound

This paper cites Federated Transfer Reinforcement Learning for Autonomous Driving.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Federated Transfer Reinforcement Learning for Autonomous Driving

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:51.461041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:1c7b2aae755cd3c1eae98bc3955f87b781aaf55e5a4311ea619fe8e4b7638853

Observation d6a73046-0631-4f66-b92c-e81d32f3f35a · outbound

This paper cites Scenario- level knowledge transfer for motion planning of autonomous driving via successor representation,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Scenario- level knowledge transfer for motion planning of autonomous driving via successor representation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:1b61c9ba5b53e055db64a0097821f99fc77fb7795a6cd5b2efa40500d7987397

Observation 9c83d9e7-d7dd-4c78-b75e-f76744b697c4 · outbound

This paper cites Self-supervised domain transfer for reinforcement learning-based autonomous driving agent,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Self-supervised domain transfer for reinforcement learning-based autonomous driving agent,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:f06ef1f7c5f0e972921714adbb8f0ef989c6dfea2eb26b9f62ff77138d0e0c66

Observation ece1c202-224d-45b4-8661-ddc911f8f7ef · outbound

This paper cites Cross-domain adaptive transfer reinforcement learning based on state-action correspondence,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Cross-domain adaptive transfer reinforcement learning based on state-action correspondence,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:357386560ce4730995023b549e72200c73a13bf3d66b55802c61ddf3e845a237

Observation c6cc2845-4143-4127-9651-c2d2b5564dfc · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:09:51.465730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:29f7fc9fd5834abc9e56d0ba07c323bb72b30c0b7bc79dc9813123f3c95929e9

Observation 51573923-6671-471a-8f60-903a588abe1a · outbound

This paper cites Policy optimization with demonstrations,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Policy optimization with demonstrations,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:95b19b01d2c8de91dabbd89d851fcd4ed58ed7c0278e78eb60f618c26fede654

Observation 7506a50e-5af8-4a9c-b942-0fa355e5c99c · outbound

This paper cites Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:09:51.463357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:f61059191db2356b5f36e00ae209f661184702d97ffabc817c809681a633c66a

Observation 5a818391-6a13-46d5-baae-a766260f1c2a · outbound

This paper cites Knowledge transfer for deep reinforcement learning with hierarchical experience replay,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Knowledge transfer for deep reinforcement learning with hierarchical experience replay,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:7c5f57be792054a694b5d7c9f2b6e9dabd7e90dbb9f6fd6d9d2c0444f73a6298

Observation dd322c82-940b-4847-9f50-009e2b09887b · outbound

This paper cites Improving reinforcement learning with confidence-based demonstrations,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Improving reinforcement learning with confidence-based demonstrations,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:544391b4f3cfed6f0bcec2fc12d56e4028ba5b82e4393181df3a05fc536c8677

Observation ac78be33-341c-4ed4-9482-7bc04e89ead8 · outbound

This paper cites An enhanced advising model in teacher-student framework using state categorization,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy An enhanced advising model in teacher-student framework using state categorization,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:0335daa5fe5bb56d6c56765469e0614593c5d641bf174319402ab31bbf1d62ba

Observation 35443eb2-b724-492f-a2a7-c1e5e229d202 · outbound

This paper cites Human as ai mentor: En- hanced human-in-the-loop reinforcement learning for safe and efficient autonomous driving,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Human as ai mentor: En- hanced human-in-the-loop reinforcement learning for safe and efficient autonomous driving,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:92df1e6059b177c688bb4b0c0edef4aa37688315f4b5976e1a7dac966b886e54

Observation 5ff7d50e-9183-4796-a818-4f3a7854372d · outbound

This paper cites Adaptive action advising with different rewards,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Adaptive action advising with different rewards,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:5cea077881ac369744e480bc878fec78c9dbc498e4a120444c9f2ba68da5d68f

Observation 44423618-ab82-467b-9f54-4ef25c0745a1 · outbound

This paper cites Safe reinforcement learning via shielding,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Safe reinforcement learning via shielding,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:d2350a35ac0cc0c74b4917bba44cbfcb20d9c048537dfe7de40889119b1129aa

Observation 9cbaf01e-e927-48f0-89f0-76929bceba3f · outbound

This paper cites Safe reinforcement learning via shielding under partial observability,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Safe reinforcement learning via shielding under partial observability,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:6f02c807fd2c65056b05c56499965e4780a592c3bb408f01b6987f0db7f2b063

Observation 2a9008c0-cc72-42e4-b01f-dee595077db0 · outbound

This paper cites Robust model predictive shielding for safe reinforcement learning with stochastic dynamics,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Robust model predictive shielding for safe reinforcement learning with stochastic dynamics,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:73f4acc05853c63b82d962c834649790bc7786552ce88210fcc645bc912ba38a

Observation 091bdc8c-596e-4a05-96d4-1c8a5475428f · outbound

This paper cites Teaching on a budget in multi-agent deep reinforcement learning,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Teaching on a budget in multi-agent deep reinforcement learning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:9b70eacc645ec170996208a6b863dd367dea7d42a6b1439f01200dc9acbe853c

Observation 09cf7116-3b99-46f5-b7f7-4c955abcde58 · outbound

This paper cites Action advising with advice imitation in deep reinforcement learning,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Action advising with advice imitation in deep reinforcement learning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:02f8505ad23422a3902204e5b1dd02b743ea35eb47a0f14326a00f66c2437122

Observation 1819784a-bfcb-4a02-a127-2312f6c6190f · outbound

This paper cites Reinforcement learning with demonstrations from mismatched task under sparse reward,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Reinforcement learning with demonstrations from mismatched task under sparse reward,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:d7189ba707724695d8da0d471e8def13847e448072f51202a7dd8629ddc05898

Observation a5e4480f-bf24-4c73-bc58-408d5628c2b3 · outbound

This paper cites Psiphi- learning: Reinforcement learning with demonstrations using successor features and inverse temporal difference learning,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Psiphi- learning: Reinforcement learning with demonstrations using successor features and inverse temporal difference learning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:4ba60a858539989c5bac43ba3a4acb22b1898f2b80d4787c6cd5a121e82fb011

Observation ab690933-2014-4bb6-b2bb-e2e0fb4f15eb · outbound

This paper cites Hybrid reinforcement learning with expert state sequences,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Hybrid reinforcement learning with expert state sequences,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:e8d5887b34c2c03552e8a66303c1d77ba89043f386f30492d263d602b1317940

Observation a1854cbe-9c52-4d5b-85a8-f8940ea56f1c · outbound

This paper cites Guided exploration with proximal policy optimization using a single demonstration,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Guided exploration with proximal policy optimization using a single demonstration,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:3175fdb158f8c7c5b9702c124d6d3247c6db74ad7759f4e0b62c718a68371d9a

Observation 251ee9eb-b1ee-4b4c-877d-25ae7f1b78be · outbound

This paper cites Hybrid rl: Using both offline and online data can make rl efficient,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Hybrid rl: Using both offline and online data can make rl efficient,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:a62577fd9bdb85973b5afbb74f1f319d421f5a4e48e452e052f56cdeeac4c57c

Observation 4041bed7-64c6-4b9e-82db-c2f9f345862b · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:be31a4749f9a84ebf3dd6d1993ec3f230efde8f5f830de290455aa018989feb0

Observation 4f2ca2ef-c774-4ebf-a232-e93879ef4e7f · outbound

This paper cites DCUR: Data Curriculum for Teaching via Samples with Reinforcement Learning.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy DCUR: Data Curriculum for Teaching via Samples with Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:51.468188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:30f1a4b1081b14db057ef7a0697839363c1f26b0f6d3e122774635e5f9ad13b0

Observation 5ab869bf-1fba-4c0a-9ffc-792e8f05443d · outbound

This paper cites An actor-critic algorithm for constrained markov decision processes,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy An actor-critic algorithm for constrained markov decision processes,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:1bbab12b5c9add0d38c1eb2e9601753e35da4e0fa61e5d062a8368d0438b299c

Observation 0578096c-f9a8-494a-928a-7a189c9dbeef · outbound

This paper cites Reinforcement learning by guided safe exploration,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Reinforcement learning by guided safe exploration,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:21f3055c94743e8c5f1b575fac5825a6d2eaffe360b81ae50c73e6c161dbf19a

Observation c2a8f051-75fd-4428-b5b5-5fb506f68ea7 · outbound

This paper cites Guarded policy optimization with imperfect online demonstrations,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Guarded policy optimization with imperfect online demonstrations,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:a25a3d6b5076891f53f61d4226434055938df2ab1cfee168e961ea711a6c5dcc

Observation 2d22defa-47d6-4af8-a64f-17b48a3c876b · outbound

This paper cites Approximately optimal approximate rein- forcement learning,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Approximately optimal approximate rein- forcement learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:98e0053c7efc6cd9ef9b4a5f273ea9e1b5b32181d7302484846b92b9a891fcdc

Observation b87f193d-c0a9-4091-a0bc-7c65dd6697f1 · outbound

This paper cites an unresolved cited work.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:b2b595ee787052ba0cd573337ff63043e3b6ef91212ef9591871557d7caa106c

Observation 3c18163e-bda3-44bd-9f4c-8dd0fd2778b3 · outbound

This paper cites an unresolved cited work.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:7c123469a6a206a7fd48e47cb02131fd3ff2632e899fff644d0b9e8896fe88f3

Observation ed8ab7f1-2817-4cbc-bfed-d933c3328404 · outbound

This paper cites An environment for autonomous driving decision-making,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy An environment for autonomous driving decision-making,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:0d8d432f208449070147a94bf9ad83c19e679144978c9fe1cc1072d187ee7098

Observation 68a96287-ca84-4736-9699-37832c968bf5 · outbound

This paper cites The kinematic bicycle model: A consistent model for planning feasible trajectories for autonomous vehicles?.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy The kinematic bicycle model: A consistent model for planning feasible trajectories for autonomous vehicles?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:71e443a73a8d1d34f25f4ae5d785abca4b1734443d2c0918aadd7ee890a19025

Observation 370964a7-e7ae-4569-ad97-1df63392fd99 · outbound

This paper cites Congested traffic states in empirical observations and microscopic simulations,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Congested traffic states in empirical observations and microscopic simulations,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:65c6b51146ea3fbc58148852764a51a66b67fe9d5e65c53d470bbb227c77ca31

Observation 15dea300-1ea8-436c-b2aa-4a247d2c1036 · outbound

This paper cites Preferred time-headway of highway drivers,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Preferred time-headway of highway drivers,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:15cc0c3647f18607bc065d3ccc4c5f0e8eeae4787362d9fc7c8c45bd135a9b99

Observation 84badfd9-f38a-4ea2-a1dd-947c2fb79cbd · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:834ef59bed92f3829fc94f71ed5f43afd8ddf34431258665bea05d0e35747591

Observation 0b27eee5-acbd-4cdb-a2e9-8d64c5f405d3 · outbound

This paper cites Responsive safety in reinforce- ment learning by pid lagrangian methods,.

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Responsive safety in reinforce- ment learning by pid lagrangian methods,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T05:24:53.844059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:24:53.844059Z digest=sha256:5002aa85170329551a3d9e4adf29f71861a199ac8556241ab763fd702a7e0fec

Pith citing papers

No inbound Pith citation observations are available.