Pith. sign in

Paper Citation Record · LEDGER

Combining Trained Models in Reinforcement Learning

As of 23 July 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2605.02159.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.02159 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T19:24:04.574344Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dcd05b3-45ed-4b75-a274-73ff85304b96 · outbound

This paper cites Human-level control through deep reinforcement learning.

Combining Trained Models in Reinforcement Learning Human-level control through deep reinforcement learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.917538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:70ff3057edcfd56b31f1b92b5e3858840e0ed8e239de05d51cc2e9c975963956

Observation 62bb8c33-8fca-4077-94f4-13a9bf917b31 · outbound

This paper cites Mastering the game of Go without human knowledge.

Combining Trained Models in Reinforcement Learning Mastering the game of Go without human knowledge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.911932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:059681c69f57ccfbd1b9656ec325761a399d568492aaa4c81739483623346ce0

Observation 898c361a-ba48-4833-9624-e925ed2c6aec · outbound

This paper cites Policy Distillation.

Combining Trained Models in Reinforcement Learning Policy Distillation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:30.231690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:c677afb383a09a6536f78383e7e8dc75b29cfc2daf7f1ea11981fb1218ba92e5

Observation 424c9021-af0d-46db-a6cd-6dd297e8f575 · outbound

This paper cites Context-aware policy reuse.

Combining Trained Models in Reinforcement Learning Context-aware policy reuse

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-08T19:29:05.471106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:798e5788b610f293a78410c339434cec88735ba754a03824162e9161dde2ab9b

Observation 09aeaef8-53e9-4d76-ab63-f1a9dfa21c9f · outbound

This paper cites Efficient bayesian policy reuse with a scalable observation model in deep reinforcement learning.

Combining Trained Models in Reinforcement Learning Efficient bayesian policy reuse with a scalable observation model in deep reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.932636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:6c2631f487a1d7ef0e6d421a9ccd2ea21cc595fc7adc457072968886282de3f6

Observation 57c7b723-d9d4-49b8-844e-418020c9cc92 · outbound

This paper cites Model-based reinforcement learning with probabilistic ensemble terminal critics for data-efficient control applications.

Combining Trained Models in Reinforcement Learning Model-based reinforcement learning with probabilistic ensemble terminal critics for data-efficient control applications

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.892264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:75cc7a25008c790d339a67891c25b170732c184d3b1f619306dd721cb6b4bde3

Observation 8a5c373b-efd2-48c4-936c-6a3152e16e0e · outbound

This paper cites FedDOVe: A federated deep Q-learning-based offloading for vehicular fog computing.

Combining Trained Models in Reinforcement Learning FedDOVe: A federated deep Q-learning-based offloading for vehicular fog computing

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.944483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:64e69d74e74e3a531820618ccf675bae297bb162deca9a03b9f469cb31aa40d3

Observation 625371cc-99db-460a-a3f8-42e2f6851413 · outbound

This paper cites Federated reinforcement learning framework for mobile robot navigation using ROS and gazebo.

Combining Trained Models in Reinforcement Learning Federated reinforcement learning framework for mobile robot navigation using ROS and gazebo

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.940208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:913feaaa508f3739bc74066ab549c366dd8821e58662b42cf8ea8809a05f93ea

Observation 84ed5440-4191-4c96-b11e-4dda6cb89933 · outbound

This paper cites Transfer learning for reinforcement learning domains: A survey.

Combining Trained Models in Reinforcement Learning Transfer learning for reinforcement learning domains: A survey

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.877506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:83cad0bf8bcec1be4a382bc6c9af6190d8d21889a2fd39b3299ef0d009e1663a

Observation f1708d71-4a22-43ac-8126-c7cdcb43a562 · outbound

This paper cites Sim-to-real transfer in deep reinforcement learning for robotics: A survey.

Combining Trained Models in Reinforcement Learning Sim-to-real transfer in deep reinforcement learning for robotics: A survey

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.904614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:b16b279196539ebed0758c916e608e7f2c21e2e86f0701c6e9294d07a28b5691

Observation b3bfbebe-4bf7-4da9-a8c2-e9169598cc6a · outbound

This paper cites A survey on transfer reinforcement learning.

Combining Trained Models in Reinforcement Learning A survey on transfer reinforcement learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.888904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:0ba82406309be8fe83168e6ae4a7eb1a02dab553a68cc021da8513fecc6c74d9

Observation ac9ec023-1889-4481-822b-083db532c5e5 · outbound

This paper cites Importance prioritized policy distillation.

Combining Trained Models in Reinforcement Learning Importance prioritized policy distillation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.959670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:ceb02d3e11bb9d2b0e5aab05e74bbe5dbe7cc46db3e85858d901f560bdb80889

Observation 91de22cc-7038-4135-b68b-3535e83e23d1 · outbound

This paper cites Online policy distillation with decision-attention.

Combining Trained Models in Reinforcement Learning Online policy distillation with decision-attention

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.895758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:f74ecad53c3111798e8ac1f304a0b54d9d919491a5bcfab56a1bf785a654e661

Observation cc44dd9b-da6e-43dc-ba9c-b81a0e6a5a61 · outbound

This paper cites Probabilistic policy reuse for safe rein- forcement learning.

Combining Trained Models in Reinforcement Learning Probabilistic policy reuse for safe rein- forcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.948925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:9e6699a9eb7c0905a911c6058bc7825bcabfbfd4a07038c44541e78e54ee0b87

Observation 72cf85ad-380f-495b-9bdd-eea41a2b3a49 · outbound

This paper cites Policy transfer via skill adaptation and composition.

Combining Trained Models in Reinforcement Learning Policy transfer via skill adaptation and composition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.956442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:191ca36e0e74dcc2ac3d6904b48e98f0d06c95679615a0945919a85ed075fc3f

Observation 9a6f5fb6-fbd1-4ec5-b5b2-e609551f54bc · outbound

This paper cites Transfer reinforcement learning based on gaussian process policy reuse.

Combining Trained Models in Reinforcement Learning Transfer reinforcement learning based on gaussian process policy reuse

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.936586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:a8db630ff5df56cd2ff5f3b86afbc45762dd4d01db19a978fbf3f28edfc6cdd7

Observation 395cfb41-45e4-45d7-8348-9a21c60215bc · outbound

This paper cites Safe adaptive policy transfer reinforcement learning for distributed multiagent control.

Combining Trained Models in Reinforcement Learning Safe adaptive policy transfer reinforcement learning for distributed multiagent control

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.927512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:a38886218037244c8348ed0c1b709623b174f1927445b31b4ac121fb8e9ad982

Observation cf5c622e-75ed-474d-a2ad-5d391c70d17c · outbound

This paper cites Combining pre-trained models for enhanced feature representation in reinforcement learning.

Combining Trained Models in Reinforcement Learning Combining pre-trained models for enhanced feature representation in reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.881768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:3680c51b6550eac9650c613b3672665ac9e465b13195b1fedaf2981844536ebf

Observation b3d3b3bd-53d2-487c-af70-ed71c8814ece · outbound

This paper cites The PRISMA 2020 statement: an updated guideline for reporting systematic reviews.

Combining Trained Models in Reinforcement Learning The PRISMA 2020 statement: an updated guideline for reporting systematic reviews

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.921673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:9bf862a1b36128e03b9dce164060d6b6ba6b495312f88fb6657564b9f7af7237

Observation 3df149d1-beea-422c-a5be-585ba198f041 · outbound

This paper cites Parallel reinforcement learning: a framework and case study.

Combining Trained Models in Reinforcement Learning Parallel reinforcement learning: a framework and case study

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.952915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:86ae0673b7f83daf6f6a82da8f0cf79917a3b4165d6df43cc99fb872fab38acd

Observation 8a93cf03-1d56-492c-89df-bc5c2e46204e · outbound

This paper cites Policy distillation and value matching in multiagent reinforcement learning.

Combining Trained Models in Reinforcement Learning Policy distillation and value matching in multiagent reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.907657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:4ccd8f5981b6cf592ca29e53891e8a26d6d96822b262eacd7c73463e91197ba2

Observation cbd5881f-28e7-42b4-8687-97a61a3a865c · outbound

This paper cites Leaders and collaborators: Address- ing sparse reward challenges in multi-agent reinforcement learning.

Combining Trained Models in Reinforcement Learning Leaders and collaborators: Address- ing sparse reward challenges in multi-agent reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.885627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:f81d0ae6a246bd3cb52755b748c1a06e9ac67dc29b6923594bf26f12f9a43ce3

Observation 05cbf563-16a2-4662-8b1c-c91d5ccebc9c · outbound

This paper cites A hybrid ensemble framework for adversarial robustness in deep reinforcement learning.

Combining Trained Models in Reinforcement Learning A hybrid ensemble framework for adversarial robustness in deep reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.900467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:ed5cf30010fe9dc2b0fd49656deeb9d23514af284b297599cc30fc8a77b02b10

Observation 44c4ab00-64e0-470c-80cb-3c77b6083ba8 · outbound

This paper cites PEARL: FPGA-based reinforcement learning acceleration with pipelined parallel environments.

Combining Trained Models in Reinforcement Learning PEARL: FPGA-based reinforcement learning acceleration with pipelined parallel environments

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:01:33.963500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:24:04.574344Z digest=sha256:46f2b27ab09218af4bae991e1e80fcb031d1a893b020aa40bab87ffbf21f43fd

Pith citing papers

No inbound Pith citation observations are available.