Pith. sign in

Paper Citation Record · LEDGER

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes

As of 13 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2501.02774.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02774 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:13:10.092408Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact6
  • verified fuzzy26
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3ce5166-102e-49cb-b77e-86b7d2e3b02e · outbound

This paper cites Human-level control through deep reinforcement learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Human-level control through deep reinforcement learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:07.874981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:07.874981Z digest=sha256:fd65e206baf718d30780274dd5227934a07653b6c0f68661ca1e2e213bcc9dad

Observation 16c667f4-1af0-449b-9d6a-6317cbd823f6 · outbound

This paper cites Continuous control with deep reinforcement learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Continuous control with deep reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:07.931810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:07.931810Z digest=sha256:b0d21e494c7a9705ab1eb422e3eda6bd4c5929b49fab4b86fc8cd527c4613e15

Observation 85a6db6c-117f-423a-9601-4f742d222a3a · outbound

This paper cites Trust Region Policy Optimization.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Trust Region Policy Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:07.973484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:07.973484Z digest=sha256:8a679e81969b481b2b18fa073f9c5efe2db13dfd62c7d9df387636649fb0ef78

Observation ebf5726b-06b7-42e8-8f35-4871cf29760a · outbound

This paper cites Deep Reinforcement Learning framework for Autonomous Driving.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Deep Reinforcement Learning framework for Autonomous Driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.021368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.021368Z digest=sha256:f28d9966e5e8bfa6110707db6ab417270864018667a1e2ad5acea76dcebd9dc5

Observation bc5a620d-d7d0-4110-9afc-d63e310814f9 · outbound

This paper cites Continuous mdp homomorphisms and homomorphic policy gradient,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Continuous mdp homomorphisms and homomorphic policy gradient,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:15.124756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.064753Z digest=sha256:4ce100ef1bdd6cfa5614d8af7dd909f7496b76aa6a73d177eb53baa28bf22312

Observation 48462379-f791-4650-bc6b-fe83478d06c5 · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.115774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.115774Z digest=sha256:d891044e93b26f35f75303b419dade8d85335e6c9d18a5cc9ee381db26074e39

Observation 773dc85a-275e-4afa-a622-a962d8bcf345 · outbound

This paper cites Movie: Visual model-based policy adaptation for view generalization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Movie: Visual model-based policy adaptation for view generalization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:15.014268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.145799Z digest=sha256:94281365f8ce785f162bca7c315defa7a352b72d5ca802246ae2939f9aad7f53

Observation c90270ed-d52e-4f73-a5af-0d61e1f651db · outbound

This paper cites Making better decision by directly planning in continuous control,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Making better decision by directly planning in continuous control,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.943242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.194612Z digest=sha256:459ac178e3bdc08aa885237e211e70fe9a2e70e39da51a59f24cc2d99ba6c6f4

Observation efd40ed2-e94b-4ee3-8e81-250cc28bf14c · outbound

This paper cites Hierarchical advantage for reinforcement learning in parameterized action space,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Hierarchical advantage for reinforcement learning in parameterized action space,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.885947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.244753Z digest=sha256:db81a2d023d864eee75d6879fa632fad28eac34813af1c16ae6684f5928d14ce

Observation dc94fb86-bcb7-49ae-af24-eb4ae35ee822 · outbound

This paper cites Deep Reinforcement Learning in Parameterized Action Space.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Deep Reinforcement Learning in Parameterized Action Space

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.294832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.294832Z digest=sha256:b0b6a25b7cfef71aec26d42e3f91fdb306b9e606c228b6a3379c23831837b4e4

Observation 58fbc369-6f81-4cd0-ac1f-861e330543f7 · outbound

This paper cites Deep Multi-Agent Reinforcement Learning with Discrete-Continuous Hybrid Action Spaces.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Deep Multi-Agent Reinforcement Learning with Discrete-Continuous Hybrid Action Spaces

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.344754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.344754Z digest=sha256:8d7cff9268ae572859b384d9588bad049011c4aaeb6e3b023354b2963bf2b648

Observation 4e967c4e-ef86-4707-83df-62f7a5cb5816 · outbound

This paper cites Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.394751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.394751Z digest=sha256:fbe2ce3206fc61d3a33cfd37ddc2b94f0c4f76497b6a1527d9c7a2d03bf5340d

Observation 214518cb-6f71-4875-bb3e-b8c8c8a3b67a · outbound

This paper cites Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.430878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.430878Z digest=sha256:93f77b8cde1e6b53ad9ef3d0216cf76637dedd654ce86cf2dab151ec03eb2e69

Observation a1030316-f089-4206-8969-67613adac404 · outbound

This paper cites HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.471598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.471598Z digest=sha256:9d993e1b98fa2ca56712b2dc13dcd965a1d231f3f573eb43c7ae87134cb12b23

Observation 023c83db-31dc-44f9-adae-bef1638bc15b · outbound

This paper cites Reinforcement learning with parameterized actions,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Reinforcement learning with parameterized actions,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.792724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.520651Z digest=sha256:1a9bd04ea5628d71a5eea88c7f6eb9a8cc56a6d2c1894a4dd29758dabeaf0d59

Observation 35cbf09d-f9b8-4de6-8456-c4e8ee3bdedf · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Dream to Control: Learning Behaviors by Latent Imagination

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.565456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.565456Z digest=sha256:d592dcab61c0c2ea58a49c09620b5d1757e6cc87511343490dd30cdfff2771bb

Observation 1f76ae8c-4ccf-4d0f-bd9a-2335b5b9b408 · outbound

This paper cites Mastering Atari with Discrete World Models.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Mastering Atari with Discrete World Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.612810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.612810Z digest=sha256:9ae0e440585872eb7e9d7539e3843d98b11140cad2b6a788857df630b3fa625f

Observation fb9e0995-93ab-4e25-bbb2-dc5d3442620a · outbound

This paper cites Privileged Sensing Scaffolds Reinforcement Learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Privileged Sensing Scaffolds Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:11.734749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.653657Z digest=sha256:e820a062086821bc50355857f7f74a6c5101f72c7e95e478143a7077fba900d9

Observation 34d92f07-a7cd-49b7-9bb0-2ca85dc350f9 · outbound

This paper cites Model-based Reinforcement Learning for Parameterized Action Spaces.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Model-based Reinforcement Learning for Parameterized Action Spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.703941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.703941Z digest=sha256:e45932dd42e6c348b4ea5e82dd37ca7f58896c00cb6f10ba8c21686d1655d856

Observation 4ef3cd29-9f52-46e3-9eaa-a75d1a33c9a7 · outbound

This paper cites The Benefits of Model-Based Generalization in Reinforcement Learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes The Benefits of Model-Based Generalization in Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.734656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.734656Z digest=sha256:d803e7c57d6c31b94a0ea0b2c75604a4b181eb06a343ee507692a6f89fd7238e

Observation a393001f-e5fd-4011-9e24-a7db96cc6e9f · outbound

This paper cites Diminishing return of value expansion methods in model-based reinforcement learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Diminishing return of value expansion methods in model-based reinforcement learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.614751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.749078Z digest=sha256:d451fa1c43e9f4456730a561e5ee9f7a66c99738e23045ef93ef1066a33b3c8d

Observation 43c3c435-2a7c-4e9e-a36b-c1bef140e093 · outbound

This paper cites Models, Pixels, and Rewards: Evaluating Design Trade-offs in Visual Model-Based Reinforcement Learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Models, Pixels, and Rewards: Evaluating Design Trade-offs in Visual Model-Based Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:11.443147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.787783Z digest=sha256:b659dbc933230dd3efdbb05fee3d211d3cb23f25dae187026b3a9f5a46d1fd0e

Observation 55f67307-798f-4012-bfd5-64cbd6767d41 · outbound

This paper cites DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:11.388643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.824907Z digest=sha256:971b802a199637518e8620668a2f2d7bbfc153cfa6a2e4f118fcb915fd8db060

Observation 57fc4af9-3968-4d15-aff1-64006f9c2c8b · outbound

This paper cites Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.859571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.859571Z digest=sha256:ad01eed0f4e63012139b696a335c37dbbfc2a788308ef49c32188ed533d544ef

Observation 61a66bf1-3717-430b-b6d1-4807988fa085 · outbound

This paper cites Learning insertion primitives with discrete-continuous hybrid action space for robotic assembly tasks,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Learning insertion primitives with discrete-continuous hybrid action space for robotic assembly tasks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.485857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.881404Z digest=sha256:a968d68be6155be7e9f475c1b4af8311c6c933afc028f96e69348fe9a3e80ee1

Observation 99753433-1d7b-4a4f-955a-f554b0454e9b · outbound

This paper cites Lipschitz continuity in model- based reinforcement learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Lipschitz continuity in model- based reinforcement learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.334749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.897601Z digest=sha256:e93aa5b0b4e6e75b8008286d55a89aec163f7113f8e016a6fc0f088b85f3c4c3

Observation 0e95acb9-b3ea-4e4b-b779-c6c713777e18 · outbound

This paper cites Markov processes over denumerable products of spaces, describing large systems of automata,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Markov processes over denumerable products of spaces, describing large systems of automata,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.268779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:08.924857Z digest=sha256:4cbbfc5621d998ae9c86c5ff222c7f181bc63f1cc17cabf6b7cb4de199757eda

Observation 055f35d8-f7a8-4de8-9abc-70969ac87f0d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Explaining and Harnessing Adversarial Examples

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.974756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.974756Z digest=sha256:b9bfcf90f5000e25ba270d44681965de6958a5ee59698e940e7f33fbc69171d3

Observation c5e37482-ebdb-46f1-afea-f32b426a8185 · outbound

This paper cites A mathematical theory of communication,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes A mathematical theory of communication,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.014851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.014851Z digest=sha256:b5d05742ff45f3b38cf65c27b1a859fe452ef632d8e2ce688c7bf838c1a3912f

Observation 8aa9f94b-ea28-4d56-957a-ae76c286218d · outbound

This paper cites The im algorithm: a variational approach to information maximization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes The im algorithm: a variational approach to information maximization,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.209556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.063516Z digest=sha256:2f2683734a817b8307c60b0c7216466d0643ba72166d1fe01d62e5e5ff9228b1

Observation 4c1b5286-a103-4fe4-ac63-cfef51d5a6bb · outbound

This paper cites Learning-based model predictive control for markov decision processes,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Learning-based model predictive control for markov decision processes,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.155133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.092320Z digest=sha256:05c05e27696df626676a04077260303c3b4ea7ed3045e81316828255027e39da

Observation 55fd597e-cee8-4b0d-92d6-e2bed94a51f7 · outbound

This paper cites Optimization of computer simulation models with rare events,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Optimization of computer simulation models with rare events,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.058320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.122904Z digest=sha256:0ec7a0e2199f0b2621a8f101b3fb3bafc1197a029b4d4fe997ed34c34a3ba99d

Observation bbf4ba56-0ff9-4a24-9d8a-20309191a49d · outbound

This paper cites Plan to predict: Learning an uncertainty-foreseeing model for model-based reinforce- ment learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Plan to predict: Learning an uncertainty-foreseeing model for model-based reinforce- ment learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.934751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.152757Z digest=sha256:3f5b312df66c0b9b8f2abd0747f2b3009861c501076d329703deb6493693c833

Observation 2ba40424-c73c-4155-b3b7-e6a2ca0c21d5 · outbound

This paper cites Choreographer: Learning and Adapting Skills in Imagination.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Choreographer: Learning and Adapting Skills in Imagination

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.179782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.179782Z digest=sha256:81b579704754a9b0091059bfcaf7b92b160b26b70d668d780ae73c7e7f858996

Observation df786e93-e2c2-480a-97e1-3820ec7f11a7 · outbound

This paper cites Mismatched no more: Joint model-policy optimization for model- based rl,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Mismatched no more: Joint model-policy optimization for model- based rl,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.784917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.224823Z digest=sha256:f8d68984f76b3323beb245b1e8b1a25245132c9266c6b7ad1b2a40b197a10d25

Observation c677bc8f-2160-41dc-b0dd-0eb365f470d0 · outbound

This paper cites Differen- tiable mpc for end-to-end planning and control,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Differen- tiable mpc for end-to-end planning and control,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.705206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.284743Z digest=sha256:1bb5d170886410151c05a6bcaa971b1cdc0ee7f3cd422806fd1f12f0e9683012

Observation 70895f64-e819-42f1-b1ac-c051685bf061 · outbound

This paper cites On information and sufficiency,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes On information and sufficiency,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.329785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.329785Z digest=sha256:c9027e439f584aad764c12cc74ca3ce3cb4ac849fbe772f4a6fbbf43e2d7edca

Observation e934efb5-2c2d-432e-a0a3-ba4114c01bd4 · outbound

This paper cites Measures of distance between probability distributions,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Measures of distance between probability distributions,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.588551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.374768Z digest=sha256:3cbfb0f5a7ec55c87da2aeab329f8c757e1b1e3a8ea7ec87f7f8278badd47567

Observation 3a1cab5a-896b-4d5f-ba88-fbc9d853d983 · outbound

This paper cites Dr. Strategy: Model-Based Generalist Agents with Strategic Dreaming.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Dr. Strategy: Model-Based Generalist Agents with Strategic Dreaming

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:11.155546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.424745Z digest=sha256:6f478096e8bcf34ff8556717a31f525504496c4d91292fafeba550f1dd2e74a5

Observation ce6fd07d-e6c5-4bb2-a43a-1b59251b0e2e · outbound

This paper cites When to trust your model: Model-based policy optimization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes When to trust your model: Model-based policy optimization,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.507143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.486673Z digest=sha256:a43f4f5ab7ffaee24a3d0d1859b18078281441c31c2c70b8d9f21e14f5dbfc84

Observation e6b81818-ff75-43bf-b7aa-7adf24c03920 · outbound

This paper cites Model-based reinforcement learning via meta-policy optimization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Model-based reinforcement learning via meta-policy optimization,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.411043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.505310Z digest=sha256:d753852218016cf83701a8036fd7c0054b158dfc47732b53860823a3448028b6

Observation 24d689a0-cebb-460a-ba81-f0f1b7f8df2c · outbound

This paper cites The virtues of laziness in model-based rl: A unified objective and algo- rithms,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes The virtues of laziness in model-based rl: A unified objective and algo- rithms,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.304750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.538880Z digest=sha256:008993eb58b16b73de841c603fd32a0c541bd0cc6cc1008642b2dc09ae48417f

Observation c3bcd5d5-fc70-48c4-ab46-475e1479b127 · outbound

This paper cites Iterative value-aware model learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Iterative value-aware model learning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.193218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.563874Z digest=sha256:59d668055e21467a61f62a2f08120b69b5ab10650ca43dc36b7dceaed05cc556

Observation 7eb6f6a4-bf98-48ea-8459-17304275825f · outbound

This paper cites Model-based value expansion for efficient model-free rein- forcement learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Model-based value expansion for efficient model-free rein- forcement learning,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.100233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.571543Z digest=sha256:02956e5edb10e7680ec0ad599fc156ac2ee485ec0cc49ca7f9b3741b55a8a6a9

Observation 7d7bce9d-9972-4c4f-8e45-2228a24eb222 · outbound

This paper cites Villani et al.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Villani et al

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.025685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.583592Z digest=sha256:c025619f774a3cd2b4efc2211a8e65d9a12a9c6032a09d9d96c697e64cf3f493

Observation 6a423560-c737-4b7b-9eb0-ba5fc99a2d78 · outbound

This paper cites Generative adversarial nets,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Generative adversarial nets,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.607965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.607965Z digest=sha256:816183817536386d26a9f50299764471f8bd34d12d90ddeddd457656e3d74693

Observation 21fad6e0-ba0c-4e0c-94a5-bf23d2e010f7 · outbound

This paper cites Towards Principled Methods for Training Generative Adversarial Networks.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Towards Principled Methods for Training Generative Adversarial Networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.634751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.634751Z digest=sha256:ffb092894003fa8de4cf1536104ac3c9006fa873cd6a74cc9a696b1f492e7869

Observation 1629a347-0aa7-448f-853a-47d4f3ec0bae · outbound

This paper cites Wasserstein generative ad- versarial networks,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Wasserstein generative ad- versarial networks,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:12.814759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.684749Z digest=sha256:9f91796bb77ff30cd6b845b6b083dc3052553ebba63d2c1d1b39761b2c9e0939

Observation 2a9e0545-bac5-4c2e-9959-ff8fe7f60c67 · outbound

This paper cites Harpy, a connected speech recognition system,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Harpy, a connected speech recognition system,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:12.704852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.734752Z digest=sha256:20c8de5b47a1789faf276e518c4c670877a6c5d364212eed30380520e777e20a

Observation 44c6830e-c21a-4e80-bd0a-d0cdc6c7c663 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.784750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.784750Z digest=sha256:d23e112a42b79163853e6724b49a65b280c55328b7bdf0d1e435ab77f6ff70d4

Observation f74544f5-fdee-48fc-bab3-6d7f128b19fd · outbound

This paper cites Walk Wisely on Graph: Knowledge Graph Reasoning with Dual Agents via Efficient Guidance-Exploration.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Walk Wisely on Graph: Knowledge Graph Reasoning with Dual Agents via Efficient Guidance-Exploration

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:10.883459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.822192Z digest=sha256:df4184457bf33bfe6a4428b1d446076d1cc77d5c90715e07b522cf076369b14f

Observation 4ae52c77-d93e-43f9-b08d-39c637e62b6b · outbound

This paper cites Spectral Norm Regularization for Improving the Generalizability of Deep Learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Spectral Norm Regularization for Improving the Generalizability of Deep Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.864761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.864761Z digest=sha256:854267a4a51abc28982c8c887d390864e63fa2590c8f8eaa357ad9033c12ee27

Observation a214a2d9-fd2c-4359-be98-2a65c8aed1b3 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Categorical Reparameterization with Gumbel-Softmax

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.893528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.893528Z digest=sha256:60d4db096a177ab6a3b508a9ba7888181e33e997698e421261fc8a998dc1db86

Observation 5a172a2a-29bd-4124-916a-e77d37a0d58b · outbound

This paper cites Introduction to online convex optimization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Introduction to online convex optimization,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:12.603409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.921439Z digest=sha256:60c83aa09bafabd9d131da90d8a2e663036926b3e0f6dfb849afa70520b283f6

Observation 6d0d3d94-3f37-47e5-841f-80391b92d272 · outbound

This paper cites Is Model Ensemble Necessary? Model-based RL via a Single Model with Lipschitz Regularized Value Function.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Is Model Ensemble Necessary? Model-based RL via a Single Model with Lipschitz Regularized Value Function

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:10.344751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:09.964752Z digest=sha256:a97bb0d4d65322219f50df92c39f27e92b0b6c3daace0542409aff074ffea557

Observation dd4be036-c677-46ec-81b8-6c3558dcd1b7 · outbound

This paper cites Visualizing data using t-sne.,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Visualizing data using t-sne.,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.994817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.994817Z digest=sha256:56ddc9b4c7515fafd931da79a4bc9e1b3cddea5af65b88dd83901f543c3c1993

Observation 10f6f2d4-73a8-4e7e-8c07-df67b7a7e86f · outbound

This paper cites General boundary conditions for denumberable markov processes,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes General boundary conditions for denumberable markov processes,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:12.519612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:10.044751Z digest=sha256:54090ce149e7e9f9428fd16dc5d37452954d9250f1ad06c433f91ab13bccdf31

Observation 51b87c37-82ee-431a-aae4-056eb7bf6398 · outbound

This paper cites an unresolved cited work.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:12.478583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:13:10.092408Z digest=sha256:99cf38264cbe0e8239368d1f57556d43befc69f74feccc69e998c965b8bfd224

Pith citing papers

No inbound Pith citation observations are available.