Pith. sign in

Paper Citation Record · LEDGER

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes

As of 13 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2501.02774.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02774 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:13:10.092408Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact6
  • verified fuzzy26
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3ce5166-102e-49cb-b77e-86b7d2e3b02e · outbound

This paper cites Human-level control through deep reinforcement learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Human-level control through deep reinforcement learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:07.874981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:07.874981Z digest=sha256:fd65e206baf718d30780274dd5227934a07653b6c0f68661ca1e2e213bcc9dad

Observation 16c667f4-1af0-449b-9d6a-6317cbd823f6 · outbound

This paper cites Continuous control with deep reinforcement learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Continuous control with deep reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:07.931810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:07.931810Z digest=sha256:b0d21e494c7a9705ab1eb422e3eda6bd4c5929b49fab4b86fc8cd527c4613e15

Observation 85a6db6c-117f-423a-9601-4f742d222a3a · outbound

This paper cites Trust Region Policy Optimization.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Trust Region Policy Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:07.973484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:07.973484Z digest=sha256:8a679e81969b481b2b18fa073f9c5efe2db13dfd62c7d9df387636649fb0ef78

Observation ebf5726b-06b7-42e8-8f35-4871cf29760a · outbound

This paper cites Deep Reinforcement Learning framework for Autonomous Driving.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Deep Reinforcement Learning framework for Autonomous Driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.021368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.021368Z digest=sha256:f28d9966e5e8bfa6110707db6ab417270864018667a1e2ad5acea76dcebd9dc5

Observation bc5a620d-d7d0-4110-9afc-d63e310814f9 · outbound

This paper cites Continuous mdp homomorphisms and homomorphic policy gradient,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Continuous mdp homomorphisms and homomorphic policy gradient,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:15.124756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.064753Z digest=sha256:3078e78eab8517fd02b7e166bce5a08e2608a05cc7046845f7826ee920ef26d9

Observation 48462379-f791-4650-bc6b-fe83478d06c5 · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.115774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.115774Z digest=sha256:d891044e93b26f35f75303b419dade8d85335e6c9d18a5cc9ee381db26074e39

Observation 773dc85a-275e-4afa-a622-a962d8bcf345 · outbound

This paper cites Movie: Visual model-based policy adaptation for view generalization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Movie: Visual model-based policy adaptation for view generalization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:15.014268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.145799Z digest=sha256:a6a4522adf468fe37eb11ba0e42583db3c6c7b43d13481943846d6dcf56a2829

Observation c90270ed-d52e-4f73-a5af-0d61e1f651db · outbound

This paper cites Making better decision by directly planning in continuous control,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Making better decision by directly planning in continuous control,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.943242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.194612Z digest=sha256:908806ba3f5d98111c693a0639992de7b0ba3c6da5246329c2e6e5cb0a3cac4d

Observation efd40ed2-e94b-4ee3-8e81-250cc28bf14c · outbound

This paper cites Hierarchical advantage for reinforcement learning in parameterized action space,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Hierarchical advantage for reinforcement learning in parameterized action space,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.885947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.244753Z digest=sha256:5de9bee60ea97383c2ee7de211a4a061f190d09cd425c24dd303d40f324833e8

Observation dc94fb86-bcb7-49ae-af24-eb4ae35ee822 · outbound

This paper cites Deep Reinforcement Learning in Parameterized Action Space.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Deep Reinforcement Learning in Parameterized Action Space

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.294832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.294832Z digest=sha256:b0b6a25b7cfef71aec26d42e3f91fdb306b9e606c228b6a3379c23831837b4e4

Observation 58fbc369-6f81-4cd0-ac1f-861e330543f7 · outbound

This paper cites Deep Multi-Agent Reinforcement Learning with Discrete-Continuous Hybrid Action Spaces.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Deep Multi-Agent Reinforcement Learning with Discrete-Continuous Hybrid Action Spaces

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.344754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.344754Z digest=sha256:8d7cff9268ae572859b384d9588bad049011c4aaeb6e3b023354b2963bf2b648

Observation 4e967c4e-ef86-4707-83df-62f7a5cb5816 · outbound

This paper cites Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.394751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.394751Z digest=sha256:fbe2ce3206fc61d3a33cfd37ddc2b94f0c4f76497b6a1527d9c7a2d03bf5340d

Observation 214518cb-6f71-4875-bb3e-b8c8c8a3b67a · outbound

This paper cites Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.430878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.430878Z digest=sha256:93f77b8cde1e6b53ad9ef3d0216cf76637dedd654ce86cf2dab151ec03eb2e69

Observation a1030316-f089-4206-8969-67613adac404 · outbound

This paper cites HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.471598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.471598Z digest=sha256:9d993e1b98fa2ca56712b2dc13dcd965a1d231f3f573eb43c7ae87134cb12b23

Observation 023c83db-31dc-44f9-adae-bef1638bc15b · outbound

This paper cites Reinforcement learning with parameterized actions,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Reinforcement learning with parameterized actions,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.792724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.520651Z digest=sha256:971c062099402ca86b4f09c54aeae500ce0c0e749da4a3410f285187d3fdd84c

Observation 35cbf09d-f9b8-4de6-8456-c4e8ee3bdedf · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Dream to Control: Learning Behaviors by Latent Imagination

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.565456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.565456Z digest=sha256:d592dcab61c0c2ea58a49c09620b5d1757e6cc87511343490dd30cdfff2771bb

Observation 1f76ae8c-4ccf-4d0f-bd9a-2335b5b9b408 · outbound

This paper cites Mastering Atari with Discrete World Models.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Mastering Atari with Discrete World Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.612810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.612810Z digest=sha256:9ae0e440585872eb7e9d7539e3843d98b11140cad2b6a788857df630b3fa625f

Observation fb9e0995-93ab-4e25-bbb2-dc5d3442620a · outbound

This paper cites Privileged Sensing Scaffolds Reinforcement Learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Privileged Sensing Scaffolds Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:11.734749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.653657Z digest=sha256:8f6f304c5b917d125e84a964087f93b1616f2e2624c5f39df4158a27675ed005

Observation 34d92f07-a7cd-49b7-9bb0-2ca85dc350f9 · outbound

This paper cites Model-based Reinforcement Learning for Parameterized Action Spaces.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Model-based Reinforcement Learning for Parameterized Action Spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.703941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.703941Z digest=sha256:e45932dd42e6c348b4ea5e82dd37ca7f58896c00cb6f10ba8c21686d1655d856

Observation 4ef3cd29-9f52-46e3-9eaa-a75d1a33c9a7 · outbound

This paper cites The Benefits of Model-Based Generalization in Reinforcement Learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes The Benefits of Model-Based Generalization in Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.734656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.734656Z digest=sha256:d803e7c57d6c31b94a0ea0b2c75604a4b181eb06a343ee507692a6f89fd7238e

Observation a393001f-e5fd-4011-9e24-a7db96cc6e9f · outbound

This paper cites Diminishing return of value expansion methods in model-based reinforcement learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Diminishing return of value expansion methods in model-based reinforcement learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.614751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.749078Z digest=sha256:5c72b7087e077ca51880ebca74c38633e9c1072be1776ab62075c728dd37bd72

Observation 43c3c435-2a7c-4e9e-a36b-c1bef140e093 · outbound

This paper cites Models, Pixels, and Rewards: Evaluating Design Trade-offs in Visual Model-Based Reinforcement Learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Models, Pixels, and Rewards: Evaluating Design Trade-offs in Visual Model-Based Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:11.443147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.787783Z digest=sha256:93ba2999460a1ea52461211a23f2a9edc35b9e25228a2e5914edd5cb2e244e70

Observation 55f67307-798f-4012-bfd5-64cbd6767d41 · outbound

This paper cites DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:11.388643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.824907Z digest=sha256:4495dfdd59ba4b999c0ac2f9ebaafa0fcbae6d2126c9ba51cdc5987e4f990c9e

Observation 57fc4af9-3968-4d15-aff1-64006f9c2c8b · outbound

This paper cites Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.859571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.859571Z digest=sha256:ad01eed0f4e63012139b696a335c37dbbfc2a788308ef49c32188ed533d544ef

Observation 61a66bf1-3717-430b-b6d1-4807988fa085 · outbound

This paper cites Learning insertion primitives with discrete-continuous hybrid action space for robotic assembly tasks,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Learning insertion primitives with discrete-continuous hybrid action space for robotic assembly tasks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.485857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.881404Z digest=sha256:dae896c2c8e8c3d1d1cbf178712c90fea71b8fec07c7abf7d8ad70c51323a146

Observation 99753433-1d7b-4a4f-955a-f554b0454e9b · outbound

This paper cites Lipschitz continuity in model- based reinforcement learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Lipschitz continuity in model- based reinforcement learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.334749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.897601Z digest=sha256:507917e0ffba757bc4fdd75b81b6198081c6b5dc56994d22d5efc981f2a7ed7c

Observation 0e95acb9-b3ea-4e4b-b779-c6c713777e18 · outbound

This paper cites Markov processes over denumerable products of spaces, describing large systems of automata,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Markov processes over denumerable products of spaces, describing large systems of automata,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.268779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:08.924857Z digest=sha256:87635ae18d4f368e6e168542a589ab0a69bc95b5afe50c5e25b3cc07953d76c2

Observation 055f35d8-f7a8-4de8-9abc-70969ac87f0d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Explaining and Harnessing Adversarial Examples

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:08.974756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:08.974756Z digest=sha256:b9bfcf90f5000e25ba270d44681965de6958a5ee59698e940e7f33fbc69171d3

Observation c5e37482-ebdb-46f1-afea-f32b426a8185 · outbound

This paper cites A mathematical theory of communication,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes A mathematical theory of communication,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.014851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.014851Z digest=sha256:b5d05742ff45f3b38cf65c27b1a859fe452ef632d8e2ce688c7bf838c1a3912f

Observation 8aa9f94b-ea28-4d56-957a-ae76c286218d · outbound

This paper cites The im algorithm: a variational approach to information maximization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes The im algorithm: a variational approach to information maximization,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.209556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.063516Z digest=sha256:beaccead5c1a8fbbb340ccd57d8eb85b4838f2d496ee070081169e3d13576e77

Observation 4c1b5286-a103-4fe4-ac63-cfef51d5a6bb · outbound

This paper cites Learning-based model predictive control for markov decision processes,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Learning-based model predictive control for markov decision processes,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.155133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.092320Z digest=sha256:48c51bbee6a4741a981e81e1287b1eac969dc761ea8c542db1f8a5f1775507a6

Observation 55fd597e-cee8-4b0d-92d6-e2bed94a51f7 · outbound

This paper cites Optimization of computer simulation models with rare events,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Optimization of computer simulation models with rare events,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:14.058320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.122904Z digest=sha256:83398680c3d1632656cb6492bb4938562c86b68f2e4a25603a1561b72038d73d

Observation bbf4ba56-0ff9-4a24-9d8a-20309191a49d · outbound

This paper cites Plan to predict: Learning an uncertainty-foreseeing model for model-based reinforce- ment learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Plan to predict: Learning an uncertainty-foreseeing model for model-based reinforce- ment learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.934751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.152757Z digest=sha256:c5dd6822a628926bcab01892cf874bddb3978706a1b3c96191c45798aca3c2f0

Observation 2ba40424-c73c-4155-b3b7-e6a2ca0c21d5 · outbound

This paper cites Choreographer: Learning and Adapting Skills in Imagination.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Choreographer: Learning and Adapting Skills in Imagination

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.179782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.179782Z digest=sha256:81b579704754a9b0091059bfcaf7b92b160b26b70d668d780ae73c7e7f858996

Observation df786e93-e2c2-480a-97e1-3820ec7f11a7 · outbound

This paper cites Mismatched no more: Joint model-policy optimization for model- based rl,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Mismatched no more: Joint model-policy optimization for model- based rl,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.784917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.224823Z digest=sha256:aae64358692401689cdc629fe735df7eb39603a7aa90837815295ad4cd1245be

Observation c677bc8f-2160-41dc-b0dd-0eb365f470d0 · outbound

This paper cites Differen- tiable mpc for end-to-end planning and control,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Differen- tiable mpc for end-to-end planning and control,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.705206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.284743Z digest=sha256:64dabce71f9b0a10eeaa2d89a3f496a8a5f5dd2d05e7f84f4ae8abde290ae41d

Observation 70895f64-e819-42f1-b1ac-c051685bf061 · outbound

This paper cites On information and sufficiency,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes On information and sufficiency,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.329785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.329785Z digest=sha256:c9027e439f584aad764c12cc74ca3ce3cb4ac849fbe772f4a6fbbf43e2d7edca

Observation e934efb5-2c2d-432e-a0a3-ba4114c01bd4 · outbound

This paper cites Measures of distance between probability distributions,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Measures of distance between probability distributions,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.588551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.374768Z digest=sha256:0ac6bbc877736e9ca257eadd1c34690c09db5b9628a2a0878ed8c14dbee7cca4

Observation 3a1cab5a-896b-4d5f-ba88-fbc9d853d983 · outbound

This paper cites Dr. Strategy: Model-Based Generalist Agents with Strategic Dreaming.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Dr. Strategy: Model-Based Generalist Agents with Strategic Dreaming

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:11.155546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.424745Z digest=sha256:581f4371bb92f2ce2b932c99c0da9d4474a1e92b07b3963996f151021412d7cc

Observation ce6fd07d-e6c5-4bb2-a43a-1b59251b0e2e · outbound

This paper cites When to trust your model: Model-based policy optimization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes When to trust your model: Model-based policy optimization,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.507143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.486673Z digest=sha256:3425c7e48a963fbf6fe844b2c969a3b31aeacec07be35077c143b9afc70499c4

Observation e6b81818-ff75-43bf-b7aa-7adf24c03920 · outbound

This paper cites Model-based reinforcement learning via meta-policy optimization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Model-based reinforcement learning via meta-policy optimization,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.411043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.505310Z digest=sha256:1f54afaf4abfed83c18418735e64c9e21d99d322db33015e8f8b0dfd18ea0c0b

Observation 24d689a0-cebb-460a-ba81-f0f1b7f8df2c · outbound

This paper cites The virtues of laziness in model-based rl: A unified objective and algo- rithms,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes The virtues of laziness in model-based rl: A unified objective and algo- rithms,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.304750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.538880Z digest=sha256:8660cbdbf1f1bf954192b6b0d95522db310ce245dcbbacb1f4e337429087f778

Observation c3bcd5d5-fc70-48c4-ab46-475e1479b127 · outbound

This paper cites Iterative value-aware model learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Iterative value-aware model learning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.193218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.563874Z digest=sha256:1ee8c65e608c2a60187f28586d115ec7d2ce565b213011cf1509ec6aa924c6df

Observation 7eb6f6a4-bf98-48ea-8459-17304275825f · outbound

This paper cites Model-based value expansion for efficient model-free rein- forcement learning,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Model-based value expansion for efficient model-free rein- forcement learning,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.100233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.571543Z digest=sha256:6797a4a5d9d68766d3e54e999cc2d542fdcb5034937103604f1bfe7ba19e54ec

Observation 7d7bce9d-9972-4c4f-8e45-2228a24eb222 · outbound

This paper cites Villani et al.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Villani et al

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:13.025685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.583592Z digest=sha256:600594ca7aa7d4425674d2731055ba34d42ab269993d3bd372528b58650b142c

Observation 6a423560-c737-4b7b-9eb0-ba5fc99a2d78 · outbound

This paper cites Generative adversarial nets,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Generative adversarial nets,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.607965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.607965Z digest=sha256:816183817536386d26a9f50299764471f8bd34d12d90ddeddd457656e3d74693

Observation 21fad6e0-ba0c-4e0c-94a5-bf23d2e010f7 · outbound

This paper cites Towards Principled Methods for Training Generative Adversarial Networks.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Towards Principled Methods for Training Generative Adversarial Networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.634751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.634751Z digest=sha256:ffb092894003fa8de4cf1536104ac3c9006fa873cd6a74cc9a696b1f492e7869

Observation 1629a347-0aa7-448f-853a-47d4f3ec0bae · outbound

This paper cites Wasserstein generative ad- versarial networks,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Wasserstein generative ad- versarial networks,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:12.814759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.684749Z digest=sha256:4296db8e457027f52bbf9f3f40211829a7c5e45ff9fa58126f1a1c59e9f26d54

Observation 2a9e0545-bac5-4c2e-9959-ff8fe7f60c67 · outbound

This paper cites Harpy, a connected speech recognition system,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Harpy, a connected speech recognition system,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:12.704852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.734752Z digest=sha256:742fca97dabc7bbd28a78d81bf7797bd57002d7c52359392d58fa017e2f622bc

Observation 44c6830e-c21a-4e80-bd0a-d0cdc6c7c663 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.784750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.784750Z digest=sha256:d23e112a42b79163853e6724b49a65b280c55328b7bdf0d1e435ab77f6ff70d4

Observation f74544f5-fdee-48fc-bab3-6d7f128b19fd · outbound

This paper cites Walk Wisely on Graph: Knowledge Graph Reasoning with Dual Agents via Efficient Guidance-Exploration.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Walk Wisely on Graph: Knowledge Graph Reasoning with Dual Agents via Efficient Guidance-Exploration

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:10.883459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.822192Z digest=sha256:6b66b1900cd58d13b8c7f98fa547e6cb9ef6064c43d9b3ffe9afacd84813068d

Observation 4ae52c77-d93e-43f9-b08d-39c637e62b6b · outbound

This paper cites Spectral Norm Regularization for Improving the Generalizability of Deep Learning.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Spectral Norm Regularization for Improving the Generalizability of Deep Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.864761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.864761Z digest=sha256:854267a4a51abc28982c8c887d390864e63fa2590c8f8eaa357ad9033c12ee27

Observation a214a2d9-fd2c-4359-be98-2a65c8aed1b3 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Categorical Reparameterization with Gumbel-Softmax

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.893528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.893528Z digest=sha256:60d4db096a177ab6a3b508a9ba7888181e33e997698e421261fc8a998dc1db86

Observation 5a172a2a-29bd-4124-916a-e77d37a0d58b · outbound

This paper cites Introduction to online convex optimization,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Introduction to online convex optimization,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:12.603409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.921439Z digest=sha256:5c8e91fc68cecd5a1dee2350d31aee5210ae2261a483e9f03c391e4d153a4393

Observation 6d0d3d94-3f37-47e5-841f-80391b92d272 · outbound

This paper cites Is Model Ensemble Necessary? Model-based RL via a Single Model with Lipschitz Regularized Value Function.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Is Model Ensemble Necessary? Model-based RL via a Single Model with Lipschitz Regularized Value Function

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:10.344751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:09.964752Z digest=sha256:d51d90384915f96b0bfdee7969b1bc1800c6d403ba8719d522eebfd535b2cc07

Observation dd4be036-c677-46ec-81b8-6c3558dcd1b7 · outbound

This paper cites Visualizing data using t-sne.,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Visualizing data using t-sne.,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:09.994817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:09.994817Z digest=sha256:56ddc9b4c7515fafd931da79a4bc9e1b3cddea5af65b88dd83901f543c3c1993

Observation 10f6f2d4-73a8-4e7e-8c07-df67b7a7e86f · outbound

This paper cites General boundary conditions for denumberable markov processes,.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes General boundary conditions for denumberable markov processes,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:12.519612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:10.044751Z digest=sha256:1dcc4853bb8670d7dbb186dd75736ae0f1c7c7ead0056035b31e1749a23467f9

Observation 51b87c37-82ee-431a-aae4-056eb7bf6398 · outbound

This paper cites an unresolved cited work.

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:12.478583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:13:10.092408Z digest=sha256:7e1697c9f0840ea7fee2ace1617ae29c21831673316fa353c350a9179af732b7

Pith citing papers

No inbound Pith citation observations are available.