Pith. sign in

Paper Citation Record · LEDGER

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.12095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12095 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:27:30.592023Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e260941e-9d6f-4510-b2ae-7ccda0e66247 · outbound

This paper cites Learning humanoid locomotion with transformers,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Learning humanoid locomotion with transformers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:32.278036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:29.499106Z digest=sha256:d9e10481ee071108fc553028afdc8f523f179ca9d853e90d5fa604e647fb1885

Observation d7889185-0ee8-4d79-918b-ee4f71d97b26 · outbound

This paper cites Real-world humanoid locomotion with reinforcement learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Real-world humanoid locomotion with reinforcement learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:32.115658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:29.537410Z digest=sha256:d0293dfc39093bbf7c90c11673cf5a53427c92e26a9465485e1b045c1ce25e11

Observation 52a8d7ac-5b49-41e5-8d74-14e8d27525b2 · outbound

This paper cites Model predictive control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Model predictive control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.932501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:29.589024Z digest=sha256:7f7f2e3ec7ad93089659eb527f8e268c65ed32115f887d54b999f2f8988ce015

Observation f4bab07a-8f8e-4064-9cf6-977f1283d99c · outbound

This paper cites Aleatoric and epistemic uncertainty with random forests,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Aleatoric and epistemic uncertainty with random forests,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.656935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:29.628274Z digest=sha256:a39a8e16b5e8691bca4c0fd086476059ac755301697f3d33f4e684e557445e06

Observation 87224d6d-60e6-4b83-a8dd-19afe8e017c6 · outbound

This paper cites Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:29.696736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:29.696736Z digest=sha256:31503af4257367d2c31d3199b58785c4b73392bcad2c55bcda41e4fa275ca884

Observation 4d584cc8-db7c-4a6e-81f4-d036127ff71f · outbound

This paper cites Stochasticity in Motion: An Information-Theoretic Approach to Trajectory Prediction.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Stochasticity in Motion: An Information-Theoretic Approach to Trajectory Prediction

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:27:30.960812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:29.748895Z digest=sha256:f12c979b761b7a075a0bedc5dae19975197c3caf45cbab039c95f84891119252

Observation 8b5c315c-d480-435d-84ac-a50798c6409d · outbound

This paper cites Learning Through Retrospection: Improving Trajectory Prediction for Automated Driving with Error Feedback.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Learning Through Retrospection: Improving Trajectory Prediction for Automated Driving with Error Feedback

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:27:30.909236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:29.806038Z digest=sha256:94aaa181df449b5c119b82ae72ef6e591432e8762ac9865f640caf93ad04a001

Observation 878a5feb-84b3-4579-beea-fa4067b7500c · outbound

This paper cites Temporal difference learning for model predictive control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Temporal difference learning for model predictive control,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:29.890578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:29.890578Z digest=sha256:5f6c22e72f75e751f1df5807acbb405ec55ea3b22cf55af814567c909de88817

Observation 0e2d9b79-dbe9-4b1f-9065-b6a28d55aaa9 · outbound

This paper cites Td-mpc2: Scalable, robust world models for continuous control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Td-mpc2: Scalable, robust world models for continuous control,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.541850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:29.976944Z digest=sha256:a632e19b9b47fd47bf35005021730c672cc2af378da16ef91b989b1036767082

Observation 56bbb6f7-6191-4408-94d9-414ebade58b5 · outbound

This paper cites Deep reinforce- ment learning in a handful of trials using probabilistic dynamics models,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Deep reinforce- ment learning in a handful of trials using probabilistic dynamics models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.489789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.035313Z digest=sha256:3553d003cc52e9f982d83338e31523eecd89f5bec9d4b73df021d3c302e6c322

Observation 70dd384c-1d61-43f9-ac3f-887c72d72b08 · outbound

This paper cites Model-Based Offline Planning.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Model-Based Offline Planning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.087984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.087984Z digest=sha256:98fb45be581d3af122c803772807dc78ffade0092cbb97706cfec1df1d7b41a7

Observation a772cadb-ca4b-44c8-844c-52e155ff9de9 · outbound

This paper cites TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.170831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.170831Z digest=sha256:90024e56d2e8798301f9974485fc12332bc7ca513d083ae28cedaf2b3fc98e9e

Observation 986fb49b-5f58-42ba-9610-d3c7d1488490 · outbound

This paper cites V ovk, A.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion V ovk, A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.198273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.198273Z digest=sha256:4b9db5d5f9020922be95d2aa2619960473b30b8900b93983284f07790fefd830

Observation 6cdcabf3-1b02-4939-94c8-9de1834b17a6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.305025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.305025Z digest=sha256:958532ee0a44193cecf72f48d753e4c2868216235a2318aa5c5607e13673f97a

Observation 8b14b05d-c8fa-4c20-b22a-d15cb39aa7d0 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.356457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.356457Z digest=sha256:2930590047c946f7a1ca5730130a903dcd9e8d4a83742acb0abae3f6a839ed6e

Observation c337b9cf-f2b7-4f20-9d21-9d6a3f968b11 · outbound

This paper cites Reinforcement learning for humanoid robotics,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Reinforcement learning for humanoid robotics,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.425332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.415385Z digest=sha256:c97fbdef0675dbf53b2f0c203740321f9cc03a521bc6f88e1b8c5722705e9757

Observation 11b382bb-3341-4f1c-90fb-c1db1f7a0b97 · outbound

This paper cites Learning off-policy with online planning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Learning off-policy with online planning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.380410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.423302Z digest=sha256:6f0a8fc7b5a79ae4665b2487c6996ed5c342af0985823010fcf9f8214b14d713

Observation f123cf6a-1eea-4d54-8c5a-9b4320964f46 · outbound

This paper cites Conformal prediction in manifold learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal prediction in manifold learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.357234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.428872Z digest=sha256:26e750ac822238b3d3c0c89d53c76e21b5abf0b8df60ee9cd3e096e62dc87cd2

Observation 614f192f-7e1f-4661-83c0-acde746bf96c · outbound

This paper cites Conformal Prediction with Learned Features.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal Prediction with Learned Features

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.435680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.435680Z digest=sha256:cebb192279fe8e12822df8ee97c7452702e35f3e32c1f6750fd549142e41e169

Observation 2d667aa7-e41b-4da9-884c-c297c762fdd9 · outbound

This paper cites Confor- mal prediction for semantically-aware autonomous perception in urban environments,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Confor- mal prediction for semantically-aware autonomous perception in urban environments,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.338215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.445140Z digest=sha256:a9ee873482ee6c67702ae09a21911506401363b4d02bc87412f5fc29df2b6907

Observation c3a52c1e-a316-4937-83dc-e277ce094e56 · outbound

This paper cites Conformal prediction for uncertainty-aware planning with diffusion dynamics model,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal prediction for uncertainty-aware planning with diffusion dynamics model,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.307658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.452592Z digest=sha256:96346187e011c1987b9248e9a49ad3803b6adb95234871a7ca6b7ad82ecef932

Observation 07eba7e6-feff-4612-8708-cba7edc3e584 · outbound

This paper cites Adaptive conformal prediction for motion planning among dynamic agents,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Adaptive conformal prediction for motion planning among dynamic agents,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.286345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.459658Z digest=sha256:075495d844e23424c3516b7b4a6e281bda99727e70776467752b27c378d53ccc

Observation 82632f0d-106c-4c04-aac0-bbe697569ced · outbound

This paper cites Safe planning in dynamic environments using conformal prediction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Safe planning in dynamic environments using conformal prediction,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.267133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.467615Z digest=sha256:0fddc5d8cb94fc9db155c1bae12f936cf313b7553d14f4b85d0cd44fab9ed92d

Observation 6f144c20-bd22-4315-af55-7676327a4b1d · outbound

This paper cites Safe perception-based control under stochastic sensor uncertainty using con- formal prediction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Safe perception-based control under stochastic sensor uncertainty using con- formal prediction,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.246167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.473135Z digest=sha256:e8d4d862a90d007cb59aa626b80c8502a2487480178fc22337f3d1dce83dc6a8

Observation 94498e85-f25f-4691-9792-b54b2eee61c2 · outbound

This paper cites Conformal decision theory: Safe autonomous decisions from imperfect predictions,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal decision theory: Safe autonomous decisions from imperfect predictions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.229092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.480932Z digest=sha256:33e87aef4c4529272a7bdc4945d557beff9c228a871de6cf5f994fcad80dc3f2

Observation 668e69c2-056f-45a2-9bbf-c26ed3c50442 · outbound

This paper cites Safe pomdp online planning among dynamic agents via adaptive conformal prediction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Safe pomdp online planning among dynamic agents via adaptive conformal prediction,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.206160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.486796Z digest=sha256:74e018ad9a73634275926242dd4b7b04582192bd087a8813c8f23521b6048789

Observation 9982d18c-a52c-4002-948c-34f7c9562c1e · outbound

This paper cites Conformal policy learning for sensorimotor control under distribution shifts,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformal policy learning for sensorimotor control under distribution shifts,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.180556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.493848Z digest=sha256:13f80336058079b63ed9099a58f5358a3cefb72a88f4a5fe053b12010a5ae672

Observation 45659071-8e67-416d-9c9d-f0cb3d6260ba · outbound

This paper cites Conformalized Teleoperation: Confidently Mapping Human Inputs to High-Dimensional Robot Actions.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conformalized Teleoperation: Confidently Mapping Human Inputs to High-Dimensional Robot Actions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.502164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.502164Z digest=sha256:27376d3eab71dddb9ab9e73693d4237d6261863b18a83c38b3b1cf658ba17e57

Observation 129a9df3-0ca0-4cd3-9f37-ab4523e42d85 · outbound

This paper cites Stabilizing off- policy q-learning via bootstrapping error reduction,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Stabilizing off- policy q-learning via bootstrapping error reduction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.157675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.512999Z digest=sha256:8e034684eb9e5122fdac42650578ece442499efd60d5dafbd1b7d229742a143f

Observation c8364e28-c826-42d7-b104-9488c5a1d8bd · outbound

This paper cites Conservative q-learning for offline reinforcement learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Conservative q-learning for offline reinforcement learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.135640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.521022Z digest=sha256:827c9ad79e4188cd16a0d7de918c52157175351e6c68b186078d5110536ebec9

Observation 03d3fb43-b9bb-442d-b1ed-df58065d65db · outbound

This paper cites A minimalist approach to offline reinforce- ment learning,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion A minimalist approach to offline reinforce- ment learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.103450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.527463Z digest=sha256:7e9db7a3724b2311ae9f709978200de8df8b84816fbb8c52b535d499cdd4634b

Observation f0913dee-c670-42da-9b0f-7160d230e533 · outbound

This paper cites Off-policy deep reinforcement learning without exploration,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Off-policy deep reinforcement learning without exploration,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.079366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.533658Z digest=sha256:108d343e15e3f06a16d67a157158585df3cd07f4d455ed9f5efa2a8463cf256d

Observation f69c47d4-639e-4bb8-a9b2-5ee79a991814 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.542382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.542382Z digest=sha256:82cfebe9d176873b464c7cb3b1b8838737309ed2240905674dbac0059fb8a8cd

Observation 6b556c63-0a75-4c9d-a391-14da58e48501 · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Extreme Q-Learning: MaxEnt RL without Entropy

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.548506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.548506Z digest=sha256:91cffae66b0f16711fcfbda32adcf3877f8ef91efe49cee3ef4375906afb962c

Observation 05760fd6-b9e7-4ee7-9971-fb4e4a86156d · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Offline Reinforcement Learning with Implicit Q-Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.555683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.555683Z digest=sha256:ec8511bf2e5b7ec6ca698924c51ad0d55d3a1555105213cf720b8655beccee7e

Observation d9f23434-f2ee-4508-8fb0-c5c7e8889472 · outbound

This paper cites Aggressive driving with model predictive path integral control,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Aggressive driving with model predictive path integral control,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.561578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.561578Z digest=sha256:23303c41fb3e6dc40a57031bacd68ee202d394e2c2fb70f0a4c5a4b1493b1b25

Observation 4e8df83b-7ac8-4750-ad88-c1415d4c9e99 · outbound

This paper cites Trust region policy optimization,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Trust region policy optimization,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.569861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.569861Z digest=sha256:01ba6f00ed88312ae5da871a03ae45bd4c207fba9ebd84d593b4e937876d62db

Observation dc279f4d-39da-4c85-b56e-c235f4528584 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:27:31.019435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:27:30.576382Z digest=sha256:58a3bedc75490ceb02b4590acbb902d270e178ded78c67d9dd897e4df4bde894

Observation f335dd01-54b4-4e0e-88e6-d9ab13ca0f86 · outbound

This paper cites Imitation is not enough: Ro- bustifying imitation with reinforcement learning for challenging driving scenarios,.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Imitation is not enough: Ro- bustifying imitation with reinforcement learning for challenging driving scenarios,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.585942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.585942Z digest=sha256:df12946fd76f008a33f241bd74cffeace05005cbfb5c1abb11e040a9e59a2058

Observation 4a59ac21-e20c-4335-982e-fa8a68de9b1a · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.592023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.592023Z digest=sha256:3b0a0f0cbe86de1770dc667d6488d711d6475252aa77de641e50f17cf0b2f16b

Pith citing papers

No inbound Pith citation observations are available.