Pith. sign in

Paper Citation Record · LEDGER

Goal-Conditioned Reinforcement Learning: Problems and Solutions

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2201.08299.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.08299 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:44:58.091912Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T03:24:28.836511Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b9f768c3-f6c6-4e19-8ced-8946163901b7 · inbound

Vision-Language Foundation Models as Effective Robot Imitators cites this paper.

Vision-Language Foundation Models as Effective Robot Imitators Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.676683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:64a04c74d9ed2d8c03d46fe9d32ce34eecb6e0c5a931300e59267569e7f307c7

Observation 4b876560-7f52-4f3a-b377-c66222a2351a · inbound

Upside-Down Reinforcement Learning for More Interpretable Optimal Control cites this paper.

Upside-Down Reinforcement Learning for More Interpretable Optimal Control Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T18:35:59.706912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:35:59.706912Z digest=sha256:c390f1ca4316130be47f93e4edc1c6c23800a544effc0f69627f924471117b35

Observation 39a3ff93-8bdf-4c5d-93fd-206c757a1307 · inbound

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering cites this paper.

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:47.893669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:28:47.893669Z digest=sha256:622fd87cdb6da12819e9ca5815e2f6ab05cb97eb72de6d9ee6541da70c5310f2

Observation 3006316e-d682-4cc8-836a-8a83f0128be4 · inbound

CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives cites this paper.

CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T05:56:42.660682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:56:42.660682Z digest=sha256:fdb66698338595952939cae860ed8351fae64c1872d57f7ca77859177434b0dc

Observation 3da55db4-9c97-4bda-94d3-71279520a723 · inbound

Goal-Conditioned Supervised Learning for Multi-Objective Recommendation cites this paper.

Goal-Conditioned Supervised Learning for Multi-Objective Recommendation Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:06.969586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:06.969586Z digest=sha256:00b29ec7d15fbc24b6d6bc5a4f42e483104bbfeee3c8ccb163c20e530533ba8e

Observation ef1f48bc-593e-4995-8e6e-c2c5dde228d8 · inbound

MGDA: Model-based Goal Data Augmentation for Offline Goal-conditioned Weighted Supervised Learning cites this paper.

MGDA: Model-based Goal Data Augmentation for Offline Goal-conditioned Weighted Supervised Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:44.169825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:02:44.169825Z digest=sha256:97173620f56c1d0f6309e0ffceb810f34c35174dec00da56365592cdec2d50fd

Observation 161940f6-fb9b-4b21-98ed-2522a694d5d1 · inbound

Physics-model-guided Worst-case Sampling for Safe Reinforcement Learning cites this paper.

Physics-model-guided Worst-case Sampling for Safe Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:51.265160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:51.265160Z digest=sha256:8ff8551d2279bae2a10015e23cb81ad0d6641d7dda9c9ea52beba6b66d4526d6

Observation 21986ddc-3d25-41ff-80b2-d8ad2f3d0980 · inbound

Generalized Back-Stepping Experience Replay in Sparse-Reward Environments cites this paper.

Generalized Back-Stepping Experience Replay in Sparse-Reward Environments Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:23:52.156041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:23:52.156041Z digest=sha256:fef30855f014882481f82947c968e849c251912db54c057d668777e2ddc0b0ee

Observation 96dcb284-a10b-410c-865a-7512b0def239 · inbound

Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning cites this paper.

Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T14:28:19.278115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:28:19.278115Z digest=sha256:c5301935e7fcbb9ef64f2f43159d9c030e1fa57c019b693414b11db336ea0c7c

Observation d5e18518-2e86-4153-bfd2-06607d1a89fd · inbound

Adviser-Actor-Critic: Eliminating Steady-State Error in Reinforcement Learning Control cites this paper.

Adviser-Actor-Critic: Eliminating Steady-State Error in Reinforcement Learning Control Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T12:49:59.106851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:49:59.106851Z digest=sha256:5e14ba251a9f6ab68af8e22d35daa5621be79ab5c1d073dfb189bdfee8863709

Observation 106e6232-3df4-4967-9c6b-b83b751e6ca6 · inbound

Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning cites this paper.

Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T11:25:49.129256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:25:49.129256Z digest=sha256:d6441ad035aba3ffec23edcec1b54dbb21b86d263ccc56ebf0fe4a10e2a4937c

Observation a4d8aa6b-f90d-481d-a873-3033fb4c3260 · inbound

GRAML: Goal Recognition As Metric Learning cites this paper.

GRAML: Goal Recognition As Metric Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:58.091912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:58.091912Z digest=sha256:fd926d74225c3e5e81795d5469d5f3eda22a8eeb9bb75b31b4eaaa4bbf8d8e34

Observation 72854cbe-c926-437f-b070-49ce31435517 · inbound

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion cites this paper.

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:52:17.883282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T12:50:17.902979Z digest=sha256:816679e568c0902ec887fa9b2c16c70d93dd2a3083fcce41a9f00a20f04217ae

Observation 1af4b4c1-774d-4db5-a01a-53e765946a7f · inbound

On the Parallels Between Evolutionary Theory and the State of AI cites this paper.

On the Parallels Between Evolutionary Theory and the State of AI Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:08.894905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:08.894905Z digest=sha256:27b0813d967dccd64364697c975f3deb4acf6a7924f449261a473d4268bd4797

Observation c5d1b977-6ba4-481b-a841-4a852e7a9bbc · inbound

Reachability Weighted Offline Goal-conditioned Resampling cites this paper.

Reachability Weighted Offline Goal-conditioned Resampling Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:24:59.397642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:24:59.397642Z digest=sha256:003c9242a686295ea63f31163dd573bd5b555f2172e13130620801cd85d3bc21

Observation fb6b9eb2-7d20-4713-b7cb-c33da20962a7 · inbound

Reward Models in Deep Reinforcement Learning: A Survey cites this paper.

Reward Models in Deep Reinforcement Learning: A Survey Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.662293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.662293Z digest=sha256:5f5187d8ca26366f85e45fa5c53543ac058799ea0b478de3710e309a4be1499a

Observation 7974a871-e96a-47a5-9f66-4fe6ab15f670 · inbound

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models cites this paper.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:30.107695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:30.107695Z digest=sha256:6b4d13bea2842d0d8f093dcc30f4c8130ca68243c1f0a68a1a3a762d04ff3c61

Observation a2a05581-2fcd-41d8-8b9b-66230649f4ba · inbound

Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning cites this paper.

Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:54:31.197090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T00:53:46.002945Z digest=sha256:c9abb50bf1c4f60e4b30a0db325e4c46a9bb3e7a59c9c7c05916cd1e4477325b

Observation 4ea218e3-983c-4711-9946-e715fba6e942 · inbound

Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving cites this paper.

Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:27:36.349835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:27:36.349835Z digest=sha256:a8694713239e412428763427e3de1ecbbc528b530f18c749a768ca5f66fbd383

Observation b1ecc38e-6980-4d53-9207-fa812b98ce3b · inbound

Equivariant Goal Conditioned Contrastive Reinforcement Learning cites this paper.

Equivariant Goal Conditioned Contrastive Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.504219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.504219Z digest=sha256:718ddab94cfac3555e987d834ee95980f396e4734f8f8790c42ee7cafe3b2981

Observation f57e4368-b5c0-4e4f-b89a-b56b308c0b8a · inbound

Self-Curriculum Model-based Reinforcement Learning for Shape Control of Deformable Linear Objects cites this paper.

Self-Curriculum Model-based Reinforcement Learning for Shape Control of Deformable Linear Objects Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T20:56:58.581176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:56:58.581176Z digest=sha256:c16c379fa7f6d4b15d4ed6175316686407c7e3966ef9b8ebefb5780b3dba42e3

Observation 3ea200fb-773a-400b-8b47-c9146e1a8b61 · inbound

Efficient Hierarchical Implicit Flow Q-learning for Offline Goal-conditioned Reinforcement Learning cites this paper.

Efficient Hierarchical Implicit Flow Q-learning for Offline Goal-conditioned Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:58.280249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:12:32.839681Z digest=sha256:97e444ce0cf4be48889f6ce1445b216db43cb935360fe4bff8537bada22514ae

Observation fe268555-00e9-4533-96a9-0d2d9d6035fc · inbound

AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking cites this paper.

AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:34:47.312157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T00:34:26.106672Z digest=sha256:14558c788cbf78ca46996e66433c8b714e5563e37856138519745681632932df

Observation 56a29c79-bb3c-482c-925d-d536bda25bad · inbound

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning cites this paper.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.179747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:d3227bc1423bcc62bbcd976f4918a320112f8c8c3f7900a1fcbb3a6285548a03

Observation 738fbe2b-4b8d-41aa-b108-1fac1b1be53e · inbound

GCImOpt: Learning efficient goal-conditioned policies by imitating optimal trajectories cites this paper.

GCImOpt: Learning efficient goal-conditioned policies by imitating optimal trajectories Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:14.892367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T11:26:22.149056Z digest=sha256:ab9c7acd6d24c5a4925f4fe496f7ba6dcc0cff7f301928ffea2eeaf8919ac0b1

Observation 14498daa-0227-42f1-a0fc-433d5d8fe9a5 · inbound

When Policies Cannot Be Retrained: A Unified Closed-Form View of Post-Training Steering in Offline Reinforcement Learning cites this paper.

When Policies Cannot Be Retrained: A Unified Closed-Form View of Post-Training Steering in Offline Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:49:16.044470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T22:15:16.221059Z digest=sha256:0eeedb40acfe3b0f0bd93eaf8157fe776f2928649c2f2a3d3b841ab02ad2c07b

Observation 2d00f843-1424-42ac-8c40-fc1ba8dcdd41 · inbound

SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning cites this paper.

SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:30.363716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T04:01:31.928056Z digest=sha256:507011836ffd704ac8a3ba1ab66e74e79be5d8c8c59f74935fba569c7c1f7074

Observation df4f7c2a-1c27-4f4d-b391-6da788d9c30e · inbound

Improving Zero-Shot Offline RL via Behavioral Task Sampling cites this paper.

Improving Zero-Shot Offline RL via Behavioral Task Sampling Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:41:18.632617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T16:27:15.347522Z digest=sha256:99cb7ded1e2fce189ca24309d5848e8cb5698efa064583fe1a78cb10d4460673

Observation dd71071a-da58-4d56-8dc2-c48e975db8a1 · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:58.762428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:a8eeaf807edf6942cffd3679bdab687e2092d72c9eb7139f4006a38c65ae96b0

Observation 6cb6eebc-8701-4752-a0d8-752b89664768 · inbound

Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation cites this paper.

Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:27.293862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-12T03:36:24.941205Z digest=sha256:d9142b3b3327553e65c176a84aa8eaf5101115f646cc2d4a72628e7c28ec5548

Observation 1b2e9950-94a1-48fc-9d8b-069ac64a58dd · inbound

Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning cites this paper.

Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:30.552367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T07:28:24.455817Z digest=sha256:7f20784543dda409e03e878c5fc348b1dd5f6c16376b003df69c0beecac3fdbb

Observation f9442c02-19e0-4661-a17f-b2f8c42540f0 · inbound

Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning cites this paper.

Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.042051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T22:48:55.661356Z digest=sha256:3673454d7cfcd7bf526257aeee24f906ad757af5d7fcfa3e84c75005f75f7b66

Observation 9e615d03-41ff-43dc-8644-a2710baab1c4 · inbound

Goal-Conditioned Supervised Learning for LLM Fine-Tuning cites this paper.

Goal-Conditioned Supervised Learning for LLM Fine-Tuning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:39:10.101400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T22:37:46.345159Z digest=sha256:1bc0954177ca717cb5ab65e39a828a1a315e90823b28f1315ad69544ea11843f

Observation 0cc70f80-1952-4bd7-b18d-166027b674c9 · inbound

CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning cites this paper.

CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:44.914592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T13:54:52.129738Z digest=sha256:0a1e2c16ee22f9bb8d688cc1885a0af5b84b86828826365f6a4aa893279b971d

Observation 37a7e5ff-c5f6-486f-a464-71898833ec2f · inbound

Decoupled Behavioral Cloning for Scalable Inductive Generalization in RL from Specifications cites this paper.

Decoupled Behavioral Cloning for Scalable Inductive Generalization in RL from Specifications Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:42:37.703603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T18:25:37.397393Z digest=sha256:168301d60d02fd142970a4e523ef1ee92a7304d5345df6a2f4bd82f1994742ec

Observation 9b1e5aaf-2f60-4689-916f-f6175e76c5b8 · inbound

Dual Advantage Fields cites this paper.

Dual Advantage Fields Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:27.870593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T10:43:13.615850Z digest=sha256:00cc91a56879d33b9ead3c5bbe909d5c336fc2cc69a0f7d29458060200d93d47

Observation 3de0bea2-bdc0-4011-88ed-02739649657b · inbound

GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation cites this paper.

GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:42.059834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T13:01:33.109009Z digest=sha256:4c8d983d4401451b55247e67d83e231b711cd39dbc12211701388524b9a4ea26

Observation 2c96bc35-26b7-44d2-8ffd-c6054725e69c · inbound

Learning Object Manipulation from Scratch via Contrastive Interaction cites this paper.

Learning Object Manipulation from Scratch via Contrastive Interaction Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.436177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T10:10:21.427118Z digest=sha256:702038bb2388ea1845a375e72d5534e6d513b1650b8afda5be9ac6e826fcf77a

Observation aa3b42f0-f16f-47a4-96c5-f77931f3d7fa · inbound

Embodiment Shapes Rolling Behavior in a Multimodal Infant Model cites this paper.

Embodiment Shapes Rolling Behavior in a Multimodal Infant Model Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.261209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T01:08:24.273577Z digest=sha256:9884d49ec5d411708fc062fef6b666678659c484aaf2d7ccba6e42e4609c51ea

Observation 1faef4cd-863f-438d-9025-a7b4142d9312 · inbound

World Models in Pieces: Structural Certification for General Agents cites this paper.

World Models in Pieces: Structural Certification for General Agents Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:00:01.600010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-25T23:04:11.950176Z digest=sha256:c2f8e6b04ad2ce542d885b4c2171dc2093e22471575be734ea5143f150dbc869

Observation 6edea12b-d8cc-420d-806a-21d0d2a678ee · inbound

Coachable agents for interactive gameplay cites this paper.

Coachable agents for interactive gameplay Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:56.521838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T12:52:05.010028Z digest=sha256:e5a812c7230db2a8c138d7e7acfb17611f127a0175dee8404b05d374806c4d8e

Observation fee86fc1-268c-4449-afde-bcd9c2db1f0b · inbound

FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games cites this paper.

FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T03:24:28.837788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-08T03:17:09.608451Z digest=sha256:d0519867ca5e8df445fa8ac3cb853c40e9f0b011df9f8206a86385a35c32a5b6

Observation 502c102c-5cce-4c05-baf4-bccce6d4c1bb · inbound

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer cites this paper.

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T08:30:10.584105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:30:10.584105Z digest=sha256:7b8aad99eb10407ba93b3ad8ab4fa9f6a313c0600f3da3275cd4648fa0be9831

Observation 8c3fc251-e5d9-4816-a034-a9522b7f4a3e · inbound

DAGR: State-Conditioned Goal Representations via Difference-Aware Goal Cross-Attention cites this paper.

DAGR: State-Conditioned Goal Representations via Difference-Aware Goal Cross-Attention Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T04:15:53.392271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:15:53.392271Z digest=sha256:8e8757ea0ab9eed34901f31bd2b6c0d23028799b5d6f72bde1d868e40732dff6