Pith. sign in

Paper Citation Record · LEDGER

Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:1909.10618.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1909.10618 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:26:50.472689Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T18:33:19.672856Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 91f1ad85-01ad-4bfd-9515-72db994af8e0 · inbound

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach cites this paper.

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.676006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T18:30:38.818711Z digest=sha256:97bb1feaa77e2b9bc619100b1f489b5bdd75bcd4905b5d868d2efd95773a380b

Observation 17db5c88-0c10-401c-a9a9-cdc1125de1a9 · inbound

Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning cites this paper.

Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:54:31.216147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:53:46.002945Z digest=sha256:e9d8ad8c8f972de9738fcdd3211779189649941fa4b006ee4842ea033ea86f94

Observation 885896cd-dfc5-465f-8fc6-464fde77d96f · inbound

Unsupervised Hierarchical Skill Discovery cites this paper.

Unsupervised Hierarchical Skill Discovery Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:50.472689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:50.472689Z digest=sha256:a68de63fb91cb30d020372b4ae1c2b15b7fdaff34f80068b51a07d2253bfbca3

Observation ee3c3102-4c4a-4b30-ba41-4e8909a376b1 · inbound

Learning to Theorize the World from Observation cites this paper.

Learning to Theorize the World from Observation Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?

Reference 197

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:21:30.269890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-07T17:15:43.429602Z digest=sha256:77b4d5da4b228f8e789a797a6ebdb68ad2fd9d62b345c030318e083cee36a28f

Observation 0b525fab-8dca-4884-9b5c-020ec92a322d · inbound

Quantum Hierarchical Reinforcement Learning via Variational Quantum Circuits cites this paper.

Quantum Hierarchical Reinforcement Learning via Variational Quantum Circuits Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:30.382581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T17:24:49.763003Z digest=sha256:917b9a4eaa4ffc071e0f7f62cb8f9054dcb75ed06cccd53a317958a0414e7650

Observation 3584eb84-14fb-4507-b12c-f1b1d92f5acb · inbound

Implicit Safety Alignment from Crowd Preferences cites this paper.

Implicit Safety Alignment from Crowd Preferences Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:31:16.861243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-22T08:28:41.652865Z digest=sha256:4ebb3b49f5e85e02940972c883a70fed49a7dcc0c66b687f103d74da262281cb

Observation 6999f00b-6592-46a9-b71b-41aec32e366f · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?

Reference 238

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.103294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.103294Z digest=sha256:1c34d81b60f3b3fbf07f283a2c632748ca52e1869d61ceb8f2acca568ae13180