Pith. sign in

REVIEW 4 cited by

Guaranteed Discovery of Control-Endogenous Latent States with Multi-Step Inverse Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.08229 v2 pith:STCIVIIT submitted 2022-07-17 cs.LG cs.ROstat.ML

classification cs.LGcs.ROstat.ML
keywords agentinformationcontrol-endogenouslatentstatediscoverymodelworld
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In many sequential decision-making tasks, the agent is not able to model the full complexity of the world, which consists of multitudes of relevant and irrelevant information. For example, a person walking along a city street who tries to model all aspects of the world would quickly be overwhelmed by a multitude of shops, cars, and people moving in and out of view, each following their own complex and inscrutable dynamics. Is it possible to turn the agent's firehose of sensory information into a minimal latent state that is both necessary and sufficient for an agent to successfully act in the world? We formulate this question concretely, and propose the Agent Control-Endogenous State Discovery algorithm (AC-State), which has theoretical guarantees and is practically demonstrated to discover the minimal control-endogenous latent state which contains all of the information necessary for controlling the agent, while fully discarding all irrelevant information. This algorithm consists of a multi-step inverse model (predicting actions from distant observations) with an information bottleneck. AC-State enables localization, exploration, and navigation without reward or demonstrations. We demonstrate the discovery of the control-endogenous latent state in three domains: localizing a robot arm with distractions (e.g., changing lighting conditions and background), exploring a maze alongside other agents, and navigating in the Matterport house simulator.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

    cs.LG 2026-07 conditional novelty 7.0 of 10

    In holonomy-cover POMDPs, the observation plus its stable quotient class is the minimal exact finite Markov state, recoverable under diagnostics and usable by standard RL.

  2. TaskSense: Focusing on What Matters in World Models

    cs.AI 2026-08 conditional novelty 6.0 of 10

    TaskSense filters visual observations with latent-conditioned stochastic spatial attention before encoding, improving world-model control under visual distractions relative to DreamerV3.

  3. How Well Do Latent World Models Understand Partially Observable Safety Constraints?

    cs.RO 2025-10 conditional novelty 6.0 of 10

    RGB-only latent safety filters fail to prevent overheating because safety-relevant temperature is unobservable in the image; a mutual-information metric predicts this failure, and multimodal supervision during world-m...

  4. Latent Action Learning Requires Supervision in the Presence of Distractors

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Latent action models need at least a small amount of action supervision to learn useful actions when observations contain distractors, as shown on the Distracting Control Suite.

Pith tools