An RL agent can earn high reward while its representation of a hidden DFA's state stays at chance; a white-box hidden-DFA instrument measures this decoupling, and permutation/group structure flags such perception gaps at 0.86 precision on 153 fresh automata.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary
An RL agent can earn high reward while its representation of a hidden DFA's state stays at chance; a white-box hidden-DFA instrument measures this decoupling, and permutation/group structure flags such perception gaps at 0.86 precision on 153 fresh automata.