Symmetric (even highly) pessimistic value functions can generalize optimally in GTI-ZSPT offline RL while mildly pessimistic asymmetric ones cannot; DAC-Output during extraction is the strongest DA method tested.
S4RL : Surprisingly Simple Self - Supervision for Offline Reinforcement Learning in Robotics
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generalization in offline RL: The structure is more important than the amount of pessimism
Symmetric (even highly) pessimistic value functions can generalize optimally in GTI-ZSPT offline RL while mildly pessimistic asymmetric ones cannot; DAC-Output during extraction is the strongest DA method tested.