REVIEW 1 cited by
Boosting Soft Q-Learning by Bounding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
An agent's ability to leverage past experience is critical for efficiently solving new tasks. Prior work has focused on using value function estimates to obtain zero-shot approximations for solutions to a new task. In soft Q-learning, we show how any value function estimate can also be used to derive double-sided bounds on the optimal value function. The derived bounds lead to new approaches for boosting training performance which we validate experimentally. Notably, we find that the proposed framework suggests an alternative method for updating the Q-function, leading to boosted performance.
Forward citations
Cited by 1 Pith paper
-
Causality-informed Anomaly Detection in Partially Observable Sensor Networks: Moving beyond Correlations
A deep Q-network that mixes causal statistics and a causality-weighted entropy term is proposed for placing sensors in partially observable anomaly detection.
Discussion (0). Sign in to comment.