Pith. sign in

REVIEW 1 cited by

Reinforcement Learning with Depreciating Assets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.14176 v1 pith:J6H3572S submitted 2023-02-27 cs.AI cs.CEmath.OC

classification cs.AIcs.CEmath.OC
keywords agentassetlearningreinforcementrewardtimevalueassets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A basic assumption of traditional reinforcement learning is that the value of a reward does not change once it is received by an agent. The present work forgoes this assumption and considers the situation where the value of a reward decays proportionally to the time elapsed since it was obtained. Emphasizing the inflection point occurring at the time of payment, we use the term asset to refer to a reward that is currently in the possession of an agent. Adopting this language, we initiate the study of depreciating assets within the framework of infinite-horizon quantitative optimization. In particular, we propose a notion of asset depreciation, inspired by classical exponential discounting, where the value of an asset is scaled by a fixed discount factor at each time step after it is obtained by the agent. We formulate a Bellman-style equational characterization of optimality in this context and develop a model-free reinforcement learning approach to obtain optimal policies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Level Strategic Classification: Incentivizing Improvement through Promotion and Relegation Dynamics

    cs.LG 2026-02 conditional novelty 7.0 of 10

    In a multi-level promotion/relegation system, thresholds placed at the leg-up steady state make honest improvement the agent's optimal long-run strategy, enabling arbitrarily high attainment.

Pith tools