{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2021:RJSCTGXDMPYLVLWRJRFUZOXPHN","short_pith_number":"pith:RJSCTGXD","schema_version":"1.0","canonical_sha256":"8a64299ae363f0baaed14c4b4cbaef3b54bc2db98e90a8870483ed85942742a2","source":{"kind":"arxiv","id":"2105.11066","version":4},"attestation_state":"computed","paper":{"title":"Policy Mirror Descent for Regularized Reinforcement Learning: A Generalized Framework with Linear Convergence","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.IT","math.IT","math.OC","stat.ML"],"primary_cat":"cs.LG","authors_text":"Baihe Huang, Jason D. Lee, Shicong Cen, Wenhao Zhan, Yuejie Chi, Yuxin Chen","submitted_at":"2021-05-24T02:21:34Z","abstract_excerpt":"Policy optimization, which finds the desired policy by maximizing value functions via optimization techniques, lies at the heart of reinforcement learning (RL). In addition to value maximization, other practical considerations arise as well, including the need of encouraging exploration, and that of ensuring certain structural properties of the learned policy due to safety, resource and operational constraints. These can often be accounted for via regularized RL, which augments the target value function with a structure-promoting regularizer.\n  Focusing on discounted infinite-horizon Markov de"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2105.11066","kind":"arxiv","version":4},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2021-05-24T02:21:34Z","cross_cats_sorted":["cs.IT","math.IT","math.OC","stat.ML"],"title_canon_sha256":"71e9bcebb098ae11002d86c23ff4eb75275adeb03f58c5cbdbab8e16ab7aca16","abstract_canon_sha256":"8e3ed1181fc73afde542f80c8842d32c7e2ee9903952e3cc9726bd7c47366d0c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T05:32:10.983945Z","signature_b64":"9Hw/9oicuAQ3ZiljLIYBHLKM69PJmppgPBR1Uk61VPKjvpTdIUg3YfEBcX1f4QOjp41oEHFLGqenjpfctaPuCg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"8a64299ae363f0baaed14c4b4cbaef3b54bc2db98e90a8870483ed85942742a2","last_reissued_at":"2026-07-05T05:32:10.983460Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T05:32:10.983460Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Policy Mirror Descent for Regularized Reinforcement Learning: A Generalized Framework with Linear Convergence","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.IT","math.IT","math.OC","stat.ML"],"primary_cat":"cs.LG","authors_text":"Baihe Huang, Jason D. Lee, Shicong Cen, Wenhao Zhan, Yuejie Chi, Yuxin Chen","submitted_at":"2021-05-24T02:21:34Z","abstract_excerpt":"Policy optimization, which finds the desired policy by maximizing value functions via optimization techniques, lies at the heart of reinforcement learning (RL). In addition to value maximization, other practical considerations arise as well, including the need of encouraging exploration, and that of ensuring certain structural properties of the learned policy due to safety, resource and operational constraints. These can often be accounted for via regularized RL, which augments the target value function with a structure-promoting regularizer.\n  Focusing on discounted infinite-horizon Markov de"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2105.11066","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2105.11066/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2105.11066","created_at":"2026-07-05T05:32:10.983510+00:00"},{"alias_kind":"arxiv_version","alias_value":"2105.11066v4","created_at":"2026-07-05T05:32:10.983510+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2105.11066","created_at":"2026-07-05T05:32:10.983510+00:00"},{"alias_kind":"pith_short_12","alias_value":"RJSCTGXDMPYL","created_at":"2026-07-05T05:32:10.983510+00:00"},{"alias_kind":"pith_short_16","alias_value":"RJSCTGXDMPYLVLWR","created_at":"2026-07-05T05:32:10.983510+00:00"},{"alias_kind":"pith_short_8","alias_value":"RJSCTGXD","created_at":"2026-07-05T05:32:10.983510+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":2,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.29092","citing_title":"Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02469","citing_title":"Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent","ref_index":55,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/RJSCTGXDMPYLVLWRJRFUZOXPHN","json":"https://pith.science/pith/RJSCTGXDMPYLVLWRJRFUZOXPHN.json","graph_json":"https://pith.science/api/pith-number/RJSCTGXDMPYLVLWRJRFUZOXPHN/graph.json","events_json":"https://pith.science/api/pith-number/RJSCTGXDMPYLVLWRJRFUZOXPHN/events.json","paper":"https://pith.science/paper/RJSCTGXD"},"agent_actions":{"view_html":"https://pith.science/pith/RJSCTGXDMPYLVLWRJRFUZOXPHN","download_json":"https://pith.science/pith/RJSCTGXDMPYLVLWRJRFUZOXPHN.json","view_paper":"https://pith.science/paper/RJSCTGXD","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2105.11066&json=true","fetch_graph":"https://pith.science/api/pith-number/RJSCTGXDMPYLVLWRJRFUZOXPHN/graph.json","fetch_events":"https://pith.science/api/pith-number/RJSCTGXDMPYLVLWRJRFUZOXPHN/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/RJSCTGXDMPYLVLWRJRFUZOXPHN/action/timestamp_anchor","attest_storage":"https://pith.science/pith/RJSCTGXDMPYLVLWRJRFUZOXPHN/action/storage_attestation","attest_author":"https://pith.science/pith/RJSCTGXDMPYLVLWRJRFUZOXPHN/action/author_attestation","sign_citation":"https://pith.science/pith/RJSCTGXDMPYLVLWRJRFUZOXPHN/action/citation_signature","submit_replication":"https://pith.science/pith/RJSCTGXDMPYLVLWRJRFUZOXPHN/action/replication_record"}},"created_at":"2026-07-05T05:32:10.983510+00:00","updated_at":"2026-07-05T05:32:10.983510+00:00"}