{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:EVXAGD2AEGI3SHJVI2ANLCBOMY","short_pith_number":"pith:EVXAGD2A","schema_version":"1.0","canonical_sha256":"256e030f402191b91d354680d5882e663a4d68bb8fa368efa0820c87fb599512","source":{"kind":"arxiv","id":"2303.08059","version":2},"attestation_state":"computed","paper":{"title":"Fast Rates for Maximum Entropy Exploration","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"stat.ML","authors_text":"Alexey Naumov, Daniele Calandriello, Daniil Tiapkin, Denis Belomestny, Eric Moulines, Michal Valko, Pierre Menard, Pierre Perrault, Remi Munos, Yunhao Tang","submitted_at":"2023-03-14T16:51:14Z","abstract_excerpt":"We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maximum entropy exploration problem of two different types. The first type is visitation entropy maximization previously considered by Hazan et al.(2019) in the discounted setting. For this type of exploration, we propose a game-theoretic algorithm that has $\\widetilde{\\mathcal{O}}(H^3S^2A/\\varepsilon^2)$ sample complexity thus improving the $\\varepsilon$-dependence upon existing results, where $S$ is a number of states, "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2303.08059","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"stat.ML","submitted_at":"2023-03-14T16:51:14Z","cross_cats_sorted":["cs.LG"],"title_canon_sha256":"f322b01768f6e8f0a3910da588763b91d1e3a8f1c16804b722c312a5c048eebb","abstract_canon_sha256":"d60469614692ec3aac5d28bf7b5f2ec99a24e07656e7c7ded6d7b1b408d437e1"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T06:17:48.394701Z","signature_b64":"yqslSSkXKrDjLtLHNBcGJbB++lkjlBjzaHsHXWYswEFAQW6B9MJqM8fIlky2OnvA7tSJdxIzq/bFj5JiOAMjBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"256e030f402191b91d354680d5882e663a4d68bb8fa368efa0820c87fb599512","last_reissued_at":"2026-07-05T06:17:48.394093Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T06:17:48.394093Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Fast Rates for Maximum Entropy Exploration","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"stat.ML","authors_text":"Alexey Naumov, Daniele Calandriello, Daniil Tiapkin, Denis Belomestny, Eric Moulines, Michal Valko, Pierre Menard, Pierre Perrault, Remi Munos, Yunhao Tang","submitted_at":"2023-03-14T16:51:14Z","abstract_excerpt":"We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maximum entropy exploration problem of two different types. The first type is visitation entropy maximization previously considered by Hazan et al.(2019) in the discounted setting. For this type of exploration, we propose a game-theoretic algorithm that has $\\widetilde{\\mathcal{O}}(H^3S^2A/\\varepsilon^2)$ sample complexity thus improving the $\\varepsilon$-dependence upon existing results, where $S$ is a number of states, "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2303.08059","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2303.08059/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2303.08059","created_at":"2026-07-05T06:17:48.394166+00:00"},{"alias_kind":"arxiv_version","alias_value":"2303.08059v2","created_at":"2026-07-05T06:17:48.394166+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2303.08059","created_at":"2026-07-05T06:17:48.394166+00:00"},{"alias_kind":"pith_short_12","alias_value":"EVXAGD2AEGI3","created_at":"2026-07-05T06:17:48.394166+00:00"},{"alias_kind":"pith_short_16","alias_value":"EVXAGD2AEGI3SHJV","created_at":"2026-07-05T06:17:48.394166+00:00"},{"alias_kind":"pith_short_8","alias_value":"EVXAGD2A","created_at":"2026-07-05T06:17:48.394166+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":1,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2412.03800","citing_title":"ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy","ref_index":12,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/EVXAGD2AEGI3SHJVI2ANLCBOMY","json":"https://pith.science/pith/EVXAGD2AEGI3SHJVI2ANLCBOMY.json","graph_json":"https://pith.science/api/pith-number/EVXAGD2AEGI3SHJVI2ANLCBOMY/graph.json","events_json":"https://pith.science/api/pith-number/EVXAGD2AEGI3SHJVI2ANLCBOMY/events.json","paper":"https://pith.science/paper/EVXAGD2A"},"agent_actions":{"view_html":"https://pith.science/pith/EVXAGD2AEGI3SHJVI2ANLCBOMY","download_json":"https://pith.science/pith/EVXAGD2AEGI3SHJVI2ANLCBOMY.json","view_paper":"https://pith.science/paper/EVXAGD2A","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2303.08059&json=true","fetch_graph":"https://pith.science/api/pith-number/EVXAGD2AEGI3SHJVI2ANLCBOMY/graph.json","fetch_events":"https://pith.science/api/pith-number/EVXAGD2AEGI3SHJVI2ANLCBOMY/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/EVXAGD2AEGI3SHJVI2ANLCBOMY/action/timestamp_anchor","attest_storage":"https://pith.science/pith/EVXAGD2AEGI3SHJVI2ANLCBOMY/action/storage_attestation","attest_author":"https://pith.science/pith/EVXAGD2AEGI3SHJVI2ANLCBOMY/action/author_attestation","sign_citation":"https://pith.science/pith/EVXAGD2AEGI3SHJVI2ANLCBOMY/action/citation_signature","submit_replication":"https://pith.science/pith/EVXAGD2AEGI3SHJVI2ANLCBOMY/action/replication_record"}},"created_at":"2026-07-05T06:17:48.394166+00:00","updated_at":"2026-07-05T06:17:48.394166+00:00"}