{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:D32H3A2ZMCTNCHTPR3EIGJPCQN","short_pith_number":"pith:D32H3A2Z","schema_version":"1.0","canonical_sha256":"1ef47d835960a6d11e6f8ec88325e2836b29889c6cba6503095704b43aa8f63a","source":{"kind":"arxiv","id":"2303.15810","version":1},"attestation_state":"computed","paper":{"title":"Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Haoran Xu, Jianxiong Li, Li Jiang, Victor Wai Kin Chan, Xianyuan Zhan, Zhaoran Wang, Zhuoran Yang","submitted_at":"2023-03-28T08:30:01Z","abstract_excerpt":"Most offline reinforcement learning (RL) methods suffer from the trade-off between improving the policy to surpass the behavior policy and constraining the policy to limit the deviation from the behavior policy as computing $Q$-values using out-of-distribution (OOD) actions will suffer from errors due to distributional shift. The recently proposed \\textit{In-sample Learning} paradigm (i.e., IQL), which improves the policy by quantile regression using only data samples, shows great promise because it learns an optimal policy without querying the value function of any unseen actions. However, it"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2303.15810","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.LG","submitted_at":"2023-03-28T08:30:01Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"a962d53ecc4304d6cf484b3a7520e8c5e16ad574fcb21052926bf3f3aa44efe5","abstract_canon_sha256":"5c751ad9c38dccb07d1712931faced392a5ad36e318d17d53a8b081aa75018ff"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T05:55:28.016456Z","signature_b64":"ltccZpQFxRvi5WT1ObWnD/YLUcM5XZXrYPhnf16w1fedk9ea4TizC9gfsvQx5LPzbh6rA/gDkfiAAS2qDFxpAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"1ef47d835960a6d11e6f8ec88325e2836b29889c6cba6503095704b43aa8f63a","last_reissued_at":"2026-07-05T05:55:28.016095Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T05:55:28.016095Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Haoran Xu, Jianxiong Li, Li Jiang, Victor Wai Kin Chan, Xianyuan Zhan, Zhaoran Wang, Zhuoran Yang","submitted_at":"2023-03-28T08:30:01Z","abstract_excerpt":"Most offline reinforcement learning (RL) methods suffer from the trade-off between improving the policy to surpass the behavior policy and constraining the policy to limit the deviation from the behavior policy as computing $Q$-values using out-of-distribution (OOD) actions will suffer from errors due to distributional shift. The recently proposed \\textit{In-sample Learning} paradigm (i.e., IQL), which improves the policy by quantile regression using only data samples, shows great promise because it learns an optimal policy without querying the value function of any unseen actions. However, it"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2303.15810","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2303.15810/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2303.15810","created_at":"2026-07-05T05:55:28.016154+00:00"},{"alias_kind":"arxiv_version","alias_value":"2303.15810v1","created_at":"2026-07-05T05:55:28.016154+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2303.15810","created_at":"2026-07-05T05:55:28.016154+00:00"},{"alias_kind":"pith_short_12","alias_value":"D32H3A2ZMCTN","created_at":"2026-07-05T05:55:28.016154+00:00"},{"alias_kind":"pith_short_16","alias_value":"D32H3A2ZMCTNCHTP","created_at":"2026-07-05T05:55:28.016154+00:00"},{"alias_kind":"pith_short_8","alias_value":"D32H3A2Z","created_at":"2026-07-05T05:55:28.016154+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.11087","citing_title":"Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning","ref_index":66,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14779","citing_title":"Peng's Q($\\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.27877","citing_title":"SPAR: Support-Preserving Action Rectification","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2506.05762","citing_title":"BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2304.10573","citing_title":"IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01663","citing_title":"Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2604.14265","citing_title":"Reinforcement Learning via Value Gradient Flow","ref_index":72,"is_internal_anchor":false},{"citing_arxiv_id":"2604.17919","citing_title":"Fisher Decorator: Refining Flow Policy via a Local Transport Map","ref_index":48,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/D32H3A2ZMCTNCHTPR3EIGJPCQN","json":"https://pith.science/pith/D32H3A2ZMCTNCHTPR3EIGJPCQN.json","graph_json":"https://pith.science/api/pith-number/D32H3A2ZMCTNCHTPR3EIGJPCQN/graph.json","events_json":"https://pith.science/api/pith-number/D32H3A2ZMCTNCHTPR3EIGJPCQN/events.json","paper":"https://pith.science/paper/D32H3A2Z"},"agent_actions":{"view_html":"https://pith.science/pith/D32H3A2ZMCTNCHTPR3EIGJPCQN","download_json":"https://pith.science/pith/D32H3A2ZMCTNCHTPR3EIGJPCQN.json","view_paper":"https://pith.science/paper/D32H3A2Z","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2303.15810&json=true","fetch_graph":"https://pith.science/api/pith-number/D32H3A2ZMCTNCHTPR3EIGJPCQN/graph.json","fetch_events":"https://pith.science/api/pith-number/D32H3A2ZMCTNCHTPR3EIGJPCQN/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/D32H3A2ZMCTNCHTPR3EIGJPCQN/action/timestamp_anchor","attest_storage":"https://pith.science/pith/D32H3A2ZMCTNCHTPR3EIGJPCQN/action/storage_attestation","attest_author":"https://pith.science/pith/D32H3A2ZMCTNCHTPR3EIGJPCQN/action/author_attestation","sign_citation":"https://pith.science/pith/D32H3A2ZMCTNCHTPR3EIGJPCQN/action/citation_signature","submit_replication":"https://pith.science/pith/D32H3A2ZMCTNCHTPR3EIGJPCQN/action/replication_record"}},"created_at":"2026-07-05T05:55:28.016154+00:00","updated_at":"2026-07-05T05:55:28.016154+00:00"}