{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:X3FYPPVPN7NQANBN3KEA2Y7NNY","short_pith_number":"pith:X3FYPPVP","schema_version":"1.0","canonical_sha256":"becb87beaf6fdb00342dda880d63ed6e3f18be26e1701fd72761023592ef4fa7","source":{"kind":"arxiv","id":"2507.07017","version":1},"attestation_state":"computed","paper":{"title":"First Return, Entropy-Eliciting Explore","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Chenghua Lin, Ge Zhang, Qian Liu, Qingshui Gu, Taoran Liang, Tianshun Xing, Tianyu Zheng, Wenhao Huang, Xingwei Qu, Xin Zhou, Yizhi Li, Zejun Ma, Zhoufutu Wen","submitted_at":"2025-07-09T16:45:48Z","abstract_excerpt":"Reinforcement Learning from Verifiable Rewards (RLVR) improves the reasoning abilities of Large Language Models (LLMs) but it struggles with unstable exploration. We propose FR3E (First Return, Entropy-Eliciting Explore), a structured exploration framework that identifies high-uncertainty decision points in reasoning trajectories and performs targeted rollouts to construct semantically grounded intermediate feedback. Our method provides targeted guidance without relying on dense supervision. Empirical results on mathematical reasoning benchmarks(AIME24) show that FR3E promotes more stable trai"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2507.07017","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-sa/4.0/","primary_cat":"cs.AI","submitted_at":"2025-07-09T16:45:48Z","cross_cats_sorted":[],"title_canon_sha256":"66deef57854505250fdfcfb710680ff60d1ebbbfab7752a6ecb4cbb24ddbfcf3","abstract_canon_sha256":"a9ec6588526ea91b1e1940d9987a3110411e9cda9c0d906c991583a902458671"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:34:27.859756Z","signature_b64":"tLqIigC0bl1WcQWTrmrB3OuEWvQnN076VZXpxg9lzdoMV2f7FkvsytklRXEOHjr+yQSIxQQkQwShYkJcIerBDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"becb87beaf6fdb00342dda880d63ed6e3f18be26e1701fd72761023592ef4fa7","last_reissued_at":"2026-07-05T11:34:27.859280Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:34:27.859280Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"First Return, Entropy-Eliciting Explore","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Chenghua Lin, Ge Zhang, Qian Liu, Qingshui Gu, Taoran Liang, Tianshun Xing, Tianyu Zheng, Wenhao Huang, Xingwei Qu, Xin Zhou, Yizhi Li, Zejun Ma, Zhoufutu Wen","submitted_at":"2025-07-09T16:45:48Z","abstract_excerpt":"Reinforcement Learning from Verifiable Rewards (RLVR) improves the reasoning abilities of Large Language Models (LLMs) but it struggles with unstable exploration. We propose FR3E (First Return, Entropy-Eliciting Explore), a structured exploration framework that identifies high-uncertainty decision points in reasoning trajectories and performs targeted rollouts to construct semantically grounded intermediate feedback. Our method provides targeted guidance without relying on dense supervision. Empirical results on mathematical reasoning benchmarks(AIME24) show that FR3E promotes more stable trai"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.07017","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.07017/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2507.07017","created_at":"2026-07-05T11:34:27.859340+00:00"},{"alias_kind":"arxiv_version","alias_value":"2507.07017v1","created_at":"2026-07-05T11:34:27.859340+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2507.07017","created_at":"2026-07-05T11:34:27.859340+00:00"},{"alias_kind":"pith_short_12","alias_value":"X3FYPPVPN7NQ","created_at":"2026-07-05T11:34:27.859340+00:00"},{"alias_kind":"pith_short_16","alias_value":"X3FYPPVPN7NQANBN","created_at":"2026-07-05T11:34:27.859340+00:00"},{"alias_kind":"pith_short_8","alias_value":"X3FYPPVP","created_at":"2026-07-05T11:34:27.859340+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":12,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.22570","citing_title":"What are Key Factors for Updates in RL for LLM Reasoning?","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2606.21943","citing_title":"Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning","ref_index":276,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18089","citing_title":"From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning","ref_index":123,"is_internal_anchor":false},{"citing_arxiv_id":"2606.20662","citing_title":"Confidence Laundering in Agent Systems: Why Uncertainty Needs a Latent Carrier","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06096","citing_title":"OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation","ref_index":114,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06080","citing_title":"On Advantage Estimates for Max@K Policy Gradients","ref_index":70,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01249","citing_title":"Trust Region On-Policy Distillation","ref_index":132,"is_internal_anchor":false},{"citing_arxiv_id":"2605.16874","citing_title":"Reasoning Can Be Restored by Correcting a Few Decision Tokens","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11505","citing_title":"Selective Off-Policy Reference Tuning with Plan Guidance","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11505","citing_title":"Selective Off-Policy Reference Tuning with Plan Guidance","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08666","citing_title":"The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08817","citing_title":"How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors","ref_index":52,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/X3FYPPVPN7NQANBN3KEA2Y7NNY","json":"https://pith.science/pith/X3FYPPVPN7NQANBN3KEA2Y7NNY.json","graph_json":"https://pith.science/api/pith-number/X3FYPPVPN7NQANBN3KEA2Y7NNY/graph.json","events_json":"https://pith.science/api/pith-number/X3FYPPVPN7NQANBN3KEA2Y7NNY/events.json","paper":"https://pith.science/paper/X3FYPPVP"},"agent_actions":{"view_html":"https://pith.science/pith/X3FYPPVPN7NQANBN3KEA2Y7NNY","download_json":"https://pith.science/pith/X3FYPPVPN7NQANBN3KEA2Y7NNY.json","view_paper":"https://pith.science/paper/X3FYPPVP","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2507.07017&json=true","fetch_graph":"https://pith.science/api/pith-number/X3FYPPVPN7NQANBN3KEA2Y7NNY/graph.json","fetch_events":"https://pith.science/api/pith-number/X3FYPPVPN7NQANBN3KEA2Y7NNY/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/X3FYPPVPN7NQANBN3KEA2Y7NNY/action/timestamp_anchor","attest_storage":"https://pith.science/pith/X3FYPPVPN7NQANBN3KEA2Y7NNY/action/storage_attestation","attest_author":"https://pith.science/pith/X3FYPPVPN7NQANBN3KEA2Y7NNY/action/author_attestation","sign_citation":"https://pith.science/pith/X3FYPPVPN7NQANBN3KEA2Y7NNY/action/citation_signature","submit_replication":"https://pith.science/pith/X3FYPPVPN7NQANBN3KEA2Y7NNY/action/replication_record"}},"created_at":"2026-07-05T11:34:27.859340+00:00","updated_at":"2026-07-05T11:34:27.859340+00:00"}