{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:LZD3HADYG4W7ELFLDWKGDUY4RK","short_pith_number":"pith:LZD3HADY","schema_version":"1.0","canonical_sha256":"5e47b38078372df22cab1d9461d31c8aa74b3d372ab8ea383c82a7d3084069d8","source":{"kind":"arxiv","id":"2508.16949","version":7},"attestation_state":"computed","paper":{"title":"Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Hengtong Lu, Jiale Zhao, Jianwei Lv, Jingwen Yang, Kongcheng Zhang, Mingli Song, Shunyu Liu, Sunzhu Li, Tongya Zheng, Wei Chen, Wenkai Fang, Yang Zhou, Yan Xie, Yihe Zhou","submitted_at":"2025-08-23T08:47:31Z","abstract_excerpt":"Recent advances in Large Language Models (LLMs) have underscored the potential of Reinforcement Learning (RL) to facilitate the emergence of reasoning capabilities. Despite the encouraging results, a fundamental dilemma persists as RL improvement relies on learning from high-quality samples, yet the exploration for such samples remains bounded by the inherent limitations of LLMs. This, in effect, creates an undesirable cycle in which what cannot be explored cannot be learned. In this work, we propose Rubric-Scaffolded Reinforcement Learning (RuscaRL), a novel instructional scaffolding framewor"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2508.16949","kind":"arxiv","version":7},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2025-08-23T08:47:31Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"98cf71f50390447705e89eb4bf5c28dcfa1c9406fc85b2313fe1a3f7634f3964","abstract_canon_sha256":"61e33567b7f2b81637571ec782f35e503117d514c783dc978a5775c9edda24a3"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-08-04T01:51:47.725226Z","signature_b64":"fL5TudENkGdjMLzZeTM47vBtwkyRnPyc4sAmm6+QFpmejou77YdUecFBDSyd88T86cy0QQpR04djtiLcDKnGBA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"5e47b38078372df22cab1d9461d31c8aa74b3d372ab8ea383c82a7d3084069d8","last_reissued_at":"2026-08-04T01:51:47.723447Z","signature_status":"signed_v1","first_computed_at":"2026-08-04T01:51:47.723447Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Hengtong Lu, Jiale Zhao, Jianwei Lv, Jingwen Yang, Kongcheng Zhang, Mingli Song, Shunyu Liu, Sunzhu Li, Tongya Zheng, Wei Chen, Wenkai Fang, Yang Zhou, Yan Xie, Yihe Zhou","submitted_at":"2025-08-23T08:47:31Z","abstract_excerpt":"Recent advances in Large Language Models (LLMs) have underscored the potential of Reinforcement Learning (RL) to facilitate the emergence of reasoning capabilities. Despite the encouraging results, a fundamental dilemma persists as RL improvement relies on learning from high-quality samples, yet the exploration for such samples remains bounded by the inherent limitations of LLMs. This, in effect, creates an undesirable cycle in which what cannot be explored cannot be learned. In this work, we propose Rubric-Scaffolded Reinforcement Learning (RuscaRL), a novel instructional scaffolding framewor"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2508.16949","kind":"arxiv","version":7},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2508.16949/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2508.16949","created_at":"2026-08-04T01:51:47.724796+00:00"},{"alias_kind":"arxiv_version","alias_value":"2508.16949v7","created_at":"2026-08-04T01:51:47.724796+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2508.16949","created_at":"2026-08-04T01:51:47.724796+00:00"},{"alias_kind":"pith_short_12","alias_value":"LZD3HADYG4W7","created_at":"2026-08-04T01:51:47.724796+00:00"},{"alias_kind":"pith_short_16","alias_value":"LZD3HADYG4W7ELFL","created_at":"2026-08-04T01:51:47.724796+00:00"},{"alias_kind":"pith_short_8","alias_value":"LZD3HADY","created_at":"2026-08-04T01:51:47.724796+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":18,"internal_anchor_count":18,"sample":[{"citing_arxiv_id":"2606.25407","citing_title":"Teach-to-Reason: Competition-Guided Reasoning with a Self-Improving Teacher","ref_index":36,"is_internal_anchor":true},{"citing_arxiv_id":"2606.23038","citing_title":"EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03968","citing_title":"QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable Rewards","ref_index":28,"is_internal_anchor":true},{"citing_arxiv_id":"2606.01249","citing_title":"Trust Region On-Policy Distillation","ref_index":182,"is_internal_anchor":true},{"citing_arxiv_id":"2605.23454","citing_title":"ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2605.29156","citing_title":"RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains","ref_index":33,"is_internal_anchor":true},{"citing_arxiv_id":"2605.30244","citing_title":"Reinforcement Learning with Robust Rubric Rewards","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2606.01091","citing_title":"Deep Research as Rubric for Reinforcement Learning","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2605.23454","citing_title":"ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2603.15646","citing_title":"Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy","ref_index":22,"is_internal_anchor":true},{"citing_arxiv_id":"2604.02795","citing_title":"Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks","ref_index":35,"is_internal_anchor":true},{"citing_arxiv_id":"2604.03472","citing_title":"Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution","ref_index":40,"is_internal_anchor":true},{"citing_arxiv_id":"2605.12474","citing_title":"Reward Hacking in Rubric-Based Reinforcement Learning","ref_index":37,"is_internal_anchor":true},{"citing_arxiv_id":"2605.08472","citing_title":"Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models","ref_index":65,"is_internal_anchor":true},{"citing_arxiv_id":"2604.20051","citing_title":"Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text","ref_index":49,"is_internal_anchor":true},{"citing_arxiv_id":"2604.13029","citing_title":"Visual Preference Optimization with Rubric Rewards","ref_index":39,"is_internal_anchor":true},{"citing_arxiv_id":"2605.08061","citing_title":"Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"2604.07941","citing_title":"Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning","ref_index":92,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/LZD3HADYG4W7ELFLDWKGDUY4RK","json":"https://pith.science/pith/LZD3HADYG4W7ELFLDWKGDUY4RK.json","graph_json":"https://pith.science/api/pith-number/LZD3HADYG4W7ELFLDWKGDUY4RK/graph.json","events_json":"https://pith.science/api/pith-number/LZD3HADYG4W7ELFLDWKGDUY4RK/events.json","paper":"https://pith.science/paper/LZD3HADY"},"agent_actions":{"view_html":"https://pith.science/pith/LZD3HADYG4W7ELFLDWKGDUY4RK","download_json":"https://pith.science/pith/LZD3HADYG4W7ELFLDWKGDUY4RK.json","view_paper":"https://pith.science/paper/LZD3HADY","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2508.16949&json=true","fetch_graph":"https://pith.science/api/pith-number/LZD3HADYG4W7ELFLDWKGDUY4RK/graph.json","fetch_events":"https://pith.science/api/pith-number/LZD3HADYG4W7ELFLDWKGDUY4RK/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/LZD3HADYG4W7ELFLDWKGDUY4RK/action/timestamp_anchor","attest_storage":"https://pith.science/pith/LZD3HADYG4W7ELFLDWKGDUY4RK/action/storage_attestation","attest_author":"https://pith.science/pith/LZD3HADYG4W7ELFLDWKGDUY4RK/action/author_attestation","sign_citation":"https://pith.science/pith/LZD3HADYG4W7ELFLDWKGDUY4RK/action/citation_signature","submit_replication":"https://pith.science/pith/LZD3HADYG4W7ELFLDWKGDUY4RK/action/replication_record"}},"created_at":"2026-08-04T01:51:47.724796+00:00","updated_at":"2026-08-04T01:51:47.724796+00:00"}