Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:52:31.630467Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 1 of 1 outbound references and 24 inbound Pith citation observations for arXiv:2508.09726.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:52:31.630467Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T04:59:13.695065Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.674459Z
1 of 1 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ef922b12-4e41-4768-bb96-cb30769cd74d · outbound
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning Multimodal Sheaf-based Network for Glioblastoma Molecular Subtype Prediction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7ca6cd1f-eb8a-45c4-b1fb-2af1848a39f1 · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0900ce65-3114-406d-b631-cf22dd7b6e5f · inbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04876c42-aeff-4d4f-a0a3-29234d105387 · inbound
CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db490758-5cb1-4c81-84ac-8cfb3c9f1d3f · inbound
Entropy After </Think> for reasoning model early exiting Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9d0c0b69-7c66-4482-bd2d-baf3af814805 · inbound
Learning to Reason Efficiently with Discounted Reinforcement Learning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 639e8c14-3980-4fb4-a234-08fec123b9f1 · inbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 681976d1-cafc-406a-81b2-c2f48d0a2e84 · inbound
On the Optimal Reasoning Length for RL-Trained Language Models Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b8e6644-4e16-4148-a510-4cf06df27157 · inbound
SaFRO: Satisfaction-Aware Fusion via Dual-Relative Policy Optimization for Short-Video Search Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2b12b8-3864-47d4-a1f7-9d783f69f61f · inbound
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 57db2f98-aae0-45df-b0ba-5df45881d80f · inbound
MEMENTO: Teaching LLMs to Manage Their Own Context Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 56a82c10-1f5d-4c11-87b8-bd938c5ff4ad · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0601cb36-b6d1-4f2f-b3b8-0ced460d87b3 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 284
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e9bbfbd2-b9f0-4854-bd3f-0f8f257169e5 · inbound
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation da80c282-8fdb-4372-9e38-5bc65c867b1c · inbound
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d9bc6d0a-3e5d-4c4c-b93a-0a7c701686d0 · inbound
LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a1e89810-fdcf-488f-9b69-72280cbe00b7 · inbound
Knowledge Graph-Driven Expert-Level Reasoning for Neuroscience Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a4ec028c-7bf4-477f-9a6d-f5b3532c8d46 · inbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation accc5c3e-6d9f-42cf-b194-b36f7fe05c6a · inbound
Trust Region On-Policy Distillation Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 252
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 459eabe6-40ea-4059-8e66-c0f25668fd16 · inbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9ddff3ca-dc50-4f7c-b576-9d6fc362d8ff · inbound
Cross-Epoch Adaptive Rollout Optimization for RL Post-Training Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 059fc053-6e17-4231-8f5c-9eb0a3866f2c · inbound
N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cd0bdc94-adc8-44e8-99f0-492753b6a445 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 177
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 769bde3a-c21a-4e9f-9df6-fd0c78feeab1 · inbound
Masked Distillation: Internalizing the Chain-of-Thought in Language Models Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e47036-5fe7-4a9e-afcb-100e785c55ca · inbound
I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.