Pith. sign in

Paper Citation Record · LEDGER

Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

As of 22 August 2026, this Paper Citation Record lists 1 of 1 outbound references and 24 inbound Pith citation observations for arXiv:2508.09726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09726 v1

Coverage vector

measured 1 of 1 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:52:31.630467Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:59:13.695065Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.674459Z

Reference resolution

1 of 1 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef922b12-4e41-4768-bb96-cb30769cd74d · outbound

This paper cites Multimodal Sheaf-based Network for Glioblastoma Molecular Subtype Prediction.

Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning Multimodal Sheaf-based Network for Glioblastoma Molecular Subtype Prediction

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:52:31.742741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T20:52:31.630467Z digest=sha256:59d6f680cd8c3d4f60358003c83bf5a3e1e251dfa10926370bfcbed770d7d4c6

Pith citing papers

Observation 7ca6cd1f-eb8a-45c4-b1fb-2af1848a39f1 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:56.804738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:138bc87c8ec3e6c95fd2ae76534a1149ed8c2fb48dbbd4ddac53f0161121388a

Observation 0900ce65-3114-406d-b631-cf22dd7b6e5f · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.525311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.525311Z digest=sha256:b3348844dfbec0a0672ad2b3bab2d153c7d963f5b5fe5083724084105de39af0

Observation 04876c42-aeff-4d4f-a0a3-29234d105387 · inbound

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning cites this paper.

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:53:34.594743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:53:34.594743Z digest=sha256:45981cff30610bb9e0053ab5bf05af7dbaf00dd06b48723287ec87a3c607631c

Observation db490758-5cb1-4c81-84ac-8cfb3c9f1d3f · inbound

Entropy After </Think> for reasoning model early exiting cites this paper.

Entropy After </Think> for reasoning model early exiting Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.589387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T11:51:58.579048Z digest=sha256:fb279003ba4eea7403468e9182d1579ac09915baad00a409348ef5257a8be354

Observation 9d0c0b69-7c66-4482-bd2d-baf3af814805 · inbound

Learning to Reason Efficiently with Discounted Reinforcement Learning cites this paper.

Learning to Reason Efficiently with Discounted Reinforcement Learning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T07:59:53.568100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:59:53.568100Z digest=sha256:fecc726d2f3bce898132fec9dc33dcf8a0ad41253f304ab8234edafd114c82ad

Observation 639e8c14-3980-4fb4-a234-08fec123b9f1 · inbound

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization cites this paper.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:55.977831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:a9433003479aed07c58972de191c79dc27986daf088f997721dc5ee885b5fcf2

Observation 681976d1-cafc-406a-81b2-c2f48d0a2e84 · inbound

On the Optimal Reasoning Length for RL-Trained Language Models cites this paper.

On the Optimal Reasoning Length for RL-Trained Language Models Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T02:53:07.787439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:53:07.787439Z digest=sha256:b3556f1403de79c495e7c51a042260a2460cf396caf15f76d1aa570d5774eeff

Observation 3b8e6644-4e16-4148-a510-4cf06df27157 · inbound

SaFRO: Satisfaction-Aware Fusion via Dual-Relative Policy Optimization for Short-Video Search cites this paper.

SaFRO: Satisfaction-Aware Fusion via Dual-Relative Policy Optimization for Short-Video Search Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T02:35:00.092811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:35:00.092811Z digest=sha256:d5bb9a688b75ee7f70528102c25e347cbbf5960f199bfdbb8d02ed3a4d10aedf

Observation 6f2b12b8-3864-47d4-a1f7-9d783f69f61f · inbound

Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks cites this paper.

Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:28:14.101220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T20:24:17.788697Z digest=sha256:05e702ce767b14263394f726f0bf37edc48c72302be6d6b61820a11f6af0a9b5

Observation 57db2f98-aae0-45df-b0ba-5df45881d80f · inbound

MEMENTO: Teaching LLMs to Manage Their Own Context cites this paper.

MEMENTO: Teaching LLMs to Manage Their Own Context Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:10:59.536875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T17:16:36.784237Z digest=sha256:ae7daaceb23939609b1edd395d873df0c7f96dbf6c0ff703b3f14108902b8490

Observation 56a82c10-1f5d-4c11-87b8-bd938c5ff4ad · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.138259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:0f62a37c4d87f7d53b52f1f54bc72ab96719e72c741eeb6245545b429541988e

Observation 0601cb36-b6d1-4f2f-b3b8-0ced460d87b3 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 284

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:09.567133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:6252a98a8082dac8b5e896d03aaa871996cefc5ef98df52f9336f943c081617b

Observation e9bbfbd2-b9f0-4854-bd3f-0f8f257169e5 · inbound

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR cites this paper.

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.144628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T12:40:53.063991Z digest=sha256:68aaaed8a85661069d2d90db931b92ecf9e77e5fffb949d98c9f73477c5c176d

Observation da80c282-8fdb-4372-9e38-5bc65c867b1c · inbound

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization cites this paper.

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:26.475965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T00:59:44.364491Z digest=sha256:215aa65f6bf35590868ff2496cddc2aed6c3487f71c8fe12514449c288e8de33

Observation d9bc6d0a-3e5d-4c4c-b93a-0a7c701686d0 · inbound

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models cites this paper.

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:31:16.791011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T02:30:42.407934Z digest=sha256:ddd6da1446c321738de5f0a6dbd60bd28475561ef15e72a7675f01ea91c84c28

Observation a1e89810-fdcf-488f-9b69-72280cbe00b7 · inbound

Knowledge Graph-Driven Expert-Level Reasoning for Neuroscience cites this paper.

Knowledge Graph-Driven Expert-Level Reasoning for Neuroscience Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:44:38.725141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T11:36:03.804989Z digest=sha256:a036be2118bc4da94983d7ce040fa06e668e965a1bf35d8ee9696bb4f9024d81

Observation a4ec028c-7bf4-477f-9a6d-f5b3532c8d46 · inbound

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training cites this paper.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.807528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:5611c303e5abcbd456aaba454de7ba6a581d82c2fb1213d40480c48d5c5b4637

Observation accc5c3e-6d9f-42cf-b194-b36f7fe05c6a · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 252

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.635585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:a860dc0e05929b592db8907bfe1c753e52b90ad71cd85dc6a318978b8b129b96

Observation 459eabe6-40ea-4059-8e66-c0f25668fd16 · inbound

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning cites this paper.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.265290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:ab603fbad577f977e9e76987dd21c692f30a0b1abf427ef0714040c8be3172da

Observation 9ddff3ca-dc50-4f7c-b576-9d6fc362d8ff · inbound

Cross-Epoch Adaptive Rollout Optimization for RL Post-Training cites this paper.

Cross-Epoch Adaptive Rollout Optimization for RL Post-Training Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.809734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T02:32:07.961430Z digest=sha256:0e10829066d79ef9485c6dc38a249774194ca3ea482a564726663992d1aec269

Observation 059fc053-6e17-4231-8f5c-9eb0a3866f2c · inbound

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization cites this paper.

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:37.320254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T14:02:05.833651Z digest=sha256:3b163f3405b2c0146f4538128ba68935be7d5655edca13868ab415ecccc7927e

Observation cd0bdc94-adc8-44e8-99f0-492753b6a445 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.676004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:53274f51faa45ea4f6da48d039dba3c745777799766e1b93303dbb4a731f8f5f

Observation 769bde3a-c21a-4e9f-9df6-fd0c78feeab1 · inbound

Masked Distillation: Internalizing the Chain-of-Thought in Language Models cites this paper.

Masked Distillation: Internalizing the Chain-of-Thought in Language Models Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:24.859537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:24.859537Z digest=sha256:3d9fd1d604f2f8903abea32c816db7f84d89d34e0aee3d49bc75a41d0998e935

Observation 31e47036-5fe7-4a9e-afcb-100e785c55ca · inbound

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning cites this paper.

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:59:13.695065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:59:13.695065Z digest=sha256:02cf77c1ed30dc7ce41f70551abf562d895c0457da79670d1b4c0c176328b9c3