Pith. sign in

Paper Citation Record · LEDGER

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 5 inbound Pith citation observations for arXiv:2505.17250.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17250 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:53:55.003012Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:53:46.148845Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T01:29:56.674717Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfee40c1-21b0-481e-a4b0-4bba34439fe1 · outbound

This paper cites an unresolved cited work.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:57.694772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.060057Z digest=sha256:a1d402a45979b8130bd9c9dca59c45183fcf333843e07a25d06d7ae04b6f979e

Observation 01ab6025-3b14-4b08-b130-730fa9bcf27c · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.592950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.592950Z digest=sha256:ea3cf62b4cbdd8a3f6e4e28c30b1abb51bb5e15bcea0145a50aa96cf6fe1aaf3

Observation 252a0634-5e80-4cfa-a6cd-97bf4f22afc5 · outbound

This paper cites Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.689958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.689958Z digest=sha256:1f8c19a337a228e38892e0f391f240c47b3e1ab76f5bb47e38c79fcae697523e

Observation 8c3d54fa-02f2-4d5e-bf42-34806e629466 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.857132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.857132Z digest=sha256:f2fc247767816e8a37dac395e0464633fa78715f382590f7ce2aacab666c855d

Observation b5a47764-a058-4ec4-8052-a3e12af3bfa9 · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:57.458987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.205339Z digest=sha256:991abcaf86b3321c600c9a94191c60c0266ace4dbe2b4a8e59f504d087fbbd72

Observation c5ef0484-65a6-41c6-8e8d-02944d49bae0 · outbound

This paper cites an unresolved cited work.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:57.282534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.369927Z digest=sha256:327d3cb13681617615e40eb919c35abb36c35e822c6d5edd2567d3989d8f47f6

Observation 20bdaeef-e216-4000-b71c-7f0103bf6dd1 · outbound

This paper cites Check \( G > B \): 26 > 9, which is true.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Check \( G > B \): 26 > 9, which is true

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:57.187488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.425961Z digest=sha256:fbdd9075083f583fc435a400dfb1f3643800268c98f2e983927048279727d834

Observation a9186dec-a6d0-4fbb-91ca-e14c666e256b · outbound

This paper cites an unresolved cited work.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:56.941051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.507187Z digest=sha256:a84ea573fe592375f139073f27ff68519c137f59b24ee693cf69aca01440204b

Observation 1ce1d444-ea9c-4a2e-9192-a7ed8fb43740 · outbound

This paper cites [...] The number of boys at the meeting is \boxed{9}.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models [...] The number of boys at the meeting is \boxed{9}

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:56.624923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.593658Z digest=sha256:bd95a3267582308d201dbf1640a20f7272c0128c7569d95ea31ff81f02008953

Observation f354b64d-04c4-465c-b955-9926bc508fe6 · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:56.445256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.659295Z digest=sha256:e8fedc45c5c1b98429798b8f79c343f9b3e51a5a1a428ec054d2e00b18f6a2a0

Observation 379006c1-2b78-4445-b93d-1e8292bb6424 · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:56.194770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.698751Z digest=sha256:a387bdf9641558be839483ca1e9a97f03abfdccd0022c809995437ccd6ebc9a6

Observation ff73a4f6-48e2-46ce-8dae-8f96ea294006 · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:56.028054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.772420Z digest=sha256:153fb7b4a560d86f3bb6cca0558c61e8cbc865c1fb0212d73e16dd60532ee52e

Observation 6470cf79-7523-4904-9497-390c3011a55c · outbound

This paper cites Thus, the result of the computation is \(\boxed{\dfrac{5}{9}}\).

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Thus, the result of the computation is \(\boxed{\dfrac{5}{9}}\)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:55.833060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.844826Z digest=sha256:ccfefb5733afc49647951ceb5ca7c7ffbcbe8499d30282b42044707993127bd7

Observation e8d56cd0-ca21-46c6-b46e-4dcaebbab35b · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:55.678609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:54.926865Z digest=sha256:2f0450ad2d91898b70420f2d0d77050d4340d3548b490aec0a001a678f7ef82b

Observation bb41f141-21cd-45da-a624-b1436980f04b · outbound

This paper cites The full reasoning traces are available at https://github.com/RazvanDu/ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models The full reasoning traces are available at https://github.com/RazvanDu/ConciseRL

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:55.546967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:55.003012Z digest=sha256:8f5cdce8fed9389812f745f0bc4a89dbd7aa4ed325b113ad25007e2fb50010b9

Observation 877a46cf-69d8-4232-87e8-fdd4b282a5b4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.407543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.407543Z digest=sha256:d6d8ab7d29689d16768c2a16561746d17349d7042c225d18a4cf6b2c4088f63b

Observation 3f8b6ab6-9178-46d5-b918-eb26621a3c06 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.957988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.957988Z digest=sha256:7f59b67c4e7f6857f8a2db03a5808058ae9407e7550657e1a3d8f905e9d62343

Observation 0e811332-6d86-4d06-b2aa-270eec8efef0 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.745836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.745836Z digest=sha256:3c65bb5ef77a1eb0def45eb11139e6dbc0f940e30e08521946fee30d0d4d24f1

Observation 33e62ba4-65ba-4b43-8254-ea5e0011726d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.478299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.478299Z digest=sha256:f4a18f62f469d5e0ef79dc60237975fac09ca4807d2c01b3774988b99fb8656f

Pith citing papers

Observation b77ec013-62b2-4281-94a6-722c927958dd · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:56.678764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:c548d0b19d101bd1aca52c4a9a51f0809372584e94234702b206aad00a96b083

Observation 47f83da4-5200-47fe-8120-9b2d8a75cc12 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:53:46.148845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:53:46.148845Z digest=sha256:591556b47418387b83888a93a2de7204028a59443fa250611b8443dddb8f11b9

Observation 9b43aca3-afee-4244-acb4-276249572b40 · inbound

Learning to Reason Efficiently with Discounted Reinforcement Learning cites this paper.

Learning to Reason Efficiently with Discounted Reinforcement Learning ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T07:59:52.455824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:59:52.455824Z digest=sha256:51d8885101618d4450b8592a69b52b4d2e21e9e63baeacb7fd142db7ad0a58c0

Observation d204015b-7dc1-438e-b5ab-c0d07dbc4a03 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 217

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.173076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:40f64c39f23a57d73c23f0bf8f74523d6415ad6f1f10a27987ad9af72e654ed2

Observation 8c70757a-9ef0-403b-b493-63747d8b6beb · inbound

Contrastive On-Policy Distillation cites this paper.

Contrastive On-Policy Distillation ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:38.524310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:38.524310Z digest=sha256:61db0d401a8cb5e18ab275a54d1fe3ea9d41ce97520ba80f994245ea5c1b47d4