Pith. sign in

Paper Citation Record · LEDGER

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training

As of 5 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2605.07316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07316 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:12:20.864362Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:39:40.459071Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact28
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 91e96891-59ce-46fb-bd14-4f1f141d98c6 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Chain-of-thought prompting elicits reasoning in large language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.528208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:4e55b78fe4c2046f8e1ddf50b81c2cc879932a6ea994b1ce34bf1538ef62b173

Observation 7c324769-3218-465d-b791-1f58c8aaf5d0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:59.752966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:d4f4424c7462cf3adcf1e41ee2de4b80abce2fc2e2b015013b4427a4cc55f76c

Observation dea649c4-6c98-4252-a624-c6b667204ecc · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:59.869378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:8640ab42314a9cae3e96f862c1f29cb7995fcccf36b57f29af0f9a5b05b3cf6b

Observation 382a8b7e-f7e9-42ba-a869-99cea33ba63d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:59.853337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:a1473dac866fe32260be18cbe6ab1f1c4ad2eb8aeeb364f9eadb9d85530af54a

Observation 4c7adc71-39bf-4e3d-b1f4-5aa3f0f4f6c8 · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.693494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:f0f235353c89742bcebfa2bb77e88e65c6553ad533ea87e90c9de79f721742a0

Observation a17cc85f-a316-4bef-adcc-640fb215f5ba · outbound

This paper cites Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.813805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:315f784c69df6c1dd8237301e964b6af25b0672e8c77a0daca26ed7a24f089a2

Observation 5e516f3c-ad6d-4590-99d6-7dcaf053f4fe · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:51:30.429014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:090ea78f10b37f8af69701d1a4ddbc5042d1705f4920a6c7536d83473085136b

Observation bc60c29f-e64e-4c1d-8d8c-ddc4ec6841ed · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:57.879683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:673f202adae73823dba7fc19472bf97255cdad7f77ac83e08b3bb64d8e4a91d5

Observation 523ea300-591f-4fe0-8279-3dd872289080 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:59.986351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:b0f6d9118626ed7c5ae8e86617b9bb014c6e583fa7c72f30bd102e727068666a

Observation 858f7524-ac87-4673-934c-914e639251a6 · outbound

This paper cites Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.784194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:95b0ae8a9b859506a13db052b2cc7ae86e1f0c08753f929b8c5a579e48d1deb4

Observation 9c3f1357-e3e3-422c-884b-cff07b2a11f8 · outbound

This paper cites Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.946039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:f13d48e8ff6b612a6be48637246e8d5ef5d5b6e5bd29467b2a0f942fb8c29bc5

Observation 2bb18992-52ae-4298-8fb8-c4e1a17e6a09 · outbound

This paper cites On the Optimal Reasoning Length for RL-Trained Language Models.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training On the Optimal Reasoning Length for RL-Trained Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-11T02:08:42.864067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:50a491e6f7b848513300a7699a7e34f386e59ce7802a4b820bc805be81e8904d

Observation d851b4fc-b9d4-4e7c-8dc5-3e7b345ffe8e · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:59.845713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:8604a13148df5d6b6e1c4e1b49e03865d6efdba953bc07e5a5364a330b1ec75b

Observation e29ec674-0fbe-463b-b15c-0e7904a9939a · outbound

This paper cites S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.831136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:29bc355d41b5f214923aa62001240f6185bd5491fdfc832c6138fcf04cb8a827

Observation 82a05dd3-6a45-44a9-834b-0ac172f5e82f · outbound

This paper cites Explore briefly, then decide: Mitigating llm overthinking via cumulative entropy regulation.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Explore briefly, then decide: Mitigating llm overthinking via cumulative entropy regulation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.877940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:4b9cd5e7917ed4f284aa6e6904af61b5a2d8cb2ab03d4f803740166537673a43

Observation 0809db36-c9f9-448e-83e8-83e98ef3a1c4 · outbound

This paper cites Optimizing Length Compression in Large Reasoning Models.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Optimizing Length Compression in Large Reasoning Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:59.740619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:3e76564bfdf973a4641b69166587194ea463e6a0d49390903ea6dfe8393fe386

Observation 30a50862-f941-4f8c-8585-f242735d6ccf · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.823484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:d6883e1d64f57b3b6b558a079a7cc5d6b5ae69e3c58c8b03e099c65ae8981080

Observation d996b2c1-faa1-4792-bec6-3842afdbf886 · outbound

This paper cites 2025 , journal =.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training 2025 , journal =

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:59.680569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:f97590aedacbaebac1e00b82f68e7ff4b701be9c837779e3d78bf1b49d3d1aee

Observation e2efe663-dd40-4762-8852-1ebd6706bb78 · outbound

This paper cites The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.710465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:14779dc91d9396f28b731d221e47a0deba4b07c3b9e52b540b351efa040ac12e

Observation 11a22fa5-d1c3-421b-85d2-0dedcfd2c553 · outbound

This paper cites Training language models to reason efficiently.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Training language models to reason efficiently

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.718679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:5c67996ad9f4fe2ae490f7c3c87fc1d761b3ec7dd891b7f6cfc9576a83b8a89b

Observation 621aa321-98df-4cf1-8d43-93fc91d81e1d · outbound

This paper cites Learn to Reason Efficiently with Adaptive Length-based Reward Shaping.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:36:00.027949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:0418f4a187553bea671af4ff1420800d5571c08e1cf3981082f0098af66d3510

Observation 902f5ff5-d894-4cdf-8cb7-c4071c5d5742 · outbound

This paper cites Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-23T01:12:02.291141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:9eb53c4ba47229d6b140dd20ee2b4feb692b01d0581da362223f86e66bc43d81

Observation e7c4f40f-4e81-4b8c-a167-3a09be53fdc0 · outbound

This paper cites Deepcompress: A dual reward strategy for dynamically exploring and compressing reasoning chains.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Deepcompress: A dual reward strategy for dynamically exploring and compressing reasoning chains

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.747885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:6c6174f045ce6613912c713980dc93cb3d6b9908dbdb17d8796870e087c7656f

Observation 5898e4d1-b6b1-4d22-af7f-e5e754e60ac1 · outbound

This paper cites SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:04:15.814003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:04a3200152322c1402b3d1b901e6125cc9086444cc9c5faf788eac809e0aa5cf

Observation 4369648a-efbc-44ca-b123-55488db6a1dc · outbound

This paper cites Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.701953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:c6697d7de3a4037163cd464228e46f5d70ab1d1d4eeaae0379bc75335498e7dd

Observation 7aac7b12-34de-4f66-a37a-85817e8be6e5 · outbound

This paper cites Adaptthink: Reasoning models can learn when to think.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Adaptthink: Reasoning models can learn when to think

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.477271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:0b36ad07c7a4bf1777bf92412d320d9735cc683926b1ae7e07069e002d4ab7fb

Observation 87f66fed-3d7f-449b-8072-97ab65fa8de9 · outbound

This paper cites Chain of Draft: Thinking Faster by Writing Less.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Chain of Draft: Thinking Faster by Writing Less

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.900134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:dd85a1545f7f706416898d7ae6d38c642c419814355c7aea58a830274ae936c5

Observation a2038e33-a1d3-4ef7-803f-99fd6fbbb808 · outbound

This paper cites Qwen3 Technical Report.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Qwen3 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:59.802958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:0c2fb68b36ead194de0520521684f2d4fc1a5e0ec1917151af5a128756eccac9

Observation 0e1542cd-3c0e-492b-aef6-270839fab4df · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.486327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:30e60c4a4cc267e7662930cf2291191102f453d84d7667000c900087433ac132

Observation 0307b7ff-7dd8-42d0-9f09-d58c28a74063 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Hybridflow: A flexible and efficient rlhf framework

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.510937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:f8b6d453cada2761b2484729567bb25ee23feb33d98d4b51e30f27ea1142cf83

Observation e0fa0a3f-c747-47a6-bb6a-3321c8fce8a3 · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:10:14.938194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:cfd0e11c051f80f0ac1eec189b36d3a02b33e8a2746b3deff214868bf6c3743f

Observation 6220c4ba-6214-40c5-b81d-b303e91ae715 · outbound

This paper cites Let's Verify Step by Step.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Let's Verify Step by Step

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:36:00.019738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:b124915d8278c017a89621738b6e9daccaa35cbf6cda1ec0185dd31394539458

Observation 1b513867-1f1a-488a-b858-ed9731639766 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Training Verifiers to Solve Math Word Problems

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:36:00.009953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:8ad2e30a645a833f6f51b058870ab3f81345b8b577a6d5f6b9c9857400e9c606

Observation d8cef271-bb77-44f9-9d9a-b0937f66f8ac · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:59.795200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:c1e27624fc6f15e8920c98b4bc98c96e9597c07a484a3f763b155ca14f5973f0

Observation fcb8b05d-5457-4aba-a0e5-5f60bc40d167 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.499355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:a91ec7ca253ac76d737af80d86919e42a1ddbebf3d518d40916e2ac0f5d8e914

Observation af6d637b-0215-42d2-a44e-eda3bd6f84d2 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 36

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T00:47:04.366365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:354a6da5e71696c90bf122b16ca9cba9b61b672800057e8528818beaf0105335

Observation 25800b19-0182-4413-b403-5c108fc0ac15 · outbound

This paper cites Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.519362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:a8fb123062f9dd6a9762092c1a20c31dacf7cb8a67b7443af88ac5c19d86bca0

Pith citing papers

Observation ecf09c03-2927-4a36-9243-d9819033cdab · inbound

Distilled Reinforcement Learning for LLM Post-training cites this paper.

Distilled Reinforcement Learning for LLM Post-training Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.459071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.459071Z digest=sha256:e748f0b59730152d210b6a1a0ec40ae01d4155b18c23ab6bdbcdfb4dff426eb3