Pith. sign in

Paper Citation Record · LEDGER

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

As of 9 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.14614.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14614 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:39:11.464437Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0689e7eb-20f2-4117-8bdf-3be15eeebe3d · outbound

This paper cites R., Geist, M., and Bachem, O.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization R., Geist, M., and Bachem, O

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.111981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.111981Z digest=sha256:85b4c94d8e6a8f44e1dfe9bb47b17ae1c932eec8b8c114049f3b6a494d7ae764

Observation c3e5efa7-533b-43f0-b14f-cbe4f735064e · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.261379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.261379Z digest=sha256:84793768cf364fa66368ebc599eb2d9cb571753a4e635b3d4a226cae81f57831

Observation 703dcbbb-8474-4336-9b52-e30328945863 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.367580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.367580Z digest=sha256:95c920d79bf4e9e3d52b5e0c071a0e799149ff182f8c0886d7e58d0d39590e8f

Observation 61b93e40-3d7f-4cdf-9e2a-120b63a30cdc · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Reasoning with Exploration: An Entropy Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.528882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.528882Z digest=sha256:8e1e85ab940d20dae941cc0465ff2409c3b1d939a1e1e2a114df9255e4d8c9ed

Observation e9dea355-746b-4272-b43b-a642bb4cdfdb · outbound

This paper cites K., Chen, G., Xu, W., Luu, A.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization K., Chen, G., Xu, W., Luu, A

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.651251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.651251Z digest=sha256:03cc7508d2bca2cef2906e60daf68dfdbaeeca4d4da2a74e971be71759011e1e

Observation b230e182-72c6-4661-8c47-925519a34555 · outbound

This paper cites Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.050935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.050935Z digest=sha256:322961c2d24634e1b2342bcf0c6b4fae284646ecab6770a546ce95c8d9161ece

Observation 20fe05de-a943-431b-a89d-ac5d7396223a · outbound

This paper cites Deepseek-r1: incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081): 633–638, September 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Deepseek-r1: incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081): 633–638, September 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.256223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.256223Z digest=sha256:96fd0d21e354f8159194044b7334bf0bee9d2aed2af629e67d0fc93cc80780cd

Observation f0e898bf-91bf-4164-b613-9c1939a72c09 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Measuring mathematical problem solving with the MATH dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.800750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.800750Z digest=sha256:e2bc82a3804ae7e57b258c7f5836dc880d8503d7f902f71d4b718278e03805a6

Observation 65be8710-17e3-4c27-89cc-916385f00bef · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Reinforcement Learning via Self-Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.961710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.961710Z digest=sha256:0c86781aa80fd7ce252a98735f4d6400bc09c9abc28313397a4c44c4c3255a40

Observation fed4f4cb-834b-4735-97b1-c9820b7609d1 · outbound

This paper cites F., and Joty, S.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization F., and Joty, S

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.080907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.080907Z digest=sha256:34214b8842964d77bee585e22674a09741fc869810df28ca6ca49b609f13e734

Observation fdd53a9a-c87d-4837-9844-6cd8ab6e4b68 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.211229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.211229Z digest=sha256:de410ccad4a29d7db4ce892d68e1549cfe409559637c655fd6e6721ae46b41b4

Observation fceea892-3f7f-4d84-bd97-d024ff84b8f4 · outbound

This paper cites V ., Jeon, M., Vu, K., Lai, V ., and Yang, E.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization V ., Jeon, M., Vu, K., Lai, V ., and Yang, E

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.336708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.336708Z digest=sha256:ccd2e2eb1fe91408646bc4433a35a110baa25e8d4c0094cabc70654821c8db5a

Observation 2d47ed9d-be70-458f-ac3e-f1b4e783acb8 · outbound

This paper cites Revisiting LLM Reasoning via Information Bottleneck.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Revisiting LLM Reasoning via Information Bottleneck

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.465608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.465608Z digest=sha256:bdf067eb4d1434915c3467b973f9bf8a86c81fba6a2462ae03fcae216edb823a

Observation a2ef72b9-6fad-4cf5-bc2f-4e97442b67c9 · outbound

This paper cites RED: Unleashing token-level rewards from holistic feedback via reward redistribution.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization RED: Unleashing token-level rewards from holistic feedback via reward redistribution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.597752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.597752Z digest=sha256:019c1475059771735a578f65a1a45b2700e92e6c2f15df0d8feed1927513f4a9

Observation 80a54d28-261d-4a24-9277-f7be0d804108 · outbound

This paper cites Can we further elicit reasoning in LLMs? critic-guided planning with retrieval-augmentation for solving challenging tasks.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Can we further elicit reasoning in LLMs? critic-guided planning with retrieval-augmentation for solving challenging tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.743422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.743422Z digest=sha256:57d4c061f2765e4ead888b041bbed82c829cc0483a08b864e3feb2bb1f30e76c

Observation 71afbd51-17a8-4d77-a281-67b2a3efd58b · outbound

This paper cites Let’s verify step by step.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Let’s verify step by step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.881210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.881210Z digest=sha256:1511777c5cda25aee83fd74fa0d7b7ff8ef74709f3f72ba75c564f93b1c91427

Observation 29d4ae00-9413-4c24-a698-8104a05dbb6b · outbound

This paper cites Ravr: Reference-answer-guided variational reasoning for large language models.arXiv preprint arXiv:2510.25206, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Ravr: Reference-answer-guided variational reasoning for large language models.arXiv preprint arXiv:2510.25206, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.013709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.013709Z digest=sha256:e33ce5af25b0eb79ec57050ffe6594e2c7551590b4a52eeba884d240d020583d

Observation 4aadada7-4f06-40c2-ae7f-1c52211b643a · outbound

This paper cites and Lab, T.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization and Lab, T

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.145093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.145093Z digest=sha256:e65af4ed28605e879ca7fddb2f763933248babbf162659f93fb63209a970d692

Observation 25553960-5153-441d-a7f7-d972e5747e04 · outbound

This paper cites American mathematics competitions (AMC), 2023.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization American mathematics competitions (AMC), 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.341996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.341996Z digest=sha256:f8165e5854cac3e7b6d53c1be45ad48f6d128132bd2be8d4e369b41993f3fb1a

Observation 378c8d48-05fa-42ab-89c3-d7fbe08828c7 · outbound

This paper cites American invitational mathematics examination (AIME),.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization American invitational mathematics examination (AIME),

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.502677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.502677Z digest=sha256:f0a5098f0a4f9638ff3cb952b3849f04548ac9dd8ebd6e0014223a81cb62174a

Observation da0d60d1-89d8-4da0-8fe0-adffee18193d · outbound

This paper cites L., Stickland, A.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization L., Stickland, A

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.860327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.860327Z digest=sha256:927337f106c138bd16a110ac441124b7bd149f537d289c0668c1d7f7dab83797

Observation 0cf673c1-2b74-44d7-a479-ee9a0a320d1f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.049351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.049351Z digest=sha256:cc185e5a24f9ef659fa0bd4c86e0c243f26dff5c909da9464ce7ec1e9262cacf

Observation 40e7aa23-9973-41a1-bbe8-68ba3ebc29eb · outbound

This paper cites Accessed: 2025-12-23.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Accessed: 2025-12-23

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.661112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.661112Z digest=sha256:a8f97d45def1fcf62084922d4ff668ce080fc69d7cd6f55799b8c3174a608520

Observation 61dfcddc-5072-4349-ab74-4c0092e78eca · outbound

This paper cites On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.405827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.405827Z digest=sha256:a6c01a5440d29bf8cbe182b9d6dc85b7cca635ebdfc7ab671a0387c107e00e7d

Observation 1f33537c-e256-496a-9950-69b95093438c · outbound

This paper cites Self-Distillation Enables Continual Learning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Self-Distillation Enables Continual Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.574318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.574318Z digest=sha256:08d0b98ce5435a3b1e2ac4588eb025af4ac63af1e5a3a74cc468ce93769d14ce

Observation e94f6c29-ce3c-4863-a626-f99a505f94d6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.238935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.238935Z digest=sha256:23cb687c90746f483e0c2a703f2b59fa9599d73551a9d00a6f394f4eae6052a0

Observation b1bec4d7-fd43-4282-aa3e-8760191cd133 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Kimi K2: Open Agentic Intelligence

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.864772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.864772Z digest=sha256:7f0e0254eb62cc8adc13db4f2c1fb247cade4b4dac1107d14e421a7b9f0354da

Observation 87738a82-72e2-41dc-bae3-09dd7cd67609 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.003116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.003116Z digest=sha256:4af172ff2e65563ecb29959a7c212bd789de0e2a0b32af5776b12169fcd379a2

Observation ce0c3b6b-14d8-4b8e-9bbf-26743f2329ba · outbound

This paper cites Espo: Entropy importance sampling policy optimization.arXiv preprint arXiv:2512.00499, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Espo: Entropy importance sampling policy optimization.arXiv preprint arXiv:2512.00499, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.745985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.745985Z digest=sha256:37095796bcdec52acf354f5dbf3924c1d84a6921f00a10521cfbe0ac1053aad3

Observation 9f56af7d-adc5-4061-ac52-35a00dfab8da · outbound

This paper cites SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.323586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.323586Z digest=sha256:6b67764e12c292ce5987b788e06efbb4a9025f4e0d67bf4c187a29dc229f0516

Observation ff2c5ad0-5fd8-464e-8202-baa1d11ad2fd · outbound

This paper cites Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.449103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.449103Z digest=sha256:2141a83ecda77032c5a26aa02f3cdd6b0dac691bdd0b5fb74dad61471c11b336

Observation 73026038-d1e2-4d6c-926f-ef4f97237b75 · outbound

This paper cites and Karkhanis, D.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization and Karkhanis, D

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.192396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.192396Z digest=sha256:ce538a49805e7c9bb348bba99d865e98214b630b1c9356c0188d830effb662e2

Observation c18f4ab1-a996-42ca-be58-bfbab0daff46 · outbound

This paper cites Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for LLM reasoning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for LLM reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.714342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.714342Z digest=sha256:60828b2c0b3650c4b127625c0597da5a4d85fad9b382a2c626921bf1621b657f

Observation 83fa065d-7d82-4a64-9b34-c917ca434a1c · outbound

This paper cites MMLU-pro: A more robust and challenging multi-task language understanding benchmark.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization MMLU-pro: A more robust and challenging multi-task language understanding benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.906301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.906301Z digest=sha256:b42b32c2049688657e2f1889a987ef0649687489b0fb4bd7d31b4d85a9be56ba

Observation 8f2f0074-1d9b-49d1-84ec-92c6fd15cd36 · outbound

This paper cites Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.605142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.605142Z digest=sha256:12affc1d0029ed899680babeb07bde345e419b70e49d23140b74a0d1bd0a89d9

Observation ddb3deaa-f872-47f7-b5f4-6ecc6f2c91f6 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-02T01:39:09.162940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.162940Z digest=sha256:2b5a617608d01d43528e32804c8de95a0ab28dac4a4251f61f55a2cddc4d85f4

Observation 52a76309-0236-4747-90da-a17746c21878 · outbound

This paper cites Quantile advantage estimation for entropy-safe reasoning.arXiv preprint arXiv:2509.22611, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Quantile advantage estimation for entropy-safe reasoning.arXiv preprint arXiv:2509.22611, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.291533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.291533Z digest=sha256:ffb2d2c7fdc5c266c4526997b302cdc9579b2418477198f70ccb6d5a250b8ce1

Observation 43079864-152a-4fda-b2cd-4706f00e4322 · outbound

This paper cites OpenClaw-RL: Train Any Agent Simply by Talking.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization OpenClaw-RL: Train Any Agent Simply by Talking

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.040248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.040248Z digest=sha256:30dfc1ee73fb0cf524da96dbec41a665b30d4726a18c3c5c335127e8484cfb08

Observation cf97f670-e25a-4dfc-bc18-fbb7f0b853f9 · outbound

This paper cites Reasons to reject? aligning language models with judgments.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Reasons to reject? aligning language models with judgments

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.686011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.686011Z digest=sha256:d95c6ad9035ac1579182767369b54cd216289ec1651335bb98a04fae76aa8190

Observation 804dc322-12c0-41bd-b001-027a18c64b29 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.838682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.838682Z digest=sha256:5b4970d34dd65bc85f7be87ed19b03410e90c2682b44205b053950cedc1a13b1

Observation 5017778e-003b-4130-a602-849d58ae4ddc · outbound

This paper cites A., Osten- dorf, M., and Hajishirzi, H.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization A., Osten- dorf, M., and Hajishirzi, H

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.482306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.482306Z digest=sha256:7cbfbb66dacba2d0b2084bb417daf601b3015e7c3605456f31d69ccfdd3cf5c5

Observation 30e4be79-6d98-45a4-b9e9-669e3bc40d89 · outbound

This paper cites Self-distillation bridges distribution gap in language model fine-tuning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Self-distillation bridges distribution gap in language model fine-tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.134964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.134964Z digest=sha256:7a03caea7834190d293ec191089f736159ece5579cb64eb77403117300e1561d

Observation d0a49c8e-a0dc-4b59-b561-71e72d61c698 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Tree of thoughts: Deliberate problem solving with large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.331443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.331443Z digest=sha256:73f03517285625ee3a1a17c979f8503b8d6888fb2603f3c00003ddbd534ee97f

Observation 6c38c0ee-0f90-401d-af90-170dc0ea9620 · outbound

This paper cites Qwen3 Technical Report.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Qwen3 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.987173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.987173Z digest=sha256:76ec2f1358052108d14dc085d50a40ebe683ff412db2da3c8eb3932ae8eee955

Observation bbdfea30-7b90-42f2-9959-1aa0c07b2398 · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.461306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.461306Z digest=sha256:043e45359b7194aeceac6faf45c7d972b6e7decddef625dc9609acab62ddc72c

Observation c46e2e8f-02b8-4482-9d30-57bcf13ee76d · outbound

This paper cites SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.555531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.555531Z digest=sha256:b71cb5613664d05cc4b16a2098d1dbf87ed16daadcaf255814a53ec73ee31533

Observation 17252ee0-4e5a-4a37-bc01-6e96b3b77dce · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization DAPO: An open-source LLM reinforcement learning system at scale

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.410737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.410737Z digest=sha256:0f3b5c009677e7d5d69e11cbfe168b9746394692a02deb44738c938369cd09f8

Observation de45ab0f-6525-4a0b-8a4a-4ac946e6c732 · outbound

This paper cites R., Kailkhura, B., Lai, F., Zhao, J., and Chen, B.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization R., Kailkhura, B., Lai, F., Zhao, J., and Chen, B

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.749976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.749976Z digest=sha256:626aed3667aee1d9b55b3382efec0dd5a3426b012f28f2cc48cb70ae542055ce

Observation 8331d4b8-dc11-4eda-b24c-c560ab2b1b34 · outbound

This paper cites First Return, Entropy-Eliciting Explore.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization First Return, Entropy-Eliciting Explore

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.853118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.853118Z digest=sha256:0344e3bcf949682eebba5caaca651c4e417e254d5949a2b593f1221de1930393

Observation ee0017c6-1e68-410b-b189-a4cd75cdf2f4 · outbound

This paper cites Group Sequence Policy Optimization.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Group Sequence Policy Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.657660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.657660Z digest=sha256:b8b33ef41d5c7c9e88596dde6373918779dd1685a62b3054f724fb57dca0b83c

Observation 2d65af65-e512-42f9-b615-bd93fd42e247 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.953149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.953149Z digest=sha256:b7ec55b1666e1f1d326831b77af2433cac6d92056d4877da03d354b8a4dea545

Observation 4291b401-e973-4b24-aaa0-f65e5ad55dce · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:11.054736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:11.054736Z digest=sha256:f058ccc4baa0f297462f9fa5f517ae225ba3cf2e29aa879c8105af0e71d2492b

Observation b3aae2d3-5a8a-4f6a-97a7-4faf062ba1d2 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:11.181475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:11.181475Z digest=sha256:74b765c89aa6e0c7c1fa0856aa1f6fe9b3c66e09508a4f89fa374c3987657793

Observation ef22e5eb-c616-46c7-90a0-0193d8386361 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:11.357786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:11.357786Z digest=sha256:3ddf0c6b2445e5ff179a7b75bd574eb52a9cb40d97643cf72d7b15186d4d099c

Observation f6baacee-6734-4eb0-a29b-3a115665d2f3 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 58

Resolution
malformed identifier
no resolver link, observed 2026-08-02T01:39:11.464437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:11.464437Z digest=sha256:312772a74e38ba3d02783c5069591253663ccc12a1fcf195d33da95a1f9ed8f4

Observation 064e667f-ead7-4b58-9de7-827eee6bce35 · outbound

This paper cites URL https://aclanthology.org/2024.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization URL https://aclanthology.org/2024

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.851537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.851537Z digest=sha256:c73bc8973148995b5f70a0202799dabfb78c485a4dafe51db4f53ab0d4f15e7b

Observation 4e3e6b8c-43f5-4ad2-a9d5-4ff845c7fd23 · outbound

This paper cites Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.601106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.601106Z digest=sha256:ee799d08f2837304037300b9145d238154264fdfac5a1894bfb0ee7ff5753207

Pith citing papers

No inbound Pith citation observations are available.