Pith. sign in

Paper Citation Record · LEDGER

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2606.08088.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.08088 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T19:56:09.820812Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch26

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0eac99e5-97d9-40c1-a7d4-60571501673e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:b3fefe53572b1b19f473476b30a184c27d161ac42af6534661e573b9724c1465

Observation 457f3a84-162b-4381-8e5e-f7a65fda66f9 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.905688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:9b110d6c131f0bb839b7cc61d9f31be3786a3e390c2495173237db4113911bd0

Observation 5e4d0d75-8a83-4d5d-8486-225e33f4ba50 · outbound

This paper cites Pan and H.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Pan and H

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.938374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:34d241c6ec8bd899abe59da4aebfdf916599368404b8f136a04bd73ca4b093be

Observation aa580c9a-1676-4f23-8d9d-52405be4a5f2 · outbound

This paper cites CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.933247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:c39f68a55164032922e4363615c0ec423004ee3b91ebfe29980825ab5ebb7240

Observation a4fbc573-08a0-4e9b-bcd5-e54607ad84ea · outbound

This paper cites arXiv preprint arXiv:2601.18533 , year=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning arXiv preprint arXiv:2601.18533 , year=

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.922659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:0fba99e590fdb8097c7966a6f8cb6f8971b73cae700bf3b44a731952e0887550

Observation f129b06e-b4d6-4fd2-bbd7-2f3857fb4541 · outbound

This paper cites Hansen and Duo Peng and Yuhui Zhang and Alejandro Lozano and Min Woo Sun and Emma Lundberg and Serena Yeung-Levy , year=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Hansen and Duo Peng and Yuhui Zhang and Alejandro Lozano and Min Woo Sun and Emma Lundberg and Serena Yeung-Levy , year=

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.925100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:f1bee888f24475af25353e674ea39af2b6bab6b17f02ad7021e4a51c27ada2be

Observation ce1eebc0-fa28-43cd-9131-117f7bf7fd6b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.927648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:f1fca8855442a35806a645850061114fd29f1099606a7ea4e715d24289c7fa07

Observation 64b39f26-c202-4303-b120-3769212a0f2d · outbound

This paper cites Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.935694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:ede0ab5c2f0267d629064aab2d86023e52f580669a52e383b72f6b44c9a0add2

Observation 5795dac7-6c37-4d11-a626-c7a151ce49f5 · outbound

This paper cites Save the good prefix: Precise error penalization via process-supervised rl to enhance llm reasoning.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Save the good prefix: Precise error penalization via process-supervised rl to enhance llm reasoning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.908613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:1f988200fd087ce8c9d076c186b614c18e0e46d2fe936eb0a4f2eec8de0151f5

Observation 96734832-56cb-42d8-90c9-2cd72e1a68b6 · outbound

This paper cites Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.914209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:68f770afa32203afbff41add84812b3ccdc0b37f09d529ec5a3965c6745c6fde

Observation a4df47c0-54a5-4439-8c05-c1f63bbcbcc6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:8aace0e6e4d42f23d083a911819e75bf52a925618ae0725c19cd7209fd66d21e

Observation 81f6aee7-5482-4a80-b541-26cb9dcd9333 · outbound

This paper cites Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:2ad32cb12874760cbf77a74d8db193a34a1e39ac412c882a48e0feff33bc89fa

Observation 718dc540-162a-4b4e-ba19-b8a81123df5c · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.900753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:2a90c8c7faed95824a73ee7385ea8fb272261e409abc86a2525d76b3bab49855

Observation fd4058a2-76b9-4c82-9ac9-bb9209ebcec3 · outbound

This paper cites Proceedings of the 2023 conference on empirical methods in natural language processing , pages=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Proceedings of the 2023 conference on empirical methods in natural language processing , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:de5fc10848e4f3ffa4de187fe9c65b77e023eee99b2cc060e5960c56f0bf4763

Observation 8c9aaf55-e41e-4f80-9b30-39471778c1b3 · outbound

This paper cites Can AI Assistants Know What They Don't Know?.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Can AI Assistants Know What They Don't Know?

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.898118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:38766e79207aed953438e10e3cbf333bd1f9d5985a9c25c72b9b6f487f38427c

Observation 2e75d40d-a96f-491f-82ca-8a19ca6aab03 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:f98f6eb135f5c215c8ab6d13bddbe73a4640dab4850c292d0dd252d37e7b1d18

Observation 1cdb34c0-bea1-4a2e-88bb-51c9e8898260 · outbound

This paper cites Enhancing Confidence Expression in Large Language Models Through Learning from Past Experience.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Enhancing Confidence Expression in Large Language Models Through Learning from Past Experience

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.894907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:7861d4ce44555636b26167ffc328118b69fee7f63215f4032b72a796804ccf3b

Observation 69a4c557-380b-43e4-82f2-9f8c87ed3f95 · outbound

This paper cites Advances in neural information processing systems , volume=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Advances in neural information processing systems , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:1aea6476115210628e3279489efc8c5ae0fbc02ba172b0d4fc275f21aa97323b

Observation be217a70-f455-46b5-a1e4-3587adc67dca · outbound

This paper cites International Conference on Learning Representations , volume=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning International Conference on Learning Representations , volume=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:c7a69b2357ba46f343e5a446cf36b72d54f44b8ea789be7a12c9a7d34c915e20

Observation c8d6598e-6707-4d0e-8607-463a32089b58 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , year=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Findings of the Association for Computational Linguistics: ACL 2024 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:c0cdf10f22a8bd094e2ee2d4a33d6b7168630bd33984d1a59cf88b06a59d1794

Observation 59cb0b50-4c60-4f3c-abb6-e8925981db91 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:a3ddb44b73fb30cb7ee59eaffed85fe5e10b7372c2b4defd6d610600644c9700

Observation 2c3a458f-77f4-4c05-af5f-d4c24edde5ab · outbound

This paper cites arXiv preprint arXiv:2503.02623 (2025).

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning arXiv preprint arXiv:2503.02623 (2025)

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.892354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:db9a601a5b87d3978b72e002eeeba2b40190d9ad3965e0ea15e6795d79817cd9

Observation 711a585e-f576-4671-8362-1e59c642e8ca · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:3736d2302cbcf9aa06f7e1e05b7da086c4e7bc3a5a186fcae7e6cef732a75fc6

Observation 526eddf0-9368-4a52-81aa-eae1d528d6d9 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:fd7690ae119eee5a33350d4373cd69147122a80cd0139c7d65b8612221b51d58

Observation c6e5bd6e-77d1-4f5d-80be-4eb688219c2b · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.903092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:d603270b43cd8770a9d090c567d253bbbd35f3656f322ab6535922d2df46e4d2

Observation caa66479-c02b-41e8-92a3-5e8caa048f64 · outbound

This paper cites Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.911497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:29e02abde7bb1948eb8132677889a4cf45a331c7f373a3deed68c522667e96f8

Observation 6acfe10f-7eb1-411e-b246-8ed429bd522f · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:7bc97882291e6f0ae2c85e12b3ced30cfa26dfc61337ea55e1b84e375bd4ab1a

Observation da3faedb-e85a-4033-a526-74de54257a75 · outbound

This paper cites Deep Think with Confidence.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Deep Think with Confidence

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.886476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:df2cede9c40b3d108831e546c91c5f1446d9289b829a0c2bcd985c515330660c

Observation fa41559d-c7d7-4b17-8d5a-b67068d91256 · outbound

This paper cites Harder is better: Boost- ing mathematical reasoning via difficulty-aware grpo and multi-aspect question reformulation.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Harder is better: Boost- ing mathematical reasoning via difficulty-aware grpo and multi-aspect question reformulation

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.881098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:22abdfe91580609f11aeccf98e675795f897d18b8a97d75e48c5b2961295622e

Observation fef274b5-193e-442b-883b-76f7d264321b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.878243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:91d92bcb6b7b850ccae1c0ff9794bb1f80770fd6192a09a7917713d1ca058dfd

Observation fc381771-0383-4da4-9c5c-5aae583d23cb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:1ce5412741524766b624abbdff2bc6fda3819f8bd00016d3ec2b61d33ea9d2c0

Observation d3d1ff10-4d27-43ec-8f1f-fc2b63b4a30a · outbound

This paper cites Group Sequence Policy Optimization.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Group Sequence Policy Optimization

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.915385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:ebebb0e4c63b2a0d9fe74ec74dd7a07a52169c482acd3d628161e573523096f2

Observation e6374cfc-faf8-4d27-bc8b-b7406f62002d · outbound

This paper cites Soft Adaptive Policy Optimization.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Soft Adaptive Policy Optimization

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.883819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:a41153abeec2f8d6c906aad160af9fcbbea5a5f96aa8f9ca7afb51b93edd33d1

Observation acf14a0f-0a73-4baa-aec8-9768f1748718 · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.875133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:00c529186dfeb7666dd78fbba8203e95e921cc50593ffeb4c1eff2a64ca44f03

Observation 4c401f2b-8499-4af7-9022-9e5fa945f24c · outbound

This paper cites Proximal Policy Optimization Algorithms.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.866648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:970e45895e993ea8865851cdb1fb47b5c71d69b5b5828fd18d032b17b67ec00d

Observation 26e49a38-037d-41f7-972b-ebddbf7287e7 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:07:23.858140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:0e2be295fd45aaf8b3dfa0e9bf84e45afaa4a2252f14b3bcc39fd72bc270f901

Observation 1934d145-c540-4c41-9459-d1bae4cdbfcb · outbound

This paper cites Proceedings of the Twentieth European Conference on Computer Systems , pages=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Proceedings of the Twentieth European Conference on Computer Systems , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:25057ac964102d9ea80701bbf2cff3a9fc7d2ea5e797c841ae15dd7372c4a298

Observation 6df45bed-9891-4866-9394-011220556066 · outbound

This paper cites Qwen3 Technical Report.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Qwen3 Technical Report

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.860897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:0df489fe6cca25803f8c172c240b0a3dcbc4061d88a6a98702137b195e2dc255

Observation 368189ad-bd44-4324-b428-4e84ad633ca6 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:07:23.863986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:0fa964c35a499f7c058cb38baca0bdf051fa0c9dbb74dcff9bfe8924f676971e

Observation 8151db6a-0ee4-488c-92c7-049b71de7618 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.848807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:af259d1928fc61b718b291a8a68cdbf40cbb7745e26ed910a1b768da8f4597ff

Observation 86a42992-4a6b-46bc-957d-59eadd424386 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Measuring Massive Multitask Language Understanding

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.851671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:8784ef38c2c52a583184c5c2b117ba235ca153976b826c311b0c3da17dd579e5

Observation cff1f99b-280b-4be0-a7dd-a7e6ac01f7b0 · outbound

This paper cites Advances in neural information processing systems , volume=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Advances in neural information processing systems , volume=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:30a9369123afa8fbddef4dac5a27f29e2a5ee6b06e88f61ffb8a709b92c579a9

Observation fe411c30-a9a1-4d1f-b629-5f0408d35319 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T19:56:09.820812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:982f7dad76cc298f41460f8553895912a9785141b4ca320f6a362ed1ab4bb8e1

Observation 7f39cbde-8113-4c3d-9d18-7008bbbd13ca · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.854950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:3673d1f5fd81f7bb690b3d56f9f5f05fa8004cfd60de84a02405db930d194794

Observation 9dfb7d77-5ad6-4732-a83e-ee13de50b8cc · outbound

This paper cites arXiv preprint arXiv:2503.17736 , year=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning arXiv preprint arXiv:2503.17736 , year=

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.869714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:6b635c04350661ddf495cae5de07279ba8a8b0b56eb950919e4bb9b4c8b5f1e2

Observation 7a0ee72f-3988-4eb9-9088-b2e897cb7516 · outbound

This paper cites VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.872673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:59a34104533c7e32e9c4788f0227d541ce640b886728a6b1985267e0aee93d35

Observation f7e8c7b0-933b-480c-a001-3ca2db9532f1 · outbound

This paper cites Vision-deepresearch benchmark: Rethinking visual and textual search for multimodal large language models.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Vision-deepresearch benchmark: Rethinking visual and textual search for multimodal large language models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.916853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:acf924e31abca0ea9afa41a8d17740b1d567ff75719f3b22cb15e0f2037cfa2c

Observation 55701cbf-2b2c-4166-b565-84f8ab479874 · outbound

This paper cites arXiv preprint arXiv:2510.01304 , year=.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning arXiv preprint arXiv:2510.01304 , year=

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.919756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:2f8b2c97896061417f4440af838a80843e9e51c65728cffffc5cd7c18cabba15

Observation ffb93de8-b762-45ef-97df-7016bb2954df · outbound

This paper cites VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.930391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:d3f725c5c01912d7bd1acd71c3e49a6c81c5b73132a26600502064d2cf3cf06b

Pith citing papers

No inbound Pith citation observations are available.