Pith. sign in

Paper Citation Record · LEDGER

Process Supervision of Confidence Margin for Calibrated LLM Reasoning

As of 2 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2604.23333.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.23333 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T08:19:09.437464Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T06:48:41.331963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact40
  • verified fuzzy43
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd1272f0-df01-4d45-bd9a-e081c8a22cd4 · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning On-policy distillation of language models: Learning from self-generated mistakes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.604886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:1de48d7110c6cbe444214ceed07190e31b35fb58f4ec02200f21c0c52e571adc

Observation 62a90c69-16d0-4d4a-a952-d2b07ca5febd · outbound

This paper cites Barber, R.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Barber, R

Reference 2

Resolution
metadata mismatch
doi, observed 2026-05-08T22:34:28.718729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:739c703e31540adb74375987db8ab9bd075cab017f4198c004f2077943d8898f

Observation 3b4e2a25-f8b2-4cb2-8cec-dd36660a7133 · outbound

This paper cites Conformal risk control.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Conformal risk control

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.632806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:af2d76c99e249bbbc6d073f444be634d10d38914b09e1aaeb3e389077d022340

Observation 225b676f-c243-467c-9aae-729aa5860759 · outbound

This paper cites Reconsidering LLM uncertainty estimation methods in the wild.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Reconsidering LLM uncertainty estimation methods in the wild

Reference 4

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.669854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:f9d5a6e02800dd6e5b14abf3a62d7c9dcbd46bbcf870235ff3deaef3e3f7af09

Observation b1cecae0-2a6c-4066-9c2f-4771d76c1a43 · outbound

This paper cites Linguistic calibration of long-form generations.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Linguistic calibration of long-form generations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.599634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:8080bdc66dc2fc897d277a032443cacf22adab5f4bc2d13cc263c4ee999ae690

Observation c6414d67-98ad-45a9-b3ca-bde9d0fccd34 · outbound

This paper cites Rewarding doubt: A reinforcement learning approach to calibrated confidence expression of large language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Rewarding doubt: A reinforcement learning approach to calibrated confidence expression of large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.628017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:22c6cdc9173fd8e12ef5d0fbc6d3a15a16cd5e5d2724c66e7ea5d7c8df779545

Observation fff23457-9f47-4199-a675-a01faf932cb1 · outbound

This paper cites Bereket and J.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Bereket and J

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.302769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b3e23f86e36328581f251ea74f866a844ed7d8a504eac866b058cd4b561a0ae3

Observation 04ee208a-c0da-4018-878d-53383f949a62 · outbound

This paper cites Do Androids Know They’re Only Dreaming of Electric Sheep?.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Do Androids Know They’re Only Dreaming of Electric Sheep?

Reference 8

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.714791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:d6dfecef9c7123ff50c79f94f80c19f2f60b9200f7a8271a8d70f6cfbd286cc8

Observation e45f9f31-1f6b-4c2a-9566-a3afb12a81c8 · outbound

This paper cites Uncertain Natural Language Inference.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Uncertain Natural Language Inference

Reference 9

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.677536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:ab6cdba052f2363d51dc3e9af3504fe46117920baac2e486f995d0236133accc

Observation 194d0830-dfe6-4e88-8a46-9f14ad0fc725 · outbound

This paper cites A close look into the calibration of pre-trained language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning A close look into the calibration of pre-trained language models

Reference 10

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.673639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:db118c54f148d637499fee20e4a60c7eee14600be273d61e6c28836bd71d3e75

Observation acad2e9a-1d91-407a-b682-c58cdc0cd67f · outbound

This paper cites Mind the confidence gap: Overconfidence, calibration, and distractor effects in large language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Mind the confidence gap: Overconfidence, calibration, and distractor effects in large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.625395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:57125483c6a7599f9e073331266d50a751f0122570f2bc701d5e583047647fca

Observation ba73bd9c-b312-46b3-a77a-b3986201bb76 · outbound

This paper cites Evaluating language models as risk scores.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Evaluating language models as risk scores

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.597097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:d13401bf3387ce8969883b5d7e2f4cdbe2337d0cb50955f074f2163861c91dae

Observation 8118ff22-c021-447c-a758-fbe4a6bca34c · outbound

This paper cites arXiv preprint arXiv:2603.09309 (2026).

Process Supervision of Confidence Margin for Calibrated LLM Reasoning arXiv preprint arXiv:2603.09309 (2026)

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:41:12.062181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:0bcead34be2efd01fc6cc81ddd60bd6adb74eedc334ba85c263d6575e0fd82df

Observation a0df8c10-e56c-4d0c-9234-65c0e2dc5bd2 · outbound

This paper cites Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:5bbfb82e558223faff48c1665720cf97f4514fbd305f66a772de0b7e60f41473

Observation 4274fba9-8265-405c-9bd1-0cd480105a2c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.070150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:4914119695ddd6aa29c894b50c78c987b9b8c63a589dd779a6aa26bfdc48db84

Observation d128a3c7-bd13-45a0-a906-8638ac224943 · outbound

This paper cites Calibration of pre-trained transformers.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibration of pre-trained transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.602127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9852fce005fa9bbe4fc7ced64bf267787fa977bc1a0671698b59294ffd13aead

Observation 5c09aea8-7652-4c17-a01f-69ddf6ec0e34 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Detecting hallucinations in large language models using semantic entropy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.607466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:aa69892fe29f3375e1fdbd1d356eb9e9590a0ecc8130027e07d0830a6d270bec

Observation af936a36-0466-4beb-b767-25f7106813be · outbound

This paper cites Deep Think with Confidence.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Deep Think with Confidence

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:30:21.518780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:2b01ba20aa1419eb843d9d318dc878293b1075ba4897f41022481a9af8fdc760

Observation c9324866-5083-4271-b0cc-967a7d707468 · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.225348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:8044c34c5a0eb26b9e7e5ca9662693c983af29e3420c9a66214bee2720ce4f5e

Observation 43982279-c7db-46c1-8238-4f257e4d80a2 · outbound

This paper cites On calibration of modern neural networks.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning On calibration of modern neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.630777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:71e8ad3be21e4e4c6f341cc6e1e4fddfce7eaf03d829a82a4f48fd4e492b23cc

Observation d904375c-697d-4c78-bfac-a464c7c4f4d1 · outbound

This paper cites Language model cascades: Token-level uncertainty and beyond.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Language model cascades: Token-level uncertainty and beyond

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.589798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:f947f698588598910858629f67d4a1b4078c2b414f949baf10e1c9255e732670

Observation c4b1cfdc-ba02-415c-bf80-965992023719 · outbound

This paper cites O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 22

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.661171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:672d166a966c3c77e8ff15bb14d5b53737177f82f68aa764b8beef02047b5021

Observation f5ec1206-ee0a-46ff-bea3-076db50247b8 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Measuring mathematical problem solving with the MATH dataset

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.592482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:0c01f8200d3b74fbadb54f0aab275ffac40a8b51e9a2cdb2ddcf0b47d5b46faf

Observation 23fb203e-c842-48a3-bbe3-dc713987d2b4 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Efficient Test-Time Scaling via Self-Calibration

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.347697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:92306e8de6ad24d3c49eda5a50cca29bf185af0abcf506dbd0427bfa1eb6a796

Observation 97eb9b03-c697-44ec-bbf5-e44f25368210 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 25

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.706923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:6a255ab312e7de552c04951c9b4cefad7d4a2380c566ea407cc02a4218388cc8

Observation 57394729-3a4e-46a8-96a9-9142b2c53f53 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Reinforcement Learning via Self-Distillation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:29:18.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9284c3a368fd64a41610c80b115c9073c71adf9fbb72cf525c4d2ff3db104f7e

Observation 209dc3a1-c0ce-493a-8d7e-bd530bf3520e · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.594691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9c74d2d3e71c228bef7cedb2702af7820be9eab5eceba780c1d48fe95d3edcf2

Observation a4d3476e-ea41-4c22-b44f-09024919246c · outbound

This paper cites Calibrating zero-shot cross-lingual (un-) structured predictions.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibrating zero-shot cross-lingual (un-) structured predictions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.609903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b2735877a1e69e066ed54bb921b425bf0ff3b44a962a090975085186066d04a3

Observation 272d0462-2136-49b8-8009-f159c4627da5 · outbound

This paper cites Addressing the Binning Problem in Calibration Assessment through Scalar Annotations.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Addressing the Binning Problem in Calibration Assessment through Scalar Annotations

Reference 29

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.665676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:8985d6be40a99246f24263ab0275c96b92988263fe29f91500cf10e4d400d641

Observation b26b2712-3bb0-462a-9470-2e43a2113618 · outbound

This paper cites Conformal linguistic calibration: Trading-off between factuality and specificity.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Conformal linguistic calibration: Trading-off between factuality and specificity

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.623048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:0f654c32f9dfef3f97f8023e90d984506090022744e223f3f86c5720a35cc3bf

Observation 17ab1cfd-b331-4c58-9cc3-cd6521422b36 · outbound

This paper cites Is that your final answer? test-time scaling improves selective question answering.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Is that your final answer? test-time scaling improves selective question answering

Reference 31

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.724549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:d31ebc3ac1923f552863855e0fe64bb3e067c7f868f4b1c8aea18cd29e06c83d

Observation ce38668a-50ac-4cdb-bc70-758d102a26fa · outbound

This paper cites Why Language Models Hallucinate.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Why Language Models Hallucinate

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:32:40.683342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:3f6ac6606125d8e851f72fbe9f8b2f03f639309872af0c8c4e72e19332933a93

Observation 8ac5acb3-0999-47cf-91ee-67809c2ad9e8 · outbound

This paper cites Scalable best-of-n selection for large language models via self-certainty.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Scalable best-of-n selection for large language models via self-certainty

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.614155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:dd76e65060009f5aa54da0e64c724aa64a41b60c5ed5de5747e9fd4f12a3c5b7

Observation 16f1ae38-4509-4703-ba91-7747fa83ceb7 · outbound

This paper cites Large language models must be taught to know what they don't know.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Large language models must be taught to know what they don't know

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.616558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:8cb3615533173df2929559a5bb4fc72be4673cacd2820027a0a81f48650b5c62

Observation bcee4fd1-bbbe-4aed-b8b1-c3d7b6ecc4da · outbound

This paper cites Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-27T03:05:05.764939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:5223e1b7586fba387fb38c689593381b033cdd47e0a04ed94ceb0a840243fc75

Observation 97afc72d-bee6-4933-96a9-1af084c9ac6d · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:52:02.773759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:bab46aa7f01eeb4da11dbdadc04d76e46bafe437a6a7c0c08ef83f47a0b11f31

Observation 1dee103a-586c-4888-8e3a-421aa1345968 · outbound

This paper cites Semantic entropy probes: Robust and cheap hallucination detection in LLM s.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Semantic entropy probes: Robust and cheap hallucination detection in LLM s

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.637437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:756a1cf3bdbeecd2276ab4b95dfa095c5319f133ca91749bf9a4893a65c5d72b

Observation 09986e8b-bda9-4e4b-a042-03608dbf0347 · outbound

This paper cites Think with moderation: Reasoning models and confidence calibration in the climate domain.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Think with moderation: Reasoning models and confidence calibration in the climate domain

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.576742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:3cc5ed760c02a507db205e3f320de6ba1968007c369d78c36b2bdd289ae5dc14

Observation b1f51ffa-1a68-4259-bf99-b87b2264c5e8 · outbound

This paper cites Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.635129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:6d3882e27f945148dea62e04bf04922a9771a4bdb74beb3bbcbb3fd23cd58ba7

Observation 77afc915-3c1a-4fc3-8d64-d3b92a4b9313 · outbound

This paper cites Taming overconfidence in LLM s: Reward calibration in RLHF.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Taming overconfidence in LLM s: Reward calibration in RLHF

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.579581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:8c85d6d817932e29633b2a73dbb6aeffda533cbf6e4f7f8f1d815547f1208730

Observation 082160ae-529d-4a88-befb-decc08150de8 · outbound

This paper cites Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:41:12.130965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9081755ff65e4cff023286300a38e43d20616f05010640981c3c45c5a43bd331

Observation dd8714b0-c0d3-408b-babd-f7873d1e2c3b · outbound

This paper cites Conftuner: Training large language models to express their confidence verbally.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Conftuner: Training large language models to express their confidence verbally

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.582151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:161b6284cf43642baa68fe4bbdfb8e8c5972a97369f33e26e8dc8d22dc5b4cab

Observation 2d407fa2-085d-4b4d-97f8-0d7402e37f9c · outbound

This paper cites Let's verify step by step.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Let's verify step by step

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.584677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:e6161188c4d07ab2706585b146bb15f4c7d074705867621e28813838f523fe30

Observation c87b0da9-7f0d-400a-9f90-cd8ba3d6a707 · outbound

This paper cites C 2gspg: Confidence- calibrated group sequence policy gradient towards self- aware reasoning.arXiv preprint arXiv:2509.23129.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning C 2gspg: Confidence- calibrated group sequence policy gradient towards self- aware reasoning.arXiv preprint arXiv:2509.23129

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.314348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:d7e7e3cdf72ee3914137f2ba4a498573c7214e49a6ad0ba29f59ded80c87aa5b

Observation 3058fae0-faa6-4bc3-b902-88b091090a92 · outbound

This paper cites C\ 2\ GSPG : Confidence-calibrated group sequence policy gradient towards self-aware reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning C\ 2\ GSPG : Confidence-calibrated group sequence policy gradient towards self-aware reasoning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.618713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:8efcb51ca9cee7500beffa16b10084e362e57641e57e516000fe8ec87cfc200e

Observation e3f73d66-006e-4162-b625-92efe674a175 · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.620953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:44d58641b2065e79c33ebaaefef7770c44257fb368b2da788ee23d3dc9093ca8

Observation 94d6c7db-a1ad-42e8-a479-bf9da999bdf4 · outbound

This paper cites Wong, Lidia S.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Wong, Lidia S

Reference 47

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.702470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:ee512a8bd146eccbe29df5caeaf21c03a1f99f0dcc291dea53d3af6a4c17cfd1

Observation f0c7e4a7-4d68-46ec-bdaa-445d282696b7 · outbound

This paper cites Uncertainty quantification and confidence calibration in large language models: A survey.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Uncertainty quantification and confidence calibration in large language models: A survey

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-08T22:34:28.695062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:f2931d0a16171582d90e90c7fee8bdb4f9d33df8d8b38baf6b9e675cca693a06

Observation c7a4a1e5-b2da-4a83-acb3-e32c4c20add6 · outbound

This paper cites Your pre-trained LLM is secretly an unsupervised confidence calibrator.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Your pre-trained LLM is secretly an unsupervised confidence calibrator

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.571118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:188fc3d7a1f3629a405d437d1c2e1d97fd49560a86ea901a738091aa6a3d8c74

Observation b11ff9b4-14cd-4bf4-af4a-9dc9f798cb6b · outbound

This paper cites Improve mathematical reasoning in language models with automated process supervision, 2025 b.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Improve mathematical reasoning in language models with automated process supervision, 2025 b

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.573802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:0fea6d1b770e033068f5321ac1267cda62a5b2d62ba956abc4667b0f482b4c37

Observation 9cce0c0d-49b6-4f68-acbd-e568ed75d00c · outbound

This paper cites Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.120451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:1958e2d0b57e7242524331755590102dd9ce93e6118efed5075b803fbb3b2b61

Observation 56b1c3ce-564e-4dd5-8e15-9a1589a609e0 · outbound

This paper cites Reasoning about uncertainty: Do reasoning models know when they don’t know? In Findings of the Association for Computational Linguistics: EACL 2026, pp.\ 3408--3458.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Reasoning about uncertainty: Do reasoning models know when they don’t know? In Findings of the Association for Computational Linguistics: EACL 2026, pp.\ 3408--3458

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.587369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:46d207317b13b1e2670b6bcf365bb7232926e502897e17717f7adf374bbca95d

Observation 042d080c-1e64-46f7-b973-e6b939b4385f · outbound

This paper cites Closing the confidence-faithfulness gap in large language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Closing the confidence-faithfulness gap in large language models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.139298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:cd427d23b9dfd63a86f23601cb648c8ab42dc7baabd3a57317f06a6b09081624

Observation 6e9acc8b-3d9e-49ec-9a54-1c0f3714c54c · outbound

This paper cites S1: Simple test-time scaling.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning S1: Simple test-time scaling

Reference 54

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.698725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:ca2e3d3206348df86f3187bfc5015fa0f4bd6d799a3b4d50ab8c214e62b6b7a6

Observation afc0299a-d722-4046-8cff-7c92ce147dc4 · outbound

This paper cites When do LLM s need retrieval augmentation? mitigating LLM s' overconfidence helps retrieval augmentation.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning When do LLM s need retrieval augmentation? mitigating LLM s' overconfidence helps retrieval augmentation

Reference 55

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.721488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:f8bc8941a0f8edc5f4849745618d2d3c44bb4b83fc534181d8829c562f9204c7

Observation 3f7fcfc4-3b30-4d0c-aa81-12789e3a1092 · outbound

This paper cites OpenAI o1 System Card.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning OpenAI o1 System Card

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.150846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:398799a59667300bd0ea05c4b652c5a533471b0214596c2e2a114ad63b2d3ff6

Observation fc5695f9-3798-490d-b24a-2275385ab356 · outbound

This paper cites Obtaining well calibrated probabilities using bayesian binning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Obtaining well calibrated probabilities using bayesian binning

Reference 57

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.650353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:c8011d95b78d66501fb9f0e5d7138eccc47827bb6cb53036cd160f1e39e2f279

Observation 836030a4-e0c3-4e41-a2fd-dfa71c0c7075 · outbound

This paper cites Optimizing anytime reasoning via budget relative policy optimization.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Optimizing anytime reasoning via budget relative policy optimization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.568389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:7e94876d5e0e36c13f68c49ea3f800862e833586421c611339395f31744ca87d

Observation f000ed63-aa98-4381-b60f-c7546516f89d · outbound

This paper cites Demystifying reasoning dynamics with mutual information: Thinking tokens are information peaks in LLM reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Demystifying reasoning dynamics with mutual information: Thinking tokens are information peaks in LLM reasoning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.612001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:3548f1ea8a32e6d6128cb127b1f61a82f43a5280402ffbbecec00400b2731b2d

Observation 63f45eef-7739-40cf-ae97-8b8b72d513f3 · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.639483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b201f456436134287f89c843d16b28a42307db6515926e2f8fe57ee320b51719

Observation dfd38fd5-50d0-4a01-b2cc-a2b5b3009daa · outbound

This paper cites Jaakkola, and Regina Barzilay.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Jaakkola, and Regina Barzilay

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.667772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:445b42619464d8f83565926502803df2d43e44156e3250c792fec3c6233a05b4

Observation 27357831-5c28-49cc-9e59-230ea2d037d3 · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.669826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:214a67209de690249ba2ce4d9926c795d5461f235be713c7dae9d01be0109294

Observation aa1f0fe6-6909-4013-956e-2eb5adae305f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Proximal Policy Optimization Algorithms

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.187773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9e2e7b27cb69f929151a3890b0f2de4830cdf6c00cf355a169523d2b82c97847

Observation b2eedcdf-8de0-4752-a704-3c85b4f9af32 · outbound

This paper cites Rewarding progress: Scaling automated process verifiers for LLM reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Rewarding progress: Scaling automated process verifiers for LLM reasoning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.672070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:5a70875686af59e239478386f783724368b0673bd36a412b2d0afced7df6bcb0

Observation 2607e0b4-fb40-4f1a-995e-5357094e4e07 · outbound

This paper cites A tutorial on conformal prediction.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning A tutorial on conformal prediction

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.674355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:c886fcbfa03e74b289b46ddebba811f372395bf4a234b2111eb0e0040a610078

Observation ce80752b-fd38-483a-8c1c-b25113cae32e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.211505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:7eb547574bec612b5892865377b628acca1ec8c5f674cf345b406fbae20d8c2e

Observation d8352934-9bbe-4f7d-9cee-2fec2acacbbd · outbound

This paper cites Wornell, and Soumya Ghosh.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Wornell, and Soumya Ghosh

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.665602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:d76ca5a520d963ac12c60db3d22a51cef32b812a462a8db704a356906e79ec32

Observation 705a36ed-57d3-4547-8498-a3a00013c4c4 · outbound

This paper cites Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.676674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:a3d728cdb298a214b4fd8bb439a459b8d6657fdcc2a3472948e999841469b9e5

Observation 20160f7c-1ef9-4585-bc4a-6178ee6e8d61 · outbound

This paper cites Alison Noble.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Alison Noble

Reference 69

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.681327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:206c4957d146657d8e29fbc97840498cbbf8406039742fb64db81f1a2d7b428a

Observation e499919b-990c-4cd5-a36b-7b3259358bef · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.330.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning doi: 10.18653/v1/2023.emnlp-main.330

Reference 70

Resolution
metadata mismatch
doi, observed 2026-05-08T22:34:28.689505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:21286bd2520c7050ad2fa448fc3634ae872c4fceb7790d4749ccddfef98b780f

Observation 3d18eb14-21ed-43f4-a7a2-9535c6f8dc6f · outbound

This paper cites Calibrating Verbalized Probabilities for Large Language Models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibrating Verbalized Probabilities for Large Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.166451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:4ab5c405416ccdb8183bd7ed9f93e32cc56b4cc505f7abb20e18105d30e96131

Observation f29053d1-b153-4b8f-9b29-2e95abf12c62 · outbound

This paper cites Always tell me the odds: Fine-grained conditional probability estimation.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Always tell me the odds: Fine-grained conditional probability estimation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.678770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:e5bec68db0e04aaedf7b53fa79dd523140f748a6a6fe59648eddd104c9c1e19d

Observation 70421b68-46a3-48ac-a7a3-2a368cf2f1de · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:34:16.071337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:c833f58d6e8d157677a2aab66038eb7b95992cb71b3b4860eed1777568b3ca88

Observation 00c0c9c8-38b9-4b76-bea2-d5827b685676 · outbound

This paper cites Calibrating verbalized confidence with self-generated distractors.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibrating verbalized confidence with self-generated distractors

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.680939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:d98517dce91f5ff214a05b80848af7951bcd0ee27314237b8e851591d11ff22c

Observation 836ac16e-1846-40c7-85a1-89adc4d56b14 · outbound

This paper cites Conformal Thinking: Risk Control for Reasoning on a Compute Budget.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Conformal Thinking: Risk Control for Reasoning on a Compute Budget

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:43:11.628743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:3ecd38536dcfa7d46c4040c8695378d22f1e871adaf62c904f57a80b7ba45959

Observation 9cc0b33f-ee76-43b3-a0f3-b0f029817e6a · outbound

This paper cites C on U : Conformal uncertainty in large language models with correctness coverage guarantees.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning C on U : Conformal uncertainty in large language models with correctness coverage guarantees

Reference 76

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.642387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:e5addf3b570c560d7605b65a6a6097016bebda058fc85d2bbcc18c8277d53247

Observation 61e175ff-4b2e-480e-a8d2-7c638f20682a · outbound

This paper cites Thought calibration: Efficient and confident test-time scaling.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Thought calibration: Efficient and confident test-time scaling

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.663418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:44d1d1bbd96231f6055ac70e58128685f84ddcd7d711d1abe0c067184bbfea78

Observation 785f3ed1-e96b-4b08-a9f1-d00e2cc542af · outbound

This paper cites Can LLM s express their uncertainty? an empirical evaluation of confidence elicitation in LLM s.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Can LLM s express their uncertainty? an empirical evaluation of confidence elicitation in LLM s

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.658902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:89d3f7ef951c54e808b45032438cafebf03a4ecde7c8167d6e965b1b00f877b8

Observation 95aa9c4f-3655-48ee-9ff5-28c2c7d55180 · outbound

This paper cites Beyond correctness: Harmonizing process and outcome rewards through rl training.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Beyond correctness: Harmonizing process and outcome rewards through rl training

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.656468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:f98da4f07d33faa541a112a5b670fa7071625331581160785f0752ce91b7ec70

Observation 166d5225-c877-4956-a144-ed5e3ffc80db · outbound

This paper cites LLM Probability Concentration: How Alignment Shrinks the Generative Horizon.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning LLM Probability Concentration: How Alignment Shrinks the Generative Horizon

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.261594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:36ca0ac543e2cfc2736acc9dc50f60fc3a36c97f9681a91817b580a08e00f91d

Observation 45eb69cf-7a3a-41b9-8ba4-87f4241e0d73 · outbound

This paper cites Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?

Reference 81

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.655110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:f9b5fac140598a4a128781a01ad2ec7bff38462ed7cdd9d77ea4567aa80dd609

Observation 4bcc19fd-c171-4037-ac66-9ee1108858cb · outbound

This paper cites OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.177273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:6f8b132344a2d92fcd025a908c0672de573f521ef14340e80bb22e8f75d6b53e

Observation b7dca2aa-de39-4794-9998-c7cb861d1a73 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.159349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:ca30f616426112671ee704e3c4e00e1ff6e5e66e65c9876d9392b661c52f3aa5

Observation d8c2b988-e882-4782-9176-d703685d228b · outbound

This paper cites Reasoning models know when they re right: Probing hidden states for self-verification.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Reasoning models know when they re right: Probing hidden states for self-verification

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.660953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:4f7a9e9bc9b744f7e3b3b3dfc7889ac2d6a8edc6ee861e3e799b5ffbea15fe0a

Observation eae8f818-50af-405d-a294-635d846bbf80 · outbound

This paper cites GRPO - LEAD : A difficulty-aware reinforcement learning approach for concise mathematical reasoning in language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning GRPO - LEAD : A difficulty-aware reinforcement learning approach for concise mathematical reasoning in language models

Reference 85

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.685372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:8ed68c975a4fc7ce752f90c9678ecab44653800000de0ca368ce9fabd9136b98

Observation ac930bb2-b14d-4541-b5b8-aea21ce3fd93 · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.648955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:707301ab1649bd6048e20f491ff536731832f451464d4cd78f20c161d8d8f206

Observation a154b8bd-8a7b-44ad-9d78-cee204d0317e · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:43:43.298530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:e9c1f4cba9ae7d467dec93e250dab6142aa61a42ffea33a0610832f576e5236b

Observation 8ba7f292-78cd-4085-b09f-68b8f4eebc34 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:54:31.187112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:5695b66ff19679e76ec7b04330b47b9e0f8e57560bcc0a7c6b85d55fd3e92bab

Observation 45abc3c0-fcdc-42c2-8ebc-8520eb949ac3 · outbound

This paper cites Calibrate before use: Improving few-shot performance of language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibrate before use: Improving few-shot performance of language models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.643897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:7be4250d97efb828840d569c22187fa0eba95a5ccffd557966034b5a47242e53

Observation 5fdcce4e-9326-46b4-bab9-8b0875ca0526 · outbound

This paper cites A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.237411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:4e55b269490e06aff24740cec7c19494a8ca60acfb0aae9245afdfb5bea6638e

Observation b1227ec7-72df-4e76-9641-87207889a73f · outbound

This paper cites write newline.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning write newline

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.646716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:863bef8abb410ce21fd475d5ded9502fe35c548f3aaf62884717cdf27ecb02c0

Observation bd5e107c-3794-471b-8f10-38d9c9321a18 · outbound

This paper cites @esa (Ref.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning @esa (Ref

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.651575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:18645eb022dad121217f404608939890be80e26ef1f8abe02ec7b3ec67172629

Observation e5932a13-9554-468e-b6c1-c8f485b73438 · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.654239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b4915be6cb6ac7fbc038b168b26e6e7087c59e49847cecd75ecb6e9992532650

Observation a834bab9-d393-4e5d-bafc-03898712ac4d · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.641516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:07f2fe4700c4f99f314b631a7d0db6bba4fa70d9a1a85520c7bcd650c226bb92

Pith citing papers

Observation bdca9a63-aad3-4285-995c-87118c8e3c97 · inbound

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute cites this paper.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Process Supervision of Confidence Margin for Calibrated LLM Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.331963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.331963Z digest=sha256:c2959c9a96c2dada890fed2233b982d62f5d350c5a9bfb96d9bca7a9b5e1ecf4