Pith. sign in

Paper Citation Record · LEDGER

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

As of 5 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 23 inbound Pith citation observations for arXiv:2507.16806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16806 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:02:40.647686Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:16:41.377531Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b7d3200-c7e3-4215-9e60-38b7914c7068 · outbound

This paper cites On Verbalized Confidence Scores for LLMs.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty On Verbalized Confidence Scores for LLMs

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T00:04:27.005902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:b80e04caa6d6dc56a6210246f095e07268887bf91bb851e8e9bc5005c76ca964

Observation 927f5937-e63b-4851-ad4d-b1db1700c5f5 · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.224097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:bf30fc7c7c391fea788006f8998cac1ea821d4f1aa958383adadae14992b7413

Observation 87ece4b9-2b42-4cb6-b3e8-6e006056ffc7 · outbound

This paper cites Examples.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Examples

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.230608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:bd185f727b0e7164179ed8c98e5de7bdba3ca56002d94e5345aece5043ecab0f

Observation a260b798-1c21-4ad6-b451-ce8f3647a295 · outbound

This paper cites We slightly modify the dataset and remove 2 non-relevant paragraphs from each question.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty We slightly modify the dataset and remove 2 non-relevant paragraphs from each question

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.234085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:466360759f6df186edc92968b2b5e3d45007b61fbac934a90d54fd81fec582fb

Observation 733a1dde-71de-4af8-8061-992413258c65 · outbound

This paper cites We measure correctness using exact-match.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty We measure correctness using exact-match

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.194680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:33ffd8278b5438b42277cefa3408d7d10a354e761c11abf6ba4cae553bc61e22

Observation 74478e29-8f43-47c8-a2c1-c829b18a38f4 · outbound

This paper cites We use the no-context split to purely test factual accuracy.We evaluate using LLM-as-a-judge.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty We use the no-context split to purely test factual accuracy.We evaluate using LLM-as-a-judge

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.211984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:b693ebe73b0946711d432c0c1a2b028112ee3737aaf7d363f2e58b869db175af

Observation 04c311f4-82ba-42e1-9a3f-891bd1cbea95 · outbound

This paper cites We evaluate using LLM-as-a-judge.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty We evaluate using LLM-as-a-judge

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.216034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:04d88d09c4fc6264443f093e742666ca421c490b91473b13ac37d58d7e278c16

Observation 9353cf4d-a698-42b9-8793-c7a50f2c785c · outbound

This paper cites We evaluate usingmath-verify, a mathematical expression evaluation system released by huggingface.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty We evaluate usingmath-verify, a mathematical expression evaluation system released by huggingface

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.203529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:2d295b33c70c5168afed60bf8f069daeb47dee8283aa59e47773aef4657abe4e

Observation 39c8870b-725a-42aa-944d-c8f4e9e2d74e · outbound

This paper cites We evaluate using math-verify.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty We evaluate using math-verify

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.175155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:b2b0dd03cc7a15d4663cede2589eb06f3dc01f8cdec72207df51a0158b6e20ef

Observation 2829eb6c-8b03-4eb9-b858-bbf1f607e683 · outbound

This paper cites We evaluate using math-verify.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty We evaluate using math-verify

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.223828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:0e1bb6afd29557daf48ccc4531777931d8438f160a475c9f851151dac51bfd5c

Observation 9a002f2f-20d0-4a75-af44-e001dd7827d2 · outbound

This paper cites We evaluate using LLM-as-a-judge.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty We evaluate using LLM-as-a-judge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.192871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:18dc332a6c5ae8e1887e3452a54aac44eaafd05470afefd99543c0294415ac42

Observation 6148607f-ee27-445c-9e8e-5273bf446419 · outbound

This paper cites We evaluate using LLM- as-a-judge.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty We evaluate using LLM- as-a-judge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.196433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:d9b1b454ed44fc55d4d4b277cd0d9238de1fead08a1b998c0fb65057caa342d8

Observation a6cc23ba-9e57-448b-9171-731d716b0fe0 · outbound

This paper cites They are evaluated with the same system prompts they are trained on.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty They are evaluated with the same system prompts they are trained on

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.198707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:43b77de3b91866393218f73f9796a651f669b34a29e1d9265f551f8bbe78fb95

Observation 1baab14e-7f37-469b-8c05-c5977fccd2cc · outbound

This paper cites It is evaluated with the same system prompt and we extract their answer from the <answer> tag.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty It is evaluated with the same system prompt and we extract their answer from the <answer> tag

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.236074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:b0face6946f73920a5a91a5331d0772807f763469b8f5a51386523961bf801fc

Observation e8790adc-3959-49af-8ec1-3c92329fef0a · outbound

This paper cites These methods thus use RLVR model as a generator and their reported accuracies in the result tables are equal.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty These methods thus use RLVR model as a generator and their reported accuracies in the result tables are equal

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.203831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:7630b6e2046d9d21aaa78b66a82cbf4481a63e35e0f8d828279c089e2baed4fe

Observation e5a7c08c-3abe-49dd-8737-8b2f20eeec97 · outbound

This paper cites In case no valid confidence can be extracted, we append ”Thinking time ended.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty In case no valid confidence can be extracted, we append ”Thinking time ended

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.231975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:52a1aec0e1de8b96307482461dae23f4fec0e1f22fb89700e8a6ff4a0f61b3bc

Observation e82acdb1-ba53-4d54-8068-740376fcb8d9 · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.205249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:ad6119890f59ba133f5b8a7c7518bc89034a89119fa312225eeb5181546cda60

Observation c86e7345-e35d-4383-8557-2661b07c5196 · outbound

This paper cites In these cases, It is also okay to have only a small number of uncertainties and then explicitly say that I am unable to spot more uncertainties.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty In these cases, It is also okay to have only a small number of uncertainties and then explicitly say that I am unable to spot more uncertainties

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.222029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:23bc1e32aff07d8eeaadcab88fdccab15a079ef80e18f83640e4a45f7f776668

Observation 67e414af-44aa-493d-82d1-2b9ae9c4fdec · outbound

This paper cites For example, uncertainties may arise from ambiguities in the question, or from the application of a particular lemma/proof.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty For example, uncertainties may arise from ambiguities in the question, or from the application of a particular lemma/proof

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.219975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:a655c6c90a32e974bfc43f8f42346edda23c7e5736ad519e8aa7b17356a4a108

Observation 50f91763-b8ea-4a38-8fe2-de0ba6de5c6b · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.241283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:e8fcbe9ec70ac476a79fa3a73ee4c4cf13beb76eecb63bba2ce578088f4ffde3

Observation a503bb22-5eff-4f77-adc8-663850b21d92 · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.219790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:bd7725ef2d3c1522f4dcd99d7dfd39467b4c379ed5f97fda163d075bbe87b323

Observation 1a2e46e7-ff74-486d-bc34-b0f6ab25489d · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.238785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:11ff6822f1d4a07a255f9fa74e1a7a816821828d21c9885fda251aa37e06b1fe

Observation c3db0f85-cc2e-4191-be37-d3ba119522fe · outbound

This paper cites The user asks a question, and the Assistant solves it.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty The user asks a question, and the Assistant solves it

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.215250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:b830fb9f2d6d02b246ff68fd7c258f440ffc9396a5b832c5a5539b536d211ce4

Observation 9280a69c-f8b5-417d-a571-4e62b7671c6b · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.217502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:29fff26ca18b1a00ec69ee3b31e75564b0fef030fffcc5a96072dd579660b01b

Observation f03f303b-2214-4bff-a1a0-ffd9802c3036 · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.229613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:0bf09b1ae5eab46f7c6b2dc3e502711c345ac4d336d9bf958d94cbe279236d94

Observation 3d049dbe-74ba-43c0-b751-ed597040aa68 · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.240109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:39ee1fd085796fb08b082e26a0e4d7001f854f2c9906b5cfc6d87ed91557f634

Observation 6352271e-cb71-4c5e-af23-c18ad64dcf1d · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.236232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:2b578f6d15325a7149c0c296cf1765dc964ff2f2fedf45cb06be386567e33894

Observation 48b4a9c2-a80d-4680-884d-9402647e7856 · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.233763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:3c58e31de532515e67b2d292586d0589dd123292806f0424241a9832b349d469

Observation 838f3ddd-94da-4d7e-90eb-ab6dbecb6e56 · outbound

This paper cites 21 D R ESULTS D.1 M ODELS TRAINED ON HOTPOT QA Method SimpleQA Trivia Acc.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty 21 D R ESULTS D.1 M ODELS TRAINED ON HOTPOT QA Method SimpleQA Trivia Acc

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.238266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:fcf0f60012f1db9e704212895bbdf79b3c1892e235692120b6c818f94435e5b6

Observation d55006d7-f05f-4a62-a4d7-9a304b79927c · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.225703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:c5a420cb4989312ffe1eaa4afa0e808c6fb55c516d63e8bba73e86ab3a4da4f8

Observation 3c68983f-6e17-4ccc-a2a3-11e3c4dfbce4 · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.227713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:60132d9b841b059d9e36caee2b24950abaaee77c3404d399252e6a94d0d41d49

Observation de9ab703-b2d2-4b71-a2bd-8b476862419a · outbound

This paper cites - Bella and Chris watched 2 more movies only with each other.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty - Bella and Chris watched 2 more movies only with each other

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.191972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:e48d06e2e3c3e14bbe5b7614f577d1761abe3f0e13986c49251ae28f95ad8573

Observation 53bffd5f-b410-4a8d-bef1-f7a3ce348eaa · outbound

This paper cites - Bella and Chris watching 2 movies only with each other are already subtracted when we subtracted the 5 movies watched together.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty - Bella and Chris watching 2 movies only with each other are already subtracted when we subtracted the 5 movies watched together

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:04:27.184442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:62ef4875bbec8215e38b93eb9040aaf5217fde3c9c8ed88978dc94fc70739023

Observation e3593f54-4b33-4ef6-8eeb-ed207a3bd368 · outbound

This paper cites an unresolved cited work.

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-22T00:04:27.228379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:02:45.493336Z digest=sha256:36fa52836d4cc8e8a81ca2d5d0bccd9af03437e9d2ae29926765f1e6d63c38db

Pith citing papers

Observation 1bd4210d-d2c6-447e-96eb-f8511bb5a4f0 · inbound

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning cites this paper.

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:02:40.647686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:02:40.647686Z digest=sha256:60d7718c73a1d9cb285778f6c1d174ecd0402312acb97a2e93e5f2b16f10a760

Observation e921c058-71ed-4487-97e0-ed7f055e6405 · inbound

Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency cites this paper.

Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T15:51:09.563645Z digest=sha256:a422b613a1419abc207a155b3964b77319f7ddb3d019f07840d04411e3c228f0

Observation 67e9d420-fd80-48b9-9d76-73a7459d3000 · inbound

NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems cites this paper.

NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T10:14:01.961040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:14:01.961040Z digest=sha256:65d7f7d0d02c9f572befebb8bdde3a19f26afdf8aeacc728e72a6bd281d42946

Observation 8f2442c1-c3df-4f79-a4df-940c482e1488 · inbound

Uncertainty-aware Generative Recommendation cites this paper.

Uncertainty-aware Generative Recommendation Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T00:05:07.481901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:05:07.481901Z digest=sha256:bd6e8d9afe7d40bba5bcba0498f265c73dd532c3fcdfb4b1e11876d0bf6eb7f5

Observation b1344618-1b6b-4c0c-9a19-baf814829d52 · inbound

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards cites this paper.

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T13:49:11.758336Z digest=sha256:b85b743f0f470a9b88b8a448f5c903bc0cef65d39bedeaa926080e425ef43201

Observation 69f1f055-77d2-4d4f-b1aa-e3eb760cc886 · inbound

Calibration-Aware Policy Optimization for Reasoning LLMs cites this paper.

Calibration-Aware Policy Optimization for Reasoning LLMs Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T15:40:24.396851Z digest=sha256:423745c28e63aef3eb41f6c55bd92368e737440ca9dfa36f8d5fe4b4ea3fc21e

Observation 41bd6740-40ce-4cfd-9295-9c9cb651bf33 · inbound

Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation cites this paper.

Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 155

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T03:13:35.541936Z digest=sha256:962e8a5ff395dcaa9982eaa364da010bd62fda0a8223ca2e248a5931c7c82201

Observation 1f5cd2f5-848a-4c68-95ff-c5e76c3d4a40 · inbound

Do Not Imitate, Reinforce: Iterative Classification via Belief Refinement cites this paper.

Do Not Imitate, Reinforce: Iterative Classification via Belief Refinement Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T21:48:15.641442Z digest=sha256:1b80fbc3d8d1a9c2d734332b5ca8992326ceb1670f714ff46004b7fade5a6180

Observation a0df8c10-e56c-4d0c-9234-65c0e2dc5bd2 · inbound

Process Supervision of Confidence Margin for Calibrated LLM Reasoning cites this paper.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:bd40c775e0151921bb826fdedf190b09ad5a3c0e4220aaab4d37f484068aed03

Observation f19aea86-4f9f-454d-9542-e32534dcbe8d · inbound

AI Observability for Large Language Model Systems: A Multi-Layer Analysis of Monitoring Approaches from Confidence Calibration to Infrastructure Tracing cites this paper.

AI Observability for Large Language Model Systems: A Multi-Layer Analysis of Monitoring Approaches from Confidence Calibration to Infrastructure Tracing Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-07T15:40:05.520653Z digest=sha256:f945ed42fae54e826e28a246baf6b0ffca13bb36650be9b3139de6c552b8167c

Observation 6d88a72d-4347-4b89-a1bc-1fe8da7854c6 · inbound

Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport cites this paper.

Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:17:01.225189Z digest=sha256:3a43ffc1cb0d3053c4876c1a39a578c0bf039c3108f27e6a39d193b6ef68527b

Observation c3fdc429-6b12-4ac1-94d3-a7fda9fc5e3e · inbound

Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport cites this paper.

Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T06:27:19.650539Z digest=sha256:d9df9b300ec86be493ac512e40bd6158497659c29d3adbdbc4bac260b7945904

Observation 06c97919-bb19-413a-b386-62b0c6bad39a · inbound

Not only where, But when: Temporal Scheduling for RLVR cites this paper.

Not only where, But when: Temporal Scheduling for RLVR Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:44:01.846492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:34:44.791173Z digest=sha256:b99e0a1d57626a179773b9a235d54b77cc4e0c842ba50bb7d9283e049adfa116

Observation 989d4b74-f5ea-47d1-9b21-4ea02128f614 · inbound

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information cites this paper.

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:43:25.798054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T12:37:16.138816Z digest=sha256:21eba25d68129a9c3b16ffb9f395b8b4c4e6ba61fd4ba0b8bc49b0a87ac3719b

Observation 621db6b5-1b2a-4b7e-a8e4-9b8b9beb9821 · inbound

GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization cites this paper.

GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:12:46.894665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T23:09:57.297404Z digest=sha256:556c7ebce3bfccbad26c23dafec9d6f621f00a48f48bcb44d6f9ae852c93ead5

Observation 297b0023-3a7d-453c-bc1a-5b71026d70f3 · inbound

Enhancing LLM Metacognition via Cognitive Pairwise Training cites this paper.

Enhancing LLM Metacognition via Cognitive Pairwise Training Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:02:33.868887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T19:01:18.153145Z digest=sha256:b941d45248958ad9a08220194fbb23f23fccc9a01d5fdacf5bea71f0aa575f81

Observation df57e23f-26b7-46f8-825e-bc09081797da · inbound

CALIBER: Calibrating Confidence Before and After Reasoning in Language Models cites this paper.

CALIBER: Calibrating Confidence Before and After Reasoning in Language Models Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:39:58.407338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T00:19:10.687624Z digest=sha256:530c16b9a8a6c3116fb02a7166b80dd14e4f9c815327b3a045554b316cc82e54

Observation 13477cd2-2cb6-4fcb-9b6e-bbdf4e762c6c · inbound

Verifiable Rewards for Calibrated Probabilistic Forecasting cites this paper.

Verifiable Rewards for Calibrated Probabilistic Forecasting Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T19:47:18.734131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T19:45:23.721172Z digest=sha256:f39b6cfd118e81bae45000f60327c2191db8f0e5c68da65e2b4185f9e85cf57e

Observation e13225bf-52b7-4951-86b6-efb0cd9277f5 · inbound

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness cites this paper.

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T01:16:41.379275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T01:09:53.133146Z digest=sha256:82fe09e697aca096178e4e55bc3351e365affd40d4e5f1e5b05f3ae5090ceb71

Observation f70029e2-a09f-404e-afdb-cc00f2c50e31 · inbound

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards cites this paper.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:1bb3081af8aa39c3e146865700d7d0cc9d2107a943159b3703cbc42005309dca

Observation 1b9d681c-d250-4dab-8d7f-ef68195f2211 · inbound

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR cites this paper.

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T07:36:47.336670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:36:47.336670Z digest=sha256:45273ec961b1fb581f8645fcb397be11d4e0671cfd30a50f8edc898208381ce2

Observation b810c039-a49e-4cf9-b93d-08407d0aa717 · inbound

Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models cites this paper.

Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:39:58.932821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:39:58.932821Z digest=sha256:87557a09cce44902915123dab1999de59674383a721b55c5473003314a24b271

Observation 57826263-223c-4742-be08-037d3faa7ccf · inbound

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning cites this paper.

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:53:26.000469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:53:26.000469Z digest=sha256:d6134550bbfc983de06ae297f77d47ddc700d6d692487188aface7b631ec9e5e