Pith. sign in

Paper Citation Record · LEDGER

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

As of 13 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 11 inbound Pith citation observations for arXiv:2505.11737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11737 v4

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:30:27.620057Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:07:50.378958Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact39
  • verified fuzzy19
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83d9ff01-8294-471e-bfbf-cc7e039f0208 · outbound

This paper cites Uncertainty quantification in fine-tuned LLMs using LoRA ensembles.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Uncertainty quantification in fine-tuned LLMs using LoRA ensembles

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.433213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:b4946189894c674f58ab5976b9792c9fee275e305fa9d0d21885e21cdd1488c4

Observation ff6f7838-a475-436e-870d-823d10ca3319 · outbound

This paper cites Area under the precision-recall curve: point estimates and confidence intervals.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Area under the precision-recall curve: point estimates and confidence intervals

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.161814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:368af0d1e4653cfe09c5e409426f7791857e8e36fdb7394c67749f7997ef5244

Observation 47a895d9-b95b-4eca-88bf-993bbd877579 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.259150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:b3fca3b14e0b72d12df15cc35cbc27aaf4a25f051538649a0850a84e8026fdde

Observation 26523ef5-fa51-4265-96f9-407490d37d4d · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.155325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:d575fc97a9e92df606ab8af8d9f850cf8f2e2f28f1f7f4ab7c07b06610478d16

Observation d2fac832-5186-4685-abd2-de240ac6bc79 · outbound

This paper cites INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.460746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:1c6d4c00eea74fbd57e1a965fd1bbd24a8495cdad0489bfc9013f3537f47d952

Observation 80a6d0b3-899b-4429-b07b-e874ba058466 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.477179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:22ff3774b8b3e9d1c7ab9a151b7ea3488fe8169dd870ef36fcbaa0daff4e3b5f

Observation 6d084ad1-c0e9-4c36-9c67-03c22ed7871f · outbound

This paper cites Understanding the uncertainty of llm explanations: A perspective based on reasoning topology.arXiv preprint arXiv:2502.17026.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Understanding the uncertainty of llm explanations: A perspective based on reasoning topology.arXiv preprint arXiv:2502.17026

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.400562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:d90c829558f0ee7c8936253310d052c8040f9a439e474bc8d24bfc1b6500bfa4

Observation 4f7834ef-d32e-485b-a25c-d4bd8304e23a · outbound

This paper cites Rainproof: An umbrella to shield text generator from out-of-distribution data.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Rainproof: An umbrella to shield text generator from out-of-distribution data

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.158559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:f26916b2a05f9e54e8af2bc42e363da1bf8f6da049a8ca31e246a1ec659a84c1

Observation 1601aa2b-3880-42d3-92ce-f2b2d5e50ec7 · outbound

This paper cites From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.443500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:25db7d15d4eeed53de52615061fa04dbac87e8ef0b512aa268efac5367b01407

Observation 31ce4579-79d7-4123-aeb7-b8fc4b580a1d · outbound

This paper cites Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.168647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:44c4e35763cf2cbc58a3012f02582dbfed60c57120d97ab636ac88e4b4edf7ff

Observation 77fd559c-ebac-4227-8ae2-39a265d0ffa0 · outbound

This paper cites Beam Search Strategies for Neural Machine Translation.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Beam Search Strategies for Neural Machine Translation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.296491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:c18e5cc5b5022f9ff4b23e27aca0b742ad14501b4b0b09a36a06b55aa8cb9d75

Observation 799f2116-42b6-41c7-b2dc-1de29f97db03 · outbound

This paper cites Deep Think with Confidence.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Deep Think with Confidence

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.412573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:e0b5c1d84b2269c77fbea9f57abbadc6069a3362b1eea081c61f9211ed97a714

Observation e1dc3c53-8cfc-41a1-ac44-427e284fefa3 · outbound

This paper cites The Llama 3 Herd of Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning The Llama 3 Herd of Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.290908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:103c52cf4abd299425c1c4937a23fa859fdbea0de621ca972c1e7659152bcac4

Observation 0ead29b4-f399-46dc-a0f7-0837fd53f05c · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.382405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:6f93e3b3fd995991fe0382ac697b273ffda2d30241f06771d5b041ef42373ab7

Observation b25019a6-9755-4bac-a1d1-80a3faea5a9b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.438351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:db42b22fbc66b868759120c63dcdd6f2255d24bdeb1a6d289efd7d0b4050b3a4

Observation 939a6910-bfb6-4513-9e36-58088ba7baba · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with Disentangled Attention.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning DeBERTa: Decoding-enhanced BERT with Disentangled Attention

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.233567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:ee9c47ca8a8b1dde67948d3ad150dd2ebdaca3e69412dae1d971042a65d8bf75

Observation f58b672c-9994-404d-8fa5-8e05ec2023c7 · outbound

This paper cites Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.464971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:afee2f58e751abbf720ceb504a7f5624e57e13b4fcc740837fa8055dbca9a25d

Observation ccd5d13c-3716-4f78-b0f8-9df0dc01609e · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T14:01:38.395278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:41d81ed62ebd588a98d701bad08c12c3749ff1330cab576740fc493fe159b2a0

Observation 49065686-f403-4f2f-869d-11f93f3935d5 · outbound

This paper cites Scalable best-of-n selection for large language models via self-certainty.arXiv preprint arXiv:2502.18581.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Scalable best-of-n selection for large language models via self-certainty.arXiv preprint arXiv:2502.18581

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.311847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:6f320dc2123f3befe6f89d0ed2c3c700a00747920293301d5fd8bc40f27066f2

Observation b5385b7a-0475-403e-ae20-a52316227076 · outbound

This paper cites Large Language Models Must Be Taught to Know What They Don't Know.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Large Language Models Must Be Taught to Know What They Don't Know

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.305531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:76fb43880230038bccdf72e9f8cf2583c1b7f8b9057e3a594faa050975b6db44

Observation 3c283511-0872-4d0b-b7f7-47baecbb0d76 · outbound

This paper cites Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.240097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:a1e03d9326449c894f58d3ed3ee67ab81315f7f3177292eb0cf94dd21ccc94de

Observation 16750add-d638-4ade-9d0a-ff98e92db484 · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.269837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:a267132ae6828987d4d82ce08f480cd27962df040b981a4732b5e0ca906c9811

Observation c01aeed3-89df-407a-a5f6-b3553f0f532f · outbound

This paper cites Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.484546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:b032055d53ae79311931b68d0165287d169ad607bbcf2503c8957a93a141e803

Observation 9f88eef8-ca47-48ee-89fd-baf571793d46 · outbound

This paper cites Uncertainty Quantification for In-Context Learning of Large Language Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Uncertainty Quantification for In-Context Learning of Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.496532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:7a6160589f5a8056f6eabb2333c33d5ebb80838624ffe0d21517139ba23be105

Observation 8376aeba-b0c0-4c09-8cd2-0a7d8e295253 · outbound

This paper cites DeLLMa: Decision Making Under Uncertainty with Large Language Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning DeLLMa: Decision Making Under Uncertainty with Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.488616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:1d1928677453bdd796581a469e01805fd44dd0cf938e51c7c184524792b10f4a

Observation 507378e6-1e64-4373-80a7-bf7f9b5d75ac · outbound

This paper cites Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.252394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:3d8e97210b94e50c44556099d30eb798961e6817bb2397ea22d0ead4149682a1

Observation eaa57b63-64b5-43c2-bf7a-bc146a5c072e · outbound

This paper cites SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.425808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:9bfd410f65500ffe7e03b699d396e49790f20dc47163cf9f31e9c112c5ad38f2

Observation 4fb7e880-cc71-45f8-8695-cc2219ce663c · outbound

This paper cites Factscore: Fine-grained atomic evaluation of factual precision in long form text generation.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Factscore: Fine-grained atomic evaluation of factual precision in long form text generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.172146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:0273068e52c31e1bd62f848358fa72af79bfc08589a1c6fceac2514e3575300d

Observation b17040dd-ffcc-4d7d-ad42-39929714d9ed · outbound

This paper cites Correcting Length Bias in Neural Machine Translation.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Correcting Length Bias in Neural Machine Translation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.473684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:ab654f4b14db83ebbde369ef807f1272437f005e78af0a27843d675857be93cf

Observation da84079e-e322-4881-b5e9-fcb073f7d014 · outbound

This paper cites Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.371687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:82b55343bd37796b6de4ef03cc0f30e0b9c96b9475ea72e3c80cd4fcad8bfbcd

Observation 58574e17-6649-4135-a92d-7d30820324f8 · outbound

This paper cites Training-free bayesianization for low-rank adapters of large language models.arXiv preprint arXiv:2412.05723.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Training-free bayesianization for low-rank adapters of large language models.arXiv preprint arXiv:2412.05723

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.364039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:74e97a8d226e7ce8ee596dfe1f4c56a5ba28b3ef6924547d501f5d38bc6f05e1

Observation 66f713f6-39b3-4e29-b8c6-c8f8ad048141 · outbound

This paper cites Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards.arXiv preprint arXiv:2505.24760.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards.arXiv preprint arXiv:2505.24760

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.456437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:efcc6b56adfec1e28897ea9cf66422047bcf5bd4f5f3eaa064384c80fbbd63be

Observation 8e0ff273-9888-466d-9a18-fa231f8bdeb5 · outbound

This paper cites Qwen2 Technical Report.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Qwen2 Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.451746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:d831f408e555f691074a72f8cf2d10f5275efea230c9fa0dc7937582b89603d3

Observation fd36d090-a143-44be-9e7a-b1bc56f09eab · outbound

This paper cites Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.165145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:75e2ae52ec4d6c8c549c554e20fbdf98331b45c49e799cb78cead39769e161f7

Observation 8158be7a-d0e1-4897-bcff-78a2e02e8328 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Solving math word problems with process- and outcome-based feedback

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T14:01:37.827431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:7a03ded71a75fc8e0010eb58bab13e842cf2acb18425a479912e30fd179fe86a

Observation 660d57c1-4417-472d-b3fc-c1aec11ec9db · outbound

This paper cites Mutual information alleviates hallucinations in abstractive summarization.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Mutual information alleviates hallucinations in abstractive summarization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.101131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:7369fc21889829755f9c2feb02139e09c0fb9eb8c10de3f9cebac0da884ddaec

Observation 74268c40-6587-4b1d-bc15-6452a413f700 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Attention is all you need.Advances in neural information processing systems, 30

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.104717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:8129d39cd615daeb55a3db31c957eb9688e06ed9a1ea12da01ee92b7f79a1149

Observation cf779a1a-eba5-4410-95a2-b6dd1dcc4e87 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.469751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:f0cb1e2a004a61ce8259dc9f637eda5579063c6907c7fb83e894717f9247342d

Observation dce899a5-6e22-4480-84f0-2c57615ddc70 · outbound

This paper cites BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.492763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:1faea0a45b28096b504162581f52018e839bee3484b3514c8975919dcc36912f

Observation deb0be3e-7613-4e73-b85d-a0b13a040b50 · outbound

This paper cites Emergent Abilities of Large Language Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Emergent Abilities of Large Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.420484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:0ffd28fa1f178acc7b30775fde0a47879e111d4f48fb489e8e3b30f6a7b4a8e2

Observation 9b9381a0-3fd0-43fd-9a48-d02d4e3a98f7 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.406957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:58c08c101011bcf6145feac1ec202dd6d5a5e50ab45b7399dbfc0e7ba5db9d05

Observation 4d3a630b-e496-423c-8b3f-61f269d95b15 · outbound

This paper cites To Believe or Not to Believe Your LLM.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning To Believe or Not to Believe Your LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.286476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:1055aae81173003a90bee45f76a9336c94c0d076b9459313cceeca2fca756db7

Observation 646bc7e1-f1cc-45f7-9e86-90a6195272b8 · outbound

This paper cites Bayesian Low-rank Adaptation for Large Language Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Bayesian Low-rank Adaptation for Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.448290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:7f9da8fb6231c979e27a63137eb9e0dc4f63c470746f637632c5cbeb2b54f68e

Observation cf372d6d-4432-42aa-b6b4-c7ce5b1dae1a · outbound

This paper cites Uncertainty-Aware Step-wise Verification with Generative Reward Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Uncertainty-Aware Step-wise Verification with Generative Reward Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.277808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:e126fdc84ac38f79d1de8d4444217879d9511a9faef5c95931ac7898fa3ee545

Observation 881969ea-f120-408b-a1e8-83b1435ee914 · outbound

This paper cites CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.220946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:02b869ee69d5387c500f0f2ba372dca79de2b45b841c6a3765f7ed0f4f9ba631

Observation b1c56a53-0093-46a3-b8b2-c47f67319986 · outbound

This paper cites Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.389956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:e9c6d3eab840ea39082960ad17fc7ebdb6c230b08e546b7f9d080464bb2ac0c3

Observation 1e660460-f535-4684-94e5-3f0020baae39 · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.246330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:50d74be921f62eb691952f30d6697c10d5aa78b561eadd5c4350a9c139eaa959

Observation 3f343d31-c1bb-4428-a353-b115c98fe112 · outbound

This paper cites A theoretical study on bridging internal probability and self-consistency for llm reasoning.arXiv preprint arXiv:2510.15444.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning A theoretical study on bridging internal probability and self-consistency for llm reasoning.arXiv preprint arXiv:2510.15444

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.227719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:55ac05a797085b5dd3473fa6027fabda121d8564442b1c532e9d63029a20b956

Observation 000480f8-4ec4-4e95-847c-0a3a0a7a5d60 · outbound

This paper cites Uncertainty-Guided Chain-of-Thought for Code Generation with LLMs.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Uncertainty-Guided Chain-of-Thought for Code Generation with LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:01:38.481134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:347714e35541df72c3fa74f798232c9cc83c4daeb0a56e2464c41e8ba77098eb

Observation fe189900-c2e7-437d-8d0b-0643774c2191 · outbound

This paper cites In Appendix B, we present the full algorithmic description of our method with low-rank weight perturbation.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning In Appendix B, we present the full algorithmic description of our method with low-rank weight perturbation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.144599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:cfe1671be802ef4f4e9b5d835146974a6a9c31b738e28b006d0f7acab69899ae

Observation f9a0ff79-822a-4c3e-b334-cfbc69b9168a · outbound

This paper cites an unresolved cited work.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-22T14:04:54.147881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:613e4538a400a3fdfb11ef84b35550ad01d9479da9a2e88af33f21e195835035

Observation 46dc3c28-f40a-40c8-aa16-00cd06d305a6 · outbound

This paper cites For Aleatoric Uncertainty (AU) and Total Uncertainty (TU) defined in Eqn.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning For Aleatoric Uncertainty (AU) and Total Uncertainty (TU) defined in Eqn

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.151848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:a71c4eaac0ed5785952e8995e752f7729d1ad7fd705797b6bce1facb53e87b09

Observation b24b033b-9725-428d-b593-fa5910346dd7 · outbound

This paper cites 4.2, we apply length normalization to TokUR to mitigate the bias introduced by varying sequence lengths.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning 4.2, we apply length normalization to TokUR to mitigate the bias introduced by varying sequence lengths

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.133057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:a88b99486239c38114bbe75681f6facfef4fcd78b9a76b94d8adf7f779bfad7b

Observation 63471ad4-10df-4091-a559-9bff0834aa25 · outbound

This paper cites True”, normalized by the sum of probabilities of token “True.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning True”, normalized by the sum of probabilities of token “True

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.136630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:6be9476b676482317ec645077833df0ef78246a3499ea8d4ed5c4e5407657ffc

Observation d21adb33-7a2f-4532-9d34-60715271a84d · outbound

This paper cites E.3 TEST-TIMESCALING VIAUNCERTAINTYESTIMATION We provide an additional visualization of the test-time scaling results in Fig.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning E.3 TEST-TIMESCALING VIAUNCERTAINTYESTIMATION We provide an additional visualization of the test-time scaling results in Fig

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.129580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:b4606fa59225fbd6e571f89f3c7073a17fc27ce4c4627d577999994ca330928a

Observation 677ba42e-3582-48f1-a742-ecd97fea6c37 · outbound

This paper cites an unresolved cited work.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-22T14:04:54.119115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:fa004a8112764d4a5eab79b47fd3ee244f4e64d4b710bdfb6dd02c560d6aff18

Observation 713f43e8-ab64-465f-95a5-021523ce6a81 · outbound

This paper cites Building upon this algorithm, we use uncertainty as the score for each particle at each step to guide the model’s generation process.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Building upon this algorithm, we use uncertainty as the score for each particle at each step to guide the model’s generation process

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.122787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:9aa457ea3de9893a4f5b5ae80c6e42fbb48ba61a9623bbc35880320640c63f52

Observation 0124b369-065b-461e-958d-2f30f6e76130 · outbound

This paper cites an unresolved cited work.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-22T14:04:54.139914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:833b395a23a22a72db0ebe651dda5218ba0e4f04752a4b2fc3d1eda12a0e275b

Observation a253ca48-fed2-4cf1-880b-255d0b04e84a · outbound

This paper cites In general, higher temperatures lead to more diverse responses.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning In general, higher temperatures lead to more diverse responses

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.126091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:e1ac029a5975df9993bd5d102400164364f000cf3f72e7fd70b2048009bf1c55

Observation cbc348a7-6b45-4136-bc12-54ea994fda1b · outbound

This paper cites In addition, we introduce a naive baseline,Negative Length, which uses sequence length alone as a confidence signal.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning In addition, we introduce a naive baseline,Negative Length, which uses sequence length alone as a confidence signal

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.175859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:7ea15502649b17a4949233109cce8e5df99c08fa3bdae932ca89eb2948b675ba

Observation f1d82121-148f-41ed-ba09-11f09c57b74f · outbound

This paper cites 9600 - 7200.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning 9600 - 7200

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.115539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:493d81029e3f7eb1faf2f0976c5eec91d00eb257e0e2069d84c743a26db86e7e

Observation 42d2b686-4dd1-4e35-8f88-f73ac6ab28ae · outbound

This paper cites 9600−7200.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning 9600−7200

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.112464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:ce301a557d5309567b99f595449cbd3c8f2a13db6553db2fa96dc412c5b730b0

Observation 2ebd2217-0c3f-47f2-96bd-9d7e2208cf54 · outbound

This paper cites 60”. In the correct solution on the left, the model had low uncertainty for the correct answer “36.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning 60”. In the correct solution on the left, the model had low uncertainty for the correct answer “36

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T14:04:54.109102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T13:58:07.913104Z digest=sha256:2542c69f45d82d424b018b466c270fc5cdbba26f39262c6abce3b34bf297f34e

Pith citing papers

Observation f7ef378b-40fd-4f59-a19e-25d06754ba4c · inbound

Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes cites this paper.

Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T16:16:29.383723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:16:29.383723Z digest=sha256:4d64aca4d78356c29fb90b6fc4df9a136b05b030ebb8ec618d87f91121e03297

Observation 0ef053b1-8b98-4ab0-8527-f7b5173820f7 · inbound

Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models cites this paper.

Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:27:44.288866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:676812aadaaf0d7066fe754b139ef7bbd302749ef207db0649b2765081c80b83

Observation 629600cc-b54d-4b96-a783-9bbc168dd5d9 · inbound

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models cites this paper.

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T20:30:59.658434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:30:59.658434Z digest=sha256:8ec1cda9378de082a80696d8e9f74580d7c43db943aa4f7268b3239cb9212a53

Observation d1a90c92-be96-4835-9c71-591e03eea92c · inbound

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence cites this paper.

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:53:11.968470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T19:48:13.133733Z digest=sha256:5e6ce3931080060a5fd97c947fb118136b3c87eb3c0609f4403967ac0b36fae1

Observation 9bc3dafb-fa52-4179-9d76-6345cf4bf1ae · inbound

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis cites this paper.

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:06:25.889812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T04:36:27.730028Z digest=sha256:aa5d347148cb75b7b258d4f868f69c1095f0a84024e4180dea3918c649d3ca23

Observation 9ec8e472-acdd-424d-b2c2-4f3b2b4b405f · inbound

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis cites this paper.

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:43:00.552622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:41:09.382812Z digest=sha256:51c7b17fef2820d1749b85dbd69408506f4faf4ee97fa93451efab5015ac890e

Observation b75d57a1-83aa-47f1-bb0c-cdd080b526a3 · inbound

Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling cites this paper.

Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T02:22:19.005834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:22:19.005834Z digest=sha256:26843b9112826524ebaaf6b1f5c23e55c4bd4fa210b744299c0bfa3a65b092a3

Observation 259e2434-5889-4dc6-891f-bf65d478325b · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 132

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:38:56.088115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:3b25507531f0ae9708b7edecae0a8d2aae04803846a308db74e2bb1cc62cab7c

Observation b7ea7294-4129-4151-8e10-c0a81d825cf1 · inbound

LLM Doesn't Know What It Doesn't Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data cites this paper.

LLM Doesn't Know What It Doesn't Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:49:19.120110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T20:55:17.597719Z digest=sha256:d775b23c6850283149476ab95107056ae12678ee40be0867f0c0712f861f14e5

Observation 56f1e266-74b9-4b47-a709-9edce26419a7 · inbound

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation cites this paper.

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:07:50.405456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-11T03:05:42.262171Z digest=sha256:8feb19c86ba60125e308c40f5c6d7d0cccd152a41a6be23032873c7c26948ba0

Observation 56c8d711-8f51-4f7f-adc4-8b65feda5946 · inbound

Attention-Path Fragility as an Uncertainty Signal in Large Language Models cites this paper.

Attention-Path Fragility as an Uncertainty Signal in Large Language Models TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T05:30:27.620057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:30:27.620057Z digest=sha256:c5ffbdbe984e2f0cd533a063b64a504252a9261fe2c564ff62eddd74fc3a51f0