Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T17:54:18.515174Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2509.09438.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T17:54:18.515174Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 08933d62-4006-400f-b7e8-14ed580ce454 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fdca9dc6-ff1a-4952-a999-6eb6eb42d824 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Claude 3: A conversational ai model
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0de05903-d348-426f-8856-70c87cdcbe15 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 109ea1f4-5ba3-4812-accd-3be8cd0c5cea · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Large language models encode clinical knowledge.Nature, 620(7972):172–180
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90f1e092-737f-441b-8a94-678715b651c4 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e860c4a1-a970-425e-a47f-018915464018 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Finben: A holistic financial benchmark for large language models.Advances in Neural Information Processing Systems, 37:95716–95743
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8cea93f-f185-4a6c-bec9-0ac77296fdcf · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Language Models (Mostly) Know What They Know
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3782065e-5476-41b7-8f83-02116060f754 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models On calibration of modern neural networks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3d3c1c7-3b44-46d5-9250-a42e2e37e314 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Calibration of pre-trained transformers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf18220f-4b9c-46e0-8967-63b23d21d73f · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Large language models are miscalibrated in-context learners
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d076c677-f83b-4c5c-881a-3780e57bf923 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c2bcd3a-b37d-413d-b568-0eb15f60cbe9 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Atomic calibration of llms in long-form generations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1301b4a4-6e97-4555-b677-726185bc6bc1 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Cali- brating large language models using their generations only
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85e46730-e15d-473b-bbe1-eb47a51880a0 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Large language models must be taught to know what they don’t know
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97667c96-49ba-4fca-858f-39d8b2d0ba52 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ffb4919-8d61-4eb2-b880-2a87652d84ba · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c81f8858-6d2f-42cb-96ce-e32e770d285a · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Linguistic calibration of long- form generations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed53c8f4-6f6f-4feb-9322-a9a4df42770d · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models LoGU: Long-form Generation with Uncertainty Expressions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cae4aef-f9af-420c-85f0-ad995bd3d50a · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Uncle: Uncertainty expressions in long-form generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01baacb6-a6f0-45cb-b072-f6ec600633be · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a730ad1d-a6e4-40f7-8261-808fa776e6ea · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Lora: Low-rank adaptation of large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89f3de43-2aa7-49b1-b27e-c74b2d8803b5 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a177a63-dd6b-45f0-adc9-b552778b2af1 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Get confused cautiously: Textual sequence memorization erasure with selective entropy maximization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6b095ae-09d1-4fb3-9116-eab805e2315c · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Softmax probabilities (mostly) predict large language model correctness on multiple-choice q&a.arXiv e-prints, pages arXiv–2402
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2190c1cd-728d-4d24-918d-0a2129f0b54f · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models The internal state of an llm knows when it’s lying
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ece7dc71-df39-4251-8a8c-feff6dad8063 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d6f1316-593c-4071-aaea-2f799821bbaf · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Enhancing language model factuality via activation- based confidence calibration and guided decoding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9223cb5-e101-4aca-a683-a8ecf4eaaa04 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models arXiv preprint arXiv:2503.14749 , year=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19bd4ae3-8f3a-441d-939a-2cde8c052cdd · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models I don’t know: Explicit modeling of uncertainty with an [idk] token.Advances in Neural Information Processing Systems, 37:10935–10958
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29798411-9659-45df-a51a-7255ac2d5da5 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8649d25b-61e3-460c-807c-3163d90334bb · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Self-consistency improves chain of thought reasoning in language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90fd3298-5194-42b6-bcda-65084ef934f4 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1f71bab-245c-47c2-b3de-02b41d321d41 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models s1: Simple test-time scaling
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b57c5af3-d79b-4191-aa1c-b425c288bd7e · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50d0baf7-7392-4ebf-a69b-f83d7a040033 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Let’s sample step by step: Adaptive- consistency for efficient reasoning and coding with llms
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee976613-51b5-4b71-8077-2c62e2cca600 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Scaling Evaluation-time Compute with Reasoning Models as Evaluators
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3d21a7a-9e4e-4a70-a6ae-5db6ad9d8974 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8716dbb-7c67-49c1-b1b2-41f33570ca58 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Confidence improves self-consistency in llms
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1558b01b-1e7e-4872-b6a6-17c016193889 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Efficient Test-Time Scaling via Self-Calibration
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 167cc81e-5d1d-452e-aded-20dbb3c36c63 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Deep Think with Confidence
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7abfb6da-f762-4287-9bdf-6537a16da157 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Think before you speak: Training language models with pause tokens
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 579c00da-3e72-447f-97b4-af2bf2aeb0df · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Guiding language model reasoning with planning tokens
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 093ae9e3-4e72-4f5b-a013-a2241fe60813 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Calibrated structured prediction.Advances in Neural Information Processing Systems, 28
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ac20496-a55c-430d-bbb9-5c782734a6b9 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Uncertainty estimation in autoregressive structured prediction
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff9f386e-761a-4aa3-860e-b5f66c36f861 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods.Advances in large margin classifiers, 10(3):61–74
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80bc7247-a912-4607-8a3e-d70ca5d3a072 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Reasoning aware self-consistency: Leveraging reasoning paths for efficient llm sampling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c875d73-d43f-4ab7-a834-60698ff28543 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96e64ac5-2b76-47a6-990c-f89ae605f7bf · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Crowdsourcing multiple choice science questions
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ca5fbe7-7ef5-4fef-b421-71c07d77b853 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Rouge: A package for automatic evaluation of summaries
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13b6c20a-52a8-48ef-95b7-ae89625033ae · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c4109c0-f65a-4f1b-936c-3f23335492b2 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24c07ed8-69f7-4b0a-ad00-488385d64e2f · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models The Llama 3 Herd of Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f2fb485-c3aa-4696-a12b-0251b5050c98 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Obtaining well calibrated probabilities using bayesian binning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80caeb5b-504d-48b7-8211-fb70a298da07 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Verification of forecasts expressed in terms of probability.Monthly weather review, 78(1):1–3
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9829f9ac-1a30-47f1-bfbe-6fd7cfe361bd · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Qwen2.5 Technical Report
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a7a77d9-755d-4bfa-b0fa-147e5c5da30e · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b7dbcc3-2b18-46a2-bed1-30c1e8c74888 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7e4c79b-102c-444a-a876-cb7428721ea2 · outbound
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.