Pith. sign in

Paper Citation Record · LEDGER

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2509.09438.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09438 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T17:54:18.515174Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact22
  • verified fuzzy34
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08933d62-4006-400f-b7e8-14ed580ce454 · outbound

This paper cites GPT-4 Technical Report.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.044711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:fbc84ad254b179ada5c24a456441afc68f5f0de4f95ff9c6fe9027fce50497c9

Observation fdca9dc6-ff1a-4952-a999-6eb6eb42d824 · outbound

This paper cites Claude 3: A conversational ai model.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Claude 3: A conversational ai model

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.347298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:1e39898a7ab04861a9dd79c692c4ac8c4379a925fa8ae696b8870e371f683a12

Observation 0de05903-d348-426f-8856-70c87cdcbe15 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.111263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:a60fc0057645e3142f530fb144cb1298896c0e802e366641ab708d57c953feaf

Observation 109ea1f4-5ba3-4812-accd-3be8cd0c5cea · outbound

This paper cites Large language models encode clinical knowledge.Nature, 620(7972):172–180.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Large language models encode clinical knowledge.Nature, 620(7972):172–180

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.355923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:47808f374fbad8426655fdbf170477ba08e3ff6e07c35feba2f5abb0989ebc82

Observation 90f1e092-737f-441b-8a94-678715b651c4 · outbound

This paper cites Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.104825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:219b8d443f5956b55f3a02d533439f7842915f00af570a0e6661f48fc92f8789

Observation e860c4a1-a970-425e-a47f-018915464018 · outbound

This paper cites Finben: A holistic financial benchmark for large language models.Advances in Neural Information Processing Systems, 37:95716–95743.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Finben: A holistic financial benchmark for large language models.Advances in Neural Information Processing Systems, 37:95716–95743

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.331089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:4f9ab7b843698cdfa395536f4bced9e2d84468d2677d83cdd97a41c2bbc0d2d0

Observation e8cea93f-f185-4a6c-bec9-0ac77296fdcf · outbound

This paper cites Language Models (Mostly) Know What They Know.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Language Models (Mostly) Know What They Know

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.097881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:c37e33253753cc2d86d8b2d4cbdebaec54273bde3e37c592287f01159f2b893e

Observation 3782065e-5476-41b7-8f83-02116060f754 · outbound

This paper cites On calibration of modern neural networks.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models On calibration of modern neural networks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.335401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:8cb846507e2df0e63ce773b29ef7a944cf353f27d794734bd8d17066a561b0df

Observation c3d3c1c7-3b44-46d5-9250-a42e2e37e314 · outbound

This paper cites Calibration of pre-trained transformers.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Calibration of pre-trained transformers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.359464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:891ee9eb88210018c33d9981dd2e24bf98eb26aea48025b33776ea22bdcca53b

Observation cf18220f-4b9c-46e0-8967-63b23d21d73f · outbound

This paper cites Large language models are miscalibrated in-context learners.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Large language models are miscalibrated in-context learners

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.339460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:ba0b69deda736f1de8c435e11cf7663902858aade12bb997fda72b29f95b4d9f

Observation d076c677-f83b-4c5c-881a-3780e57bf923 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.118255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:85211543ccf39fb1b2a7cdbb82296a6c5c423fdfa4c8457c50dce2cd5e686a67

Observation 0c2bcd3a-b37d-413d-b568-0eb15f60cbe9 · outbound

This paper cites Atomic calibration of llms in long-form generations.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Atomic calibration of llms in long-form generations

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.091914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:bbf4f846863696a72c6921d8dd6b74e9a97a63923dce1b6055d0bc617fe0564e

Observation 1301b4a4-6e97-4555-b677-726185bc6bc1 · outbound

This paper cites Cali- brating large language models using their generations only.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Cali- brating large language models using their generations only

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.343511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:4754e226b3e522fa46551b80d149246349fcb9f59fb6b69e61212a8d5f068e1f

Observation 85e46730-e15d-473b-bbe1-eb47a51880a0 · outbound

This paper cites Large language models must be taught to know what they don’t know.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Large language models must be taught to know what they don’t know

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.326127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:5dfe1d20b410dd4051970380b32b64abfc266b857411ca0c1ac2da42f8f0e9a5

Observation 97667c96-49ba-4fca-858f-39d8b2d0ba52 · outbound

This paper cites Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.317474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:d2886b10f3f775d52cb92ee33787606a984e612439b9921bef8a980d1c5e5a9c

Observation 4ffb4919-8d61-4eb2-b880-2a87652d84ba · outbound

This paper cites Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.292848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:93dfdfc77fb6bdde5944556be169b8cd698299469d6f8a17b7416d23a5b4ad1a

Observation c81f8858-6d2f-42cb-96ce-e32e770d285a · outbound

This paper cites Linguistic calibration of long- form generations.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Linguistic calibration of long- form generations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.285841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:67767758c3390b995b16ef76c85bddd27f54a45337baa8e8d08c3d6e1333955d

Observation ed53c8f4-6f6f-4feb-9322-a9a4df42770d · outbound

This paper cites LoGU: Long-form Generation with Uncertainty Expressions.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models LoGU: Long-form Generation with Uncertainty Expressions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.130120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:0da57f985a61233ccf9a79dfbd055dc6b791081ddf885ec4610d4245a039f16b

Observation 3cae4aef-f9af-420c-85f0-ad995bd3d50a · outbound

This paper cites Uncle: Uncertainty expressions in long-form generation.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Uncle: Uncertainty expressions in long-form generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.124226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:f543920ed4724e55f3e9fc19d228d9d4889f7cb277e9e38c020b626c3027cd92

Observation 01baacb6-a6f0-45cb-b072-f6ec600633be · outbound

This paper cites Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T17:56:42.157411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:1079839ce1622daa77dba2060e7c9a7f7948cc2be16f847222342ba0111359de

Observation a730ad1d-a6e4-40f7-8261-808fa776e6ea · outbound

This paper cites Lora: Low-rank adaptation of large language models.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Lora: Low-rank adaptation of large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.275498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:87b58f26da10f5919ff18eab55920827e2f7b07c5910ba1a9f7ced578c6174fc

Observation 89f3de43-2aa7-49b1-b27e-c74b2d8803b5 · outbound

This paper cites Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.271450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:b05889579d1dc4cb1ec0b1d4f3dff68613e5e0735cf71999f977388e190768cb

Observation 7a177a63-dd6b-45f0-adc9-b552778b2af1 · outbound

This paper cites Get confused cautiously: Textual sequence memorization erasure with selective entropy maximization.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Get confused cautiously: Textual sequence memorization erasure with selective entropy maximization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.263942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:87df679105d8bf57d363157bd41dc5aefafef12967fd56ff4da274f789ec49c2

Observation b6b095ae-09d1-4fb3-9116-eab805e2315c · outbound

This paper cites Softmax probabilities (mostly) predict large language model correctness on multiple-choice q&a.arXiv e-prints, pages arXiv–2402.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Softmax probabilities (mostly) predict large language model correctness on multiple-choice q&a.arXiv e-prints, pages arXiv–2402

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.279253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:65c75661ef9c444920a8a1c93a82d50b429e997440b64c18f02021c91deabff3

Observation 2190c1cd-728d-4d24-918d-0a2129f0b54f · outbound

This paper cites The internal state of an llm knows when it’s lying.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models The internal state of an llm knows when it’s lying

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.296344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:656e0e52ae9f2ad1405ad7b880575fd3477b9bed2bf8a901a01419f11b77dab0

Observation ece7dc71-df39-4251-8a8c-feff6dad8063 · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.145469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:e3b36c27f2ca90cf5aac9fb456773caf25d9880c4dd30f9c089e8213c9f34fcf

Observation 7d6f1316-593c-4071-aaea-2f799821bbaf · outbound

This paper cites Enhancing language model factuality via activation- based confidence calibration and guided decoding.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Enhancing language model factuality via activation- based confidence calibration and guided decoding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.267670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:df3d4a094f377d9f767b0eb6680b0b53240ce9663d30e53f48f85a64238405ac

Observation e9223cb5-e101-4aca-a683-a8ecf4eaaa04 · outbound

This paper cites arXiv preprint arXiv:2503.14749 , year=.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models arXiv preprint arXiv:2503.14749 , year=

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.151417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:c287c9a0afcd2590b6fe0c000b51a5e71495a2b4adcc911a6acdd617fe384a18

Observation 19bd4ae3-8f3a-441d-939a-2cde8c052cdd · outbound

This paper cites I don’t know: Explicit modeling of uncertainty with an [idk] token.Advances in Neural Information Processing Systems, 37:10935–10958.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models I don’t know: Explicit modeling of uncertainty with an [idk] token.Advances in Neural Information Processing Systems, 37:10935–10958

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.234081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:dfb8a06028b1a7481808e351603ba6e4477ff69c7e72c23e645b8505ccce06e4

Observation 29798411-9659-45df-a51a-7255ac2d5da5 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.056469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:cc301e8b207c3c6f806944a190cb514bee6091d1019a8e76d1964a1a385a9647

Observation 8649d25b-61e3-460c-807c-3163d90334bb · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Self-consistency improves chain of thought reasoning in language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.238227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:9bfc07ba145d05a72b1632e97fff67a78799a9ce2025bf9479b112d7d1c97d2b

Observation 90fd3298-5194-42b6-bcda-65084ef934f4 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.066613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:f4b8d1bdfa010d274a389de6bd007484e16d4b4ad99b4a53f73a021a4ab49933

Observation b1f71bab-245c-47c2-b3de-02b41d321d41 · outbound

This paper cites s1: Simple test-time scaling.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models s1: Simple test-time scaling

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.061587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:e0dcfa2d2faac3b69343f89cfceb0babdcae191e3a08bc63a7d59c570f14a334

Observation b57c5af3-d79b-4191-aa1c-b425c288bd7e · outbound

This paper cites Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.249238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:f8044353811430a88eed19d4eb14ec9903ddf53fc2b000cdd43eecb49b0f3b94

Observation 50d0baf7-7392-4ebf-a69b-f83d7a040033 · outbound

This paper cites Let’s sample step by step: Adaptive- consistency for efficient reasoning and coding with llms.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Let’s sample step by step: Adaptive- consistency for efficient reasoning and coding with llms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.241800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:cf34bc96af9b232f5327937e037b647bcc8987318bbab8a99bd60797bbf4c273

Observation ee976613-51b5-4b71-8077-2c62e2cca600 · outbound

This paper cites Scaling Evaluation-time Compute with Reasoning Models as Evaluators.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Scaling Evaluation-time Compute with Reasoning Models as Evaluators

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:04:41.392178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:6994d1056a0dffed8551fc5599d3c548b18c56e3946bb3d25dbe52b5a24feb89

Observation f3d21a7a-9e4e-4a70-a6ae-5db6ad9d8974 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.245509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:9f06ecc1ba2b2e71f806aabb2eb3cd7daccefbce542c956ae62318a3268c4b06

Observation c8716dbb-7c67-49c1-b1b2-41f33570ca58 · outbound

This paper cites Confidence improves self-consistency in llms.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Confidence improves self-consistency in llms

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.084833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:fee19658b49090dffbceb56435165a4fc2d58f7a590bc51ea62e78039f28c6e4

Observation 1558b01b-1e7e-4872-b6a6-17c016193889 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Efficient Test-Time Scaling via Self-Calibration

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.072706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:207ad9a41a4891438fc29db821803eb91e50393bbe64851f3e9c988d8b4aaaa2

Observation 167cc81e-5d1d-452e-aded-20dbb3c36c63 · outbound

This paper cites Deep Think with Confidence.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Deep Think with Confidence

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.078309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:89e94beefd0aeefeeb75a8c892e1cfdfcf89b913aca8ca1b7069c834ebc42df2

Observation 7abfb6da-f762-4287-9bdf-6537a16da157 · outbound

This paper cites Think before you speak: Training language models with pause tokens.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Think before you speak: Training language models with pause tokens

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.229923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:fb4f30ea2276292fad64d395783929dcc91e696705dbdda53dbb1da6c8522e3c

Observation 579c00da-3e72-447f-97b4-af2bf2aeb0df · outbound

This paper cites Guiding language model reasoning with planning tokens.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Guiding language model reasoning with planning tokens

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.252859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:205601dc8d5311c9dcb66de2f5746704a48a91cbe3519987877862067b16a334

Observation 093ae9e3-4e72-4f5b-a013-a2241fe60813 · outbound

This paper cites Calibrated structured prediction.Advances in Neural Information Processing Systems, 28.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Calibrated structured prediction.Advances in Neural Information Processing Systems, 28

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.303661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:2db8ebac1dd972fc80bd48e3c535f5ea4d98610692cc5405c2c7d9951c3c6436

Observation 0ac20496-a55c-430d-bbb9-5c782734a6b9 · outbound

This paper cites Uncertainty estimation in autoregressive structured prediction.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Uncertainty estimation in autoregressive structured prediction

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.351910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:d6d0e0c96bb1d714d98a7dc5122131ec1d1f33651f76e06bc582acd3aba98483

Observation ff9f386e-761a-4aa3-860e-b5f66c36f861 · outbound

This paper cites Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods.Advances in large margin classifiers, 10(3):61–74.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods.Advances in large margin classifiers, 10(3):61–74

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.260423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:82be9e79b901fb8ea90a04cb72a56d993aaaec9ee1e414664461faa99ba16b55

Observation 80bc7247-a912-4607-8a3e-d70ca5d3a072 · outbound

This paper cites Reasoning aware self-consistency: Leveraging reasoning paths for efficient llm sampling.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Reasoning aware self-consistency: Leveraging reasoning paths for efficient llm sampling

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.256637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:734ed7026d6e5f7ade24afd05f102b5523b5486f5299cf8c8a5a89401ff047f7

Observation 0c875d73-d43f-4ab7-a834-60698ff28543 · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.282470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:6f541a576340874a2037341a2627ef51df0f97bf361fef328f90cec2e0af3b13

Observation 96e64ac5-2b76-47a6-990c-f89ae605f7bf · outbound

This paper cites Crowdsourcing multiple choice science questions.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Crowdsourcing multiple choice science questions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.289130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:7fc1f91199bb1d1db79dfa2c61c5dd5d4912adc32554af81c8255997a08f51e2

Observation 7ca5fbe7-7ef5-4fef-b421-71c07d77b853 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Rouge: A package for automatic evaluation of summaries

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.313162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:f7ca27a536bdc7b6fe22ed276deddb7809ca1f486154e43e98ccaf273d4c7baf

Observation 13b6c20a-52a8-48ef-95b7-ae89625033ae · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.178435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:ce5a1f87a411146ec23f4ec6f3e67d56cceb1c7ec9f559540b3f093812912410

Observation 0c4109c0-f65a-4f1b-936c-3f23335492b2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.168290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:70d8d1483226bbbb426e638fa9bf8c1097a3939169daabb9d2aa4a2887059199

Observation 24c07ed8-69f7-4b0a-ad00-488385d64e2f · outbound

This paper cites The Llama 3 Herd of Models.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models The Llama 3 Herd of Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.173070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:8bbdf6b1d1e05b5dfb326f4f6a8b5bf1a6c57789b91598e8ef9ad20c3adf7d7d

Observation 6f2fb485-c3aa-4696-a12b-0251b5050c98 · outbound

This paper cites Obtaining well calibrated probabilities using bayesian binning.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Obtaining well calibrated probabilities using bayesian binning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.321835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:8fe36a525f2aeac2efa92a4530bb21578e7c7632c74c5c7437b33ab5bbfb2b58

Observation 80caeb5b-504d-48b7-8211-fb70a298da07 · outbound

This paper cites Verification of forecasts expressed in terms of probability.Monthly weather review, 78(1):1–3.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Verification of forecasts expressed in terms of probability.Monthly weather review, 78(1):1–3

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.308941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:7b931cb81119c0983d4b5f0b72cb5e2d42bf7ef7dfeaaa4c107d0ab7b31d8225

Observation 9829f9ac-1a30-47f1-bfbe-6fd7cfe361bd · outbound

This paper cites Qwen2.5 Technical Report.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Qwen2.5 Technical Report

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.140456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:feecdff7fb9631c94461fa39ff267ff7563ad6720533be2ba3391fd39d06e48b

Observation 5a7a77d9-755d-4bfa-b0fa-147e5c5da30e · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Mathqa: Towards interpretable math word problem solving with operation-based formalisms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T17:56:43.300307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:4028902c2d569288cd16fd02903bcf4e3fe8851c0d456c8aaedf46151013a346

Observation 2b7dbcc3-2b18-46a2-bed1-30c1e8c74888 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.163103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:1f0cd0fe5c431a8eac9abe87051d5d507b85fac3b5177c39b89cd0c7f16484df

Observation f7e4c79b-102c-444a-a876-cb7428721ea2 · outbound

This paper cites Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs.

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs

Reference 59

Resolution
malformed identifier
arxiv_id, observed 2026-05-18T17:56:42.135178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:54:18.515174Z digest=sha256:549a9d665cab5ac0c556f928ecda297b779d1b35339e2c1b8126e454c60b6b59

Pith citing papers

No inbound Pith citation observations are available.