Pith. sign in

Paper Citation Record · LEDGER

Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2407.16221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.16221 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:22.449090Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T19:55:33.960847Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1555fe95-6146-4ff3-bf68-eb5919e48133 · inbound

Towards Harmonized Uncertainty Estimation for Large Language Models cites this paper.

Towards Harmonized Uncertainty Estimation for Large Language Models Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.449090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:22.449090Z digest=sha256:f3bbd830faa4dffe94c70a3ac546a64f7f3282f07fb88a3b40c9f75fe8c50879

Observation ddee5a53-226a-47fc-9163-449121d25e26 · inbound

CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention cites this paper.

CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:37.100267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:08:37.100267Z digest=sha256:960b16f001739deae68e672f14b892d580ad879357dc343a724f76a608f90cc0

Observation 4f7557c1-9c72-41b1-9710-27353326a2ed · inbound

Fine-Tuning and Prompt Engineering of LLMs, for the Creation of Multi-Agent AI for Addressing Sustainable Protein Production Challenges cites this paper.

Fine-Tuning and Prompt Engineering of LLMs, for the Creation of Multi-Agent AI for Addressing Sustainable Protein Production Challenges Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:23.531540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:23.531540Z digest=sha256:d31b4f78829daae904f0ba738a131a693167143ae7505b529085f211aab37361

Observation 5aa7c029-4be7-4f30-81db-8b7951dcb419 · inbound

VisionTrap: Unanswerable Questions On Visual Data cites this paper.

VisionTrap: Unanswerable Questions On Visual Data Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:08.233389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:08.233389Z digest=sha256:d693cac3d5c90bdb282488605aae2c5f643eadfb4062d0044c3a8566ba25a2f7

Observation a58d7397-4577-4cc5-9ca4-b5aab31dee69 · inbound

Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control cites this paper.

Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:23:47.223084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:23:47.223084Z digest=sha256:86ee2faffd1eb9f9dc6075aab2a12c6a147321271c3dbf9483ed45e77781190e

Observation 2418c629-0792-4380-a047-235bada9e429 · inbound

Scaling Truth: The Confidence Paradox in AI Fact-Checking cites this paper.

Scaling Truth: The Confidence Paradox in AI Fact-Checking Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T20:09:16.900567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:09:16.900567Z digest=sha256:02a17c748e629d644a62846efbeab7283a870b80d3620070e50339cd1e573922

Observation f091fd4c-f072-403c-a475-1c593a7ef720 · inbound

Causal Evidence that Language Models use Confidence to Drive Behavior cites this paper.

Causal Evidence that Language Models use Confidence to Drive Behavior Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:44:05.697748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T09:43:05.524088Z digest=sha256:d2c7c74fabc525896886daf7126660057077a9b144815756eb7aa0a521c284e9

Observation a5f1f369-6d9a-4308-a79e-158928104152 · inbound

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence cites this paper.

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:53:11.971607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T19:48:13.133733Z digest=sha256:1b4de97435402eb2f39bcde5db2f15357e9fc7132a63c940ca3547091c13e180

Observation c22f51ba-b88e-46f3-8a36-4a272bcd8648 · inbound

Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents cites this paper.

Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:57:15.529664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:55:07.930578Z digest=sha256:ea6b2df8169ef7d9e99914695015fb2bf50eef87c55c393dd7764706f1799f42

Observation 2ef30ac9-5799-416e-bae7-7a688aaa7a63 · inbound

Partial Evidence Bench: Benchmarking Authorization-Limited Evidence in Agentic Systems cites this paper.

Partial Evidence Bench: Benchmarking Authorization-Limited Evidence in Agentic Systems Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.879312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:28:51.943383Z digest=sha256:f55b6b730b9e93e3369976c39b1b036decfda87f1c233145f842f788ab381f28

Observation f21201c5-3228-4cca-be59-44147ff29bac · inbound

When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents cites this paper.

When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:19:49.568196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T08:04:57.334657Z digest=sha256:65609464d067efa4c5d433476d42c35ecd13d3c982fa5278a958931497c1aa53

Observation 7cbb30f9-6e33-4599-8211-fc5229f886a2 · inbound

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation cites this paper.

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.962308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-08T19:45:46.002113Z digest=sha256:f18f892242cf52322a80fc0b9ce9cf4b6355c13a140283e310981fdf741cf35c