Pith. sign in

Paper Citation Record · LEDGER

Measuring short-form factuality in large language models

As of 24 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 100 inbound Pith citation observations for arXiv:2411.04368.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04368 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T06:45:50.219157Z

measured 119 of 119 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 100 of 147 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:09:31.122730Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact11
  • verified fuzzy5
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b073f0d9-5944-412f-9adf-8a89a7b45b32 · outbound

This paper cites Do Language Models Know When They're Hallucinating References?.

Measuring short-form factuality in large language models Do Language Models Know When They're Hallucinating References?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.241868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:51468c23346af8b1c87bcba334f7730bbf4eb598dba61a3cb8196697b68d9a3e

Observation db243db8-6c9c-4349-a301-7f9cb93a2cbb · outbound

This paper cites Anthropic.

Measuring short-form factuality in large language models Anthropic

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:45:50.313374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:6b337160253fa4c92156b66c895461effa9a8a12b25f37f33a36d85a63965271

Observation e7946d0d-d737-409a-9651-ab722df88e72 · outbound

This paper cites Evaluating Hallucinations in Chinese Large Language Models.

Measuring short-form factuality in large language models Evaluating Hallucinations in Chinese Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.302052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:0333fe3d0772e8470124607e05d5f6d52b9d85ec405e63aa16bf9aa10b130310

Observation 39e907eb-f1fa-420d-9054-080ae0e0ce06 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Measuring short-form factuality in large language models TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:45:50.309098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:3efa2061284f20757b173012320d0daa5afab63ef3ad213202b93dd8cf35c1a5

Observation 6799a419-3a33-45e1-9eb7-41b49cee82ca · outbound

This paper cites Language Models (Mostly) Know What They Know.

Measuring short-form factuality in large language models Language Models (Mostly) Know What They Know

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:45:50.247729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:0fc5ed95c092bca9c4e1a4d823d8fb35c5e917eedaac194ace95553ca2178df2

Observation f1c95610-5e4f-4f56-82b0-251e11647666 · outbound

This paper cites an unresolved cited work.

Measuring short-form factuality in large language models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-15T06:45:50.330120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:d49b77e31e24ad4571c3902c5750d6b5c6296562a361acf8ba1de28b1eedecb9

Observation ed2ca68b-042e-41d5-bc2f-332a8fd42f93 · outbound

This paper cites Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation.

Measuring short-form factuality in large language models Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.254884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:25e60cb8572084062ff29ddf461028c3f48e4e21cb925d7f34952bef70a23649

Observation bde3cc3d-94ff-426a-89e5-9b645bf69cef · outbound

This paper cites Kwiatkowski, J.

Measuring short-form factuality in large language models Kwiatkowski, J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:45:50.339191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:a9abc5aab34e6d274be9adec8e55777e587b54eab5b81978ce43cdfb32746578

Observation d3b46f2b-f045-40d8-9e11-be3075ccf31c · outbound

This paper cites HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models.

Measuring short-form factuality in large language models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.262718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:f9a22911251c9bb7290b1abcde9639e41684e030234fe51f0b01f35ef478e548

Observation ac5c9d88-26cb-4779-99f6-4ab53a4d63d8 · outbound

This paper cites an unresolved cited work.

Measuring short-form factuality in large language models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-15T06:45:50.317471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:091f660700403c38b9b1b2439f90b936a4ff0ce96514d0d409352001e3a68c88

Observation 70495176-badf-4cfd-ac03-e572429f0681 · outbound

This paper cites Teaching Models to Express Their Uncertainty in Words.

Measuring short-form factuality in large language models Teaching Models to Express Their Uncertainty in Words

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:36:08.649609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:6e6183b10faac75cc1d281cb369de8c5b1808fde78273d76738294ca91349c99

Observation ae740448-1b5c-44fa-af53-df06ce47ebbf · outbound

This paper cites an unresolved cited work.

Measuring short-form factuality in large language models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-15T06:45:50.326041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:1f7d27b2ef70e63f840d86801df9b8884620c3906262ad1c0c3bcb8ea0b5d412

Observation 5b06c113-8476-44f4-8286-ce49947e43a6 · outbound

This paper cites Hello gpt-4o, 2024 a.

Measuring short-form factuality in large language models Hello gpt-4o, 2024 a

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:45:50.334798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:a0c6570c4946c749d887cb4b33884f8ec41e03c5f5613372933b7b362f33a604

Observation da15d8a0-8180-4e76-89d1-a37120a204da · outbound

This paper cites Openai o1-mini, 2024 b.

Measuring short-form factuality in large language models Openai o1-mini, 2024 b

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:45:50.343851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:ee798d3cb0eaea6b86042470dc44a8f574938b642312e6b5c3f67bc46a0dd7a7

Observation c6c47803-e762-410b-a0ed-a6e74bc60ae2 · outbound

This paper cites Learning to reason with llms, 2024 c.

Measuring short-form factuality in large language models Learning to reason with llms, 2024 c

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:45:50.321566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:8d2ae9d7cb180c232e6a7f4ba9d4afaf17531a543c2517d003ff656df3a465da

Observation 1437c2c2-9997-49b4-80fa-c6512fe92af5 · outbound

This paper cites FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation.

Measuring short-form factuality in large language models FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.275719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:adf2dd3bda6b7ac1d27d977df340e57c4faa7a472932e79c8dfc1784a6f6e806

Observation 9470a51c-6d82-44e1-9256-b44614e48517 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Measuring short-form factuality in large language models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:45:50.282772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:611b2225fe832737ba1805e5f392064098e6a7288aaaf2eb2bcd174b9e740c13

Observation 06b1d0f2-7278-4747-b188-f7c739200a4b · outbound

This paper cites Long-form factuality in large language models.

Measuring short-form factuality in large language models Long-form factuality in large language models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.288896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:66262aae29bea5f2786f26e8cd0f036908bcfd566cb2c9385d8d05fe557280ec

Observation f67ea904-82ad-45f6-96fe-888ce408dc48 · outbound

This paper cites FELM: Benchmarking Factuality Evaluation of Large Language Models.

Measuring short-form factuality in large language models FELM: Benchmarking Factuality Evaluation of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.295212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:e7c709804dd8c7af5215b321d425e4003e03c014dc9c3e1e1a5639a232b60f34

Pith citing papers

Observation e34ef943-f0b5-4d86-bc6f-10cae8658b79 · inbound

O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson? cites this paper.

O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson? Measuring short-form factuality in large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:07:48.894841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:07:48.894841Z digest=sha256:a44e1352f4b8f92b61eb1430fbd9dcac94128d93a2dae2b2fa88903dbd5ac419

Observation 3c931fb1-a8f3-46e2-9890-e16ce7c9656d · inbound

100% Elimination of Hallucinations on RAGTruth for GPT-4 and GPT-3.5 Turbo cites this paper.

100% Elimination of Hallucinations on RAGTruth for GPT-4 and GPT-3.5 Turbo Measuring short-form factuality in large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:55:26.901084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:55:26.901084Z digest=sha256:f4ccc2bfc223cc19fa84dbea30a901381b3c160ff1fa4531c4e60bebe272c18d

Observation 7cb769b5-1664-425b-950f-9cc7ed93c96b · inbound

Deliberative Alignment: Reasoning Enables Safer Language Models cites this paper.

Deliberative Alignment: Reasoning Enables Safer Language Models Measuring short-form factuality in large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:44:16.347418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:44:16.347418Z digest=sha256:497253557661bf8f9178d1b850c0bc59623382360acf4403f24d2eeadf109b52

Observation 6bdfc51e-a3d6-4f79-aa35-b1d2a30439c4 · inbound

Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking cites this paper.

Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking Measuring short-form factuality in large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:49.095505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:49.095505Z digest=sha256:7b61bb4013508924521fec135cec17d3851302844476622bf0a80bda4b424193

Observation c8ecd44b-6284-4d06-998a-a87c5b07c9d4 · inbound

Decoding Knowledge in Large Language Models: A Framework for Categorization and Comprehension cites this paper.

Decoding Knowledge in Large Language Models: A Framework for Categorization and Comprehension Measuring short-form factuality in large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:33:24.532912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:33:24.532912Z digest=sha256:156853b9c76af1ec2ab8bf836829be10d416c2056a61a8c5ff3b2f20d8588c48

Observation 830fab3b-2009-431d-be43-48d6d4ded6b2 · inbound

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input cites this paper.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Measuring short-form factuality in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.916094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.916094Z digest=sha256:7360705c9f1ef710c28379e53a87e250462f293e2b2338b6cd8e62efdf4634f5

Observation 0219d421-062f-42b2-9758-a7ad296a4264 · inbound

TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time cites this paper.

TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time Measuring short-form factuality in large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:43:45.409278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:43:45.409278Z digest=sha256:825e9f8127ba6ac65367d8c375846b6a0e4e851f41b374b23d66342a2a5dca63

Observation 250d4475-5baa-45c9-b35d-d1c508a22886 · inbound

Predictions as Surrogates: Revisiting Surrogate Outcomes in the Age of AI cites this paper.

Predictions as Surrogates: Revisiting Surrogate Outcomes in the Age of AI Measuring short-form factuality in large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T19:54:58.050866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:54:58.050866Z digest=sha256:1fbe1e482f6d264c25bff5438c1df5e29c6649cb6b1e181211f45f23ecc73a96

Observation ea073248-6907-44e2-b96d-353a4f384ba4 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Measuring short-form factuality in large language models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:67aa9328b047d435bb76b7eb1e596d34f61a8d23fea29f2e56e907018b732952

Observation 9d1158f4-6348-4a56-97f5-b98797720e26 · inbound

Trading Inference-Time Compute for Adversarial Robustness cites this paper.

Trading Inference-Time Compute for Adversarial Robustness Measuring short-form factuality in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T22:19:41.395082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:19:41.395082Z digest=sha256:f7a6abc56e97da50422bb9e0005ca9f210ea587bef63aa436b3661860317b02a

Observation 77d63d76-0b38-4929-95cb-12c593631201 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning Measuring short-form factuality in large language models

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:11:37.251634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:c5502be0bb80a7dbfa105cdbe390cd71147632d90ff633c681c120f701b3f2ba

Observation 1bb2c226-877e-4409-9356-16042d1529cf · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Measuring short-form factuality in large language models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.479245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.479245Z digest=sha256:b0fc85c043ac81698481cb9d4296b4290a7a03c87c84f9a2304cad908f2459ee

Observation 75b9796e-604a-43ed-a377-36da011eea82 · inbound

Learning to Reason at the Frontier of Learnability cites this paper.

Learning to Reason at the Frontier of Learnability Measuring short-form factuality in large language models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.155841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:33e4c95f5b9069446febfcf17bb88d6e0453d55a78b124bebf5bc2db9db0364d

Observation c873965e-5314-42be-9784-99015204d247 · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Measuring short-form factuality in large language models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:27:21.310735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:f1a20d0673f1f16a5c7fdb5813fdc1e50cfdf724ffb10dedd53f525bc1945c98

Observation 5f317ab6-a637-4ac1-8003-7c7c8dea20fa · inbound

Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results cites this paper.

Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results Measuring short-form factuality in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:09:31.122730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:09:31.122730Z digest=sha256:7fa364f7f2462307467b410f605fa66bece8247672cedc3b0d64ca6677540e96

Observation 539bbfea-6fcc-4932-b7bf-33ec54fcd598 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark Measuring short-form factuality in large language models

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.757700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.757700Z digest=sha256:8dfba67b6ffa3871e57baf687160dbc3d92da951bbbd7494c37bfd5f43acca08

Observation e10034c9-f11a-42d8-986e-814e5055f879 · inbound

aiXamine: Simplified LLM Safety and Security cites this paper.

aiXamine: Simplified LLM Safety and Security Measuring short-form factuality in large language models

Reference 128

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:11.275027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:11.275027Z digest=sha256:f67e7345f50d38ddc47068751f0d7172122ee39ccb74eca081100f7051584cc6

Observation ca7a9cbc-c633-41bf-8faf-dbd2ad02ca5e · inbound

HalluLens: LLM Hallucination Benchmark cites this paper.

HalluLens: LLM Hallucination Benchmark Measuring short-form factuality in large language models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:19.340665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:19.340665Z digest=sha256:4c2d2cb839047b58e272341e5277973df54be5337da76be1a91d6c0409df7830

Observation 6231d726-a778-47b9-b3f3-b19517bcab23 · inbound

Investigating task-specific prompts and sparse autoencoders for activation monitoring cites this paper.

Investigating task-specific prompts and sparse autoencoders for activation monitoring Measuring short-form factuality in large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:38:12.232963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:38:12.232963Z digest=sha256:a207bd5bf1b2b84368a5a527482f4af5b96fcd0120b103966080aa8b7e6abc70

Observation b5aa205e-0465-4378-8d75-1f4719ce8f88 · inbound

Evaluating LLM Metrics Through Real-World Capabilities cites this paper.

Evaluating LLM Metrics Through Real-World Capabilities Measuring short-form factuality in large language models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T22:03:20.469685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:03:20.469685Z digest=sha256:00b544af076a40868df510f7dcb2b178a1312810a74e0c689fef77be5ab492a0

Observation 9bfd8c1c-7a42-490f-8262-aa9cc102b295 · inbound

Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs cites this paper.

Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs Measuring short-form factuality in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:38:44.694284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:38:44.694284Z digest=sha256:74b4c1c997f7f58271fb1b18b8aade3c756b0f37f070e74ff9a2c460a3cb0af2

Observation d3db9def-01cd-4792-bf47-2a938be442c0 · inbound

Phare: A Safety Probe for Large Language Models cites this paper.

Phare: A Safety Probe for Large Language Models Measuring short-form factuality in large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:23.215085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:23.215085Z digest=sha256:0cfffade86244ef71c4f160a95986c3a73c02d83586b338dc955402dfec76aa6

Observation 2a38d3bc-a9c3-4402-a5b7-0b8712479fa4 · inbound

AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning cites this paper.

AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning Measuring short-form factuality in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:49.430503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:49.430503Z digest=sha256:d831b8f27fa31180e862e218d6bb42f5e7daf24530c48a97ab91998f49f56bb4

Observation cee2d8f7-c7d7-4872-8ffd-955bf72fc76f · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Measuring short-form factuality in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.846476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.846476Z digest=sha256:b6a8f65ddf5000ee102f01abe97c1ee669e65d3c04bd65da593ff64df6b6eecd

Observation beeab439-929c-4b2b-aba6-664376dde5d1 · inbound

MedBrowseComp: Benchmarking Medical Deep Research and Computer Use cites this paper.

MedBrowseComp: Benchmarking Medical Deep Research and Computer Use Measuring short-form factuality in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:51.016582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:30:51.016582Z digest=sha256:6d41737b307a5b7dac54508ccb6bd5c0199c02f6e3e59221a81a4b1749ce805f

Observation 29f35dad-8602-4e0c-91fc-4b05780a2a6c · inbound

An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents cites this paper.

An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Measuring short-form factuality in large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:53.898164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:53.898164Z digest=sha256:8b364851db669465dbb169afa6e58fc2313929d7d84ec71565ab43c2b2bb4998

Observation 2769800b-4406-4c1e-8e46-04605919d1df · inbound

Adaptive Plan-Execute Framework for Smart Contract Security Auditing cites this paper.

Adaptive Plan-Execute Framework for Smart Contract Security Auditing Measuring short-form factuality in large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:14.383511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:26:14.383511Z digest=sha256:753d46682981d0b28bceeb3ad99b9a3cdb724acc0afcdd46146a584070940fa7

Observation 02815d34-6320-4292-86cb-a907077c71bd · inbound

Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs cites this paper.

Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs Measuring short-form factuality in large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:52.150225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:52.150225Z digest=sha256:2989bae9753490338ff8214cfe3bdf6cc9933217367f0dff8deaef7eb8fc3034

Observation 22190d48-0f88-4f5e-bb12-f47a37c0b9ae · inbound

Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA cites this paper.

Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA Measuring short-form factuality in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:27.490649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:27.490649Z digest=sha256:507b0dff49f4ec7079a422160fef20557def816f9b3fe15bfb8060f18c2af729

Observation acbd60b6-dba9-4d11-8652-d52550417321 · inbound

RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models cites this paper.

RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models Measuring short-form factuality in large language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:50.289751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:32:50.289751Z digest=sha256:b4347c405bffc414f317ef70971782327590ba745261b095231aeecc5a70c104

Observation a1abe4f0-9b5d-4aef-ba9a-c2fba24a4f3d · inbound

EvolveSearch: An Iterative Self-Evolving Search Agent cites this paper.

EvolveSearch: An Iterative Self-Evolving Search Agent Measuring short-form factuality in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.778096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.778096Z digest=sha256:15204067b893e2915fe2da20ff55fa9c15d24bb24da97848c7452f2c434fa1a5

Observation fc6fa950-9580-4289-92bf-3011befb0c22 · inbound

Are Reasoning Models More Prone to Hallucination? cites this paper.

Are Reasoning Models More Prone to Hallucination? Measuring short-form factuality in large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:32.707752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:32.707752Z digest=sha256:766ae385ff671c440ff976201a4b4168cc13f0caf2d26aa6c0c401e95270c3be

Observation fa3b2187-3169-49b8-be3f-f2d630da372e · inbound

Reconsidering LLM Uncertainty Estimation Methods in the Wild cites this paper.

Reconsidering LLM Uncertainty Estimation Methods in the Wild Measuring short-form factuality in large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:51.492314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:51.492314Z digest=sha256:a7262b5c0c2f0c56358b0d1da7cb5db8dbd7c228ab6da8c033fe3653a207c27e

Observation ce2080a2-9371-4e44-8d81-267562a7ef10 · inbound

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows cites this paper.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Measuring short-form factuality in large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.669134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.669134Z digest=sha256:f1731fdc9a20e292594be48dbb42ee133142f8b1065125e3e14357facb43a40e

Observation b0e4bb2c-45b7-4e5d-8659-41979ecf42d5 · inbound

Quantifying Cross-Modality Memorization in Vision-Language Models cites this paper.

Quantifying Cross-Modality Memorization in Vision-Language Models Measuring short-form factuality in large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:49.586262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:49.586262Z digest=sha256:1b4027ec15b46dd32eb2b37bdcd5620e47ea5f92f042b3218fb8042a172186e9

Observation 62e5ed4e-0cb3-4a94-a314-773e22ff5983 · inbound

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems cites this paper.

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems Measuring short-form factuality in large language models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:22:15.910547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T09:20:12.827871Z digest=sha256:0bcff37c7478c10133b964c8735024df802d2ca91059b2b01cd5aadcf6ae1051

Observation a67319c4-9aba-48ee-a765-8ba08e47cab0 · inbound

AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models cites this paper.

AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models Measuring short-form factuality in large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:18.755967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:18.755967Z digest=sha256:57fa07fb181718cbd23eb7c71b8e25880512adf74a1e2fd03cf548dc699c8a1c

Observation aa04c82a-eca2-449a-a320-c53ea51779a6 · inbound

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention cites this paper.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Measuring short-form factuality in large language models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:f52da6510b6e292381d09d3221befaf321694d57f3855c798de148787a597909

Observation 048aff7e-4691-4cbb-89fe-1d70bc06fde9 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap Measuring short-form factuality in large language models

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.940654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.940654Z digest=sha256:22e18a7dc35fe24c77d40771320a80d22ae14821206e7ade583cd591390afb41

Observation 4089ce7a-91ae-4ff1-a2f0-0ea4f506acc8 · inbound

Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know? cites this paper.

Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know? Measuring short-form factuality in large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:22.109878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:22.109878Z digest=sha256:bf22c0452dd77d7e3304d6550639e3226f617fc331d17b166ebae8e8f579f170

Observation 8a25a484-6c09-4d7d-823c-3e1bdf29f7df · inbound

Jan-nano Technical Report cites this paper.

Jan-nano Technical Report Measuring short-form factuality in large language models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:33.833633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:33.833633Z digest=sha256:9494a808b0bf5b73826760a953720f0ca94d79ceda8a60fc14f04f8473a10be8

Observation e091268f-41eb-4008-bc86-6a39117b640e · inbound

L0: Reinforcement Learning to Become General Agents cites this paper.

L0: Reinforcement Learning to Become General Agents Measuring short-form factuality in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.718465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.718465Z digest=sha256:96abdf00658f39cb31313b1832c972913a215c97dc33744d1b52f0ecf36fc711

Observation 587eb515-049d-4032-a808-c06af7a2e0ba · inbound

WebSailor: Navigating Super-human Reasoning for Web Agent cites this paper.

WebSailor: Navigating Super-human Reasoning for Web Agent Measuring short-form factuality in large language models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:37:09.727529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T15:37:09.572241Z digest=sha256:4d5890f63ff6353ce20d665c150cea18982ae6dcfa0ab50658f30e00dad3938c

Observation 7e007b0c-0b86-4d42-94b7-84dd4f755e1e · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring short-form factuality in large language models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.263615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.263615Z digest=sha256:9360bde57616ad70b1eb467fce5e5b008010fccb5a61fbdaf63e1086983d10c0

Observation 949d3b7c-2feb-468e-ab0e-542558767ab6 · inbound

How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models cites this paper.

How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models Measuring short-form factuality in large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:55.431654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:55.431654Z digest=sha256:d00cab427631f4f5fb501ef3106a4da9cc99eb6517b52a890e3e4f36823f16bd

Observation 9fbcc871-6593-478b-b69d-5ef38f14d250 · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Measuring short-form factuality in large language models

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:52:07.841600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:af737e33e5690289b214e9818f72bb68dbbf1e68b713e8dba26d6fa5cc044fed

Observation 3c8a7c98-2fc2-429a-84e7-18a47b7acd64 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Measuring short-form factuality in large language models

Reference 199

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.842336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.842336Z digest=sha256:6b39f5fbb74f0e6bfac7df3588b1bfe41722ad963982e08511e91ffbcf4c33a6

Observation aea26196-b430-4ef0-bc43-bcd375bec63c · inbound

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization cites this paper.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Measuring short-form factuality in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.587682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.587682Z digest=sha256:40492e75df644e4420bc42c22802209ac59b72989926d766eb4f06af97885571

Observation 469bc87b-11d9-4711-8630-64820a91cc2e · inbound

RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation cites this paper.

RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation Measuring short-form factuality in large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:29.262992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:29.262992Z digest=sha256:b770de4f27166972d6f1323e24aef818731eda14e6aa380afa8796d4e2bdf34e

Observation 02b01546-0239-435a-a083-f9976411cbf7 · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence Measuring short-form factuality in large language models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:59100039cb19a6080e27c9f3a25250ce6f63f2b6b5d00b0f7563f076fa2b3ab3

Observation 3d62732f-eedf-4992-90d0-6df84b72ea84 · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Measuring short-form factuality in large language models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.277099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.277099Z digest=sha256:82e7ab9e8015d6ed5ad7117957b543acdfe59b842638f9f4fc7dcb2cb8a3fbd9

Observation 7c39e5ee-23a0-4633-9785-eaa81364b574 · inbound

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution cites this paper.

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution Measuring short-form factuality in large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:25.574191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:25.574191Z digest=sha256:cb3bdfe792c4e48c93c7917da8847e5dc94ff047e10e16ca78ab1d6707064677

Observation ef6ec15e-a4ab-4d7a-8305-06bccb3321e2 · inbound

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty cites this paper.

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty Measuring short-form factuality in large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:36:49.146226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:36:49.146226Z digest=sha256:d412be3af9fd69572431740b56f017a4b536d9380d7b146108b2e46430075627

Observation 914e4dc4-33d4-4f3d-9017-87ea06285be2 · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning Measuring short-form factuality in large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.901494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.901494Z digest=sha256:34cd5f9cec73e4dea570605e05cc9890090369b27f90aa7e05818ccccccfa506

Observation 4aeba49c-029e-42aa-bf11-55bf1c781dce · inbound

Search-Time Data Contamination cites this paper.

Search-Time Data Contamination Measuring short-form factuality in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T21:08:06.044958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:08:06.044958Z digest=sha256:2716bd63f563925a6be272512aac4d9a858f4e0916242a144f3921eadc4d66fc

Observation 2cc5c784-86cf-42bb-8fe8-8a4622da8dc3 · inbound

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents cites this paper.

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents Measuring short-form factuality in large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:18:58.169572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:18:58.169572Z digest=sha256:4d582c1c22fa77c14cedfea98dfd897c05072c97e7b86b38501462ff1965a1da

Observation c6bd90e3-7244-4700-9910-041f1c89fc39 · inbound

Hallucinations in medical devices cites this paper.

Hallucinations in medical devices Measuring short-form factuality in large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:17.413030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:20:17.413030Z digest=sha256:c98bce41708bc5db4475b9f06bc78a462d2c243bda97e8bc7ea69c97993571d9

Observation c9f21bba-2b91-4da7-a383-e50b111deb34 · inbound

UQ: Assessing Language Models on Unsolved Questions cites this paper.

UQ: Assessing Language Models on Unsolved Questions Measuring short-form factuality in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.840580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.840580Z digest=sha256:1d61af165b657af8d32c4462e371e90742f8623267f205bee2d58527c69514cc

Observation 2368bc24-3900-4dc0-b95f-49b04e3dff55 · inbound

Hermes 4 Technical Report cites this paper.

Hermes 4 Technical Report Measuring short-form factuality in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:54.645316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:32:54.645316Z digest=sha256:6461eeef0289fc3dbb9c7817bce2ec5b41a250a3db34944177b0f5b01daba73e

Observation 9f988073-9cb7-44ec-84ec-b22e5cb22be2 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring short-form factuality in large language models

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.083052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.083052Z digest=sha256:657b8592ffeb108498cff68c56a9dee8b06a536f5d39555d11fd7ca9a2979d41

Observation 798a97a3-3e8d-496f-93ff-410cba03e8bd · inbound

Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference cites this paper.

Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference Measuring short-form factuality in large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:03:02.053966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:03:02.053966Z digest=sha256:08af7c646d1ae2a022700326176eb1f1fc733e19bc7696c3b26d842815f59ee7

Observation e905bff6-e959-490e-b8ce-1316b2089016 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Measuring short-form factuality in large language models

Reference 184

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:43.663112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:43.663112Z digest=sha256:7b621b77088fec829709d50a0001e9fb27639dc610efa7c4895c6697cab8a8ac

Observation 4020a24b-881d-4602-b6f1-a03f5b034ae3 · inbound

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework cites this paper.

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework Measuring short-form factuality in large language models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:36:24.970158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T13:34:22.790199Z digest=sha256:e27f940cca9718ff87ea0538ca8826aae09393892d505b66c2b86a019f0469ee

Observation f5b3aea6-6b39-4bab-8574-fed944a3d09e · inbound

Evidence for Limited Metacognition in LLMs cites this paper.

Evidence for Limited Metacognition in LLMs Measuring short-form factuality in large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:51:23.756907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:51:23.756907Z digest=sha256:c40ff350e03498757ed4aa469d826e970fb4f899c7b805bce6e82eeec7d53282

Observation c1364c78-bb7a-499e-8e71-cc8a28e1d0b7 · inbound

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents cites this paper.

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents Measuring short-form factuality in large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:54.380028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:54.380028Z digest=sha256:cac1707159fef646f5fec0f93235605dd53975e4dcaf7e947e4b6c5b5a3156b6

Observation fbbefafa-2783-4855-b415-739d359bdbca · inbound

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations cites this paper.

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations Measuring short-form factuality in large language models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:56:24.448221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T12:54:01.717015Z digest=sha256:30f952b9a03c0702addf62d2ed2e9c792f6eca1ccecd9289f3f45d9a83d2c487

Observation 3ac914c7-b7e5-420e-b1ce-7faa668ab44e · inbound

Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity cites this paper.

Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity Measuring short-form factuality in large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T13:23:15.694081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:23:15.694081Z digest=sha256:12c7d3f8ebb17173d76918b3eb47ac76a2342620526ca5025de846fcb8e745d4

Observation 97876d2a-cfe0-490c-ac01-79b929618d3d · inbound

Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation cites this paper.

Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation Measuring short-form factuality in large language models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T09:21:10.034000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T09:18:48.453505Z digest=sha256:3b69a9461086aa97df40322d9ea3f7a02dcea81d0c4a5210c0a86b0a97ad87c0

Observation f511b76f-0bb3-42c2-b79c-fa37532d1217 · inbound

MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning cites this paper.

MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning Measuring short-form factuality in large language models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:00:34.635162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T00:57:25.902674Z digest=sha256:d336f9ee86521cdb12f8dfaf8af71e933a994efce844b85e03cd6f2d11274fda

Observation 93d143fb-9a73-43a2-99fa-750dbff1eeb8 · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Measuring short-form factuality in large language models

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:48:51.115403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:7c6373f2c4dcbe0d0740a75bf55d0d2a93b06a1324af9e338cba9d6e386c0b63

Observation 55bdf1dc-13d7-4401-86db-cdc67205a95f · inbound

FaithLens: Detecting and Explaining Faithfulness Hallucination cites this paper.

FaithLens: Detecting and Explaining Faithfulness Hallucination Measuring short-form factuality in large language models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T20:48:32.840198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T20:44:22.891896Z digest=sha256:985058dd0b35f9d9f3ead1a68392f444fbe6cc283859be659d8ac1a507ebe066

Observation a7d60c3f-986a-4584-8259-9f8bc4885468 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning Measuring short-form factuality in large language models

Reference 143

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:44.748564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:44.748564Z digest=sha256:15d88f6a617e4b669d9981b7219d1ae1c646deaa467b29732edfdb45229117f6

Observation 6b357cec-c250-49d0-888a-0fc47f7957ec · inbound

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection cites this paper.

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection Measuring short-form factuality in large language models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:40:42.368628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T06:39:17.785715Z digest=sha256:8e5eaca6eafd3e467535b25855d635d39de0a492e9d4fac47e5a6f7cfb4b8e32

Observation a297ef69-a840-4ca6-819d-3cb75e057804 · inbound

Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality cites this paper.

Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality Measuring short-form factuality in large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:47.218315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:47.218315Z digest=sha256:a2c283d355b351ae0c2a90b7e9fa73aca58fb53ee69344952dab533a16b8b5db

Observation 7113c071-720d-410b-98c3-0ac1ed7a4add · inbound

VeRO: A Harness for Agents to Optimize Agents cites this paper.

VeRO: A Harness for Agents to Optimize Agents Measuring short-form factuality in large language models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:06:30.922205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T19:01:50.547095Z digest=sha256:eeaad959a3db00e85c5049a62e381350a3e6995cbbcff2b9c5cf2e03e814ea90

Observation 386a46fc-88ef-475e-9678-b9d8aa9a06a8 · inbound

VeRO: A Harness for Agents to Optimize Agents cites this paper.

VeRO: A Harness for Agents to Optimize Agents Measuring short-form factuality in large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T20:45:12.868399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:45:12.868399Z digest=sha256:909914d68bf856e7b5d13407e07c19569d9e8df844b9e921900ff9ea05b1a5f0

Observation 1b68bec0-229f-4a02-9f4f-391effcf86e5 · inbound

Evaluating the Search Agent in a Parallel World cites this paper.

Evaluating the Search Agent in a Parallel World Measuring short-form factuality in large language models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:06:19.298161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T17:02:10.963839Z digest=sha256:2d5825a00c6d2eedda82b4a6cf3eeab4385e6fc5c2aa7a9983f7e44d069c0bcf

Observation 32cf9a80-4bcd-4f27-a22d-0d418fa8025e · inbound

CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering cites this paper.

CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering Measuring short-form factuality in large language models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T10:45:28.268236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T10:43:33.300825Z digest=sha256:093d1258ee3814d6417412b1f102f3f1bf2f8439852f558ce1d25260f201551e

Observation 788adcf4-e859-4442-9671-ee9ca90b2319 · inbound

CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering cites this paper.

CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering Measuring short-form factuality in large language models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T10:40:00.640012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:38:10.070691Z digest=sha256:c5e89d28c95249688a461b10db3652e2e24b5bff3deee37145d66fd8189643ab

Observation a2229f91-e9b7-4dba-96af-229625db583b · inbound

Causal Evidence that Language Models use Confidence to Drive Behavior cites this paper.

Causal Evidence that Language Models use Confidence to Drive Behavior Measuring short-form factuality in large language models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:44:05.701288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T09:43:05.524088Z digest=sha256:1ab5d93b9b0f6f90f6b5c188b5a8979a0c64ec3eb7bf45cef52198c00fbdf562

Observation 4b1d43f9-c86c-4653-a431-9f3582fddad2 · inbound

JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency cites this paper.

JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency Measuring short-form factuality in large language models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T19:26:38.134505Z digest=sha256:37d9a896a4f66910f0a513dce563ec608c24cfeabc9484bb04d1f95c75f46330

Observation add04dd0-fd88-49ac-be44-615bf3e0d692 · inbound

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence cites this paper.

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence Measuring short-form factuality in large language models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T19:48:13.133733Z digest=sha256:ad97a727d3f9a6cc039db266a54a7b48084c2b9309f6c7f6fd1a0dc0ef297415

Observation 6f1f69cb-644f-4225-a192-52cc1e7c0bac · inbound

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts cites this paper.

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts Measuring short-form factuality in large language models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:44:26.931344Z digest=sha256:c1b32ea444fe23e228d9e3955c8219dc67a149778ba0e3e8a1ab51d2bfc46336

Observation 229362ad-4062-43b4-8c5e-90a43240ac8c · inbound

A 4.5-s Quasiperiodic Spectral Oscillation in GRB 230307A: Evidence for Free Precession of a Post-Merger Magnetar? cites this paper.

A 4.5-s Quasiperiodic Spectral Oscillation in GRB 230307A: Evidence for Free Precession of a Post-Merger Magnetar? Measuring short-form factuality in large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T08:50:25.297894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:50:25.297894Z digest=sha256:9beafa4f9c351ae6ebbafd2a20a266943a3f944caf4fcb28c58c40ff71f5e328

Observation 5238890d-d710-452f-9bb8-ca9e4d8c3946 · inbound

WRAP++: Web discoveRy Amplified Pretraining cites this paper.

WRAP++: Web discoveRy Amplified Pretraining Measuring short-form factuality in large language models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:23:14.376066Z digest=sha256:cc771f319d3345e6ccbbbe30aeec4cd66654b35e20d42af3565d723873fe6232

Observation 723a4754-b539-4ee9-bc35-f07de6b59ffe · inbound

EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools cites this paper.

EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools Measuring short-form factuality in large language models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:36:45.297356Z digest=sha256:0b6db18a4f67d558b77f4e2f10087871a13b87e4ef29e1f8e954191403191424

Observation 2ae13732-9bc1-4e47-80d3-7f75f4dc0e1e · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts Measuring short-form factuality in large language models

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:54430d4c28c9ca1830b83fdba9c524ab836f701e408ff5e8f72328c8d75dfcef

Observation fde4d85a-3be9-45c8-b692-2c6466ec65b3 · inbound

NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data cites this paper.

NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data Measuring short-form factuality in large language models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:39:06.587762Z digest=sha256:1cdd0c8614f5d2ba1246016783b3b79c5f4e240a5da39cc3660ee4869d6c1a41

Observation ef92f343-9186-4aa2-9ca9-cc9c782cc1c9 · inbound

Evaluation of Agents under Simulated AI Marketplace Dynamics cites this paper.

Evaluation of Agents under Simulated AI Marketplace Dynamics Measuring short-form factuality in large language models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T12:33:04.984953Z digest=sha256:887f49709545209ae85c99dbf8e5d2ff9bad5c24ee3d11f2524ffd4713b890d3

Observation 028e7c5d-9756-45f9-a819-3c3b406f4482 · inbound

Purging the Gray Zone: Latent-Geometric Denoising for Precise Knowledge Boundary Awareness cites this paper.

Purging the Gray Zone: Latent-Geometric Denoising for Precise Knowledge Boundary Awareness Measuring short-form factuality in large language models

Reference 2

Resolution
malformed identifier
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:50:25.128014Z digest=sha256:889e8bc61632480449919d4853e4e8d2e1b596982f59c75db3acfbdfbedeabf7

Observation 9482e1b5-5561-4b94-a68c-53a3a6ac47a5 · inbound

Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization cites this paper.

Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization Measuring short-form factuality in large language models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T10:57:25.591595Z digest=sha256:ede562938b55819180baa9733b65ac812e2ea2823b02d1eda1db4330e39e4747

Observation f5489cfc-c2c2-4505-be68-a2ec1ca315ca · inbound

Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification cites this paper.

Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification Measuring short-form factuality in large language models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T06:29:24.974157Z digest=sha256:35ddff56f368fdc7958762e5a92e1513439596a2c3a65be9d0a0e32814051b01

Observation 5ee285da-fd08-480d-8e6a-cbff23dd7b43 · inbound

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts cites this paper.

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts Measuring short-form factuality in large language models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T05:52:28.822723Z digest=sha256:e61dfa7b7244a2ec81371aee1c8cc8e1123a56e305dd76e568b21b4ef6db7a27

Observation 5b4f3514-d817-4d72-941f-a328a3f31c3f · inbound

Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs cites this paper.

Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs Measuring short-form factuality in large language models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:26:21.545889Z digest=sha256:cdecce299b2b1b719aa61583a5afe1e4f12d3b34497159df329d700c15910083

Observation 8e7acf7a-549b-4526-87ad-78e0b6b4aa1a · inbound

Decomposing and Steering Functional Metacognition in Large Language Models cites this paper.

Decomposing and Steering Functional Metacognition in Large Language Models Measuring short-form factuality in large language models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:57:09.180530Z digest=sha256:bd3b1b6135a0c0a8a42a15e1310f451d1042f11c22e9d320bb79e1991c556229

Observation a57e9cde-7286-4fc0-ba53-d4e9438e3776 · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs Measuring short-form factuality in large language models

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:50:17.399580Z digest=sha256:9a475326b1064c095c50f29434d7403f65e971dcf7e4082ee8466ebfac90ede0

Observation a6656881-6328-40e3-83a1-43915dcbad09 · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs Measuring short-form factuality in large language models

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:20:32.494840Z digest=sha256:fdf89f0e645d0239b75d0283cdac32933f6a918b5e9cf4268d318f3132c5a38e

Observation 8b047e32-67c3-40f2-9d40-2476e6d911a7 · inbound

Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery cites this paper.

Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery Measuring short-form factuality in large language models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:55:22.596920Z digest=sha256:07431eb29564dfd163542367786be140ac7f37916256951a517bd99ac9c16846

Observation 0c08e022-3462-46bc-86c8-7584a815ff6a · inbound

Improving Multi-turn Dialogue Consistency with Self-Recall Thinking cites this paper.

Improving Multi-turn Dialogue Consistency with Self-Recall Thinking Measuring short-form factuality in large language models

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T20:25:02.538160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T20:22:11.252657Z digest=sha256:1669c14bc5f395283c587428f2ed0304075fbc1aa724dcf45edd6cc55af1b093

Observation 33c0abc2-b2b2-4c64-a5b4-055b5769488f · inbound

AgentStop: Terminating Local AI Agents Early to Save Energy in Consumer Devices cites this paper.

AgentStop: Terminating Local AI Agents Early to Save Energy in Consumer Devices Measuring short-form factuality in large language models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:37:39.619062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T16:36:02.233255Z digest=sha256:10de9f9b10af1b309aae243469eee5e28377f5f017d1b8e6bf513c20e7f59470