Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:45:50.219157Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 100 inbound Pith citation observations for arXiv:2411.04368.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:45:50.219157Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:09:31.122730Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
19 of 19 outbound references displayed
External citation measurements
12
pith, observed 2026-08-05T02:28:24.338817Z
Observation b073f0d9-5944-412f-9adf-8a89a7b45b32 · outbound
Measuring short-form factuality in large language models Do Language Models Know When They're Hallucinating References?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation db243db8-6c9c-4349-a301-7f9cb93a2cbb · outbound
Measuring short-form factuality in large language models Anthropic
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e7946d0d-d737-409a-9651-ab722df88e72 · outbound
Measuring short-form factuality in large language models Evaluating Hallucinations in Chinese Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 39e907eb-f1fa-420d-9054-080ae0e0ce06 · outbound
Measuring short-form factuality in large language models TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6799a419-3a33-45e1-9eb7-41b49cee82ca · outbound
Measuring short-form factuality in large language models Language Models (Mostly) Know What They Know
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f1c95610-5e4f-4f56-82b0-251e11647666 · outbound
Measuring short-form factuality in large language models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ed2ca68b-042e-41d5-bc2f-332a8fd42f93 · outbound
Measuring short-form factuality in large language models Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bde3cc3d-94ff-426a-89e5-9b645bf69cef · outbound
Measuring short-form factuality in large language models Kwiatkowski, J
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d3b46f2b-f045-40d8-9e11-be3075ccf31c · outbound
Measuring short-form factuality in large language models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ac5c9d88-26cb-4779-99f6-4ab53a4d63d8 · outbound
Measuring short-form factuality in large language models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 70495176-badf-4cfd-ac03-e572429f0681 · outbound
Measuring short-form factuality in large language models Teaching Models to Express Their Uncertainty in Words
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ae740448-1b5c-44fa-af53-df06ce47ebbf · outbound
Measuring short-form factuality in large language models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5b06c113-8476-44f4-8286-ce49947e43a6 · outbound
Measuring short-form factuality in large language models Hello gpt-4o, 2024 a
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation da15d8a0-8180-4e76-89d1-a37120a204da · outbound
Measuring short-form factuality in large language models Openai o1-mini, 2024 b
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c6c47803-e762-410b-a0ed-a6e74bc60ae2 · outbound
Measuring short-form factuality in large language models Learning to reason with llms, 2024 c
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1437c2c2-9997-49b4-80fa-c6512fe92af5 · outbound
Measuring short-form factuality in large language models FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9470a51c-6d82-44e1-9256-b44614e48517 · outbound
Measuring short-form factuality in large language models Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 06b1d0f2-7278-4747-b188-f7c739200a4b · outbound
Measuring short-form factuality in large language models Long-form factuality in large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f67ea904-82ad-45f6-96fe-888ce408dc48 · outbound
Measuring short-form factuality in large language models FELM: Benchmarking Factuality Evaluation of Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e34ef943-f0b5-4d86-bc6f-10cae8658b79 · inbound
O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson? Measuring short-form factuality in large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c931fb1-a8f3-46e2-9890-e16ce7c9656d · inbound
100% Elimination of Hallucinations on RAGTruth for GPT-4 and GPT-3.5 Turbo Measuring short-form factuality in large language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb769b5-1664-425b-950f-9cc7ed93c96b · inbound
Deliberative Alignment: Reasoning Enables Safer Language Models Measuring short-form factuality in large language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bdfc51e-a3d6-4f79-aa35-b1d2a30439c4 · inbound
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking Measuring short-form factuality in large language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ecd44b-6284-4d06-998a-a87c5b07c9d4 · inbound
Decoding Knowledge in Large Language Models: A Framework for Categorization and Comprehension Measuring short-form factuality in large language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 830fab3b-2009-431d-be43-48d6d4ded6b2 · inbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Measuring short-form factuality in large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0219d421-062f-42b2-9758-a7ad296a4264 · inbound
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time Measuring short-form factuality in large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250d4475-5baa-45c9-b35d-d1c508a22886 · inbound
Predictions as Surrogates: Revisiting Surrogate Outcomes in the Age of AI Measuring short-form factuality in large language models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea073248-6907-44e2-b96d-353a4f384ba4 · inbound
Humanity's Last Exam Measuring short-form factuality in large language models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9d1158f4-6348-4a56-97f5-b98797720e26 · inbound
Trading Inference-Time Compute for Adversarial Robustness Measuring short-form factuality in large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d63d76-0b38-4929-95cb-12c593631201 · inbound
LIMO: Less is More for Reasoning Measuring short-form factuality in large language models
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1bb2c226-877e-4409-9356-16042d1529cf · inbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Measuring short-form factuality in large language models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b9796e-604a-43ed-a377-36da011eea82 · inbound
Learning to Reason at the Frontier of Learnability Measuring short-form factuality in large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c873965e-5314-42be-9784-99015204d247 · inbound
LLM-Safety Evaluations Lack Robustness Measuring short-form factuality in large language models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5f317ab6-a637-4ac1-8003-7c7c8dea20fa · inbound
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results Measuring short-form factuality in large language models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 539bbfea-6fcc-4932-b7bf-33ec54fcd598 · inbound
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark Measuring short-form factuality in large language models
Reference 132
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10034c9-f11a-42d8-986e-814e5055f879 · inbound
aiXamine: Simplified LLM Safety and Security Measuring short-form factuality in large language models
Reference 128
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca7a9cbc-c633-41bf-8faf-dbd2ad02ca5e · inbound
HalluLens: LLM Hallucination Benchmark Measuring short-form factuality in large language models
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6231d726-a778-47b9-b3f3-b19517bcab23 · inbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Measuring short-form factuality in large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5aa205e-0465-4378-8d75-1f4719ce8f88 · inbound
Evaluating LLM Metrics Through Real-World Capabilities Measuring short-form factuality in large language models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bfd8c1c-7a42-490f-8262-aa9cc102b295 · inbound
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs Measuring short-form factuality in large language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3db9def-01cd-4792-bf47-2a938be442c0 · inbound
Phare: A Safety Probe for Large Language Models Measuring short-form factuality in large language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a38d3bc-a9c3-4402-a5b7-0b8712479fa4 · inbound
AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning Measuring short-form factuality in large language models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee2d8f7-c7d7-4872-8ffd-955bf72fc76f · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Measuring short-form factuality in large language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beeab439-929c-4b2b-aba6-664376dde5d1 · inbound
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use Measuring short-form factuality in large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29f35dad-8602-4e0c-91fc-4b05780a2a6c · inbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Measuring short-form factuality in large language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2769800b-4406-4c1e-8e46-04605919d1df · inbound
Adaptive Plan-Execute Framework for Smart Contract Security Auditing Measuring short-form factuality in large language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02815d34-6320-4292-86cb-a907077c71bd · inbound
Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs Measuring short-form factuality in large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22190d48-0f88-4f5e-bb12-f47a37c0b9ae · inbound
Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA Measuring short-form factuality in large language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acbd60b6-dba9-4d11-8652-d52550417321 · inbound
RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models Measuring short-form factuality in large language models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1abe4f0-9b5d-4aef-ba9a-c2fba24a4f3d · inbound
EvolveSearch: An Iterative Self-Evolving Search Agent Measuring short-form factuality in large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc6fa950-9580-4289-92bf-3011befb0c22 · inbound
Are Reasoning Models More Prone to Hallucination? Measuring short-form factuality in large language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa3b2187-3169-49b8-be3f-f2d630da372e · inbound
Reconsidering LLM Uncertainty Estimation Methods in the Wild Measuring short-form factuality in large language models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce2080a2-9371-4e44-8d81-267562a7ef10 · inbound
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Measuring short-form factuality in large language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e4bb2c-45b7-4e5d-8659-41979ecf42d5 · inbound
Quantifying Cross-Modality Memorization in Vision-Language Models Measuring short-form factuality in large language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e5ed4e-0cb3-4a94-a314-773e22ff5983 · inbound
Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems Measuring short-form factuality in large language models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a67319c4-9aba-48ee-a765-8ba08e47cab0 · inbound
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models Measuring short-form factuality in large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa04c82a-eca2-449a-a320-c53ea51779a6 · inbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Measuring short-form factuality in large language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 048aff7e-4691-4cbb-89fe-1d70bc06fde9 · inbound
Deep Research Agents: A Systematic Examination And Roadmap Measuring short-form factuality in large language models
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4089ce7a-91ae-4ff1-a2f0-0ea4f506acc8 · inbound
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know? Measuring short-form factuality in large language models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a25a484-6c09-4d7d-823c-3e1bdf29f7df · inbound
Jan-nano Technical Report Measuring short-form factuality in large language models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e091268f-41eb-4008-bc86-6a39117b640e · inbound
L0: Reinforcement Learning to Become General Agents Measuring short-form factuality in large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 587eb515-049d-4032-a808-c06af7a2e0ba · inbound
WebSailor: Navigating Super-human Reasoning for Web Agent Measuring short-form factuality in large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7e007b0c-0b86-4d42-94b7-84dd4f755e1e · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring short-form factuality in large language models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 949d3b7c-2feb-468e-ab0e-542558767ab6 · inbound
How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models Measuring short-form factuality in large language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fbcc871-6593-478b-b69d-5ef38f14d250 · inbound
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Measuring short-form factuality in large language models
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c8a7c98-2fc2-429a-84e7-18a47b7acd64 · inbound
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Measuring short-form factuality in large language models
Reference 199
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea26196-b430-4ef0-bc43-bcd375bec63c · inbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Measuring short-form factuality in large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 469bc87b-11d9-4711-8630-64820a91cc2e · inbound
RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation Measuring short-form factuality in large language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b01546-0239-435a-a083-f9976411cbf7 · inbound
Kimi K2: Open Agentic Intelligence Measuring short-form factuality in large language models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d62732f-eedf-4992-90d0-6df84b72ea84 · inbound
Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Measuring short-form factuality in large language models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c39e5ee-23a0-4633-9785-eaa81364b574 · inbound
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution Measuring short-form factuality in large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6ec15e-a4ab-4d7a-8305-06bccb3321e2 · inbound
Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty Measuring short-form factuality in large language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 914e4dc4-33d4-4f3d-9017-87ea06285be2 · inbound
SSRL: Self-Search Reinforcement Learning Measuring short-form factuality in large language models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aeba49c-029e-42aa-bf11-55bf1c781dce · inbound
Search-Time Data Contamination Measuring short-form factuality in large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc5c784-86cf-42bb-8fe8-8a4622da8dc3 · inbound
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents Measuring short-form factuality in large language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6bd90e3-7244-4700-9910-041f1c89fc39 · inbound
Hallucinations in medical devices Measuring short-form factuality in large language models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f21bba-2b91-4da7-a383-e50b111deb34 · inbound
UQ: Assessing Language Models on Unsolved Questions Measuring short-form factuality in large language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2368bc24-3900-4dc0-b95f-49b04e3dff55 · inbound
Hermes 4 Technical Report Measuring short-form factuality in large language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f988073-9cb7-44ec-84ec-b22e5cb22be2 · inbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring short-form factuality in large language models
Reference 231
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798a97a3-3e8d-496f-93ff-410cba03e8bd · inbound
Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference Measuring short-form factuality in large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e905bff6-e959-490e-b8ce-1316b2089016 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Measuring short-form factuality in large language models
Reference 184
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4020a24b-881d-4602-b6f1-a03f5b034ae3 · inbound
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework Measuring short-form factuality in large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5b3aea6-6b39-4bab-8574-fed944a3d09e · inbound
Evidence for Limited Metacognition in LLMs Measuring short-form factuality in large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1364c78-bb7a-499e-8e71-cc8a28e1d0b7 · inbound
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents Measuring short-form factuality in large language models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbbefafa-2783-4855-b415-739d359bdbca · inbound
ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations Measuring short-form factuality in large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3ac914c7-b7e5-420e-b1ce-7faa668ab44e · inbound
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity Measuring short-form factuality in large language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97876d2a-cfe0-490c-ac01-79b929618d3d · inbound
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation Measuring short-form factuality in large language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f511b76f-0bb3-42c2-b79c-fa37532d1217 · inbound
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning Measuring short-form factuality in large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 93d143fb-9a73-43a2-99fa-750dbff1eeb8 · inbound
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Measuring short-form factuality in large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 55bdf1dc-13d7-4401-86db-cdc67205a95f · inbound
FaithLens: Detecting and Explaining Faithfulness Hallucination Measuring short-form factuality in large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a7d60c3f-986a-4584-8259-9f8bc4885468 · inbound
Toward Efficient Agents: Memory, Tool learning, and Planning Measuring short-form factuality in large language models
Reference 143
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b357cec-c250-49d0-888a-0fc47f7957ec · inbound
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection Measuring short-form factuality in large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a297ef69-a840-4ca6-819d-3cb75e057804 · inbound
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality Measuring short-form factuality in large language models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7113c071-720d-410b-98c3-0ac1ed7a4add · inbound
VeRO: A Harness for Agents to Optimize Agents Measuring short-form factuality in large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 386a46fc-88ef-475e-9678-b9d8aa9a06a8 · inbound
VeRO: A Harness for Agents to Optimize Agents Measuring short-form factuality in large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b68bec0-229f-4a02-9f4f-391effcf86e5 · inbound
Evaluating the Search Agent in a Parallel World Measuring short-form factuality in large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 32cf9a80-4bcd-4f27-a22d-0d418fa8025e · inbound
CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering Measuring short-form factuality in large language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 788adcf4-e859-4442-9671-ee9ca90b2319 · inbound
CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering Measuring short-form factuality in large language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a2229f91-e9b7-4dba-96af-229625db583b · inbound
Causal Evidence that Language Models use Confidence to Drive Behavior Measuring short-form factuality in large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4b1d43f9-c86c-4653-a431-9f3582fddad2 · inbound
JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency Measuring short-form factuality in large language models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation add04dd0-fd88-49ac-be44-615bf3e0d692 · inbound
BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence Measuring short-form factuality in large language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6f1f69cb-644f-4225-a192-52cc1e7c0bac · inbound
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts Measuring short-form factuality in large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 229362ad-4062-43b4-8c5e-90a43240ac8c · inbound
A 4.5-s Quasiperiodic Spectral Oscillation in GRB 230307A: Evidence for Free Precession of a Post-Merger Magnetar? Measuring short-form factuality in large language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5238890d-d710-452f-9bb8-ca9e4d8c3946 · inbound
WRAP++: Web discoveRy Amplified Pretraining Measuring short-form factuality in large language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 723a4754-b539-4ee9-bc35-f07de6b59ffe · inbound
EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools Measuring short-form factuality in large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2ae13732-9bc1-4e47-80d3-7f75f4dc0e1e · inbound
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts Measuring short-form factuality in large language models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fde4d85a-3be9-45c8-b692-2c6466ec65b3 · inbound
NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data Measuring short-form factuality in large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ef92f343-9186-4aa2-9ca9-cc9c782cc1c9 · inbound
Evaluation of Agents under Simulated AI Marketplace Dynamics Measuring short-form factuality in large language models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 028e7c5d-9756-45f9-a819-3c3b406f4482 · inbound
Purging the Gray Zone: Latent-Geometric Denoising for Precise Knowledge Boundary Awareness Measuring short-form factuality in large language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9482e1b5-5561-4b94-a68c-53a3a6ac47a5 · inbound
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization Measuring short-form factuality in large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5489cfc-c2c2-4505-be68-a2ec1ca315ca · inbound
Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification Measuring short-form factuality in large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ee285da-fd08-480d-8e6a-cbff23dd7b43 · inbound
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts Measuring short-form factuality in large language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5b4f3514-d817-4d72-941f-a328a3f31c3f · inbound
Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs Measuring short-form factuality in large language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8e7acf7a-549b-4526-87ad-78e0b6b4aa1a · inbound
Decomposing and Steering Functional Metacognition in Large Language Models Measuring short-form factuality in large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a57e9cde-7286-4fc0-ba53-d4e9438e3776 · inbound
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs Measuring short-form factuality in large language models
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a6656881-6328-40e3-83a1-43915dcbad09 · inbound
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs Measuring short-form factuality in large language models
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8b047e32-67c3-40f2-9d40-2476e6d911a7 · inbound
Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery Measuring short-form factuality in large language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0c08e022-3462-46bc-86c8-7584a815ff6a · inbound
Improving Multi-turn Dialogue Consistency with Self-Recall Thinking Measuring short-form factuality in large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 33c0abc2-b2b2-4c64-a5b4-055b5769488f · inbound
AgentStop: Terminating Local AI Agents Early to Save Energy in Consumer Devices Measuring short-form factuality in large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.