Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:57:11.951271Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 25 inbound Pith citation observations for arXiv:2501.03200.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:57:11.951271Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:58:39.544437Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T14:07:06.683844Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c763b47e-0141-4132-ac44-55f29c0dd1d9 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22c7df76-4787-4494-b360-eea5eba58d55 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input The Claude 3 model family: Opus , Sonnet , Haiku , 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e69daaf4-9c60-4610-85ae-ec3efb997447 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 238f238c-0534-48c7-82b3-8fbde7c33425 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input BooookScore: A systematic exploration of book-length summarization in the era of LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71d7b26e-de38-45c0-99da-ff2674ec62b5 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input TrueTeacher: Learning Factual Consistency Evaluation with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84dee49c-0c99-404f-9077-36132d935343 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Introducing gemini 2.0: our new ai model for the agentic era, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2a32f5d8-fe4f-4144-a0a0-3b0065aff8ca · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Gemini: A Family of Highly Capable Multimodal Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1812b83c-93ce-422d-9256-43a92dc3e0da · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input TRUE: Re-evaluating Factual Consistency Evaluation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a65f861e-3fb6-4f46-bb1e-a86baea32953 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input FactAlign: Long-form Factuality Alignment of Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c0a98ade-5822-4545-9d8b-1cf5bfbe2d25 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input CoverBench: A Challenging Benchmark for Complex Claim Verification
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2dbb18-77d2-4014-af30-9289e0cf2f93 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation df790ae7-e49c-4bea-bce0-36c21edbd9c9 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input One Thousand and One Pairs: A "novel" challenge for long-context language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e59414a3-384f-4234-b60f-033eb067ac0c · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input FABLES: Evaluating faithfulness and content selection in book-length summarization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a4cdf4-5af0-424a-a6d9-638144dcd43f · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac12361-d93c-4b7f-aaf8-2212b0eb57ee · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5511854b-c4a0-4b0d-a537-82d544f5cb66 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2ad42837-cbbe-4133-bebc-a7cd4ddc5028 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5255417e-c9b9-4616-ba7e-2d9188ee10a8 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db1e092-abc2-479a-862d-5483187577e6 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc75e232-d566-4e48-979b-866ee58a9553 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Learning to reason with LLMs , 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3ae4f0f3-e226-4a67-8a3a-c0427cbaa142 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Fact-Checking Complex Claims with Program-Guided Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19822b6c-de57-4309-8687-e321c1394bd2 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Ramprasad and B
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf15b7a3-1029-4889-a531-b3721cc696bd · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1b0a7c-8ba6-46ee-adbf-a06ab0f03231 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Rashkin, V
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f5978bf4-fef7-41c7-87c0-fe6233318f02 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f5e0ed0-f702-4cf2-8a19-85588c069e2b · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Data Contamination Report from the 2024 CONDA Shared Task
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaf1367d-41a2-472f-b0ae-68a9d2853f69 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa05910-0983-4e0a-8c94-669ef8a68a54 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ccc6744-7ac2-43f7-be71-f473dc05d537 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6488028e-f276-4c55-a34e-933129a7f92a · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Hallucination evaluation model (revision 7437011), 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a2d13a3f-f9e0-4878-9701-26724eda8e4f · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Self-Preference Bias in LLM-as-a-Judge
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 830fab3b-2009-431d-be43-48d6d4ded6b2 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Measuring short-form factuality in large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435aeff9-4beb-40f2-abd4-64dea21e67fe · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Long-form factuality in large language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd320f1-405f-4b18-9105-c8a107d6d4c6 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 201db2a0-a18d-471e-92ca-2fae0bd12284 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1001b3f-22d4-4d86-a73b-b926abfb9af1 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc46dfc-5a6f-4083-abaa-774052874713 · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 042ffe57-bf12-439a-bbd9-c38be20c2a3e · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Zheng, W.-L
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 077cf027-d212-48de-8c96-b6a39e749b9f · outbound
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24580c8-082e-4d85-a512-7d38e063c98a · inbound
ConSens: Assessing context grounding in open-book question answering The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7fb3a0c-6d0d-452f-93d8-23d5fd6eb86c · inbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89b41f11-e77d-41f8-acbe-59c284267aae · inbound
Evaluating LLM Metrics Through Real-World Capabilities The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0ac765-2a37-4dd1-9ad1-3c987da4f752 · inbound
LIFEBench: Evaluating Length Instruction Following in Large Language Models The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 934ffd7b-4604-4766-802b-a12986f35cfa · inbound
Towards Large Reasoning Models for Agriculture The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b91d1bc-198e-4186-a41a-9aa4a291001e · inbound
How Does Response Length Affect Long-Form Factuality The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d36c2852-fd98-410d-88c0-c664a7268218 · inbound
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5da0c4c-1331-4f40-83ee-f911d5eaefbd · inbound
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a1c2b675-ad1d-4723-be6b-8c543833809a · inbound
Kimi K2: Open Agentic Intelligence The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d493dbc0-17da-45ef-b449-36fe420c341e · inbound
StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7fa6e91-bedc-4052-8660-5c9a2e750fb5 · inbound
A Neurosymbolic Approach to Natural Language Formalization and Verification The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5b6b225-781e-4335-b577-abe861a02ca9 · inbound
Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2d6b555-5801-431f-b0a6-01762320e696 · inbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7bfa351-3ceb-4b6a-989b-528fd538842d · inbound
TRACE: Tourism Recommendation with Accountable Citation Evidence The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2435c00e-d144-4f5a-a439-7cd2a045fa22 · inbound
Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e77a223b-c5ed-443f-9b4e-32f3f381f802 · inbound
OpenAaaS: An Open Agent-as-a-Service Framework for Distributed Materials-Informatics Research The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ecb089f6-a38f-41f2-8914-f6be9cf60cce · inbound
Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ef784c11-4a7d-4d8f-a5d4-1a663846657d · inbound
Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eae57c51-4248-4afd-a4b8-e559119de2f6 · inbound
APEX: Automated Prompt Engineering eXpert with Dynamic Data Selection The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 444372f9-e854-470e-a635-d21d90470b77 · inbound
WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 253e777e-ccbd-492d-9550-19d9751f0ef7 · inbound
ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 440c548c-f473-46dc-b9e1-df1700d1a73e · inbound
Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c7092507-80ac-406d-bd35-f09354737179 · inbound
Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 563b240d-ab6f-4515-b819-eb27b438c232 · inbound
AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6f5f99d-e6f3-4fe3-b737-588b8ce3664d · inbound
Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.