Pith. sign in

Paper Citation Record · LEDGER

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

As of 22 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 25 inbound Pith citation observations for arXiv:2501.03200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03200 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:57:11.951271Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:58:39.544437Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T14:07:06.683844Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c763b47e-0141-4132-ac44-55f29c0dd1d9 · outbound

This paper cites GPT-4 Technical Report.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.752194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.752194Z digest=sha256:7de25e34349468fef10e6a6830adfb1e24e79aeeb208d20a4a256027d4a8991e

Observation 22c7df76-4787-4494-b360-eea5eba58d55 · outbound

This paper cites The Claude 3 model family: Opus , Sonnet , Haiku , 2024.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input The Claude 3 model family: Opus , Sonnet , Haiku , 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:12.786085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.758139Z digest=sha256:5e3eb70c1e61f2af766ad4ab0601ed0f52d25a52a6912a545f175f18c98c42db

Observation e69daaf4-9c60-4610-85ae-ec3efb997447 · outbound

This paper cites LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.763682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.763682Z digest=sha256:d1b5450841be6c3e64c7217e591c5a6e02ce9be6bc6d86c7965dea6cfc5e7403

Observation 238f238c-0534-48c7-82b3-8fbde7c33425 · outbound

This paper cites BooookScore: A systematic exploration of book-length summarization in the era of LLMs.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input BooookScore: A systematic exploration of book-length summarization in the era of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.769165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.769165Z digest=sha256:18fe89549b031fec637e97106deaf5b41892d776ac3c7981f6df602f513d6917

Observation 71d7b26e-de38-45c0-99da-ff2674ec62b5 · outbound

This paper cites TrueTeacher: Learning Factual Consistency Evaluation with Large Language Models.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input TrueTeacher: Learning Factual Consistency Evaluation with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.774378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.774378Z digest=sha256:fb50c9bebe595bd4a7133b6f99406533516a499f77c510fbeff7ed849b4bd3e1

Observation 84dee49c-0c99-404f-9077-36132d935343 · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era, 2024.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Introducing gemini 2.0: our new ai model for the agentic era, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:12.767474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.780723Z digest=sha256:d8daecec541f69c6ae74c5391d17393efcf33fe64655ab82122bfb448e6d8e4f

Observation 2a32f5d8-fe4f-4144-a0a0-3b0065aff8ca · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Gemini: A Family of Highly Capable Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.787147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.787147Z digest=sha256:cfeda5382f8d608ec33c4e65df2ee6ec7249579606aaa4ad30993cb4103c33aa

Observation 1812b83c-93ce-422d-9256-43a92dc3e0da · outbound

This paper cites TRUE: Re-evaluating Factual Consistency Evaluation.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input TRUE: Re-evaluating Factual Consistency Evaluation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.793041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.793041Z digest=sha256:a1f6f92f779e92e7098409cfa4fc5ae9fd7e02b47808c880651c24bbd3b9d6c7

Observation a65f861e-3fb6-4f46-bb1e-a86baea32953 · outbound

This paper cites FactAlign: Long-form Factuality Alignment of Large Language Models.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input FactAlign: Long-form Factuality Alignment of Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:57:12.518666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.798483Z digest=sha256:f8ae2be2c70b597af10e51877d863a893b06068298ac0708addb6bbed4eb2ad9

Observation c0a98ade-5822-4545-9d8b-1cf5bfbe2d25 · outbound

This paper cites CoverBench: A Challenging Benchmark for Complex Claim Verification.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input CoverBench: A Challenging Benchmark for Complex Claim Verification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.804116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.804116Z digest=sha256:08584844c02ca2af0ef7a8eb842465ac6feb620ce54f1ad28b73e91b0e54e624

Observation fc2dbb18-77d2-4014-af30-9289e0cf2f93 · outbound

This paper cites an unresolved cited work.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:57:12.750203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.809569Z digest=sha256:9e76e148e9a8d98c7dd6259a249e685f64eb8b466e325173549e78f7c1cf05b0

Observation df790ae7-e49c-4bea-bce0-36c21edbd9c9 · outbound

This paper cites One Thousand and One Pairs: A "novel" challenge for long-context language models.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input One Thousand and One Pairs: A "novel" challenge for long-context language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.814838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.814838Z digest=sha256:625f30d490d49a6378d30fada83e12aac817a69cdc81d6e17a54a84d55ad097c

Observation e59414a3-384f-4234-b60f-033eb067ac0c · outbound

This paper cites FABLES: Evaluating faithfulness and content selection in book-length summarization.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input FABLES: Evaluating faithfulness and content selection in book-length summarization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.819492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.819492Z digest=sha256:d0f7871db128bacfdfa4ab5cceb41184277a2fe791710b98c830c8ae648f0b36

Observation 77a4cdf4-5af0-424a-a6d9-638144dcd43f · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.824253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.824253Z digest=sha256:be6a3f670b40f2f35b5a7f38288822328d572b39fd607e7ee2605ee1b1fbfb9f

Observation 1ac12361-d93c-4b7f-aaf8-2212b0eb57ee · outbound

This paper cites an unresolved cited work.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:57:12.731343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.828822Z digest=sha256:2f477dc0718b7047cb6c926d86f513a1f82fabccc3fdf354930ce7c6b04d3575

Observation 5511854b-c4a0-4b0d-a537-82d544f5cb66 · outbound

This paper cites an unresolved cited work.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-10T21:57:12.419995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.833610Z digest=sha256:416820e802a71b637b62c0ff6ed226cf30499d128ff58f68f4e36313bb3690ca

Observation 2ad42837-cbbe-4133-bebc-a7cd4ddc5028 · outbound

This paper cites HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.838528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.838528Z digest=sha256:8825269f89bfd45a04535373c4265728319d9211754bb3925f10cee2b2c7dd9f

Observation 5255417e-c9b9-4616-ba7e-2d9188ee10a8 · outbound

This paper cites an unresolved cited work.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.844230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.844230Z digest=sha256:169c0c383631cce3c6809c1f9ec591949da03766b33cf5a7a0ae711cbb3dddcc

Observation 5db1e092-abc2-479a-862d-5483187577e6 · outbound

This paper cites FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.849479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.849479Z digest=sha256:11086372e501c4543c453b42028e565ddbc73c14d97249e24d1eeb6979e9ac9b

Observation bc75e232-d566-4e48-979b-866ee58a9553 · outbound

This paper cites Learning to reason with LLMs , 2024.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Learning to reason with LLMs , 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:12.711791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.854786Z digest=sha256:521fade86879aeaab5d8f3df9b09d45126c52e83c5670e717fd04f9dd4882865

Observation 3ae4f0f3-e226-4a67-8a3a-c0427cbaa142 · outbound

This paper cites Fact-Checking Complex Claims with Program-Guided Reasoning.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Fact-Checking Complex Claims with Program-Guided Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.859626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.859626Z digest=sha256:0ffd4f5fe1b5689503ab9d4379722badc9b3230c337e70b1ad9955b1481d6d84

Observation 19822b6c-de57-4309-8687-e321c1394bd2 · outbound

This paper cites Ramprasad and B.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Ramprasad and B

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.864849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.864849Z digest=sha256:36f24dc307419f1a57a0aeffeb9bc16b2ecfdbc57cc4a68b0e906591b98cdd03

Observation bf15b7a3-1029-4889-a531-b3721cc696bd · outbound

This paper cites Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.869602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.869602Z digest=sha256:3a9d10fb14de29566071e55d0bb7b3afd09a5492036515f66410d26bae895bca

Observation 6a1b0a7c-8ba6-46ee-adbf-a06ab0f03231 · outbound

This paper cites Rashkin, V.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Rashkin, V

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:12.693866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.874818Z digest=sha256:e54149efe5185f7797e78cd7cf429c002e28fd9533a7157e5f4d7ccd57fb927c

Observation f5978bf4-fef7-41c7-87c0-fe6233318f02 · outbound

This paper cites Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.880011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.880011Z digest=sha256:9747c5168ac9da3c93fd9d99c7993298ca6b65419eadc9c280d6a138ba9c8a58

Observation 5f5e0ed0-f702-4cf2-8a19-85588c069e2b · outbound

This paper cites Data Contamination Report from the 2024 CONDA Shared Task.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Data Contamination Report from the 2024 CONDA Shared Task

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.885384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.885384Z digest=sha256:0b70e442e660fa223c97c46414e5694deb98e93f87fdccf39708137597b41de0

Observation aaf1367d-41a2-472f-b0ae-68a9d2853f69 · outbound

This paper cites VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.891838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.891838Z digest=sha256:02797f39d7a2dabefb00916294d974c252fbbb8e0422388d0a5d113da1c5f1df

Observation 4aa05910-0983-4e0a-8c94-669ef8a68a54 · outbound

This paper cites Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.897120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.897120Z digest=sha256:cc6b35f419e35e494b95c69cb5d101b738a7ed787ffabc8bde0c6cb383efb996

Observation 0ccc6744-7ac2-43f7-be71-f473dc05d537 · outbound

This paper cites MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.902334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.902334Z digest=sha256:8de49007a1dfbc6c8de0b8dfd360e58782740e369d94c515ace81ce95553adb9

Observation 6488028e-f276-4c55-a34e-933129a7f92a · outbound

This paper cites Hallucination evaluation model (revision 7437011), 2024.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Hallucination evaluation model (revision 7437011), 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:12.676609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.907270Z digest=sha256:4a1c9bd84fcbcfd8b790acf9aed8c6d979e2b8a190dc7046dfc440c295bb2433

Observation a2d13a3f-f9e0-4878-9701-26724eda8e4f · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Self-Preference Bias in LLM-as-a-Judge

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.911664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.911664Z digest=sha256:774e7f14d583b14c7836f3e0f86d904be3edb7b64076ffe97c429b066efabc7f

Observation 830fab3b-2009-431d-be43-48d6d4ded6b2 · outbound

This paper cites Measuring short-form factuality in large language models.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Measuring short-form factuality in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.916094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.916094Z digest=sha256:7360705c9f1ef710c28379e53a87e250462f293e2b2338b6cd8e62efdf4634f5

Observation 435aeff9-4beb-40f2-abd4-64dea21e67fe · outbound

This paper cites Long-form factuality in large language models.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Long-form factuality in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.920951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.920951Z digest=sha256:3a88eb5a93b42882016de6d62ff5fcd9e976c41af6f00c9112e5ae7227e34b6a

Observation 9dd320f1-405f-4b18-9105-c8a107d6d4c6 · outbound

This paper cites an unresolved cited work.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.925367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.925367Z digest=sha256:0a2318a9c227fa80e8f96cff541d9813b3be698040f7308c8fb000997b7822a3

Observation 201db2a0-a18d-471e-92ca-2fae0bd12284 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.930851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.930851Z digest=sha256:e8f7e8c1999a0f9da508bcf2db6762d64816b2acc00f974d1e2a983a587d900f

Observation b1001b3f-22d4-4d86-a73b-b926abfb9af1 · outbound

This paper cites WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.935950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.935950Z digest=sha256:1d3790abd0162c892a5ec28f3930b140cb85b1749847388d252f79e13a02bd82

Observation 4bc46dfc-5a6f-4083-abaa-774052874713 · outbound

This paper cites an unresolved cited work.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:57:12.658972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.941189Z digest=sha256:0be433df5fd6b1c654c41c7fcc6f6fb5f5d2a0f3fc4827d857bc24cd69a6ca36

Observation 042ffe57-bf12-439a-bbd9-c38be20c2a3e · outbound

This paper cites Zheng, W.-L.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Zheng, W.-L

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:12.641918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T21:57:11.946531Z digest=sha256:f4d0fec6928d4117d6840a03d2117b4182faff2fb3945ec0f5e63db95b7e51ef

Observation 077cf027-d212-48de-8c96-b6a39e749b9f · outbound

This paper cites HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.951271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.951271Z digest=sha256:d5020892cc1b760e32a23eb04e0dbb507cb1609607969dea8222187ef4dc3469

Pith citing papers

Observation f24580c8-082e-4d85-a512-7d38e063c98a · inbound

ConSens: Assessing context grounding in open-book question answering cites this paper.

ConSens: Assessing context grounding in open-book question answering The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:58:39.544437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:58:39.544437Z digest=sha256:2b790512294060e2dc114776cf3513b41db276e6c72b38a283b75df610cd08ff

Observation f7fb3a0c-6d0d-452f-93d8-23d5fd6eb86c · inbound

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation cites this paper.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.238393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.238393Z digest=sha256:0b26f56137c7620b2ff7868bb4570a9cbf08851ea9784b6b2e00760bc0dd1d8d

Observation 89b41f11-e77d-41f8-acbe-59c284267aae · inbound

Evaluating LLM Metrics Through Real-World Capabilities cites this paper.

Evaluating LLM Metrics Through Real-World Capabilities The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T22:03:20.453455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:03:20.453455Z digest=sha256:ad16f6e4f1e3aa7df19378b1d0f156a38403581ef6c5fce8a1a0b6d39f025b90

Observation 9b0ac765-2a37-4dd1-9ad1-3c987da4f752 · inbound

LIFEBench: Evaluating Length Instruction Following in Large Language Models cites this paper.

LIFEBench: Evaluating Length Instruction Following in Large Language Models The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:04.558589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:08:04.558589Z digest=sha256:5eae03bc3b0bf5a73c7ac2c58dfcfb74f1caa89848eeeb778f8d562c6e870291

Observation 934ffd7b-4604-4766-802b-a12986f35cfa · inbound

Towards Large Reasoning Models for Agriculture cites this paper.

Towards Large Reasoning Models for Agriculture The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:43.463461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:43.463461Z digest=sha256:63bb13c6b791a334771e053e268b7745e20feb0b2356954e558dbad9cc11e2af

Observation 9b91d1bc-198e-4186-a41a-9aa4a291001e · inbound

How Does Response Length Affect Long-Form Factuality cites this paper.

How Does Response Length Affect Long-Form Factuality The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:38.496979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:38.496979Z digest=sha256:8d38c7a8c9d043286f2ae3f543db38f2561ada350188c524caa1ea286dea2788

Observation d36c2852-fd98-410d-88c0-c664a7268218 · inbound

DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs cites this paper.

DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:14.770034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:12:14.770034Z digest=sha256:1aa3b93326f6689d0ef9b085dff53bc82f18e52331072ae0db95782b40743e02

Observation b5da0c4c-1331-4f40-83ee-f911d5eaefbd · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:52:07.726059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:36bb810aa9514e96a7b6b3feee3244f53777d1724e65f7e249bb48962a87a8b6

Observation a1c2b675-ad1d-4723-be6b-8c543833809a · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:49:28.148937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:3f02c67335898466fcf812d909380fa8c284c32f401fc6407a8672ddcc6fd03d

Observation d493dbc0-17da-45ef-b449-36fe420c341e · inbound

StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation cites this paper.

StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:57:05.522192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:57:05.522192Z digest=sha256:62c4a51c1c4d943945adb98f45315a22ff80b4570035a31bffe3e5b1d4efb10b

Observation e7fa6e91-bedc-4052-8660-5c9a2e750fb5 · inbound

A Neurosymbolic Approach to Natural Language Formalization and Verification cites this paper.

A Neurosymbolic Approach to Natural Language Formalization and Verification The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T22:49:01.771793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:49:01.771793Z digest=sha256:2f5fe6d641a32784654d10c8f8510d3fa5b110a7fa271a39c10670f32c875ede

Observation d5b6b225-781e-4335-b577-abe861a02ca9 · inbound

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models cites this paper.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:48.472543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:48.472543Z digest=sha256:d2b47d8229c0e9021a8452599ee37534eb5e7cefe32a999fca2deaac209fd503

Observation a2d6b555-5801-431f-b0a6-01762320e696 · inbound

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation cites this paper.

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:49.587447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:49.587447Z digest=sha256:9429cf279dcc2e83305c7db618a8e8e2550eb292df09c98104c0ec4708d2df71

Observation c7bfa351-3ceb-4b6a-989b-528fd538842d · inbound

TRACE: Tourism Recommendation with Accountable Citation Evidence cites this paper.

TRACE: Tourism Recommendation with Accountable Citation Evidence The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:30:57.973089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:27:26.082625Z digest=sha256:308a7b0051ec092a03683a2baa50a0841d5d531955e5a2496e7477c7c39848b3

Observation 2435c00e-d144-4f5a-a439-7cd2a045fa22 · inbound

Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights cites this paper.

Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:37:03.597207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:34:43.464899Z digest=sha256:579add10596193e00e19586f37b45c4a4a43d87cf43f0547590f34d32919adf9

Observation e77a223b-c5ed-443f-9b4e-32f3f381f802 · inbound

OpenAaaS: An Open Agent-as-a-Service Framework for Distributed Materials-Informatics Research cites this paper.

OpenAaaS: An Open Agent-as-a-Service Framework for Distributed Materials-Informatics Research The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:52:35.650120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T18:50:27.219978Z digest=sha256:74fc7c55315f530ce1451efbe5d9df5805e5b13b964267366d0e8292dfb20995

Observation ecb089f6-a38f-41f2-8914-f6be9cf60cce · inbound

Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification cites this paper.

Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:50.649461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T18:19:53.820922Z digest=sha256:f3632a62eba8c439a1d5b262669a8e10f28b5723d8142fcd50ca3217f71ac781

Observation ef784c11-4a7d-4d8f-a5d4-1a663846657d · inbound

Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring cites this paper.

Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:37:09.586047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T22:29:01.722453Z digest=sha256:564714f320bb2e2bc3260e2660668512ab3c99c49900f998817991cb9f33464f

Observation eae57c51-4248-4afd-a4b8-e559119de2f6 · inbound

APEX: Automated Prompt Engineering eXpert with Dynamic Data Selection cites this paper.

APEX: Automated Prompt Engineering eXpert with Dynamic Data Selection The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:47:42.008859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T13:02:34.623221Z digest=sha256:38b3d419fa56f8a6bbbc2fafa0956b89e948bdb16313e21a9b8ec906c406e72b

Observation 444372f9-e854-470e-a635-d21d90470b77 · inbound

WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning cites this paper.

WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.663170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T09:47:04.122464Z digest=sha256:2d391c4ecf5f03327bfa9db034261bdbb0c84712c3b37eccb0f059d1d5e22657

Observation 253e777e-ccbd-492d-9550-19d9751f0ef7 · inbound

ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence cites this paper.

ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:49:58.064081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T01:15:53.269926Z digest=sha256:ace63ddce2b3765146756392df93777e686dd6b5d6e94b7d63ac5590f5da24d4

Observation 440c548c-f473-46dc-b9e1-df1700d1a73e · inbound

Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search cites this paper.

Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:29:51.712710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T05:10:59.825246Z digest=sha256:daeba51fe7e27103d892c853ee1c13dd03d6bbce34a6db9c44c5106007ed4d11

Observation c7092507-80ac-406d-bd35-f09354737179 · inbound

Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator cites this paper.

Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T14:07:06.685896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T14:06:11.853570Z digest=sha256:2ead69c1ab3d122155099732ffb0eb7270242187f3945bd764be2202d05cf171

Observation 563b240d-ab6f-4515-b819-eb27b438c232 · inbound

AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge cites this paper.

AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T13:42:51.563514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T13:42:51.563514Z digest=sha256:1fe4b3a8a6bde9a778809d3470d9e7be1e62dcba5e1f5e57351d8d37433d4f08

Observation e6f5f99d-e6f3-4fe3-b737-588b8ce3664d · inbound

Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge cites this paper.

Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T00:53:53.961682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:53:53.961682Z digest=sha256:0c05b88f90ada6a380c9c0a391d4ed774d6c75aaf1c059746294dd570d7b1ea6